# Abhinandan > Inference engineer. RL post-training on reasoning models and the serving stack that runs them: vLLM, SGLang, TensorRT-LLM, INT8 KV caching. Site: https://abhinandan.one Index (short): https://abhinandan.one/llms.txt Profile page: https://abhinandan.one/about ## At a glance - Name: Abhinandan - Roles: Inference Engineer, Reinforcement Learning Engineer, ML Engineer - Currently: Founding Software Engineer at Browzer (since 2025-09) - Looking for: full-time inference and ML engineering roles. Available immediately. - Based in: India (IST, UTC+5:30), works remotely - Education: B.Tech Computer Science, Dr. APJ Abdul Kalam Technical University, 2026 - Recognition: International Youth Math Challenge — Gold Honour; Top 1% TypeScript Engineer Globally (Algora); Amazon ML Summer School 2025 ## FAQ **Who is Abhinandan?** An inference engineer based in India, working remotely. He does RL post-training on reasoning models and builds the inference systems that serve them — vLLM, SGLang, TensorRT-LLM, INT8 KV caching, and the surrounding serving stack. **What is Abhinandan working on now?** Founding Software Engineer at Browzer since September 2025, building a browser-agent runtime with a streaming ReAct loop and zero-LLM replay. In the open, he maintains RolloutCore, an RL rollout control plane for vLLM, and HiQCache, an INT8 host-tier KV cache for SGLang HiCache. **What roles is Abhinandan looking for?** Full-time inference and ML engineering roles: LLM serving and runtime work, KV cache and quantization, and RL post-training for reasoning models. He can join immediately. **What is Abhinandan's inference stack?** Inference: vLLM, SGLang, TensorRT-LLM, NCCL, quantization. ML: PyTorch, Hugging Face, LoRA/PEFT, Qdrant. Engineering: Python, C++, TypeScript, FastAPI, Node.js, Redis, PostgreSQL, AWS, GCP, Docker, GitHub Actions. **Has Abhinandan contributed to open source?** Yes. He has merged pull requests and open work across vLLM, SGLang, HeroUI, Mooncake, ai-dynamo and others. Sixteen merged pull requests to HeroUI (then NextUI) led to a personal offer from the CEO. The full, live list is at abhinandan.one/contributions. **What is Abhinandan's education?** B.Tech in Computer Science from Dr. APJ Abdul Kalam Technical University, 2026. **What recognition has Abhinandan received?** International Youth Math Challenge Gold Honour, Top 1% TypeScript Engineer Globally on Algora, and selection for Amazon ML Summer School 2025. **How do I contact Abhinandan?** Email abhinandan@abhinandan.one, or reach him on GitHub at awesome-pro, LinkedIn at abhibuilds, or X at abhibuilds. # Projects ## RolloutCore - tag: RL Rollout Runtime - date: 2026-09-23 - stack: PyTorch, NCCL, vLLM, RL - keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput - GitHub: https://github.com/awesome-pro/rolloutcore - case study: none Versioned RL rollout runtime for vLLM with NCCL hot weight updates, version-pure rollouts, and cache-coherent transitions. ## HiQCache - tag: Quantized SGLang HiCache - date: 2026-09-23 - stack: PyTorch, SGLang, Quantization - keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput - GitHub: https://github.com/awesome-pro/hiqcache - case study: none INT8 hierarchical KV caching for SGLang HiCache, delivering 43.75% lower host KV memory and 1.78× more L2 cache capacity for Qwen3-8B. ## MiniServe - tag: LLM Inference Runtime - date: 2026-09-23 - stack: Python 3.12, PyTorch, Paged KV Cache, Continuous Batching, Prefix Caching - keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput - GitHub: https://github.com/awesome-pro/miniserve - case study: none a from-scratch LLM serving runtime with continuous batching, chunked prefill, block-based KV management, preemption, and prefix caching. ## AgentFlow-Pro - tag: Agentic RL Research - date: 2026-05-20 - stack: PyTorch, TRL, DAPO, PRM, PEFT / LoRA, Qwen3-8B, Ollama, FastMCP - keywords: AgentFlow, process reward model, DAPO, reinforcement learning, agentic reasoning, Qwen3, LoRA fine-tuning, GPQA, AIME, RLHF - GitHub: https://github.com/awesome-pro/agentflow-pro - Original AgentFlow paper: https://arxiv.org/abs/2510.05592 - case study: https://abhinandan.one/agentflow-pro Process-supervised RL that made an 8B model reason better, and the gain carried over to a domain it never trained on. ## SmartMemo - tag: Semantic LLM Cache - date: 2026-03-15 - stack: FAISS, SentenceTransformers, PyTorch, SQLite, Pydantic - keywords: semantic cache, LLM cache, FAISS, embeddings, sentence-transformers, pairwise classifier, semantic equivalence, prompt caching, MiniLM - GitHub: https://github.com/awesome-pro/smartmemo - PyPI: https://pypi.org/project/smartmemo/ - case study: https://abhinandan.one/smartmemo A semantic cache for LLM agents where a trained classifier decides if a cached answer is safe to reuse, instead of raw cosine similarity. ## Orchflow - tag: Agent Orchestration Framework - date: 2026-02-20 - stack: AsyncIO, LiteLLM, Pydantic - keywords: multi-agent, orchestration, asyncio, pipeline, LiteLLM, workflow framework, checkpoint resume, dependency-free - GitHub: https://github.com/awesome-pro/orchflow - PyPI: https://pypi.org/project/orchflow/ - case study: https://abhinandan.one/orchflow A dependency-free Python framework for readable multi-agent pipelines: sequential, parallel, conditional, and resumable flows. ## agenteval - tag: LLM Evaluation Tooling - date: 2026-01-25 - stack: AsyncIO, OpenAI SDK, Anthropic SDK, LangChain, Typer - keywords: agent evaluation, LLM testing, pass rate, CI, behavioral assertions, agent tracing, non-deterministic testing, regression tracking - GitHub: https://github.com/awesome-pro/agenteval - PyPI: https://pypi.org/project/agenteval-py/ - case study: https://abhinandan.one/agenteval Behavioral eval for agents: replaces brittle exact-match asserts with repeated-run pass-rate scoring for CI gates. # Artifacts ## 05 — Measuring Adaptive Speculative Decoding in SGLang with EAGLE3 on Triton - canonical: https://abhinandan.one/artifacts/measuring-adaptive-speculative-decoding-in-sglang - published: 2026-09-24T08:00:00+00:00 - updated: 2026-09-24T10:14:34.144+00:00 - Harness Github: https://github.com/awesome-pro/heterospec I do not know if it is my poor luck or my mind, but after failing to measure `spec-decoding` myself, I realized that SGLang has already merged adaptive speculative decoding. Technically, the controller takes the accepted draft count from every request in the batch, averages them into one number, smooths it into an EMA, and picks K from that. and that one sentence is the whole design, so let me say it in slower words. imagine 32 requests running together. and A few of them are chewing through very predictable text and accepting four or five drafts every round. A few others are writing something open ended and accepting almost nothing. SGLang adds all of that up, divides by 32, and gets one number. That number picks the depth for everybody in the batch. actually I kept coming back to that. If every request has its own curve of how likely position k is to be accepted, why throw the curve away and keep a mean? The answer we give ourselves is that the curve belongs to the request, so the mean is a fair summary of something stable. It feels good & easy when you say it like that. I wanted to know if it is actually true. - If the curve belongs to the request, why does one number work at all? - What if the curve itself moves when you change K? - And how would you even know, if you never measured it? These are some questions I had while reading `adaptive_spec_params.py`. I did not want to argue about them, so I built a harness and rented one A6000 to go find out. Let me first tell you what the harness actually does. It runs the same workload at a fixed K, across a few concurrency levels. K = 1, 3, 5, 7 and concurrency = 1, 8, 32. That is 16 runs, plus one run with speculation switched off completely, so I have an honest floor to compare everything against. Two controls I want to mention, because without them the whole grid is junk. Every server gets warmed up before I start timing it. The first requests into a fresh server are slower, Triton is still compiling kernels, the allocator is still growing, and at the few percent level we are chasing here, that warmup would show up as signal. And every server runs with `--disable-radix-cache`. Without that flag, the grid reuses the same seeded prompt pool across concurrencies on one server, so later runs would look cheaper than earlier ones purely because the cache got warm. I am not studying prefix caching, so I do not want it quietly moving my batch size axis. Now came the tough part. SGLang already hands you `spec_correct_drafts_histogram` for every request. It tells you how many rounds accepted 0 drafts, how many accepted 1, and so on up the chain. From that you can compute P(position k was accepted), which is exactly the curve I wanted to look at. But that number only means something if K was held still. Here is the problem. Say a request ran under adaptive K, and for most of its life the controller gave it depth 2. It accepts something at position 1, accepts at position 2, and then it finishes. Positions 3 through 7 are sitting there as zeros in that histogram. But the draft at position 3 was never proposed. Nobody asked the model that question. If I average those zeros into my curve, I am not measuring rejection, I am measuring absence, and those are two very different things. And you cannot repair it afterwards. The information is not in the file. It was never written. So I patched the fork to also emit an iteration level trace. Every decode iteration writes a line with the active K and the rids that were in that batch. Now I can bound K per request from the trace instead of assuming it, and I can throw away any request whose curve I do not trust. It writes about 31,000 of those lines across the three adaptive runs, and it makes the adaptive capture readable for the first time. Then there is the oracle. Four levels of it, all scored on the deepest static capture, and all with a train and test split. That split is not decoration. A scalar policy with fine enough bins can put every batch in its own bin, memorise it, and report a gap of exactly zero no matter what the truth is. If you do not hold batches back, you will measure your own overfitting and call it a result. But before I looked at any gap number, I wrote a gate. Let me explain why, because the gate is the part I am happiest with. Every level of that oracle needs to know what a request would have done at a depth it never actually ran. That is a guess. It is a fair guess only if the acceptance curve does not depend on K, which is the assumption everybody in this space quietly makes and almost nobody checks. So I wrote the rule down first. Run the same workload, same seed, same batching, at K = 1, 3, 5, 7. Compare P(A at k) across those depths. If the curve moves when you move K, then the gap does not get reported at all. I wrote that down before I had a single number from the GPU. Which is the only time you can write a rule like that honestly. It failed. ![The gate firing on the GPU host](https://mehqnbdhdsgejnmdvvjc.supabase.co/storage/v1/object/public/artifact-images/artifacts/measuring-adaptive-speculative-decoding/1790237864493-9usqa73t.png) P(accepted at position 1) is 0.440 at K=1. At K=3, 5 and 7 it is 0.376, 0.360 and 0.362. The same split shows up at every concurrency I ran, so it is not one noisy cell. And K=3, 5 and 7 agree with each other well inside the tolerance, which tells me the wobble is not random at all. It sits at the shallowest depth, and it is consistent. ![Per-position acceptance depends on K](https://mehqnbdhdsgejnmdvvjc.supabase.co/storage/v1/object/public/artifact-images/artifacts/measuring-adaptive-speculative-decoding/1790237409458-e7l4e11m.png) I did not argue with my own rule. The gap is not reportable, and I am not going to show you a number I cannot defend. But the same session answered two other questions, and honestly those are worth more to me than the gap would have been. The first one is that speculation lost to not speculating. At concurrency 32, the run with speculation switched off did 1049 tokens per second. The best speculative cell did 753. Turning speculation off made the server about 39% faster. ![Throughput by K and concurrency](https://mehqnbdhdsgejnmdvvjc.supabase.co/storage/v1/object/public/artifact-images/artifacts/measuring-adaptive-speculative-decoding/1790237409459-twtikr3l.png) That surprised me, and then it made sense. The draft model is not broken, the arithmetic just does not work out. The number that matters is how many tokens you get for each position you push through the target model. At K=0 that number is 1.000, one token per position, by definition. At K=7 it is 0.239. So you are shoving 8 positions through a forward pass to collect 1.9 tokens at the end of it. Every extra draft position costs close to a full pass at this batch size, because the GPU is already busy, and the acceptance rate is nowhere near high enough to pay that back. The second thing is that the heterogeneity I was actually hunting for is real. I just could not use my oracle to price it. ![Per-class acceptance](https://mehqnbdhdsgejnmdvvjc.supabase.co/storage/v1/object/public/artifact-images/artifacts/measuring-adaptive-speculative-decoding/1790237409460-tggi0aqo.png) Same batch, same concurrency, same everything. The repetitive prompts hit 1.66 accepted drafts per round at K=7. The open ended ones sit at 0.50. That is a 3.3x spread inside a single batch, which is exactly the thing that made me think a single mean was a bad idea. The premise survived. What did not survive is the tool I built to measure it. Before I trusted any of those numbers, I ran one correctness check, and I want to put it here because it bounds everything above it. At temperature 0, speculative decoding has to produce the same tokens as plain decoding. That is the deal. So I compared the completion length for every prompt across every K. 768 prompts at concurrency 32, and every single one matched exactly. The pipeline is lossless. So the throughput numbers are a real cost outcome and not a symptom of something being broken underneath. Two bugs still ate real time, and I think they are the most useful part of this whole thing. The first one I still think about. Every measured request came back as HTTP 500. Not some of them. Every one. And warmup passed every single time, on the same server, with the same prompts, hitting the same endpoint. So I stared at that for a while. What is different between warmup and the real run? Same model, same temperature, same max tokens. The only difference was that my warmup code does not send a seed, and my measured code does. It turns out the field in this SGLang is called `sampling_seed`, not `seed`. The server builds `SamplingParams(**kwargs)` from whatever you put in `sampling_params`, and it does not filter out keys it does not recognise. So my `seed` raised a TypeError on the server, and all the client ever saw was an empty 500 with no body. The actual error was sitting in the server log the whole time. I found it by reading the log instead of guessing, which is a lesson I had to learn twice in this project. The second one was an exit code. The launcher died with `-9`, and if you have ever debugged a process on a rented box, `-9` means the OOM killer and nothing else. That is what I assumed. I went looking at memory. It was not memory at all. SGLang had killed its own process tree after a `ValueError` in the draft model's context length check, and the launcher died as a side effect of that cleanup. The real error was three lines above the kill message in the log. So both bugs had the same shape. The thing that killed me was not the thing that was reported. The part that bothers me most is that over 400 tests passed on my laptop while the real server was rejecting every single request. My mock server accepted whatever payload I handed it, so nothing local ever checked whether I was speaking the same language as the real thing. It does not accept that anymore. If this project left one useful thing behind, it is probably that. I stopped here. The gate said the method cannot answer the question, and I would rather write that down than publish a number I cannot stand behind. If I run this again, the first thing I change is the attention backend. SGLang's own EAGLE3 tests run flashinfer. I ran triton, because it kept flashinfer off the attention path and saved me pulling a wheel. I never wrote down what that convenience cost me, and that was careless. The verify pass attends over K+1 positions at once, which is exactly the kind of shape a backend handles well or badly. So my cost numbers describe EAGLE3 on triton, and not EAGLE3. That is the honest boundary of what I measured here. Which is the artifact for me. Not "adaptive K does not work." I wanted to know what one scalar throws away, and I found out the harder way, that the curve it summarises does not hold still long enough for anyone to compare it across K. --- ## 04 — Building a Mini Inference Engine from Scratch - canonical: https://abhinandan.one/artifacts/building-a-mini-inference-engine-from-scratch - published: 2026-09-21T10:38:43.277+00:00 - updated: 2026-09-23T11:25:27.735+00:00 - demo video: https://youtu.be/0-gmk9ND6yc - GitHub: https://github.com/awesome-pro/miniserve - YT Series: https://www.youtube.com/playlist?list=PLIi0G24WvbAA Most of us know that LLMs are auto-regressive, i.e., they generate token one by one. We also know that they have a concept of KV caching and can process multiple requests together, requests of different sizes, producing different amounts of output tokens. Reading these concepts feels good & easy because we can just say they work like this or they work that, but how does it actually happen? - How do they produce tokens one by one? - How do they decide that this is the next token? - How does KVCache actually get implemented? - How do to process batching? These are some questions which I have had in mind for a long time. And I wanted to realize them by implementing them myself. So in this artifact, we are going to build a mini inference system from scratch, which will help us answer these questions from the first principles :) So lets get started, but before that, I already created a proper [YouTube Series](https://www.youtube.com/playlist?list=PLIi0G24WvbAA) on building the LLM inference engine from scratch. So, if you want a visual explanation and if you understand the concepts better by videos, I hope going through this series will be much helpful for you to visualise and understand the whole implementation. Or you may follow your own custom combination, its all upto you. Ok, let us first understand what actually happens when we provide the prompt or input tags to the model.So first step, that your token goes through: you gave a prompt to ChatGPT, Grok, Cloude, or whatever LLM of your choice, assuming that it is an auto-regressive LLM. The first step it goes through is the **tokenizer**, which converts input text into token IDs and matrices. It cannot process text like us, so it needs the input to be present in the form of numbers. So to make it understandable for the LLM, the input gets converted into token IDs. If you want to know more or understand, you may go to the [OpenAI tokenizer](https://platform.openai.com/tokenizer) to get a real feel of how it works in GPT. And token ID is basically referred to the tokens present in the vocabulary set of the model. So essentially, we have converted the text into tokens which the model can understand. After this step, we convert the tag tokens to vector embeddings. What does that mean? Vector is basically a mathematical entity having features. And in case of a model, it is simply converting into a row all columns of numbers. And this collection of all the input vectors is something we represent by **`X`**. > will be completing it very soon. have something urgent on plate at the moment --- ## 03 — Building a Semantic Cache that Does Not Trust Cosine - canonical: https://abhinandan.one/artifacts/building-a-semantic-cache-that-does-not-trust-cosine - published: 2026-07-03T07:37:17.016+00:00 - updated: 2026-09-23T08:34:02.344+00:00 - demo video: https://youtu.be/UoHwsRx7J-I - GitHub: https://github.com/awesome-pro/smartmemo - PyPi: https://pypi.org/project/smartmemo Usually semantic caches for LLMs work on embeddings and cosine similarity. You take the new prompt, turn it into a vector, find the closest old prompt, and if the score is high enough you reuse the old answer. Reading this feels good & easy because we can just say similar prompts should have similar answers, but how true is that actually? - What if two prompts use the same words and talk about the same object, but ask for the opposite thing? - What if "approve the refund" and "deny the refund" sit next to each other in embedding space? - Who decides that a high cosine score means it is safe to reuse the cached response? These are some questions which I had in mind when I started building SmartMemo. Cosine is good at finding neighbours. It is not the same thing as semantic equivalence. And I wanted to realize that difference by implementing it myself. Ok, let us first understand what actually goes wrong. You gave a prompt to your agent. The cache embeds it. Somewhere in the store there is an old prompt that looks almost the same. The cosine is 0.92, threshold was 0.90, so it is a hit. The old answer comes back. The problem is that the two prompts were never the same request. One was approve, one was deny. Same customer, same refund, opposite action. The vectors do not care about that, because the words around the action are shared and the embedding model collapses them. So cosine can select candidates. It should not be the final judge. What SmartMemo does is split that into two steps. First step, same as everyone else. The prompt is converted into an embedding. A vector is basically a row of numbers that tries to hold the meaning. That vector goes into FAISS and we take the nearest cached prompts. These are only candidates. Nothing is returned yet. Second step is the part I actually cared about. A small classifier looks at the pair: this new prompt and that old prompt. It has been trained to say whether they are asking for the same thing, or they just look similar. If the score is high enough, we reuse the cached response. If not, we call the LLM, store the new prompt with its answer, and move on. The classifier that ships with it is a small MLP on top of MiniLM embeddings. I trained it on a lot of prompt pairs across a few domains, including the hard ones, the same-object opposite-action kind, and also negated versions of the same request. On a small gold set I kept aside, cosine at the same recall was at 0.53 precision. The classifier went to 0.83. Same recall, fewer wrong hits. That +30 is the whole reason this project exists. It is still not magic. On some high-stakes medical or legal opposite-action pairs it is better than cosine, but it still gets a few wrong. A generic classifier cannot know your domain perfectly. So there is a feedback path. If a hit was bad, you can report it. There is also an optional implicit version: if the same prompt comes again quickly after a cache hit, treat that as a signal that the earlier answer was not useful. Those pairs can be exported and used later to train a classifier for your own domain. The thing you actually call is one async function, `get_or_call`. You pass the prompt and the function that would call the LLM. You get back the response, whether it was a hit, and the classifier score. Storage is just SQLite. Nothing distributed. It is meant to sit inside one agent process and not pretend to be a cluster. That is basically it. Embedding search to find likely matches. A classifier to decide if they are actually the same thing. And a way to learn when they were not. Repo: [github.com/awesome-pro/smartmemo](https://github.com/awesome-pro/smartmemo) --- ## 02 — Running a Health Voice: All Private on Your Mac - canonical: https://abhinandan.one/artifacts/running-a-health-voice-all-private-on-your-mac - published: 2026-06-26T18:19:07.916+00:00 - updated: 2026-09-23T08:25:36.27+00:00 - demo video: https://youtu.be/KEieS276Rrw - GitHub: https://github.com/awesome-pro/health_voice - Case Study: https://www.aidaddy.tech/learn/16-case-studies/13-voice-ai-healthcare There is this promise that a doctor can just talk, and a note will write itself. It feels good & easy when you say it like that. Talk, note, done. But the patient's voice has to go somewhere, and most of these products send it to a cloud box you are supposed to trust. That part never sat right with me. You are in a room with a person. Their audio leaving that room is not a small detail. So I picked up the "Voice AI Assistant for Healthcare" case study from ombharatiya's system design guide. I did not want another slide with boxes and arrows. I wanted to know, if I actually sit and build this, how much of it is real. Press record, talk, and at the end a signed note is sitting in a real FHIR server. That was the itch. A few questions I had: - Can the audio stay on my machine and still become a proper note? - If the case study says real time, what does that even mean on a laptop? - Who is allowed to invent a sentence here, and who is not? The rule I gave myself: the voice never leaves. Transcription, who is speaking, medical terms, all of that has to happen locally. The only thing that can go out is a finished transcript, and even that comes back as a draft. A human has to look at it, change it, put a name on it. Otherwise I am just making another box to trust. Ok, let us walk through it the way it actually runs, not the way the diagram looks. You open the console. First you record about ten seconds of the nurse. That is just to make a voiceprint, so later I can tell who is who. Then the encounter starts. The browser takes the mic, squashes it down to 16kHz mono, and pushes it over a WebSocket to FastAPI. Silence is the first problem. If I send every quiet frame into Whisper, Whisper starts making things up. So silero VAD cuts the stream into utterances. Only speech goes forward. Then there are two Whisper models, and this is where I stopped pretending. The case study talks about real time as if the model just keeps up. On this Mac, large-v3-turbo needs around 2 seconds before it even starts, every single call, even for a short clip. There is no trick around that locally. Their numbers assume a datacenter GPU. Mine does not have one. So base.en talks first, fast and messy, just so the screen feels alive. The big model comes behind it and writes the line I actually keep. I did not like this at first. It felt like cheating. Then I realised the other option was to lie about latency, and I liked that even less. Each finished line gets a speaker label from the voiceprint, and a biomedical NER pass for symptoms, meds, vitals, the usual. I also run three checks that are just code. Safety. Spoken self-correction. Later, did the note forget something that was said. I did not want a model doing those. If something is wrong I want to point at the exact phrase, not at a score. When you hit stop, the whole speaker-labeled transcript goes to a SOAP endpoint. What is SOAP here? Basically the shape of a clinical note. Subjective, objective, assessment, plan. gpt-5.5 fills that shape from the transcript, with a strict JSON schema, and I tell it in the prompt: do not invent a fact. It comes back as a draft. Someone edits it. Types their name. Hits approve. Only then I build a FHIR bundle myself, in code, and POST it to a local HAPI server, and write one boring line in an audit file. FHIR is just the hospital's way of storing the record. I am not going to let a model be creative at the door of the medical record. That idea made me uncomfortable from day one. A few moments that stayed with me. I first put the SOAP call inside the WebSocket, right when the encounter ended. Felt tidy. Then I turned on pyannote to refine speakers. On CPU that can sit there for 30, 40, 90 seconds. The frontend gave up waiting. The note never arrived. No crash. Just an empty screen, and me staring at it thinking the model had failed, when it was my own wiring. I pulled SOAP out into `POST /soap`. Suddenly I could also regenerate a note from an old transcript without recording again. That bug annoyed me and then I was glad it happened. The one that still makes me laugh is Whisper on silence. Near-quiet room, and it starts speaking YouTube. "Thanks for watching." "Please subscribe." "Consult a qualified healthcare professional." One afternoon it wrote "love love love" sixteen times. The NER, very sincerely, marked "love" as a lab value. Then the completeness check got angry that "love" was missing from the note. I sat there looking at this thing I had built, a clinical pipeline arguing about love. The fix was ugly and I like that it is ugly. Drop the high no-speech segments. Drop the one-word-repeated nonsense. Regex for the stock phrases. If the utterance is weak, do not even send it to NER. Real speech stays. Garbage does not get a second chance to become a vital. Safety I refused to outsource. Regex. If it fires I can show you the words. A false alert is a glance. A miss is the actual failure. Red banner, no FHIR push, someone has to check a box. I do not want to argue with a model about whether a sentence was dangerous. Even the SOAP call had a small stupid moment. I send temperature because that is what I have always sent. gpt-5.5 says no, reasoning models only want the default. So I catch that complaint and retry without it. Nothing deep. Just the kind of thing that ruins a demo at 2am if you assume the old API still behaves. I keep coming back to the same feeling. Listening can be a model. Drafting can be a model. Deciding that this note is true enough to file cannot be. That last step needs a name on it. That is Health Voice for me. Not a product pitch. A way to sit with that discomfort and see how far I could take it on one machine, without sending the room away to someone else. --- ## 01 — Building a RAG based Search System for an Enterprise - canonical: https://abhinandan.one/artifacts/building-a-rag-based-search-system-for-an-enterprise - published: 2026-06-25T18:19:07.916+00:00 - updated: 2026-09-23T08:31:13.386+00:00 - demo video: https://youtu.be/00RRYFfVe_Y - B_GitHub: https://github.com/awesome-pro/enterprise_rag_backend - F_GitHub: https://github.com/awesome-pro/enterprise_rag_frontend There will be only 0.000001% people in tech who have not heard the word `RAG`. Embed a few documents, find the close ones, paste them into a prompt. The answer comes back. It feels good & easy. You can say it works like this or it works like that. Then I think about the same thing inside a company, and that feeling does not survive. Same question, two people. They are not allowed to see the same files. If I retrieve first and hide later, the model has already looked. That sat wrong with me. Vector search also misses the boring exact things, error codes, project names, acronyms. And the closest chunk is often not the one that actually answers you. And none of this matters if the model just invents a sentence and puts a citation under it. So I took a published system design case study. It was written for 5000 users and 500k documents. I did not have that. I had a laptop and free tiers. I still wanted every stage they had. Smaller rooms, same house. I needed to feel the whole path, not a slide of it. Questions I kept circling: - How do you make sure a document above someone's clearance never even enters the search? - Why do we pretend cosine is relevance? - If I show a live pipeline on the right side of the screen, will I still believe my own architecture? Ok, let us walk through what a question actually does here. It first hits some guardrails. Prompt injection, PII, the obvious stuff. If there is chat history, an LLM rewrites the follow-up into a standalone question. Only then I look at cache. Exact cache first, because that one is just a hash of role plus question, no embedding, no drama. Same person, same words, same answer, costs nothing. If that misses, I embed the question once. One vector, reused everywhere after that. Semantic cache is next, still namespaced by role, so one team's answer cannot leak into another team's mouth. The case study said use 0.95 as the similarity cutoff. With this embedding model, real paraphrases were landing around 0.71. 0.95 almost never fired. I sat with that for a bit. The number on the page was confident. The number in my logs was not. I moved it to 0.90 and stopped pretending their threshold was a law. Then the search. Two of them, in parallel. Qdrant for meaning. Elasticsearch for the exact ugly strings that vectors forget. Both filters see the user's access levels before they run. Not after. A document above your clearance is not retrieved, not reranked, not shown to Claude. It does not exist for you. That was the part I cared about more than fancy fusion. The two lists get fused with reciprocal rank fusion, 0.7 on vector, 0.3 on keyword. What does that mean, basically? Both rankings get a vote, closer to the top counts more. Then a cross encoder actually reads the question against the top 50 and keeps 10. This is the slow part. I left it slow on purpose and put the latency on the screen. Hiding it would have made the demo prettier and the system less honest. Claude writes the answer only from those 10, with citations. Output guardrails check that a citation points at a real chunk. If it does not, it does not ship. Every query is logged. The UI is a split. Chat on the left. On the right you can watch the path: role, access levels, each stage, how the reranker reordered things, whether cache saved you. I built that right side because I did not want this to be a story I tell. I wanted to look at it while it was happening. One bug ate a day and I still think about it. The UI was blank. Network tab was full. Events were arriving. I blamed the backend, then the model, then the network. It was my parser. The stream used CRLF, carriage return plus newline, twice between frames. I split on newline newline. They never matched. The whole stream went in. Zero events came out. One line to fix it. The feeling was stupid in a very specific way. Data was right there and my screen was still empty. For the public URL I use Claude Haiku, rate limit per IP, cap the day, because I do not want a link that quietly burns money. The reranker needs torch and a big local model, so the public build skips it and lives on fusion order. The demo I record has the reranker on. I do not love that split. It is the honest one I could afford. I keep returning to the same discomfort. A notebook RAG feels finished when the answer sounds right. Inside a company the answer sounding right is the last thing, not the first. Who was allowed to see the file. Did we find the error code or only a cousin of it. Did the model point at a chunk that exists. If those are not true, the sentence does not matter. That is this artifact for me. Not "I built RAG." I wanted to sit with the parts the demo skips, at a size I could hold, and see which of them still hurt when they were real. --- # Contact - Email: mailto:abhinandan@abhinandan.one - GitHub: https://github.com/awesome-pro - LinkedIn: https://linkedin.com/in/abhibuilds - X: https://x.com/abhibuilds - YouTube: https://youtube.com/@0xAbhinandan