Abhinandan
I do RL post-training on reasoning models, & build the inference systems that serve them.
i like working when the world is sleeping. my work cycle is generally 12pm to 4am. And I usually write my thoughts in my artifacts .
I live the part after the launch day. When the reasoning breaks, when the cost skyrockets, when the first traffic hits - all the similar thrills :)
Artifacts
- 15
Understanding Attention From Scratch
Oct 202621viewsI think if you have ever spent even a single day learning about current models, then attention is most probably the first term you may have heard. And when it comes to the…
https://abhinandan.one/artifacts/understanding-attention-from-scratch · serial 15 · published 2026-10-06T17:30:00.202+00:00 · updated 2026-10-06T17:42:39.065+00:00 - 14
Prefill & Decode made Crystal Clear
Oct 202666viewsThese two terms are very common. And it's very possible that you may have heard them before, or maybe you even know what they mean. But I feel that even though earlier I knew…
https://abhinandan.one/artifacts/prefill-decode-made-crystal-clear · serial 14 · published 2026-10-02T13:10:55.521+00:00 · updated 2026-10-02T13:58:38.825+00:00 - 05
I do not know if this was my poor luck or me being late, but while trying to understand and measure speculative decoding myself, I realized that SGLang had already merged adaptive…
https://abhinandan.one/artifacts/measuring-adaptive-speculative-decoding-in-sglang · serial 05 · published 2026-09-24T08:00:00+00:00 · updated 2026-10-03T11:46:38.499+00:00
Open source
Projects
Versioned RL rollout runtime for vLLM with NCCL hot weight updates, version-pure rollouts, and cache-coherent transitions.
keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput
date: 2026-09-23 · case study: none
INT8 hierarchical KV caching for SGLang HiCache, delivering 43.75% lower host KV memory and 1.78× more L2 cache capacity for Qwen3-8B.
keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput
date: 2026-09-23 · case study: none
a from-scratch LLM serving runtime with continuous batching, chunked prefill, block-based KV management, preemption, and prefix caching.
keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput
date: 2026-09-23 · case study: none
Process-supervised RL that made an 8B model reason better, and the gain carried over to a domain it never trained on.
keywords: AgentFlow, process reward model, DAPO, reinforcement learning, agentic reasoning, Qwen3, LoRA fine-tuning, GPQA, AIME, RLHF
date: 2026-05-20 · case study: https://abhinandan.one/agentflow-pro
A semantic cache for LLM agents where a trained classifier decides if a cached answer is safe to reuse, instead of raw cosine similarity.
keywords: semantic cache, LLM cache, FAISS, embeddings, sentence-transformers, pairwise classifier, semantic equivalence, prompt caching, MiniLM
date: 2026-03-15 · case study: https://abhinandan.one/smartmemo
A dependency-free Python framework for readable multi-agent pipelines: sequential, parallel, conditional, and resumable flows.
keywords: multi-agent, orchestration, asyncio, pipeline, LiteLLM, workflow framework, checkpoint resume, dependency-free
date: 2026-02-20 · case study: https://abhinandan.one/orchflow
Behavioral eval for agents: replaces brittle exact-match asserts with repeated-run pass-rate scoring for CI gates.
keywords: agent evaluation, LLM testing, pass rate, CI, behavioral assertions, agent tracing, non-deterministic testing, regression tracking
date: 2026-01-25 · case study: https://abhinandan.one/agenteval
Work
Amazon ML Summer School 2025 · IYMC Gold Honour · Top 1% TypeScript Engineer, Algora