abhinandan

Abhinandan

I do RL post-training on reasoning models, & build the inference systems that serve them.

i like working when the world is sleeping. my work cycle is generally 12pm to 4am. And I usually write my thoughts in my artifacts .

I live the part after the launch day. When the reasoning breaks, when the cost skyrockets, when the first traffic hits - all the similar thrills :)

Artifacts

  1. 15

    I think if you have ever spent even a single day learning about current models, then attention is most probably the first term you may have heard. And when it comes to the…

    https://abhinandan.one/artifacts/understanding-attention-from-scratch · serial 15 · published 2026-10-06T17:30:00.202+00:00 · updated 2026-10-06T17:42:39.065+00:00
  2. 14

    These two terms are very common. And it's very possible that you may have heard them before, or maybe you even know what they mean. But I feel that even though earlier I knew…

    https://abhinandan.one/artifacts/prefill-decode-made-crystal-clear · serial 14 · published 2026-10-02T13:10:55.521+00:00 · updated 2026-10-02T13:58:38.825+00:00
  3. 05

    I do not know if this was my poor luck or me being late, but while trying to understand and measure speculative decoding myself, I realized that SGLang had already merged adaptive…

    https://abhinandan.one/artifacts/measuring-adaptive-speculative-decoding-in-sglang · serial 05 · published 2026-09-24T08:00:00+00:00 · updated 2026-10-03T11:46:38.499+00:00

Open source

Projects

RolloutCoreRL Rollout Runtime

Versioned RL rollout runtime for vLLM with NCCL hot weight updates, version-pure rollouts, and cache-coherent transitions.

stack: PyTorch, NCCL, vLLM, RL
keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput
date: 2026-09-23 · case study: none
HiQCacheQuantized SGLang HiCache

INT8 hierarchical KV caching for SGLang HiCache, delivering 43.75% lower host KV memory and 1.78× more L2 cache capacity for Qwen3-8B.

stack: PyTorch, SGLang, Quantization
keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput
date: 2026-09-23 · case study: none
MiniServeLLM Inference Runtime

a from-scratch LLM serving runtime with continuous batching, chunked prefill, block-based KV management, preemption, and prefix caching.

stack: Python 3.12, PyTorch, Paged KV Cache, Continuous Batching, Prefix Caching
keywords: LLM inference, serving runtime, continuous batching, chunked prefill, paged attention, KV cache, prefix caching, preemption, request scheduling, vLLM, TTFT, throughput
date: 2026-09-23 · case study: none
AgentFlow-ProAgentic RL Research

Process-supervised RL that made an 8B model reason better, and the gain carried over to a domain it never trained on.

stack: PyTorch, TRL, DAPO, PRM, PEFT / LoRA, Qwen3-8B, Ollama, FastMCP
keywords: AgentFlow, process reward model, DAPO, reinforcement learning, agentic reasoning, Qwen3, LoRA fine-tuning, GPQA, AIME, RLHF
date: 2026-05-20 · case study: https://abhinandan.one/agentflow-pro
SmartMemoSemantic LLM Cache

A semantic cache for LLM agents where a trained classifier decides if a cached answer is safe to reuse, instead of raw cosine similarity.

stack: FAISS, SentenceTransformers, PyTorch, SQLite, Pydantic
keywords: semantic cache, LLM cache, FAISS, embeddings, sentence-transformers, pairwise classifier, semantic equivalence, prompt caching, MiniLM
date: 2026-03-15 · case study: https://abhinandan.one/smartmemo
OrchflowAgent Orchestration Framework

A dependency-free Python framework for readable multi-agent pipelines: sequential, parallel, conditional, and resumable flows.

stack: AsyncIO, LiteLLM, Pydantic
keywords: multi-agent, orchestration, asyncio, pipeline, LiteLLM, workflow framework, checkpoint resume, dependency-free
date: 2026-02-20 · case study: https://abhinandan.one/orchflow
agentevalLLM Evaluation Tooling

Behavioral eval for agents: replaces brittle exact-match asserts with repeated-run pass-rate scoring for CI gates.

stack: AsyncIO, OpenAI SDK, Anthropic SDK, LangChain, Typer
keywords: agent evaluation, LLM testing, pass rate, CI, behavioral assertions, agent tracing, non-deterministic testing, regression tracking
date: 2026-01-25 · case study: https://abhinandan.one/agenteval

Work

BrowzerFounding Software EngineerSF
Sept 2025 to Oct 2026
Cynos NexusFounding Software EngineerNoida
2025
Etkin.aiContract Software EngineerTürkiye
2024
HeroUI (YC S24)Open Source ContributorRemote
2024

Amazon ML Summer School 2025 · IYMC Gold Honour · Top 1% TypeScript Engineer, Algora