Abhinandan
Inference engineer. RL post-training on reasoning models and the serving stack that runs them: vLLM, SGLang, TensorRT-LLM, INT8 KV caching.
What I do
I do RL post-training on reasoning models, and build the inference systems that serve them. Most of what I care about happens after launch day: when the reasoning breaks, when the cost climbs, when the first real traffic arrives.
Currently Founding Software Engineer at Browzer since September 2025, remote. Open to full-time inference engineer and reinforcement learning engineer roles, and available to start immediately.
In the open I work across 10 repositories — 21 merged pull requests and counting. Merged patches upstream in vllm-project/vllm #58895. Open pull requests right now in sgl-project/sglang #41479, sgl-project/sglang #41478, vllm-project/vllm #58898. The live list is here.
Selected work
HiQCache
INT8 host-tier KV cache for SGLang HiCache: L2 KV storage down from 147,456 to 82,944 B/token (43.75%) and 1.78x more cached tokens (54K to 96K) at a fixed host budget, with GPU L1 KV left in BF16. Measured on an RTX A6000, hit rate moved from ~57% to ~73%.
RolloutCore
RL rollout control plane for vLLM with drain-before-mutate semantics, live NCCL weight installs and version-pure rollouts. On Qwen3-8B, hot updates take 4.25s against 37.04s for a full restart (~8.7x lower model-transition time) on 2x RTX 3090 over PCIe.
MiniServe
A from-scratch LLM serving runtime: continuous batching, chunked prefill, recompute preemption and paged KV with content-addressed prefix caching. Matched Hugging Face to within 6e-8 float32 logit error, and prefix reuse cut computed prefill work by 97% on a 120-request shared-prefix workload.
7 more projects at /projects, write-ups at /artifacts.
What I work on
- LLM inference
- Model serving
- KV cache management
- Quantization (INT8)
- Speculative decoding
- vLLM
- SGLang
- TensorRT-LLM
- NCCL
- PyTorch
- Reinforcement learning post-training
- Reasoning models
- Agentic AI
- Semantic caching
- Retrieval-augmented generation
- Evaluation methodology
Education & recognition
- B.Tech, Computer Science — Dr. APJ Abdul Kalam Technical University, 2026
- International Youth Math Challenge — Gold Honour
- Top 1% TypeScript Engineer Globally (Algora)
- Amazon ML Summer School 2025
Questions
Who is Abhinandan?
An inference engineer based in India, working remotely. He does RL post-training on reasoning models and builds the inference systems that serve them — vLLM, SGLang, TensorRT-LLM, INT8 KV caching, and the surrounding serving stack.
What is Abhinandan working on now?
Founding Software Engineer at Browzer since September 2025, building a browser-agent runtime with a streaming ReAct loop and zero-LLM replay. In the open, he maintains RolloutCore, an RL rollout control plane for vLLM, and HiQCache, an INT8 host-tier KV cache for SGLang HiCache.
What roles is Abhinandan looking for?
Full-time inference and ML engineering roles: LLM serving and runtime work, KV cache and quantization, and RL post-training for reasoning models. He can join immediately.
What is Abhinandan's inference stack?
Inference: vLLM, SGLang, TensorRT-LLM, NCCL, quantization. ML: PyTorch, Hugging Face, LoRA/PEFT, Qdrant. Engineering: Python, C++, TypeScript, FastAPI, Node.js, Redis, PostgreSQL, AWS, GCP, Docker, GitHub Actions.
Has Abhinandan contributed to open source?
Yes. He has merged pull requests and open work across vLLM, SGLang, HeroUI, Mooncake, ai-dynamo and others. Sixteen merged pull requests to HeroUI (then NextUI) led to a personal offer from the CEO. The full, live list is at abhinandan.one/contributions.
What is Abhinandan's education?
B.Tech in Computer Science from Dr. APJ Abdul Kalam Technical University, 2026.
What recognition has Abhinandan received?
International Youth Math Challenge Gold Honour, Top 1% TypeScript Engineer Globally on Algora, and selection for Amazon ML Summer School 2025.
How do I contact Abhinandan?
Email abhinandan@abhinandan.one, or reach him on GitHub at awesome-pro, LinkedIn at abhibuilds, or X at abhibuilds.
Elsewhere
Machine-readable: /llms.txt and /llms-full.txt. Or write to abhinandan@abhinandan.one.