abhinandan

Abhinandan

Inference engineer. RL post-training on reasoning models and the serving stack that runs them: vLLM, SGLang, TensorRT-LLM, INT8 KV caching.

What I do

I do RL post-training on reasoning models, and build the inference systems that serve them. Most of what I care about happens after launch day: when the reasoning breaks, when the cost climbs, when the first real traffic arrives.

Currently Founding Software Engineer at Browzer since September 2025, remote. Open to full-time inference engineer and reinforcement learning engineer roles, and available to start immediately.

In the open I work across 10 repositories — 21 merged pull requests and counting. Merged patches upstream in vllm-project/vllm #58895. Open pull requests right now in sgl-project/sglang #41479, sgl-project/sglang #41478, vllm-project/vllm #58898. The live list is here.

Selected work

HiQCache

INT8 host-tier KV cache for SGLang HiCache: L2 KV storage down from 147,456 to 82,944 B/token (43.75%) and 1.78x more cached tokens (54K to 96K) at a fixed host budget, with GPU L1 KV left in BF16. Measured on an RTX A6000, hit rate moved from ~57% to ~73%.

RolloutCore

RL rollout control plane for vLLM with drain-before-mutate semantics, live NCCL weight installs and version-pure rollouts. On Qwen3-8B, hot updates take 4.25s against 37.04s for a full restart (~8.7x lower model-transition time) on 2x RTX 3090 over PCIe.

MiniServe

A from-scratch LLM serving runtime: continuous batching, chunked prefill, recompute preemption and paged KV with content-addressed prefix caching. Matched Hugging Face to within 6e-8 float32 logit error, and prefix reuse cut computed prefill work by 97% on a 120-request shared-prefix workload.

7 more projects at /projects, write-ups at /artifacts.

What I work on

  • LLM inference
  • Model serving
  • KV cache management
  • Quantization (INT8)
  • Speculative decoding
  • vLLM
  • SGLang
  • TensorRT-LLM
  • NCCL
  • PyTorch
  • Reinforcement learning post-training
  • Reasoning models
  • Agentic AI
  • Semantic caching
  • Retrieval-augmented generation
  • Evaluation methodology

Education & recognition

  • B.Tech, Computer Science — Dr. APJ Abdul Kalam Technical University, 2026
  • International Youth Math Challenge — Gold Honour
  • Top 1% TypeScript Engineer Globally (Algora)
  • Amazon ML Summer School 2025

Questions

Who is Abhinandan?

An inference engineer based in India, working remotely. He does RL post-training on reasoning models and builds the inference systems that serve them — vLLM, SGLang, TensorRT-LLM, INT8 KV caching, and the surrounding serving stack.

What is Abhinandan working on now?

Founding Software Engineer at Browzer since September 2025, building a browser-agent runtime with a streaming ReAct loop and zero-LLM replay. In the open, he maintains RolloutCore, an RL rollout control plane for vLLM, and HiQCache, an INT8 host-tier KV cache for SGLang HiCache.

What roles is Abhinandan looking for?

Full-time inference and ML engineering roles: LLM serving and runtime work, KV cache and quantization, and RL post-training for reasoning models. He can join immediately.

What is Abhinandan's inference stack?

Inference: vLLM, SGLang, TensorRT-LLM, NCCL, quantization. ML: PyTorch, Hugging Face, LoRA/PEFT, Qdrant. Engineering: Python, C++, TypeScript, FastAPI, Node.js, Redis, PostgreSQL, AWS, GCP, Docker, GitHub Actions.

Has Abhinandan contributed to open source?

Yes. He has merged pull requests and open work across vLLM, SGLang, HeroUI, Mooncake, ai-dynamo and others. Sixteen merged pull requests to HeroUI (then NextUI) led to a personal offer from the CEO. The full, live list is at abhinandan.one/contributions.

What is Abhinandan's education?

B.Tech in Computer Science from Dr. APJ Abdul Kalam Technical University, 2026.

What recognition has Abhinandan received?

International Youth Math Challenge Gold Honour, Top 1% TypeScript Engineer Globally on Algora, and selection for Amazon ML Summer School 2025.

How do I contact Abhinandan?

Email abhinandan@abhinandan.one, or reach him on GitHub at awesome-pro, LinkedIn at abhibuilds, or X at abhibuilds.

Elsewhere

Machine-readable: /llms.txt and /llms-full.txt. Or write to abhinandan@abhinandan.one.