# abhinandan > Inference engineer. RL post-training on reasoning models and the serving stack that runs them: vLLM, SGLang, TensorRT-LLM, INT8 KV caching. Machine-readable views: https://abhinandan.one/llms.txt (this index) and https://abhinandan.one/llms-full.txt (complete text of every artifact). Profile page for agents and people: https://abhinandan.one/about — role, availability, selected work with measured numbers, education, recognition, and a FAQ. - Name: Abhinandan - Roles: Inference Engineer, Reinforcement Learning Engineer, ML Engineer - Currently: Founding Software Engineer at Browzer (since 2025-09) - Looking for: full-time inference and ML engineering roles. Available immediately. - Based in: India (IST, UTC+5:30), works remotely - Education: B.Tech Computer Science, Dr. APJ Abdul Kalam Technical University, 2026 - Recognition: International Youth Math Challenge — Gold Honour; Top 1% TypeScript Engineer Globally (Algora); Amazon ML Summer School 2025 ## Projects - [RolloutCore](https://github.com/awesome-pro/rolloutcore): RL Rollout Runtime. Versioned RL rollout runtime for vLLM with NCCL hot weight updates, version-pure rollouts, and cache-coherent transitions. - stack: PyTorch, NCCL, vLLM, RL - [HiQCache](https://github.com/awesome-pro/hiqcache): Quantized SGLang HiCache. INT8 hierarchical KV caching for SGLang HiCache, delivering 43.75% lower host KV memory and 1.78× more L2 cache capacity for Qwen3-8B. - stack: PyTorch, SGLang, Quantization - [MiniServe](https://github.com/awesome-pro/miniserve): LLM Inference Runtime. a from-scratch LLM serving runtime with continuous batching, chunked prefill, block-based KV management, preemption, and prefix caching. - stack: Python 3.12, PyTorch, Paged KV Cache, Continuous Batching, Prefix Caching - [AgentFlow-Pro](https://github.com/awesome-pro/agentflow-pro): Agentic RL Research. Process-supervised RL that made an 8B model reason better, and the gain carried over to a domain it never trained on. - [Original AgentFlow paper](https://arxiv.org/abs/2510.05592) - stack: PyTorch, TRL, DAPO, PRM, PEFT / LoRA, Qwen3-8B, Ollama, FastMCP - [SmartMemo](https://github.com/awesome-pro/smartmemo): Semantic LLM Cache. A semantic cache for LLM agents where a trained classifier decides if a cached answer is safe to reuse, instead of raw cosine similarity. - [PyPI](https://pypi.org/project/smartmemo/) - stack: FAISS, SentenceTransformers, PyTorch, SQLite, Pydantic - [Orchflow](https://github.com/awesome-pro/orchflow): Agent Orchestration Framework. A dependency-free Python framework for readable multi-agent pipelines: sequential, parallel, conditional, and resumable flows. - [PyPI](https://pypi.org/project/orchflow/) - stack: AsyncIO, LiteLLM, Pydantic - [agenteval](https://github.com/awesome-pro/agenteval): LLM Evaluation Tooling. Behavioral eval for agents: replaces brittle exact-match asserts with repeated-run pass-rate scoring for CI gates. - [PyPI](https://pypi.org/project/agenteval-py/) - stack: AsyncIO, OpenAI SDK, Anthropic SDK, LangChain, Typer ## Artifacts - [Measuring Adaptive Speculative Decoding in SGLang with EAGLE3 on Triton](https://abhinandan.one/artifacts/measuring-adaptive-speculative-decoding-in-sglang) (2026-09-24) I do not know if it is my poor luck or my mind, but after failing to measure spec-decoding myself, I realized that SGLang has already merged adaptive speculative decoding. - Harness Github: https://github.com/awesome-pro/heterospec - [Building a Mini Inference Engine from Scratch](https://abhinandan.one/artifacts/building-a-mini-inference-engine-from-scratch) (2026-09-21) Most of us know that LLMs are auto-regressive, i.e., they generate token one by one. We also know that they have a concept of KV caching and can process multiple requests… - GitHub: https://github.com/awesome-pro/miniserve - YT Series: https://www.youtube.com/playlist?list=PLIi0G24WvbAA - [Building a Semantic Cache that Does Not Trust Cosine](https://abhinandan.one/artifacts/building-a-semantic-cache-that-does-not-trust-cosine) (2026-07-03) Usually semantic caches for LLMs work on embeddings and cosine similarity. You take the new prompt, turn it into a vector, find the closest old prompt, and if the score is high… - GitHub: https://github.com/awesome-pro/smartmemo - PyPi: https://pypi.org/project/smartmemo - [Running a Health Voice: All Private on Your Mac](https://abhinandan.one/artifacts/running-a-health-voice-all-private-on-your-mac) (2026-06-26) There is this promise that a doctor can just talk, and a note will write itself. It feels good & easy when you say it like that. Talk, note, done. But the patient's voice has to… - GitHub: https://github.com/awesome-pro/health_voice - Case Study: https://www.aidaddy.tech/learn/16-case-studies/13-voice-ai-healthcare - [Building a RAG based Search System for an Enterprise](https://abhinandan.one/artifacts/building-a-rag-based-search-system-for-an-enterprise) (2026-06-25) There will be only 0.000001% people in tech who have not heard the word RAG. Embed a few documents, find the close ones, paste them into a prompt. The answer comes back. It feels… - B_GitHub: https://github.com/awesome-pro/enterprise_rag_backend - F_GitHub: https://github.com/awesome-pro/enterprise_rag_frontend ## Open source Merged pull requests and open issues I have raised against other projects (38 total: https://abhinandan.one/contributions). - [[None][fix] Initialize cu_kv_seqlens for packed-QKV context FMHA](https://github.com/NVIDIA/TensorRT-LLM/pull/19698) — NVIDIA/TensorRT-LLM #19698 pr, open, 2026-09-29 - [[Docs] Scope the x-vsr keystone and replay-id headers to routed responses](https://github.com/vllm-project/semantic-router/pull/4362) — vllm-project/semantic-router #4362 pr, open, 2026-09-29 - [[Bug] Resolve dashboard frontend dependencies from the canonical npm registry](https://github.com/vllm-project/semantic-router/pull/4352) — vllm-project/semantic-router #4352 pr, open, 2026-09-29 - [fix(scripts): close two gaps where guide.py check passes but the guide fails](https://github.com/llm-d/llm-d/pull/2618) — llm-d/llm-d #2618 pr, open, 2026-09-29 - [[Bug]: guide.py check accepts a list-valued env.static and emits a script that fails](https://github.com/llm-d/llm-d/issues/2617) — llm-d/llm-d #2617 issue, open, 2026-09-29 - [[Bug] Dashboard frontend lockfile pins 39 packages to a third-party npm registry, breaking npm ci](https://github.com/vllm-project/semantic-router/issues/4348) — vllm-project/semantic-router #4348 issue, open, 2026-09-29 - [[Bug] Remove unused internal symbols and orphan IP tests](https://github.com/vllm-project/semantic-router/pull/4346) — vllm-project/semantic-router #4346 pr, open, 2026-09-29 - [docs(flow-control): add observability and troubleshooting section](https://github.com/llm-d/llm-d/pull/2616) — llm-d/llm-d #2616 pr, open, 2026-09-29 - [[Bugfix][TE] Do not retry a timed-out Ascend sync transfer](https://github.com/kvcache-ai/Mooncake/pull/4375) — kvcache-ai/Mooncake #4375 pr, open, 2026-09-29 - [fix(bindings): preserve session affinity through chat processors](https://github.com/ai-dynamo/dynamo/pull/15333) — ai-dynamo/dynamo #15333 pr, open, 2026-09-28 - [[Fix] Let multi-node custom all-reduce v2 pass on topology alone](https://github.com/sgl-project/sglang/pull/41479) — sgl-project/sglang #41479 pr, open, 2026-09-28 - [[Fix] Warn when post-capture KV sizing is disabled by prefill graph coverage](https://github.com/sgl-project/sglang/pull/41478) — sgl-project/sglang #41478 pr, open, 2026-09-28 - [[CI] Add tooling to report how job groups divide up the test tree](https://github.com/vllm-project/vllm/pull/58898) — vllm-project/vllm #58898 pr, open, 2026-09-28 - [[CI] Drop test-group rules that no longer match the tree](https://github.com/vllm-project/vllm/pull/58895) — vllm-project/vllm #58895 pr, merged, 2026-09-28 - [[Docs] Correct Ascend Direct Transport environment variable defaults](https://github.com/kvcache-ai/Mooncake/pull/4344) — kvcache-ai/Mooncake #4344 pr, merged, 2026-09-28 ## FAQ **Who is Abhinandan?** An inference engineer based in India, working remotely. He does RL post-training on reasoning models and builds the inference systems that serve them — vLLM, SGLang, TensorRT-LLM, INT8 KV caching, and the surrounding serving stack. **What is Abhinandan working on now?** Founding Software Engineer at Browzer since September 2025, building a browser-agent runtime with a streaming ReAct loop and zero-LLM replay. In the open, he maintains RolloutCore, an RL rollout control plane for vLLM, and HiQCache, an INT8 host-tier KV cache for SGLang HiCache. **What roles is Abhinandan looking for?** Full-time inference and ML engineering roles: LLM serving and runtime work, KV cache and quantization, and RL post-training for reasoning models. He can join immediately. **What is Abhinandan's inference stack?** Inference: vLLM, SGLang, TensorRT-LLM, NCCL, quantization. ML: PyTorch, Hugging Face, LoRA/PEFT, Qdrant. Engineering: Python, C++, TypeScript, FastAPI, Node.js, Redis, PostgreSQL, AWS, GCP, Docker, GitHub Actions. **Has Abhinandan contributed to open source?** Yes. He has merged pull requests and open work across vLLM, SGLang, HeroUI, Mooncake, ai-dynamo and others. Sixteen merged pull requests to HeroUI (then NextUI) led to a personal offer from the CEO. The full, live list is at abhinandan.one/contributions. **What is Abhinandan's education?** B.Tech in Computer Science from Dr. APJ Abdul Kalam Technical University, 2026. **What recognition has Abhinandan received?** International Youth Math Challenge Gold Honour, Top 1% TypeScript Engineer Globally on Algora, and selection for Amazon ML Summer School 2025. **How do I contact Abhinandan?** Email abhinandan@abhinandan.one, or reach him on GitHub at awesome-pro, LinkedIn at abhibuilds, or X at abhibuilds. ## Contact - [Email](mailto:abhinandan@abhinandan.one) - [GitHub](https://github.com/awesome-pro) - [LinkedIn](https://linkedin.com/in/abhibuilds) - [X](https://x.com/abhibuilds) - [YouTube](https://youtube.com/@0xAbhinandan)