Nikhil
Agrawal
Applied AI Engineer · RAG, Agents & AI Evaluation
I build production retrieval and agent systems, and measure — with statistics, not vibes — whether they actually got better. That's an evidence-grounded RAG and agent platform orchestrated with LangGraph and governed through MCP, a canonical data platform spanning 1.7M+ records, and local LLM & GPU systems work below the API layer.
Featured Work
Production-grade AI systems built for scale, reliability, and real-world impact.
→ Agentic retrieval with governed tools, human approval, and enforced abstention
→ Statistical rigor as a release gate: bootstrap CIs, ablations, calibrated LLM-judge
→ 1.7M+ profiles · 20,060 new profiles reconciled, 1 ambiguous identity isolated
→ Live dashboard used in a real campaign, with optimized multi-stop routing
→ 3rd Prize — MCP × A2A Hackathon
→ Six experiments: inference, memory, scheduling & fine-tuning
→ 0→1 hardware + edge AI product, from founder concept to working prototype
How I Think About AI
Principles that guide every system I build.
Systems, not demos
I focus on AI systems that survive real-world messiness: noisy inputs, partial failures, edge cases, and changing data formats.
Feedback loops matter
The best AI systems are not one-shot generators. They generate, evaluate, improve, and become more reliable over time.
AI is infrastructure
The next frontier is better systems around models: observability, reliability, cost control, and trust at scale.
Currently interested in
"I build AI systems that don't just work once —
they keep getting better."