Open to new opportunities

Nikhil
Agrawal

Applied AI Engineer · RAG, Agents & AI Evaluation

I build production retrieval and agent systems, and measure — with statistics, not vibes — whether they actually got better. That's an evidence-grounded RAG and agent platform orchestrated with LangGraph and governed through MCP, a canonical data platform spanning 1.7M+ records, and local LLM & GPU systems work below the API layer.

~1,600
civic meetings indexed & evaluated
retrieval and evaluation in production
6
production systems shipped at OpGov
owned end-to-end, from ingestion to evaluation
1.7M+
profiles unified
canonical public-record platform
3rd
MCP × A2A Hackathon
prize winner
Selected Projects

Featured Work

Production-grade AI systems built for scale, reliability, and real-world impact.

Evidence-Grounded Civic Intelligence Platform (ASK NOHA)
Hybrid RAG, knowledge-graph and agent retrieval over ~1,600 civic meetings — orchestrated with LangGraph, exposed as governed tools through MCP, with human approval required on sensitive actions.
Python
LangGraph
MCP
FastAPI
MongoDB
Neo4j
Weaviate
FAISS
Gemini
CrossEncoder

Agentic retrieval with governed tools, human approval, and enforced abstention

AI Evaluation & Regression Harness
A versioned evaluation harness with real statistics behind it — bootstrap confidence intervals on retrieval deltas, ablation studies, and an LLM-as-judge calibrated against human labels via Cohen's kappa.
Python
Bootstrapping
Ablation Studies
LLM-as-Judge
Cohen's Kappa
Langfuse
OpenTelemetry

Statistical rigor as a release gate: bootstrap CIs, ablations, calibrated LLM-judge

Multi-County Public Records Data Platform
A canonical MongoDB platform unifying 1.7M+ public-record profiles and ~4M historical participation records from incompatible county schemas.
Python
MongoDB
PyMongo
ETL
Identity Resolution
Temporal Data

1.7M+ profiles · 20,060 new profiles reconciled, 1 ambiguous identity isolated

Geospatial Field Operations Dashboard
A live field dashboard over a structured voter database — GPS tracking, nearest-neighbor search over 2dsphere-indexed GeoJSON, and optimized multi-stop routing. Used in a real campaign.
Next.js
TypeScript
MongoDB
GeoJSON
Google Routes API
Python

Live dashboard used in a real campaign, with optimized multi-stop routing

MCP × A2A Hackathon — Multimodal Agent System
A multimodal autonomous agent system for complex task automation, built against the Model Context Protocol and agent-to-agent communication. 3rd Prize.
MCP
A2A
Multimodal
Autonomous Agents

3rd Prize — MCP × A2A Hackathon

GPU & Local Inference Lab
A systems-focused lab of six experiments on local LLM inference, GPU memory, scheduling and fine-tuning — measuring what most demos skip.
PyTorch
CUDA
Local LLMs
Quantization
QLoRA
GPU Profiling
Benchmarking

Six experiments: inference, memory, scheduling & fine-tuning

Pet POV Wearable — Edge AI Device
A founder's early concept turned into a working prototype: a wearable capturing point-of-view video with on-device YOLO object detection and multimodal emotion inference.
Raspberry Pi
Pi AI Camera
YOLO
TensorFlow Lite
Python
React

0→1 hardware + edge AI product, from founder concept to working prototype

Philosophy

How I Think About AI

Principles that guide every system I build.

Systems, not demos

I focus on AI systems that survive real-world messiness: noisy inputs, partial failures, edge cases, and changing data formats.

Feedback loops matter

The best AI systems are not one-shot generators. They generate, evaluate, improve, and become more reliable over time.

AI is infrastructure

The next frontier is better systems around models: observability, reliability, cost control, and trust at scale.

Currently interested in

Retrieval quality
Structured outputs
AI evaluation
Observability
Local LLM inference
GPU memory & quantization
Scalable backend design
Full-stack AI products
"I build AI systems that don't just work once — they keep getting better."