Open to new opportunities

Nikhil
Agrawal

Founding AI Engineer · Production LLM Systems · Retrieval & Full-Stack AI

I build production LLM systems, retrieval infrastructure, backend services, and full-stack AI products — spanning semantic search, RAG pipelines, document intelligence, and local LLM & GPU systems work.

1M+
records/week
processed at OpGov.ai
4+
multi-agent systems
shipped to production
3rd
MCP × A2A Hackathon
prize winner
6+
years
building AI systems
Selected Projects

Featured Work

Production-grade AI systems built for scale, reliability, and real-world impact.

Multi-Agent System for Technology & Market Evaluation
Multi-agent platform where autonomous agents collaborate to perform real-time evaluation of business and technology proposals — Flask backend, React frontend.
CrewAI
GPT-4o
Flask
React
Python
AI Agents

Full-stack AI-powered decision support

Agentic City Data Automation System
Agentic pipeline that ingests city email notifications, scrapes city websites, extracts meeting details, and generates structured calendars, agendas, and summaries.
Python
FastAPI
LLMs
Web Scraping
MongoDB
Agentic Workflows
Docker

1M+ records/week · ~30% accuracy improvement

KG-RAG Hybrid Retrieval System
Hybrid retrieval system combining Neo4j knowledge graphs with vector databases (Weaviate, FAISS) for multi-hop reasoning and context-aware search.
Neo4j
Weaviate
FAISS
RAG
Python
LLMs
Knowledge Graphs

Improved retrieval relevance by ~30%

Local LLM & GPU Systems Lab
A systems-focused project exploring how local LLMs behave under different memory, inference, and fine-tuning conditions — GPU memory allocation, quantization, context length, and benchmarking across models.
PyTorch
CUDA
Local LLMs
Quantization
QLoRA
Benchmarking
Llama / Mistral / Qwen / Phi

VRAM, tokens/sec & load-time benchmarks across models

Mass Communication Platform
A full-stack communication platform (FastAPI, Next.js, MongoDB) for scalable outreach — contact management, API integrations, and reliable message-delivery workflows.
FastAPI
Next.js
MongoDB
REST APIs
Full-Stack
Python

Scalable outreach & message delivery

AI Search & Knowledge Platform
A document search and retrieval system for natural-language questions over large collections — combining vector search, metadata filtering, and LLM summarization behind backend APIs.
RAG
Vector Search
Embeddings
LLMs
FastAPI
Document Processing

Natural-language search over large document sets

Philosophy

How I Think About AI

Principles that guide every system I build.

Systems, not demos

I focus on AI systems that survive real-world messiness: noisy inputs, partial failures, edge cases, and changing data formats.

Feedback loops matter

The best AI systems are not one-shot generators. They generate, evaluate, improve, and become more reliable over time.

AI is infrastructure

The next frontier is better systems around models: observability, reliability, cost control, and trust at scale.

Currently interested in

Retrieval quality
Structured outputs
AI evaluation
Observability
Local LLM inference
GPU memory & quantization
Scalable backend design
Full-stack AI products
"I build AI systems that don't just work once — they keep getting better."