Nikhil
Agrawal
Applied AI Engineer · RAG, Agents & AI Evaluation
I build production retrieval and agent systems, and the evaluation infrastructure that keeps them honest — evidence-grounded RAG and knowledge graphs, a canonical data platform spanning 1.7M+ records, and local LLM & GPU systems work below the API layer.
Featured Work
Production-grade AI systems built for scale, reliability, and real-world impact.
→ Claim-level citation validation + enforced abstention across ~1,600 meetings
→ Golden datasets & hard negatives as a versioned release gate
→ 1.7M+ profiles · 20,060 new profiles reconciled, 1 ambiguous identity isolated
→ Live geospatial nearest-record search with real-time device location
→ 3rd Prize — MCP × A2A Hackathon
→ Six experiments: inference, memory, scheduling & fine-tuning
How I Think About AI
Principles that guide every system I build.
Systems, not demos
I focus on AI systems that survive real-world messiness: noisy inputs, partial failures, edge cases, and changing data formats.
Feedback loops matter
The best AI systems are not one-shot generators. They generate, evaluate, improve, and become more reliable over time.
AI is infrastructure
The next frontier is better systems around models: observability, reliability, cost control, and trust at scale.
Currently interested in
"I build AI systems that don't just work once —
they keep getting better."