Open to new opportunities

Nikhil
Agrawal

Applied AI Engineer · RAG, Agents & AI Evaluation

I build production retrieval and agent systems, and the evaluation infrastructure that keeps them honest — evidence-grounded RAG and knowledge graphs, a canonical data platform spanning 1.7M+ records, and local LLM & GPU systems work below the API layer.

1.7M+
profiles unified
canonical public-record platform
~1,600
civic meetings
evidence-grounded RAG + KG
3rd
MCP × A2A Hackathon
prize winner
M.S.
Statistics
CSU East Bay, GPA 3.9
Selected Projects

Featured Work

Production-grade AI systems built for scale, reliability, and real-world impact.

Evidence-Grounded Civic Intelligence Platform
Hybrid and graph RAG over ~1,600 civic meetings, with claim-level citation validation and enforced abstention when evidence is insufficient.
Python
FastAPI
MongoDB
Neo4j
Weaviate
FAISS
Gemini
BGE
CrossEncoder
Docker
Azure

Claim-level citation validation + enforced abstention across ~1,600 meetings

AI Evaluation & Regression Harness
A versioned evaluation harness measuring retrieval, grounding, attribution, conversation and security behavior — used as a release gate rather than a one-off report.
Python
Evaluation Pipelines
Langfuse
OpenTelemetry
Retrieval Metrics
Statistical Inference

Golden datasets & hard negatives as a versioned release gate

Multi-County Public Records Data Platform
A canonical MongoDB platform unifying 1.7M+ public-record profiles and ~4M historical participation records from incompatible county schemas.
Python
MongoDB
PyMongo
ETL
Identity Resolution
Temporal Data

1.7M+ profiles · 20,060 new profiles reconciled, 1 ambiguous identity isolated

Geospatial Field Operations Application
A Next.js and MongoDB geospatial field app using GeoJSON, 2dsphere indexing and live device location to surface the nearest eligible records with contact and navigation actions.
Next.js
TypeScript
MongoDB
GeoJSON
Google Maps
Python

Live geospatial nearest-record search with real-time device location

MCP × A2A Hackathon — Multimodal Agent System
A multimodal autonomous agent system for complex task automation, built against the Model Context Protocol and agent-to-agent communication. 3rd Prize.
MCP
A2A
Multimodal
Autonomous Agents

3rd Prize — MCP × A2A Hackathon

GPU & Local Inference Lab
A systems-focused lab of six experiments on local LLM inference, GPU memory, scheduling and fine-tuning — measuring what most demos skip.
PyTorch
CUDA
Local LLMs
Quantization
QLoRA
GPU Profiling
Benchmarking

Six experiments: inference, memory, scheduling & fine-tuning

Philosophy

How I Think About AI

Principles that guide every system I build.

Systems, not demos

I focus on AI systems that survive real-world messiness: noisy inputs, partial failures, edge cases, and changing data formats.

Feedback loops matter

The best AI systems are not one-shot generators. They generate, evaluate, improve, and become more reliable over time.

AI is infrastructure

The next frontier is better systems around models: observability, reliability, cost control, and trust at scale.

Currently interested in

Retrieval quality
Structured outputs
AI evaluation
Observability
Local LLM inference
GPU memory & quantization
Scalable backend design
Full-stack AI products
"I build AI systems that don't just work once — they keep getting better."