LLM & retrieval systems
Grounded assistants, document ingestion, chunking, embeddings, vector search, citations, structured outputs, and evaluation-aware RAG workflows.
Open to senior AI/ML and GenAI engineering roles
Senior AI/ML Engineer with 7+ years across LLM applications, retrieval systems, retail forecasting, MLOps, distributed data pipelines, and cloud-native deployment across GCP and AWS.
01 / Expertise
I work across the full engineering path: retrieval and model orchestration, data pipelines, APIs, infrastructure, deployment, monitoring, and optimization.
Grounded assistants, document ingestion, chunking, embeddings, vector search, citations, structured outputs, and evaluation-aware RAG workflows.
Repeatable training and inference systems with CI/CD, containers, autoscaling services, GPU workloads, observability, and reliability controls.
Retail demand forecasting, time-series modelling, distributed feature pipelines, customer analytics, churn and delinquency models, and production inference.
02 / Selected work
Selected work across production AI, data movement, retrieval, and deployment. Client-sensitive implementations are described without exposing private source or data.
Implemented a production-oriented FastAPI service for document retrieval and grounded generation on AWS, with S3 ingestion, Bedrock Knowledge Bases, least-privilege IAM, container delivery, structured logging, citations, tests, and deployment guidance. This is independent portfolio engineering and is not presented as client work.
View repositoryExtended a Cloud Run-ready MCP gateway with OIDC authentication, role-based authorization, constrained BigQuery and PostgreSQL tools, governed REST connectors, PII masking, rate limits, and structured audit logging. Caller identity remains separate from the cloud workload identity.
View repositoryBuilt a full-stack RAG workbench with swappable OpenAI, Claude, and Gemini providers; pgvector hybrid retrieval; reranking; citations; abstention; and latency, token, and cost traces. Its evaluation harness compares configurations on recall@k, MRR, keyword coverage, and abstention using the same pipeline as the live API.
View repositoryBuilt a router/supervisor graph with knowledge and order agents, persistent state, structured outputs, human approval for risky actions, low-confidence escalation, retries, timeouts, circuit breaking, full-round-trip tracing, and scenario-based evaluation through the compiled graph.
View repositoryEngineered a production platform for exploring policy records through semantic retrieval, combining a Dash and Flask application with Vertex AI embeddings, ScaNN search, and Cloud SQL. Delivered the supporting API and scheduled data workflows on Cloud Run with private database connectivity, Secret Manager, Artifact Registry, and workload identity federation.
Source private by design · Architecture discussable in interviewsDesigned document-aware assistants spanning ingestion, chunking, embedding generation, retrieval, citations, response structure, and cloud deployment. Focused on grounded answers and operational reliability rather than demo-only prompting.
Built a Cloud Run service that reads source records from Cloud SQL, generates embeddings with Vertex AI, and writes the enriched records back to a dedicated database table.
View repositoryAutomated delivery of Python notebooks and ingestion assets across development, pre-production, and production GCP environments using service accounts and a GitHub pipeline.
View repositoryA growing set of interview problems documented with intuition, dry runs, complexity analysis, clean Python, and automated tests—not answer dumps.
View repository03 / Experience
Upwork · Remote
Building GPT-powered automation, RAG platforms, assistants, GCP pipelines, GPU workloads, analytics integrations, and production Python services.
Open Insights · US, remote
Contributed to the Open ML 1.5 forecasting platform; built PySpark feature pipelines and TensorFlow training and inference workflows across BigQuery, Dataproc, and GCP infrastructure.
Capgemini · Bengaluru
Deployed ML systems on GCP, moved architectures toward serverless delivery, and introduced carbon-emission tracking for ML workflows.
Hughes Systique · Gurugram
Led GCP and AWS model deployment, developed vision, NLP, and tabular models, optimized TensorFlow for mobile edge use, and maintained internal CI/CD tooling.
Integration Wizards · Bengaluru
Trained TensorFlow and PyTorch models, prototyped pose estimation, and helped deploy a chemical-fluid leak detection system.
04 / Credentials
B.Tech in Computer Science & Engineering from Shoolini University, with an OGPA of 8.14.
05 / Recommendation
A LinkedIn recommendation from a former manager who worked with me directly at Open Insights.
Bill Karr · Chief Scientist, Open Insights
Managed Raghu directly on the Open ML 1.5 platform
Raghu was a core contributor to Open ML 1.5, our hierarchical probabilistic forecasting platform serving national retail clients at scale. Over the course of our work together he became genuinely well-rounded, comfortable moving between training and deploying ML models, writing and optimizing PySpark jobs, and managing cloud infrastructure on GCP, including BigQuery and Dataproc.
06 / Contact
I’m interested in senior AI/ML, GenAI, MLOps, and cloud engineering opportunities where reliability matters as much as model capability.