
Sarvesh Kumar
About Candidate
AI Engineer with 2+ years designing and shipping production Generative AI and agentic systems for enterprise use. Built a LangGraph ReAct agent that coordinates 8 tools — RAG, natural-language-to-SQL, code execution, and document generation — inside a multi-modal assistant used across the company. Comfortable across the stack: LLM orchestration (LangGraph, LangChain), vector search
(ChromaDB, FAISS, pgvector), and backend engineering (Python, Flask, FastAPI, async and streaming APIs). Experienced moving systems from prototype to production, including real-time token streaming, failover-resilient model serving on AWS Bedrock, and reliability work on LLM-driven routing and tool use.
Location
Education
Work & Experience
•Architected and own Atul Agent, a full-stack multi-modal enterprise AI assistant (Flask + React) covering 8 capabilities — general
chat, knowledge base Q&A, document Q&A, document generation, NL-to-SQL, code interpretation, app scaffolding, and web search
— used across the company.
• Redesigned the core orchestration layer as a LangGraph ReAct agent, replacing a monolithic if/elif intent-dispatch system by
registering all 8 capabilities as LangGraph tools, which improved routing accuracy and made the system extensible without rewriting
dispatch logic.
• Implemented token-level streaming and live tool-status updates using astream_events (v2) and adispatch_custom_event, giving the
React frontend real-time visibility into agent execution.
• Built a sync-to-async bridge to plug legacy blocking generator services into the agent's async execution loop without a full service
rewrite.
• Engineered a multi-module NL-to-SQL pipeline (Sales, Purchase, Finance) on live MySQL, combining Schema RAG via
ChromaDB, per-module few-shot stores, LLM query rewriting, SQL validation, and self-healing retry logic.
• Built a sandboxed code interpreter modeled on ChatGPT's Advanced Data Analysis — schema extraction on upload, LLM-generated
pandas execution with safe builtins, automatic retry on errors, and multi-file support.
• Deployed a company-wide pgvector knowledge base (PostgreSQL + HNSW) with adaptive chunking that auto-detects Q&A-style
spreadsheets, supporting PDF, DOCX, XLSX, PPTX, and CSV ingestion with upsert-by-title.
• Built a multi-engine OCR pipeline (PaddleOCR primary, EasyOCR/Tesseract fallback) integrated with Gemini 2.5 Flash for
structured extraction, and a Shipping Bill Verification System that pulls 50+ fields with cross-document validation — cutting manual
verification effort by 80%.
• Delivered the company's first RAG-based enterprise chatbot using LangGraph and ChromaDB with semantic search, cutting manual
document lookup effort by 60% across large internal corpora.
• Architected a hybrid Sentence-Transformers + LLM re-ranking pipeline for employee-to-department mapping, and migrated the
LLM backend to a self-hosted Qwen2.5-7B-Instruct-AWQ model on vLLM, resolving GPU out-of-memory and concurrency issues
through chunked batching.
• Optimized ETL pipelines in Pentaho, cutting data processing time by 35% through query restructuring and streamlined
transformations.
• Delivered 8+ UiPath RPA bots across enterprise workflows, saving 120+ man-hours annually and cutting manual errors by 70%.
• Built an intelligent data-extraction system combining multi-engine OCR (Gemini, Tesseract, PaddleOCR) for GST and invoice
processing; schema-constrained prompts achieved near-100% output reliability, and ThreadPool concurrency boosted API throughput
by 25%.