Forward Deployed Engineer, Observe.AI/M.S. Applied ML, University of Maryland

I build voice agents that hold up in production.

I work where LLM reasoning meets real systems: telephony, APIs, databases, and the evaluations that show whether any of it actually works. Before this, I built recommendation and forecasting pipelines at Flipkart, RAG systems at EY, and multimodal clinical ML as a research assistant at MIT Manipal.

NowOwner of voice AI agent and Companion Agent (real-time agent assist) deployments for 5 enterprise customers at Observe.AI, all taken from scoping to go-live.

Selected work

3 case studies
01

SLA-Aware Inference Gateway

A gateway that keeps image-classification latency inside a 300 ms SLA by routing each request between an accurate ResNet-50 (int8 ONNX) and a fast MobileNetV3, based on live p95 latency. Built from scratch and measured end to end under load.

Problem

At twice its capacity a single accurate model collapses: p95 latency of 2.7 s, 27% of requests failing, and 0.4% answered within the SLA. Always using the small model avoids that at a permanent 8-point accuracy cost.

What I built

The whole system: ONNX export and int8 quantization of 8 candidate models evaluated on 9,500 held-out images, the model servers, the gateway and its control loop, canary rollouts, an open-loop load generator, Docker Compose, kind and GKE deployments, and the benchmark suite (34 tests).

Result

2,721 → 192 msp95 at 2× capacity, with 99.8% of requests within the SLA and zero errors, while keeping accuracy 4 points above the always-fast option. A faulty canary was rolled back in 16 s with no client-visible errors. Found and fixed a CFS-throttling issue in ONNX Runtime that cut p95 by 3.8×. On GKE, a 6-minute 4× spike stayed 99.8% within SLA while the HPA and cluster autoscaler added pods and a node in 81 s. Measured energy per inference with powermetrics: ONNX Runtime int8 ResNet-50 at 0.79 J vs 3.75 J for Core ML fp32 on the same CPU, and 0.15 J for the fast tier's MobileNetV3.

  • Python
  • FastAPI
  • ONNX Runtime
  • int8 quantization
  • Kubernetes (GKE)
  • Prometheus
  • OpenTelemetry
02
2025 · Agentic LLMs

Scientific Research Analyzer

A multi-agent LLM workflow that reads research papers, retrieves supporting context, and writes grounded summaries with citation and hallucination checks.

Problem

LLM literature summaries sound confident but often cite the wrong source, or no source at all.

What I built

Agent workflows built with LangChain Graphs and transformer models. TGI handles serving, Dask handles scale, and LangFuse gives visibility into multi-agent coordination.

Result

+24%retrieval performance after tuning agent coordination.

  • LangChain
  • Hugging Face TGI
  • Dask
  • LangFuse
  • RAG
03
2024–2025 · Research, MIT Manipal

Multimodal ML for Myeloma Treatment Response

Predicting how multiple myeloma patients will respond to treatment by combining clinical tables, flow cytometry, bone-marrow variables, and PET/CT imaging, and explaining what drives each prediction.

Problem

Response signals are spread across modalities, and clinicians need to know why a model predicts what it does.

What I built

Transformer and CNN architectures for each modality, GPU training sped up with CUDA, PyTorch, Spark, and MLflow, and graph-based modeling for explainable decision support.

Result

+12%AUC from the SHAP-identified IBV biomarker. Co-authored and presented a technical report.

  • PyTorch
  • TabTransformer
  • ClinicalBERT
  • 3D U-Net
  • SHAP
  • CUDA

Experience

May 2026 – PresentRedwood City, CA

Forward Deployed Engineer · Observe.AI

Voice-AI and real-time agent-assist deployments for 5 enterprise customers

  • Own voice-AI and real-time agent-assist deployments for 5 enterprise customers, from scoping and integrations through testing and go-live.
  • Led a real-time agent-assist pilot for a US regional bank: English and Spanish support across 5 customer-call scenarios, with live prompts and step-by-step guidance for agents, integrated with the bank's Cisco phone system and taken from scoping to a live customer demo.
  • Co-led a US healthcare contact centre's move to the Companion Agent platform (Genesys phone integration, separate flows for new and experienced agents), passing a 45-case regression clean before customer testing.
  • Built an LLM-as-judge regression harness in Python (Typer, Pydantic, Azure OpenAI) that scores the AI's behaviour against 400+ rules; raised the pass rate from 77.9% to 95.8% across five scenarios.
  • Caught an LLM summariser inventing "booked" outcomes on 12 of 25 calls before they reached the customer's CRM; replaced an unreliable LLM compliance check (9 out of 9 production false positives) with a deterministic rule in n8n.
  • Traced a silent, multi-release failure to 16 misconfigured AI tools; one config fix took two compliance alerts from never firing to firing every time.
  • Designed a Snowflake secure-share data pipeline (n8n batch loads every ~15 min) delivering call summaries to a large US fintech.
  • Built the team's agentic developer toolkit (16 Claude Code subagents, 31 skills, MCP integrations) and onboard new AI Agent Engineers through build reviews, 6 so far.
  • Python
  • n8n
  • Azure OpenAI
  • Genesys
  • Cisco PCCE
  • Snowflake
  • ClickHouse
Feb – Jun 2025Bangalore

Data Intern · Flipkart

Recommendations & demand forecasting

  • Built end-to-end ML pipelines over 120M+ e-commerce transactions (PySpark, SQL, Pandas), lifting recommendation click-through rate by 14%.
  • Developed LightGBM and Prophet demand-forecasting models (MAE < 5%) to optimize inventory across 12 warehouses.
  • Containerized models with Docker and deployed FastAPI inference services at 20–40 ms latency in staging.
  • Automated workflows with Airflow and PostgreSQL, cutting manual analytics by 90% and runtime by 60%. SQL-driven persona segmentation raised notification engagement by 22%.
  • PySpark
  • LightGBM
  • Prophet
  • FastAPI
  • Airflow
  • Docker
Jan 2024 – May 2025Manipal

Research Assistant · MIT Manipal

Clinical AI Lab · multimodal clinical ML

  • Built transformer and CNN models to predict myeloma treatment response from multimodal clinical data. SHAP analysis identified IBV as a novel biomarker (+12% AUC). See case study
Jun – Aug 2024Noida

Generative AI Intern · Ernst & Young

RAG & LLM systems

  • Designed a RAG chatbot with Flan-T5 and Mistral-7B for Excel workflow automation, extracting insights from 10K+ corporate files.
  • Engineered hybrid retrieval with LangChain and FAISS: <400 ms latency, 99.3% precision.
  • A/B tested 400+ prompts and embedding setups (Instructor-XL, OpenAI), improving query accuracy by 25%.
  • Integrated retrieval graph reasoning to improve multi-document linking and semantic search.
  • Mistral-7B
  • Flan-T5
  • LangChain
  • FAISS
  • RAG

Earlier internships · 2024

May – Jul 2024Remote
Cisco · AICTE Cybersecurity Internship. Python intrusion detection pipelines (Snort, TF-IDF): 93% F1, 28% fewer false alerts.
May 2024Gurugram
Akzo Nobel · IT Intern. Anomaly detection pipelines (Isolation Forest, ARIMA) and Grafana dashboards that cut defects by 12%.
Apr 2024Gurugram
Hero MotoCorp · Data Science Intern. Sales forecasting with Random Forest and Prophet on 5M+ records, improving MAE by 18%.
Aug 2025 – May 2027 (expected)

University of Maryland

M.S. Applied Machine Learning · College Park, MD

Deep Learning · Reinforcement Learning · Advanced Machine Learning · Computing Systems for ML · Algorithms & Data Structures for ML · Introduction to Optimization · Principles of Machine Learning · Principles of Data Science · Probability & Statistics

Oct 2021 – May 2025

Manipal Institute of Technology

B.Tech Electronics & Communication, Minor in Data Science

Data Structures & Algorithms · Object-Oriented Programming (C++) · Computer Vision · Statistical Inference & Regression · Practical Machine Learning · Database Systems · Computer Organization & Architecture · Communication Networks

Publications

Published · Vol. 27(2), 2026 · online Oct 2025

Digital Sentiment and the Retail Crowd: How Finfluencers Shape IPO Valuations ↗

Kavitha Bharath Raja Guru, Krishna Prasad, Sandhya Parasnath Dubey, Simran Kharbanda. Journal of Behavioral Finance, Taylor & Francis.

Across 395 Indian IPOs (2014–2025), sentiment and engagement modeling (VADER, TextBlob, BERT) found that finfluencer endorsements are associated with 7.2% more underpricing, 6.8% higher initial returns (p < 0.05), and 10.9× higher retail demand (p < 0.01).

doi:10.1080/15427560.2025.2566736 ↗
Under review

Digital Signals and Market Discipline: Finfluencer Sentiment under Regulatory Signaling

Decisions in Economics and Finance

How regulatory signals shape finfluencer sentiment, and how digital financial communication influences investor behavior and market response.

More projects

  • 2026

    A model-serving platform on Kubernetes (AWS EKS) with a high-accuracy and a fast model variant and an SLA-aware router that shifts traffic to the fast model when the latency SLA is at risk, plus HPA autoscaling, Prometheus/Grafana monitoring and canary rollouts. Under high load, p95 latency fell from 30+ s to 5 s and errors from 5–10% to 0–2%. The Inference Gateway (case study 01) is a from-scratch rebuild of this idea with reproducible measurements.

    Python · FastAPI · ONNX · Kubernetes · AWS EKS · Locust
  • 2026

    Full-stack app that turns receipt photos into a per-store price database (LLM vision extraction with structured outputs, human review, rules + fuzzy product matching, unit-price normalization), then solves a MILP to pick which stores to visit, trading item prices against trip costs. Beats single-store and per-item greedy baselines in benchmarks. Also has price-drop alerts, spending insights and a recipe-to-list assistant built on LLM tool calling. Shipped with Docker, GitHub Actions CI, a Helm chart, plan-only Terraform (EKS, RDS, ElastiCache), OpenTelemetry tracing, Prometheus metrics and JSON logs; 158 tests.

    Python · FastAPI · SQLAlchemy · PuLP (MILP) · React / Next.js · TypeScript · Docker · Helm · Terraform · OpenTelemetry
  • 2026

    Dimensional warehouse built with dbt that runs unchanged on DuckDB locally and on BigQuery (98 of 98 checks on both, the BigQuery build billing 1.29 GiB at $0.008): a star schema with SCD Type 2 customer history, enforced model contracts, PII tagging and masking with role-based analyst views, freshness checks and a per-build data-quality log, all over 1.19M order lines with 84 tests passing in under 5 seconds. Alongside it, a Kafka to PyFlink streaming pipeline computes one-minute order metrics with event-time watermarks and lands them in the same warehouse: 2,000 events/s sustained with results queryable 2.3 s after each window closes, and 18,600 events/s into Kafka in a burst. The same stream also runs on Google Cloud as Pub/Sub → Dataflow (Beam) → BigQuery, measured at the same rate with the latency difference explained. Synthetic data, every number from a real run.

    dbt · DuckDB · BigQuery · SQL · Kafka · Flink (PyFlink) · Pub/Sub · Dataflow (Beam) · Docker · GitHub Actions · Python
  • 2026

    Fraud detection on 284,807 real card transactions (0.17% fraud) with a strict time-based split. XGBoost, Random Forest, LightGBM and logistic regression across three imbalance strategies; unsupervised Isolation Forest and autoencoder detectors; K-means fraud segments (one holds 44% of frauds at 16.6x the base rate); isotonic calibration with a cost-based threshold that cuts expected loss from $2,633 to $646 per 10,000 transactions; SHAP; 5-seed stability. Differentially private training (DP-SGD) with a measured privacy/utility curve; ONNX, int8 and Core ML export at p95 0.04 ms per transaction; Kafka to Spark Structured Streaming scoring at 2,000 events/s that reproduces the batch alerts exactly; PSI drift monitoring. The whole lifecycle also runs as a Vertex AI Pipeline (validate, train, evaluate with a promotion gate, Model Registry, endpoint, batch prediction, weekly schedule): one measured run at test PR-AUC 0.80, endpoint p50 68 ms, about $0.40. The stream also runs on Google Cloud as Pub/Sub → Dataflow (Beam, ONNX Runtime) → BigQuery with per-minute score-drift windows, flagging exactly the same 136 transactions as the batch evaluation.

    Python · XGBoost · PyTorch · Opacus · ONNX Runtime · Core ML · Kafka · Spark · Pub/Sub · Dataflow · BigQuery · Vertex AI Pipelines · Docker · GitHub Actions
  • 2026

    Retrieval-augmented question answering over 2,067 Wikipedia paragraphs (SQuAD dev) with an evaluation harness that scores answers against gold spans, not just an LLM judge. Compared BM25, dense embeddings (Vertex AI text-embedding-005 in BigQuery vector search) and a hybrid fusion: hybrid retrieves the right paragraph in the top 5 for 94.6% of questions, and Gemini 2.5 Flash reaches EM 0.78 / F1 0.86 with it against 0.09 / 0.14 with no retrieval, at $0.29 per 1,000 questions. Served from FastAPI on Cloud Run; a load test found the async route serialising requests and the fix took throughput from 2.4 to 12 requests/s on one instance. Every number from a real run.

    Python · Vertex AI · Gemini · BigQuery · FastAPI · Cloud Run · rank-bm25 · pytest · GitHub Actions
  • 2026

    A Java 21 / Spring Boot 3 service that runs YAML evaluation suites against AI agent outputs: strategy-pattern evaluators (contains, regex, JSON Schema, must-not-fire negative assertions, LLM judge), a sealed outcome type, repeat-N banded scoring, per-check precision and recall of alerts, and run-to-run regression diffs, persisted in PostgreSQL. 90 JUnit 5 tests with Mockito and a Testcontainers integration test, 84.7% line coverage. Benchmarked virtual threads against fixed pools on a simulated 50 to 200 ms judge: 7,205 cases/s at a 1,000 bound versus 247 for a 32-thread pool. Every number from a committed results file.

    Java 21 · Spring Boot · Spring Data JPA · PostgreSQL · Flyway · JUnit 5 · Mockito · Testcontainers · JaCoCo · Docker · GitHub Actions
  • 2026

    Formulated depth completion as a constrained optimization over the dense depth map (data fidelity, smoothness, edge-aware and coarse-consistency terms, non-negativity by projection). Over 5,000 NYU Depth V2 samples at 5% observed depth, the full objective cut RMSE 33% vs data-only; L-BFGS matched Adam's accuracy at 3.2x the speed. MSML604 team project.

    PyTorch · Adam · L-BFGS · NYU Depth V2
  • 2026
    Agentic Voice Assistant for Insurance Claims

    Vapi + n8n voice agent that authenticates callers, retrieves real-time claim status, and logs every call through Google Sheets webhooks, routing callers by claim outcome with stateful tool calls.

    Vapi · n8n · OpenAI APIs · REST · STT/TTS
  • 2026

    Frames grocery shopping as a 0/1 knapsack problem (price = weight, nutrition = value) solved with dynamic programming on Open Food Facts data. Live Streamlit app; code.

    Python · Dynamic programming · Pandas · Streamlit
  • 2026

    Hash-table indexes over a 1.2M-song Spotify dataset for fast artist, song, and album lookup, with a Streamlit search UI.

    Python · Hashing · Streamlit
  • 2025

    End-to-end pipeline on FAOSTAT food-balance data (12 countries, 2016–2023): feature engineering, Random Forest / Logistic Regression / XGBoost comparison with stratified CV, and SHAP analysis.

    Pandas · scikit-learn · XGBoost · SHAP
  • 2025
    RL for Warehouse Optimization

    Custom single- and multi-agent warehouse environments with DQN, PPO, and MAPPO policies for placement, picking, and restocking. Reward shaping cut path inefficiency by 18%.

    DQN · PPO · MAPPO · Multi-agent RL
  • 2024
    Multi-Object Tracking in Video

    YOLOv5 detection with Kalman filtering and Hungarian matching, reaching 82% MOTA and 76% IDF1 on the MOT dataset.

    YOLOv5 · Kalman filter · OpenCV
  • 2024
    Movie Recommendation System

    Hybrid ALS + autoencoder recommender on 15M+ interactions: +12% CTR, +15% watch time, −15% RMSE with Bayesian tuning.

    ALS · Autoencoders · PySpark
  • 2024
    Customer Segmentation & Recommendations

    K-means customer segmentation plus a collaborative-filtering recommender for online retail, boosting sales by 20% through personalized suggestions.

    scikit-learn · K-means · Collaborative filtering
  • 2024
    Abstractive Text Summarization

    Attention-based Seq2Seq Bi-LSTM on CNN/DailyMail, benchmarked against BERT and T5 with ROUGE, BLEU, and perplexity.

    PyTorch · Bi-LSTM · T5
  • 2024
    Smart Shopping Assistant

    OCR + RAG pipeline that extracts structured data from grocery receipts for expense categorization and semantic search.

    OCR · RAG · LLMs
  • 2024
    Real-Time Disaster Response System

    LSTM forecasting and CNN satellite-image analysis with Kafka streaming and live Streamlit heatmaps for resource allocation.

    LSTM · CNN · Kafka · Streamlit

Leadership & volunteering

Leadership

Aug 2023 – Sep 2024
Head of Logistics · Leaders of Tomorrow, MIT Manipal. Managed 40+ members and coordinated large-scale university events.
Mar 2023
Head of Organizing Committee · Revels cultural festival, MIT Manipal. Led 90+ members and streamlined resources and event planning.
Oct 2022
Head of Organizing Committee · TechTatva technical festival, MIT Manipal.

Volunteering

UMD
Volunteer Tutor · Maryland Mentor Corps. Math education and tutoring.
UMD
Volunteer · Terps for Change, Kids Achieve Club. Youth development and community outreach.
UMD
Member · Active Minds. Mental health awareness and campus engagement.
Sep 2024
Volunteer · Suvidha NGO. Fundraising for underprivileged children.
Nov 2023 – Jan 2024
Volunteer · Kerala State AIDS Control Society. Organized HIV/AIDS prevention and awareness programs.
May – Jul 2022
Volunteer Teacher · PARIVAR Society. Taught mathematics and led self-help groups for the elderly.

Certifications

Deep Learning Specialization · DeepLearning.AI, Aug 2024

Data Science Professional Certificate · IBM, Nov 2024

Data Science Specialization · Johns Hopkins University, Nov 2024

Awards

4th place, Cisco Forecast League · Feb 2024

Toolkit

LLMs & agents

RAG, LLM evaluation (LLM-as-judge), prompt engineering, fine-tuning, LangChain, FAISS, OpenAI / Azure OpenAI, Claude Code (subagents, skills, MCP), Vapi, n8n, Hugging Face TGI, LangFuse

Machine learning

PyTorch, TensorFlow, Keras, scikit-learn, XGBoost, LightGBM, Prophet, SHAP, OpenCV, spaCy, NLTK, ONNX

MLOps & backend

FastAPI, Flask, Docker, Kubernetes, AWS, Airflow, MLflow, Kafka, PySpark, REST APIs, webhooks, Pydantic, Git, Linux

Data & analytics

PostgreSQL, MySQL, MongoDB, Snowflake, ClickHouse, Metabase, Superset, Grafana, Power BI

Enterprise integrations

Genesys, Cisco PCCE / SIPREC, Salesforce, SSO / SCIM

Languages

Python, SQL, JavaScript, C++, R, Go, MATLAB, Bash

Methods

Deep learning, NLP, reinforcement learning, time-series forecasting, causal inference, survival analysis, A/B testing, recommender systems

Security & hardware

Kali Linux, Nessus, Wireshark, Nmap · Qt/QML, Simulink, HFSS

Let's talk agents, evals, or ML systems.

I'm open to AI/ML engineering roles, research collaborations, and conversations about building LLM systems that hold up in the real world.