Staff Data Scientist, Bangalore, India
- Cut median forecast error 21% (holdout WAPE 0.574 to 0.451) on a global consumer-devices manufacturer's weekly retail demand portfolio by rebuilding the model-selection criterion: WAPE-primary, gated by event-window WAPE at or below 1.2x WAPE, with a 1.5x per-series degradation cap, after proving the incumbent in-sample criterion disagreed with the true out-of-sample winner on 69 of 76 series.
- Classified 40,447 spares materials for a mining and industrial operator by Syntetos-Boylan demand type and routed each class to its own statsforecast family (AutoETS with forced-seasonal ZZA, AutoTheta, Croston-SBA, TSB) against a pooled global LightGBM, then reframed evaluation from statistical accuracy to inventory outcome: test fill rate 97.7% against the incumbent policy's 74.5% at lower cost, decided under a nested train, validation and test backtest with leakage tripwires and per-series conformal coverage audits.
- Shipped a proprietary tariff-classification agent, now in production at a North American customs brokerage: LLM description enrichment, hybrid BM25 and dense retrieval over pgvector and LanceDB, a Neo4j knowledge graph of customs rulings, and a written justification per code, served as a containerised FastAPI API on GCP Cloud Run. Cut cost to $0.09 per classification from $0.13 by re-architecting the monolithic ReAct agent into a LangGraph StateGraph of 7 nodes and 9 edges. Built its 2,603-item evaluation harness and instrumented the stack with Langfuse tracing; fine-tuned an open-source 8B model on a 2-node SLURM cluster as a cost alternative to frontier models.
- Built the rest of the modelling stack for the planning product: multi-echelon deep-learning forecasters in PyTorch for a global PC manufacturer, a multi-level LSTM predicting channel, distribution-centre and manufacturing demand jointly across 80 SKUs over 5 geographic regions, their countries and in-country distribution centres; a multi-agent demand-planning workflow coordinating root, anomaly, causal and forecasting agents under an Observe-Orient-Decide-Act loop; and a Bayesian anomaly service on rolling Normal-Inverse-Gamma regression with a Student-t posterior predictive, live as a FastAPI container on GCP Cloud Run.
- Profiled 36 SAP MM parquet tables covering 130M+ rows across 15 plants and recovered five years of demand history an in-house extract filter had silently dropped (51,461 of 60,022 lines) by running the complement of the team's own filter; caught a covariate leak worth 9.9 WMAPE points in a published benchmark, retracted the claim and recomputed all 16 sweep configurations leak-free; authored the forecasting methodology for a planning discipline that had no owner and is building the team, five engineers recruited to date.