CV

Rajesh Rajendran, Staff Data Scientist

7 years building systems, 6+ years building models.

At Avathon I build demand forecasts, inventory policy, production plans and LLM agents for enterprise supply chains. The roots: robotics state estimation, signalling software, then causal inference on rail data.

Read the CVDownload PDF

Bangalore, India

Now, September 2026:multi-echelon inventory optimisation for a mining and industrial operator, a graph against relational database benchmark on a US electric vehicle manufacturer's bill of materials, hiring, and Avathon's new AI Center of Excellence, as one of its three seeding members.

Selected work

Eleven pieces of work, each with its number and its basis

detect

01

Unusual, given the drivers

Bayesian anomaly detection on demand signals

  • I built a service that asks of demand what I first asked of train telemetry: is this point out of line with what I expected? A rolling Normal-Inverse-Gamma regression on promotion, price, seasonality and weather gives a Student-t posterior predictive at each step. A two-sided posterior predictive p-value flags the tail.
  • The expectation moves with the drivers. A promotion week expects more demand, so only the part the promotion does not explain gets flagged. The service runs as a FastAPI container on GCP Cloud Run, and I tagged it v1 on 8 September 2025 with the method written up on Confluence.

Basis Build and deployment facts from my journal, August to October 2025, tested on generated demand signals. I marked the detection evaluation done on 15 October 2025 without recording a figure, so no detection rate is quoted.

deployed as a serviceon GCP Cloud Run, in Avathon's demand planning platform

Python, Normal-Inverse-Gamma updating, scikit-learn BayesianRidge, PyMC, FastAPI, Docker, GCP Cloud Run

02

What the rules never fired on

Unsupervised anomaly detection on train telemetry

  • I developed unsupervised anomaly detection on train odometry and ATC subsystem signals using autoencoder architectures. It reached over 99% recall and surfaced failures that rule-based monitoring missed.
  • A statistical fault detection model that I packaged and deployed into the fleet management system improved maintenance efficiency by 30%.

Basis Measured by the Alstom team, 2020 to 2023.

in productionon a fleet management system

Python, PyTorch, TensorFlow, autoencoders

explain

03

Which driver moved demand

Attributing a flagged anomaly inside a multi-agent planning co-pilot

  • At Alstom the question was which fault caused a failure. Here it is which driver caused a flagged anomaly. I drop each candidate in turn (promotion, price, weather, competitor stockout), refit on the anomaly window and rank them by ΔBIC or Bayes factor. The posterior coefficient and its credible interval size each effect.
  • Between 29 September and 1 October 2025 I built a planning co-pilot around it. A root agent passes a planner's request to anomaly, causal and forecasting agents through structured tool calls, under an Observe Orient Decide Act loop, and GPT-4o returns the recommendation with its explanation.

Basis Method from my journal, September 2025; the four agents built 29 September to 1 October 2025. Tested on generated demand in which a weather shock hit only some states. The attribution evaluation was still open on 15 October 2025, so no accuracy is quoted.

presented internallythe attribution method, at an Avathon monthly all-hands

Python, scikit-learn BayesianRidge, ΔBIC, Bayes factors, LangChain, GPT-4o, structured tool calls

04

A ranked cause, not a correlation

Causal Bayesian networks on urban transit failures

  • I built a root cause system using causal discovery and a Causal Bayesian Network. It answers interventional and counterfactual queries over Bayesian posteriors, so engineers get causes ranked by probability instead of correlations.
  • Troubleshooting time fell 86%, and the method became a UAI 2024 paper on time-dependent counterfactual root cause analysis.

Basis Measured by the Alstom team, 2023 to 2025.

in productionon an urban transit fleet

Python, causal discovery, Causal Bayesian Networks, do-calculus, counterfactual queries

Read the paper

forecast

05

Select on the holdout

What the fit-time winner could not tell me

  • On 69 of 76 series (90.8%), the model that won at fit time was not the model that won on the holdout, so I retired the in-sample criterion.
  • I rebuilt the criterion as WAPE primary, gated by event window WAPE at or below 1.2x WAPE and capped at 1.5x per-series degradation. Median holdout WAPE went from 0.574 to 0.451 across three pipeline reruns.

Basis Median across 76 weekly retail series, holdout window, single split.

in productionfor a global consumer-devices manufacturer

Python, LightGBM, XGBoost quantile objective, statsmodels ETS and SARIMAX, pandas, parquet

06

Products that inherit demand

Lending a new product the history of the one it replaces

  • I mapped each PC product family to the models it replaces, including many-to-many transitions, and fed the predecessor in as a feature beside product age and days left in the lifecycle. A new model with little history borrows signal from the old one. A family with a single record and no predecessor on file has nothing to learn from and nothing to borrow, so when I reviewed results across all families I marked it as not forecastable from history.
  • The same question comes up wherever new products replace old ones. At a global consumer-devices manufacturer, one launch series overshot about 12 times over (WAPE 11.76) and another was forecast at zero because of pre-release clipping. I concluded that capped and launch series should be scored by cap correctness and launch-ramp shape, not by WAPE.

Basis Multi-echelon forecast of 80 SKUs over 5 regions, 2025, a scale fact: no accuracy was recorded for the transition features. Launch series: holdout WAPE, 2026.

deliveredto a global PC manufacturer

Python, pandas, XGBoost, LightGBM, SARIMAX, Prophet, PyTorch LSTM

07

Forty thousand materials, four demand classes

Routing a spares catalogue by its time-series properties

  • I classified 40,447 maintenance and spares materials by Syntetos-Boylan demand type, with quantity share smooth 34%, lumpy 22%, intermittent 21% and erratic 18%. I flagged the 2,888 of 40,447 (7.1%) that sit on a class boundary as unstable rather than settled.
  • AutoETS was collapsing to a level model, so I added a forced seasonal AutoETS entry to the model menu. It took one material from MASE 0.77 to 0.51 and another from 1.65 to 1.31.

Basis Per-material panel, one site, holdout MASE per series.

shipped to customera mining and industrial operator, where production model selection reads it

Python, statsforecast with AutoETS including forced seasonal ZZA, AutoTheta, Croston-SBA and TSB, conformal prediction intervals, LightGBM, Plotly

decide

08

Fill rate is the verdict

Judging a forecast by the inventory it produces

  • I moved evaluation from forecast error to inventory outcome and shipped the verdict: deploy on 9 smooth series, where test fill rate was 97.7% against the incumbent policy's 74.5% at lower inventory cost, and keep the incumbent on intermittent and lumpy series.
  • A pre-deployment guardrail caught a 54 percentage point fill rate regression on one material before rollout. I traced it to a reorder point that had collapsed from 3,618 to 489.

Basis Test window 2026-H1, nested train, validation and test backtest with leakage tripwires, 57 tests.

decision deliveredto a mining and industrial operator

Python, statsforecast, LightGBM, NumPy and SciPy lognormal fitting, pytest

09

What to build this week, and how to ship it

One program for the production plan and the freight plan

  • I built a two-stage plan. A demand forecast by SKU, channel, geography and week feeds a mixed integer linear program that plans the Master Production Schedule and the pack-out jointly, with one colour per production week and weeks-of-supply targets of 12 for channels 1 and 2 and 13 for channel 3. Each week it picks air, fast boat or ocean under a 7-week ocean lead time and 14,409 units a week of pre-build capacity.
  • Demand averaged 13,033 units a week against that cap (90.5%), but 23 of 91 weeks ran over it, peaking at 80,059 in the Presidents Day week, so when to build ahead is the decision. Replaying 52 weeks of actual demand, the plan cost $508K in freight against $1.53M all air and $763K all fast boat. That equals all ocean: with the pre-build scheduled, ocean could carry the year.

Basis 52-week backcast replay of actual demand, channel 3 (about 95% of sales) in one region, run 21 April 2026.

proof of conceptbackcast replay for a global consumer-devices manufacturer

Python, PuLP, CBC, pandas, XGBoost, SARIMAX

verify

10

The leak I found in my own benchmark

Feature attribution against a published result

  • A feature importance export showed a lookahead aggregate at 68.8% of model gain against lag_1 at 0.08%. That made the model a scale estimator rather than a model of demand dynamics.
  • I measured the leak at 9.9 and 6.2 WMAPE points on intermittent and lumpy demand, retracted the published claim in writing, and recomputed all 16 sweep configurations without the leak. The ranking held.

Basis Intermittent and lumpy quadrants, walk-forward backtest, one site.

shipped to customercorrected findings reissued

Python, LightGBM, feature attribution, pandas

orchestrate

11

Ten digits, one defensible answer

Hierarchical tariff classification under customs rules that keep changing

  • A tariff code is a walk down a hierarchy, from chapter to heading to subheading to the full ten digits. Each step is governed by the General Rules of Interpretation and by thousands of binding customs rulings, and the schedule and the rulings change every year. I built a proprietary classification agent that enriches a sparse part description, retrieves candidate headings with hybrid BM25 and dense search over pgvector and LanceDB, walks the subheadings against a Neo4j knowledge graph of rulings, and writes a justification with every code it assigns.
  • Re-architecting it from a monolithic ReAct loop into a LangGraph StateGraph of 7 nodes and 9 edges took cost per classification from $0.13 to $0.09. A 2,603-item evaluation harness, sampled from three months of live traffic with Langfuse tracing on every run, decides whether a change ships. One trace showed the model spending its full 4k token allowance on thinking, which a lean prompt and a tool fetch fixed.

Basis Cost per classification measured after the change; harness size as sampled. No accuracy figure is published.

in productionat a North American customs brokerage

LangGraph StateGraph, LangChain, Langfuse, Gemini 2.5 Pro, GPT-4o, pgvector, LanceDB, Neo4j, FastAPI, Docker, GCP Cloud Run; an 8B open model fine-tuned on a 2-node SLURM cluster as a cost alternative

More work

Range beyond the eleven, including the software years

  1. Five years of history the filter had hidden

    • An in-house extract filter had silently dropped 51,461 of 60,022 lines, five years of demand history back to 2021-05. Running the complement of the filter recovered all of it.
    • From the recovered 7.97M-row extract I classified 6,148 materials by demand visibility, separating 930 requires-forecast materials holding 60.3% of value from 5,218 order-on-demand materials.

    Shipped to a mining and industrial operator. The incident became a standing rule for the group.

    Python, pandas, SAP AFKO, RESB, MATDOC and MARC tables

  2. The review the green suite could not do

    • A pre-merge review of a new production forecast engine, 50 files and +6,440 lines, returned 18 verified findings, 6 of them critical, and 0 refuted.
    • All 6 critical findings were silent divergences from the source notebook. A passing suite of 78 tests could not see them, because its fixtures were dense and synthetic. The first fix I recommended was a golden-run parity test.

    Gate applied before merge.

    Python, pytest, uv, git worktrees

  3. One network, three levels of demand

    • I built multi-echelon deep learning forecasters in PyTorch: a multi-level LSTM that predicts channel, distribution centre and manufacturing demand jointly rather than as three separate models, across 80 SKUs over 5 geographic regions, their countries and in-country distribution centres.
    • On the same dataset I benchmarked SARIMAX, Prophet, NeuralProphet, gradient boosted trees, Bayesian neural networks and the Chronos foundation model. MAPE for the machine learning tier was 39.9 at shipdate, rising to 61.7 at the deepest disaggregation, against 46.4 to 72.4 for SARIMAX.

    Delivered to a global PC manufacturer.

    PyTorch, LSTM, MLP, torchbnn, XGBoost, LightGBM, CatBoost, Chronos, AutoGluon

  4. Linear programs for the parts that are not forecasting

    • For a bulk logistics operator I delivered set partitioning and facility assignment formulations in PuLP, integrated with an existing Scala scheduler.

    Logistics optimizer integrated.

    Python, PuLP, Scala

  5. Replenishment as a control problem

    • As a proof of concept I framed inventory replenishment as sequential decision making. A simulation API plays out demand, lead time, holding cost and stockout cost week by week; a newsvendor rule and a heuristic reorder rule are the baselines; and a reinforcement learning policy trains against the same simulator.
    • The public demo is on this page. The drawer on the fill rate card runs newsvendor, an (s,S) rule and a Q-learning benchmark on one synthetic part across 500 simulated years, and shows fill rate against cost with its spread.

    Proof of concept and benchmark only, not deployed. The next step is a proof of concept on a live spares catalogue.

    Python, PyTorch, NumPy, simulation API, newsvendor, (s,S), Q-learning

  6. Before the models, the simulators

    • For four years I built simulation and test software for railway signalling. The main piece was a simulator for iVPI-based interlocking systems that replicated logic evaluation, sensor input, relay activation and vital to non-vital parameter exchange, later extended with failover, hot standby and synchronised variable states.
    • I also built automated factory acceptance testing and the protocol interfaces behind it, including a WCF service that exposes data through a SCADA OPC UA server, and a portable diagnostic application that reads datalogs, event faults and snapshots off a train car ATC module over direct serial and Ethernet TCP/IP.

    Shipped inside Alstom Transport, 2015 to 2020.

    C++ with the Microsoft Foundation Class framework, C#, object-oriented design, socket programming, WCF, SCADA OPC UA

Method

Three rules, each backed by a card above

  1. Select on the holdout.

    A model earns its place on data it has not seen. When the fit-time winner and the holdout winner disagreed on 69 of 76 series, I retired the fit-time criterion.

    See card 05
  2. Retract in public.

    When I find an error in my own published result, I say so in writing and recompute. The covariate leak in my benchmark was worth 9.9 WMAPE points on intermittent demand, so I retracted the claim and reran all 16 configurations.

    See card 10
  3. Judge by the decision.

    I score a forecast by the inventory it produces. On the spares work that meant test fill rate and cost, and the incumbent policy stayed on the intermittent and lumpy series.

    See card 08
Rajesh Rajendran

About

Rajesh Rajendran

I studied electronics and instrumentation in Chennai and process automation at TU Dortmund. Then came state estimation for humanoid robots at the German Aerospace Center, and MATLAB simulations for a cancer diagnostic device at Blue Triangle.

At Alstom I wrote railway signalling simulators, then turned to the data: anomaly detection on train telemetry, and a causal root cause system that became a UAI 2024 paper. I was a founding member of a data science team that grew to 20.

At Avathon I forecast demand across a 40,447-material spares catalogue, judge inventory policy by fill rate, and plan production and freight with mixed integer programs. I built a proprietary tariff classification agent now in production, and Bayesian anomaly detection on Cloud Run under a multi-agent planning co-pilot. I fine-tuned an 8B model on a 2-node SLURM cluster as a cost alternative to frontier models. I recruited five engineers, run weekend AI workshops and brought Claude Code into the team.

The journey continues: I am one of three seeding members of Avathon's AI Center of Excellence in Bangalore, set up in August 2026 for Physical AI and frontier models in autonomous industrial operations, and I am reading reinforcement learning.

Timeline

Every role, 2012 to now

  1. Feb 2025 to present

    Staff Data Scientist, Avathon (formerly SparkCognition)

    Demand forecasting, inventory policy, production and freight planning, and LLM agents for enterprise supply chain.

    • Built the SAP MM data foundation for a mining and industrial operator: 36 parquet tables, 130M+ rows, 15 plants and about 5.7 years of history.
    • Defined the Clear-to-Build layer for a US electric vehicle manufacturer: bill of materials explosion, probabilistic ETA, runout projection and a cutover readiness score, taken to a CEO-level review.
    • Fine-tuned an open 8B model on a 2-node SLURM cluster as a cost alternative to frontier models.

    Bangalore, India

  2. Apr 2023 to Feb 2025

    Data Scientist Sr, Alstom Transport

    Causal root cause analysis for urban transit failures, and anomaly detection on real-time transit data with Bayesian networks and statistical process control.

    Bangalore, India

  3. Jan 2020 to Mar 2023

    Data Scientist, Alstom Transport

    Anomaly detection on train odometry and ATC telemetry, and log analytics on Spark with Scala that reduced failure detection time by 53%.

    Bangalore, India

  4. Apr 2018 to Jan 2020

    Software Architect, Alstom Transport

    Diagnostic and simulation tools for train control subsystems in C++ and C#, including a portable diagnostic application for the ATC module and a shared signalling layout editor.

    Bangalore, India

  5. Sep 2015 to Mar 2018

    Software Designer, Alstom Transport

    Simulators and automated factory acceptance testing for railway signalling equipment.

    Bangalore, India

  6. Jul 2014 to Jul 2015

    Product Engineer, Blue Triangle Innovations

    MATLAB simulations supporting the prototyping of a cancer diagnostic device.

    Ruhr Region, Germany

  7. Dec 2012 to Jun 2014

    Research Assistant, German Aerospace Center (DLR)

    State estimation of the under-actuated degrees of freedom in humanoid robots, in MATLAB and Simulink.

    Munich, Germany

  • M.Sc. Process Automation, Technical University of Dortmund, Germany2010 to 2013
  • B.E. Electronics & Instrumentation, MIT, Anna University, India2006 to 2010

Recognition

A paper, a seeding membership, two awards

  • PaperUncertainty in Artificial Intelligence (UAI), 2024, Barcelona

    Industrial-Grade Time-Dependent Counterfactual Root Cause Analysis through the Unanticipated Point of Incipient Failure

    Co-author. The paper sets out the causal method behind the root cause card above.

  • MembershipAvathon, Aug 2026, Bangalore

    Seeding member, AI Center of Excellence

    One of three seeding members of the centre, announced on 27 August 2026 for research and development on Physical AI and frontier models for autonomous industrial operations. The announcement confirms the centre; the membership is my own statement.

    Read the announcement
  • AwardAlstom, Mar 2024

    World Class Expert

    Recognised for developing data solutions.

  • AwardAlstom, Aug 2023

    Winner, SoftWar 2.0

    Applied AI to estimate odometry system tuning parameters. I drove the product design and the presentations.

CertificationsMachine Learning Specialization, DeepLearning.AIFoundations of Causality, causaLensMathematics for Machine Learning and Data Science, DeepLearning.AIData Science Specialization, Johns Hopkins University

Teaching

Trainings, workshops and white papers

I teach as much as I build. At Alstom I owned the causal inference methodology and presented it to cross-functional teams.

At Avathon I run the weekend AI workshops for colleagues who do not work in AI. I brought Claude Code into the team by researching each release and showing the team how to work with it, and I write internal white papers on what we learn.

Stack

What I build with

Forecasting

  • statsforecast
  • AutoETS
  • AutoTheta
  • Croston-SBA
  • TSB
  • SARIMAX
  • Prophet
  • Chronos
  • conformal prediction
  • WAPE
  • MASE

Machine learning

  • LightGBM
  • XGBoost
  • CatBoost
  • scikit-learn
  • PyTorch
  • LSTM
  • autoencoders

Causal and Bayesian

  • Causal Bayesian Networks
  • causal discovery
  • do-calculus
  • PyMC
  • Normal-Inverse-Gamma updating
  • CausalImpact

Agents and retrieval

  • LangGraph
  • LangChain
  • Langfuse
  • pgvector
  • LanceDB
  • Neo4j
  • hybrid BM25 and dense retrieval

Decisions and serving

  • (s,S) policy
  • newsvendor
  • PuLP
  • FastAPI
  • Docker
  • GCP Cloud Run
  • Airflow
  • Python
  • SQL

ReadingSutton and Barto, through the University of Alberta reinforcement learning specialization.

The full list is on the CV page.

Contact

Talk to me

I am open to staff and principal level applied science roles, research collaboration on causal and probabilistic forecasting, and consulting on supply chain AI.

Read the CVDownload PDF

Code I write for learning is on GitHub.

What the rules never fired on: the working

Illustrative reconstruction on synthetic data.

A ranked cause, not a correlation: the working

Illustrative reconstruction on synthetic data.

Select on the holdout: the working

One synthetic part, mechanism demo. Not the 76-series portfolio quoted above.

Fill rate is the verdict: the working

One synthetic part, mechanism demo. Not the production fill rate result quoted above.

Ten digits, one defensible answer: the working

Scripted trace. Node topology is real; per-step tokens, latency and cost are chosen to land on the measured $0.13 and $0.09 totals.