Which Deploy Moved the Metric?
Conversion dropped 12% last week. Correlating a metric's step-change against your commit history turns a two-day investigation into a five-minute Slack thread, if you build the detector correctly.
Conversion rate dropped 12% last week. Every dashboard I’ve ever used will tell you that much. Not one of them will tell you it started Tuesday the 14th, and that the only thing that shipped Tuesday the 14th was a change to the checkout address form.
That gap is where analysts lose their afternoons. The workflow is always the same: notice the weekly number moved, pull the daily series, eyeball where the level shifted, open the repo, scroll git log around that date, squint at forty commit subjects, ask three engineers which one it was. Steps two through four are mechanical. I built stepchange to do them.
The detection side has three failure modes I kept hitting, and each one has a specific fix.
A single bad day in your baseline wrecks a naive threshold. Mean-plus-k-standard-deviations is the textbook approach, and it’s wrong in a specific, expensive way: one outlier day inflates the standard deviation and buries the real step underneath it. Median and MAD (median absolute deviation) instead of mean and stdev fixes this, and I have a test case in the suite where the naive version misses a real 15% step and the robust version catches it clean.
A perfectly flat baseline breaks the math a different way. Spread goes to roughly zero, and now every test day flags at infinite sigma. This is the single most common false positive I saw in practice, and the fix is almost embarrassingly simple: floor the spread at 2% of the baseline center.
Day-of-week cycles will eat you alive if you don’t handle them explicitly. Weekly seasonality is the biggest source of false “step changes” in business metrics, full stop. The tool defaults to windows sized in multiples of seven days; there’s also a mode that fits weekday factors over a longer history and reports them so you can sanity-check what it assumed.
One more distinction that matters more than it sounds: a step and a spike are different findings and need different confidence. A single bad day is real information, but correlating a one-day blip against your commit history the same way you’d correlate a sustained level shift produces confident-sounding nonsense. The tool requires three-plus consecutive flagged days before it calls something a step; anything shorter goes in a separate spikes list.
None of this proves causation, and the tool says so in its own output. What it gives you is a ranked shortlist of candidate commits, scored on proximity to the pivot date, keyword overlap with the metric, touched file paths, and blast radius, with an honest confidence label attached. “Strong” or “weak,” not a made-up percentage. That’s the whole point: turn a two-day investigation into a five-minute Slack thread where someone says “yeah, that’s probably it” and goes and checks.
Measured against synthetic ground truth (fixed seeds, thirty trials per cell), recall on a clean 10% step with no noise is 100%. Recall on a 5% step with typical day-of-week noise drops to single digits. The tool is honest about that gap too: small, noisy steps are genuinely hard, and pretending otherwise is how you end up debugging the wrong deploy.