Proving Ground — Ship a Model Only When the Evidence Says So
A template repository for releasing a model through explicit, config-driven gates, a shadow then canary stage, drift monitoring, a hash-chained audit log, and a tested one-command rollback. A worked motor-insurance claim-frequency example and an offline report viewer ship with it.
The problem & requirements
- Train a constant baseline, a Poisson GLM and LightGBM, then run every candidate through gates G1–G8 before it can be registered
- Promote through shadow and a 10% canary using stable-hash routing, with no promotion unless a passing gate report exists for that exact candidate
- Monitor windows of served traffic for drift, move ok → watch → alert, and write a rollback recommendation to the audit log
- Roll back to the previous champion with one command, verified through the serving health endpoint
- Render all release evidence into a single static, offline HTML report
- No force flag: every stage transition re-checks the gate report, the candidate hash, the data version and the current gates.yaml hash
- Tamper-evident audit: append-only log where each entry hashes the previous one, serialized canonically so Python and the viewer agree byte for byte
- Holdout discipline: the holdout is scored once per candidate, tuning uses cross-validation inside the training split only, and a ledger records it
- Honest data labeling: synthetic traffic and drift are always flagged as synthetic in every artifact and in the viewer banner
Scale & constraints
A template, not a production system, so the numbers below are what the evidence rests on: a real public dataset, a grouped holdout, and a demo cheap enough to run on every CI push.
Interfaces & Endpoints
Data model
| Entity | Key fields |
|---|---|
| Candidate | Project, model family, params, seed, data version, and a prediction fingerprint. Its hash is SHA-256 over the canonical JSON of those fields. |
| GateReport | Per-gate pass/fail/warn/skip plus the gates.yaml hash it ran under. Changing a threshold invalidates older reports. |
| Registry | File-based. Aliases candidate, champion and previous_champion, plus a release stage from registered through stable or aborted. |
| AuditEntry | Canonical JSON with prev_entry_hash and entry_hash. The first entry chains from 64 zeros. |
| MonitorWindow | Per-feature PSI, KS or chi-squared, new-category and out-of-range rates, prediction PSI, validation-failure rate, and label drift. Schema-versioned and flagged synthetic. |
The registry is a plain file rather than MLflow on purpose. Rollback is a two-field edit and the demo needs no database. MLflow is an opt-in mirror, never the source of truth.
Architecture
Every transition re-checks the evidence. A candidate cannot skip a stage, and there is no override.
Fig. 2a — The gates run in order and all must pass. A rejected candidate leaves an audit entry naming the gates it failed.
Drift detection feeds the same audit log that rollback and the report viewer read from.
Fig. 2b — The viewer is rendered by Python, escapes all input, and only uses JavaScript for tabs, sorting and optional in-browser chain verification.
Key decisions & trade-offs
Lessons learned
- Declaring each drift scenario's expected outcome before running it made the scenario matrix a real test. The recorded run met every declared expectation, and the misses that remained are listed as limits.
- Exercising rollback in CI turned it from a runbook claim into a tested path, including the idempotent second call.
- Hashing the gates config into every report closed the quiet loophole of loosening a threshold and promoting on an old green report.
- The GitHub Actions workflows ran locally before they ran on GitHub. I would push a minimal workflow on day one so CI is a real signal rather than something I wrote and have not watched execute.
- Detection thresholds were set from one dataset's reference noise. They are sensitive to window size, and I would test them against more than one dataset before treating the numbers as anything but a starting point.
- Exposure is not proportional to claims in this data, so it is also used as a covariate. That is a pragmatic fix, documented, but not a pricing-grade treatment.