Projects›Proving Ground — Gated Model Release Template
Case Study · MLOps · Release Engineering

Proving Ground — Ship a Model Only When the Evidence Says So

A template repository for releasing a model through explicit, config-driven gates, a shadow then canary stage, drift monitoring, a hash-chained audit log, and a tested one-command rollback. A worked motor-insurance claim-frequency example and an offline report viewer ship with it.

PythonLightGBMFastAPIPanderaDVCEvidentlyGitHub Actionspytest
Role
Solo design & implementation
Domain
MLOps · insurance claim frequency
Primary Services
FastAPI · LightGBM · GitHub Actions
01

The problem & requirements

Functional
  • Train a constant baseline, a Poisson GLM and LightGBM, then run every candidate through gates G1–G8 before it can be registered
  • Promote through shadow and a 10% canary using stable-hash routing, with no promotion unless a passing gate report exists for that exact candidate
  • Monitor windows of served traffic for drift, move ok → watch → alert, and write a rollback recommendation to the audit log
  • Roll back to the previous champion with one command, verified through the serving health endpoint
  • Render all release evidence into a single static, offline HTML report
Non-functional
  • No force flag: every stage transition re-checks the gate report, the candidate hash, the data version and the current gates.yaml hash
  • Tamper-evident audit: append-only log where each entry hashes the previous one, serialized canonically so Python and the viewer agree byte for byte
  • Holdout discipline: the holdout is scored once per candidate, tuning uses cross-validation inside the training split only, and a ledger records it
  • Honest data labeling: synthetic traffic and drift are always flagged as synthetic in every artifact and in the viewer banner
02

Scale & constraints

A template, not a production system, so the numbers below are what the evidence rests on: a real public dataset, a grouped holdout, and a demo cheap enough to run on every CI push.

8
Gates, G1 data validation through G8 model card, all of which must pass
57
Automated tests covering core, claims example and viewer, including rollback exercised in CI
84s
End-to-end demo on a laptop: train, gate, shadow, canary, promote, drift, rollback, report
102,035
Policies in the grouped holdout, scored with exposure weighting and 95% bootstrap CIs
03

Interfaces & Endpoints

ModelProject (protocol)
The one interface a fork implements. The template core never imports examples/, which a structure test enforces.
make demo
Runs the full release story: train, gate, reject a degraded challenger, shadow, canary, promote, inject drift, roll back, verify the audit chain, build the viewer.
POST /predict
Validates the record against the project Pandera schema. Invalid records get a structured error, are counted, and become a monitor signal.
GET /health
Reports the serving model hash. Rollback calls it to verify the restored champion is actually the one answering.
POST /admin/reload
Local-only, token-guarded hot reload used by rollback. No image redeploy, and a second rollback call is a no-op.
04

Data model

EntityKey fields
CandidateProject, model family, params, seed, data version, and a prediction fingerprint. Its hash is SHA-256 over the canonical JSON of those fields.
GateReportPer-gate pass/fail/warn/skip plus the gates.yaml hash it ran under. Changing a threshold invalidates older reports.
RegistryFile-based. Aliases candidate, champion and previous_champion, plus a release stage from registered through stable or aborted.
AuditEntryCanonical JSON with prev_entry_hash and entry_hash. The first entry chains from 64 zeros.
MonitorWindowPer-feature PSI, KS or chi-squared, new-category and out-of-range rates, prediction PSI, validation-failure rate, and label drift. Schema-versioned and flagged synthetic.

The registry is a plain file rather than MLflow on purpose. Rollback is a two-field edit and the demo needs no database. MLflow is an opt-in mirror, never the source of truth.

05

Architecture

Release pipeline

Every transition re-checks the evidence. A candidate cannot skip a stage, and there is no override.

🧪
Train candidates
baseline · GLM · LightGBM
↓
#
Candidate hash
params · seed · data version · fingerprint
↓
Gates G1–G8 · one holdout evaluation per candidate
📋
G1 Data
schema · no id in two splits
📏
G2 Floors
absolute metrics
📉
G3 Regression
paired bootstrap vs champion
🎯
G4 Calibration
🍰
G5 Slices
age · region
⏱
G6 Budget
p95 latency · size
🔁
G7 Repro
same seed, same hash
📄
G8 Model card
↓
✗
Rejected, audited
named failing gates
or
✓
Registered
gate report attached
↓
👥
Shadow
champion answers, candidate logged
→
🐤
Canary 10%
sha256 bucket routing
→
🏆
Champion
promoted

Fig. 2a — The gates run in order and all must pass. A rejected candidate leaves an audit entry naming the gates it failed.

Monitor, rollback & evidence

Drift detection feeds the same audit log that rollback and the report viewer read from.

🌐
Serving
FastAPI · /predict · /health
↓
📈
Windowed drift monitor
PSI · KS · chi² · label drift
↓
👀
watch
surfaced everywhere
or
🚨
alert
hold + rollback recommendation
↓
↩
Rollback
alias flip · hot reload · health check
↓
Evidence
🔗
Hash-chained audit log
🧾
Incident before/during/after
Output
🖥
Static report viewer
🔐
CSP + script hash, offline
9 drift scenariosEvidently (optional)GitHub PagesMLflow mirror (opt-in)

Fig. 2b — The viewer is rendered by Python, escapes all input, and only uses JavaScript for tabs, sorting and optional in-browser chain verification.

06

Key decisions & trade-offs

Trap. The first drift alert fired on noise. Alerting whenever the actual-to-expected claims ratio moved 20% raised an early false alert on silent_decay. With roughly 130 claims per window, noise alone moves it about 10%. Label drift now requires the 99% bootstrap CI to exclude the reference value, which costs sensitivity to small base-rate shifts and is stated as a limit.
Invariant. No promotion without evidence, and no force flag. Promotion requires a passing gate report that names this candidate hash, this data version and the current gates.yaml hash. Editing a threshold invalidates every older report, so gates must re-run rather than being quietly re-interpreted.
Trade-off. File registry first, MLflow as a mirror. Rollback is a two-field file edit with a hot reload and a health check, and the whole demo needs no database. The cost is that the registry is single-writer and local.
Constraint. The dataset has no timestamps. freMTPL2freq has no date column, so the split is grouped by policy and risk profile rather than by time. Every artifact says temporal generalization was not evaluated, and no policy id appears in two splits.
Detail. Canonical serialization shared across two languages. The audit chain uses sorted keys, fixed separators and no floats so the Python writer and the viewer JavaScript compute identical hashes, and a test checks that they do. A hostile-input test once caught U+2028 splitting an entry when the log was read with splitlines().
Honesty. Limits are written down, not implied. The audit log is tamper-evident, not tamper-proof. Traffic and drift are simulated. Driver age and region are kept but reported by slice, with the proxy risk stated in the model card. Fairness and compliance are explicitly not claimed.
07

Lessons learned

💡 What actually held up
  • Declaring each drift scenario's expected outcome before running it made the scenario matrix a real test. The recorded run met every declared expectation, and the misses that remained are listed as limits.
  • Exercising rollback in CI turned it from a runbook claim into a tested path, including the idempotent second call.
  • Hashing the gates config into every report closed the quiet loophole of loosening a threshold and promoting on an old green report.
↺What I'd do differently
  • The GitHub Actions workflows ran locally before they ran on GitHub. I would push a minimal workflow on day one so CI is a real signal rather than something I wrote and have not watched execute.
  • Detection thresholds were set from one dataset's reference noise. They are sensitive to window size, and I would test them against more than one dataset before treating the numbers as anything but a starting point.
  • Exposure is not proportional to claims in this data, so it is also used as a covariate. That is a pragmatic fix, documented, but not a pricing-grade treatment.
🔁If I started this over: put the audit chain and the canonical serialization in first, then build every other stage as a function that appends to it, so the evidence trail is the spine of the system rather than something the viewer reconstructs afterward.
System
Proving Ground — Gated Model Release Template
Primary services
FastAPI · LightGBM · GitHub Actions
Status
Open source template, live sample report on GitHub Pages
Type
MLOps release engineering

Interested in the architecture behind this or another project?