Question → grain → transformation → test → API → Evaluation
fct_player_week retains model and candidate identity; fct_weekly_eval aggregates only outcome-bearing rows at an explicit evaluation-window grain.
Explore how raw statistics and model artifacts become tested player-week facts, versioned decision metrics, and useful football features.
Live values from the API; MAE and fold bars come from evaluation artifact ….
DATA PLATFORM
Start with a player-week key, inspect the SQL that computes error, open the reconciliation test, and see which API and football component consumes the result—even when the live prediction API is unavailable.
fct_player_week retains model and candidate identity; fct_weekly_eval aggregates only outcome-bearing rows at an explicit evaluation-window grain.
Inspect the four-period incremental lookback, the honest pre-capture fallback, and both decision-policy versions without mistaking exported v2 for served v1.
Read the decisionsTHE CASE STUDY
Fantasy lineup selection is a safe analogue for real-world risk work: incomplete information, asymmetric errors, shifting populations, and a hard decision deadline every week.
A point estimate does not answer this. The useful product combines an expected value, an interval around it, and an auditable trail from the published number back to the data and code that produced it.
SYSTEM DESIGN
nflverse weekly player data, the only source
As-of signals from each player’s own prior weeks
RandomForest champion, XGBoost challenger, per position
Gaussian-mixture preseason tiers from prior-season aggregates
Weekly GitHub Actions job: publish, hold, or promote
A forward-time holdout, rolling-origin folds, and a causal baseline are fixed before any model is compared.
MAE, median AE, RMSE, and within-band rates are stored next to their formulas in the evaluation artifact.
Hashed inputs, versioned features, a manifest that names the champion, and a policy that can refuse to publish.
A challenger is trained beside the champion every run and promoted only when the policy says so.
ENGINEERING EVIDENCE
Floors, ceilings, and an explicit policy over them
Forward-time splits, rolling-origin folds, causal baselines
Versioned artifacts, hashed inputs, metric definitions
FastAPI, containers, committed manifests
Every figure on this site is read from a committed evaluation artifact. Nothing is typed into the UI by hand, and the API names the model, feature set, and evaluation behind each response so a reader can trace it to repository evidence.