Applied data science + analytics engineering

Football predictions, backed by inspectable data models.

Explore how raw statistics and model artifacts become tested player-week facts, versioned decision metrics, and useful football features.

Modelloading…
Data through—
StandardNo metric without provenance
FLOOR POLICY
FROM API
MAE · CHAMPION—reading evaluation artifacts
WITHIN ±3 POINTS—
0 pts
More candidatesSafer floors
…Loading the latest predictions

Live values from the API; MAE and fold bars come from evaluation artifact ….

PYTHON SCIKIT-LEARN XGBOOST FASTAPI SQL · DBT · DUCKDB NEXT.JS GITHUB ACTIONS
DP

DATA PLATFORM

Follow one metric all the way into the product.

Start with a player-week key, inspect the SQL that computes error, open the reconciliation test, and see which API and football component consumes the result—even when the live prediction API is unavailable.

METRIC TRACE

Question → grain → transformation → test → API → Evaluation

fct_player_week retains model and candidate identity; fct_weekly_eval aggregates only outcome-bearing rows at an explicit evaluation-window grain.

Trace within ±3
MODELING JUDGMENT

Corrections, captured history, and compatible contracts

Inspect the four-period incremental lookback, the honest pre-capture fallback, and both decision-policy versions without mistaking exported v2 for served v1.

Read the decisions
01

THE CASE STUDY

From a ranking model to a measurable decision strategy.

Fantasy lineup selection is a safe analogue for real-world risk work: incomplete information, asymmetric errors, shifting populations, and a hard decision deadline every week.

THE BUSINESS QUESTION

Who should we act on, and what evidence would change that decision?

A point estimate does not answer this. The useful product combines an expected value, an interval around it, and an auditable trail from the published number back to the data and code that produced it.

decision = f(prediction, floor, ceiling) · every input versioned
SUCCESS CRITERIA
  • Beat the baselineLower error than a causal trailing mean of the player's own points
  • Leak nothingFeatures use only weeks that were already scored
  • Publish on evidenceA weekly job publishes, holds, or promotes based on the evaluation
  • Trace every figureModel version, feature version, and eval id on every response
02

SYSTEM DESIGN

An observable path from raw data to a published projection.

01

Ingest

nflverse weekly player data, the only source

02

Features

As-of signals from each player’s own prior weeks

03

Model

RandomForest champion, XGBoost challenger, per position

04

Tiers

Gaussian-mixture preseason tiers from prior-season aggregates

05

Publish

Weekly GitHub Actions job: publish, hold, or promote

Evaluation before optimization

A forward-time holdout, rolling-origin folds, and a causal baseline are fixed before any model is compared.

Metrics with definitions attached

MAE, median AE, RMSE, and within-band rates are stored next to their formulas in the evaluation artifact.

Guardrails over magic

Hashed inputs, versioned features, a manifest that names the champion, and a policy that can refuse to publish.

Built for iteration

A challenger is trained beside the champion every run and promoted only when the policy says so.

03

ENGINEERING EVIDENCE

Inspect the contracts behind the published output.

01Decision science

Floors, ceilings, and an explicit policy over them

02Modeling

Forward-time splits, rolling-origin folds, causal baselines

03Analytics engineering

Versioned artifacts, hashed inputs, metric definitions

04Production

FastAPI, containers, committed manifests

PUBLICATION PRINCIPLE

Every published figure needs provenance.

Every figure on this site is read from a committed evaluation artifact. Nothing is typed into the UI by hand, and the API names the model, feature set, and evaluation behind each response so a reader can trace it to repository evidence.