The Brier score, made into a running habit

The Calibration Gym

There's a real line between having an opinion and having a checkable one: state an honest probability before you know the answer, not just a pick, and you can actually measure how well-calibrated your confidence really is — whether your "60%" picks hit close to 60% of the time, or closer to a coin flip. That's what this logs, using the same Brier scoring rule weather services and election modelers are graded on.

Log a forecast before you know the answer. Come back and resolve it once you do. The stated probability can never be edited after logging — only resolved or deleted — because a number quietly revised after the fact defeats the entire point of tracking this.

Log a new forecast

This is a record of what you believe right now, not a bet. Nothing here is sent anywhere — it's saved only in this browser.

Pending — awaiting a result

Resolved

Your calibration, so far

Your mean Brier score across every resolved forecast — lower is better. Needs at least one resolved forecast.

A flat, uninformative 50%-every-time forecaster on coin-flip events averages 0.25. A perfect forecaster averages 0.

Reliability curve

Each dot is a bucket of your stated probabilities (0–10%, 10–20%, …). Its horizontal position is your average stated confidence in that bucket; its vertical position is how often those forecasts actually happened. A dot sitting on the diagonal line means that confidence level is well-calibrated. Dots below the line mean overconfidence in that range; above it, underconfidence. Small, faint dots are thin-sample buckets — don't read much into a bucket with only one or two forecasts yet.

Every forecast you log is stored only in this browser's local storage — nothing is sent to a server, and there's no account. That also means it only lives on this device in this browser; export a backup if you want to move it or keep one safe.