An evaluation is useful only when the question, forecast, outcome and scoring rule have compatible meanings. The current rule supports unconditional binary forecasts with a declared probability; it reports exclusions for unsupported or incomplete candidates.
Eligibility before arithmetic
The evaluator checks the pinned question and its completeness, declared probability, required timing, outcome and applicable revisions before calculating a loss. Keep the information cutoff separate from the target period and the grading deadline.
| Candidate | Current treatment |
|---|---|
| Eligible binary probability with a resolved 0/1 outcome | Scored by binary-brier-v1 |
| Forecast awaiting an outcome | Unresolved when appropriate; no invented result |
| Narrative without declared probability | Named exclusion |
| Conditional scenario | Named exclusion; not scored as an unconditional event |
| Point or interval output | Unsupported by this binary rule |
| Missing time or incomplete question | Named exclusion, preserving the reason |
Understand the Brier formula
For a single binary outcome, loss = (p − y)², where p is the declared probability and y is 0 or 1. Lower loss is better for the realized outcome. This worked example shows the arithmetic only; it does not run the SDK's eligibility checks.
Declared probability: 0.70. Observed outcome: 1. Loss = (0.70 − 1)² = 0.09.
If the outcome is corrected to 0, the new loss is (0.70 − 0)² = 0.49. Keep both evaluations with their different outcome evidence.
The quickstart's loss of 0.09 does not show that a forecasting system is well calibrated. Comparative skill needs a declared cohort, suitable baselines, enough resolved independent questions and uncertainty analysis. These aggregate analyses are outside the 0.1 evaluator.
Run a correction and reproduce both versions
The complete Tutorial B is in the example bundle:
./.venv/Scripts/python.exe ./tutorial_b_forecast.py --fixtures ./fixtures/synthetic --work ./corrected-forecast-run
It records the original fictional outcome, evaluates it, then appends a corrected outcome and a new evaluation. The original pinned evaluation remains reproducible.
| Evaluation | Probability | Outcome | Displayed loss |
|---|---|---|---|
| Original synthetic outcome | 0.70 | 1 | 0.09 |
| Corrected synthetic outcome | 0.70 | 0 | 0.49 |
The changing result is expected: it is now evaluated against different outcome evidence. Neither evaluation should silently replace the other.
Read the report honestly
report.counts distinguishes candidates, forecasts, independent questions, forecast families, eligible, scored, unresolved and excluded. Count independent questions separately from multiple forecasts about the same question. Report the excluded and unresolved counts alongside the scores.
Use evaluation_records to create the pinned evaluation bodies, then append them explicitly. store.verify() checks storage identity and references; reproduce() checks the saved evaluation against its pinned inputs and rule. One does not substitute for the other.
What reproduction can return
| Result | Meaning | Response |
|---|---|---|
match |
Recomputed result matches the retained evaluation | Retain both inputs and result |
mismatch |
Recomputed result differs | Investigate the exact pinned inputs and rule; preserve the original |
unavailable |
Reproduction cannot be completed | Surface the reason; do not claim successful verification |
See the core API reference for exact signatures and result fields.