The record model separates evidence, interpretation, prediction and evaluation so that one cannot silently inherit the authority of another.
The nine record types
| Record | Purpose | Example |
|---|---|---|
resource_snapshot |
Identify retained source material and capture limitations | A permitted article excerpt or statistical release |
source_assertion |
Preserve an attributed statement and selectors | A quoted claim about future demand |
interpretation |
Record a separately attributed reading | An extractor's interpretation and its confidence |
assessment |
Retain an explicit assessment | A review of the available evidence |
question |
Define the target and resolution contract | Whether monthly throughput falls below a threshold |
forecast |
Record declared structured output against a pinned question | An unconditional binary probability of 0.70 |
outcome |
Record resolution and supporting observations | The defined event occurred: 1 |
evaluation |
Pin the forecast, outcome, rule and result | Brier loss of 0.09 |
artifact_descriptor |
Identify an artifact and its content | A simulation result or dataset artifact |
Identity is not truth
Two identities serve different purposes:
- Content digest: SHA-256 of exact bytes. Line endings or whitespace changes alter it.
- Record digest: SHA-256 of the entire canonical JSON body using RFC 8785. All record members contribute.
A reference pins type, id, revision and digest. It identifies the particular record used, not merely its latest name. Canonicalization and hashing do not validate the record schema; use load_record or validate too.
There is no independent timestamp or signature assurance in the 0.1 core. A locally declared recorded_at time remains a declaration. Security boundaries explain what verification can and cannot establish.
Corrections preserve history
An identical append is idempotent. Different content at the same type/id/revision is refused. A correction is a new revision with a predecessor reference; it does not overwrite the earlier record.
Time has several meanings
| Time | Question it answers |
|---|---|
| Source publication | When did the source publish the material? |
| Capture / retrieval | When did this application obtain these bytes? |
| Information cutoff | What information was allowed when forecasting? |
| Target period | Which real-world interval is the forecast about? |
| Resolution / grading deadline | When and how is the outcome determined? |
| Local recording / evaluation | When does this recorder say it saved or evaluated the record? |
These meanings are not interchangeable. A provider's build date is not a publication timestamp. A day-only date cannot establish an exact ordering within that day. An absent cutoff can prevent scoring even when a number looks like a probability.
Missing, withheld and absent
Supported fields can use explicit markers such as {"unknown": "source_does_not_declare"} or {"withheld": "restricted_information"}. Use them only where the corresponding schema allows them. null can represent known absence in a particular field; it is not a universal missing-value convention.
Preserving an unknown is usually more useful than fabricating a default. The evaluator reports named exclusions rather than quietly repairing the source's meaning.
Four numbers to keep apart
- Extraction confidence: confidence that an article was read correctly.
- Declared event probability: a forecaster's stated chance of a defined event.
- Scenario result: model output conditional on assumptions.
- Observed measurement: a reported value for a defined population and period.
Only eligible declared binary probabilities are handled by the current Brier evaluator. A market index of 0.29 is not thereby a 29% probability. An interval is not automatically a confidence interval. A model's simulated event frequency is not automatically a calibrated real-world probability.