Version 0.1.0 alpha · published on PyPI 2026-10-01; pin the version.

`comuvia` provides offline evidence records, identity checks, an append-only store and binary evaluation. It has no runtime dependencies. Package metadata requires Python 3.11 or later; measured support covers CPython 3.11–3.13 on Windows and Linux x86-64.

Import the package before using the examples. Functions in the signatures below are public names on `comuvia`; use `comuvia.validate(...)`, for example. `RecordStore` is also imported directly here. `Mapping` and `Iterable` in signatures describe ordinary Python mapping and iterable inputs.

```python
import comuvia
from comuvia import RecordStore
```

For the data model and its boundaries, see [evidence concepts](/sdk/docs/concepts). For the optional provider reader, see [foreglass](https://foreglass.ai/developers/reference/).

## Parse and validate records

```python
parse_json(data: bytes | str)
load_record(data: bytes | str) -> dict
validate(body) -> list[Issue]
schema_files()
is_marker(value) -> bool
```

| Function | Result and behavior |
|---|---|
| `parse_json` | Strictly parses a JSON value; rejects duplicate members and unsupported or nonfinite numbers. It does not validate a record schema. |
| `load_record` | Parses and validates one record, returning a dictionary or raising `RefusalError`. |
| `validate` | Returns named issues. An empty list means no schema or semantic refusal was found. |
| `schema_files` | Returns the packaged schema resource tree for inspection through Python's resource API. |
| `is_marker` | Recognizes the shape of an `unknown` or `withheld` marker; use record validation to check whether that marker is permitted in a particular field. |

```python
from pathlib import Path

try:
    record = comuvia.load_record(Path("record.json").read_bytes())
except comuvia.RefusalError as exc:
    for issue in exc.issues:
        print(issue.pointer, issue.code, issue.message)
```

`Issue` has `pointer`, `code` and `message`; pointers use RFC 6901 JSON Pointer. `RefusalError` is a `ValueError` subclass with `.issues` and `.codes`. Examples of named issues include `unsupported_schema_version`, `unknown_record_type` and `missing_required_field`. Normal Python exceptions such as `OSError`, `TypeError` and `KeyError` can also apply; not every failure is a `RefusalError`.

## Canonical identity and references

```python
canonicalize(value) -> bytes
content_digest(data: bytes) -> str
record_digest(body: Mapping) -> str
reference_to(body: Mapping) -> dict
verify(body: Mapping, *, expected_digest: str | None = None,
       content: bytes | None = None, records: Iterable[Mapping] = ()) -> Verification
check_selector(assertion: Mapping, content: bytes | None,
               *, snapshot: Mapping | None = None) -> SelectorCheck
revision_conflicts(bodies: Iterable[Mapping]) -> list[Issue]
latest_revisions(records: Iterable[Mapping]) -> dict[tuple[str, str], Mapping]
```

`content_digest` hashes exact bytes. Changing a line ending changes that digest. `canonicalize` produces RFC 8785 canonical UTF-8; `record_digest` hashes the entire canonical record body, excluding no member. Both digests use the `sha256:` prefix. **Hashing does not validate a record**: use `load_record` or `validate` first.

`reference_to(record)` returns its `record_type`, `id`, `revision` and record `digest`. `revision_conflicts` reports different bodies sharing the same type/id/revision. `latest_revisions` returns the latest view keyed by `(record_type, id)`.

For an already validated resource snapshot:

```python
snapshot = comuvia.load_record(Path("snapshot.json").read_bytes())
check = comuvia.verify(snapshot, content=Path("retained-source.txt").read_bytes())
print(check.content_identity)
print(check.timestamp_assurance)  # "unavailable"
```

`Verification` exposes `record_digest`, `record_identity`, `content_identity`, `references`, `issues`, `timestamp_assurance` and `signature`. Without `expected_digest`, record identity is `not_checked`. Reference checks can be `match`, `mismatch` or `unresolved`: inspect these results as well as issues. `verify` never fetches a locator. `check_selector` separately checks assertion selectors against retained content and returns `result`, `checked` and `issues`.

These operations check retained identities and selected relationships. They do not establish factual truth, authorship, data-use rights or independent publication time. Timestamp assurance is `unavailable`; signature is `none`.

## RecordStore

```python
RecordStore(path: str | os.PathLike, *, writer: bool = False)
RecordStore.create(path: str | os.PathLike) -> RecordStore
store.append(body: Mapping) -> StoreEntry
store.entries() -> tuple[StoreEntry, ...]
store.records() -> list[dict]
store.latest() -> dict[tuple[str, str], dict]
store.get(record_type: str, record_id: str, revision: int | None = None) -> dict | None
store.get_by_digest(digest: str) -> dict | None
store.verify() -> StoreReport
store.close() -> None
```

`create` initializes a new or empty directory and returns a reader. Open a writer explicitly; the context manager releases its OS lock on exit. This example assumes `records` is a list of validated bodies, ordered so reference targets are appended first.

```python
RecordStore.create("evidence-store")
with RecordStore("evidence-store", writer=True) as store:
    for record in records:
        entry = store.append(record)
        print(entry.digest, entry.appended)

reader = RecordStore("evidence-store")
print(reader.verify().ok)
```

The store supports one writer at a time. Append validates bodies and resolves references against records already stored. Re-appending identical content is idempotent: `StoreEntry.appended` is false. Conflicting content at the same type/id/revision is refused as `revision_conflict`; unresolved or changed targets produce `unresolved_reference` or `reference_digest_mismatch`. A second writer is refused as `store_locked`.

`records()` includes all revisions in append order; `latest()` is a derived view. Corrections append a successor revision whose `predecessor` references the prior revision. Reads do not rewrite history. `StoreEntry` contains `sequence`, `record_type`, `id`, `revision`, `digest` and `appended`.

`StoreReport` contains `entries`, `bodies`, `lock_holder`, `issues`, `detection_limit` and the `.ok` property. Verification checks stored identity and references, **not evaluation arithmetic**. Run it after a writer closes to avoid transient in-flight findings. Complete index-tail removal or whole-store replacement is not detectable without an independent checkpoint; version 0.1 provides none.

Share exported records and permitted source bytes rather than a raw live store: its lock file can contain machine metadata. The advanced recovery method `store.truncate_torn_tail() -> int` removes only an incomplete final index line and requires a writer; it is not a general record deletion API.

## Evaluate and reproduce

```python
load_rule(rule_id: str = "binary-brier-v1", version: str | None = None) -> ScoringRule
evaluate(records: Iterable[Mapping], *, rule: ScoringRule | None = None) -> EvaluationReport
evaluation_records(report: EvaluationReport, records: Iterable[Mapping], *,
                   evaluated_at: str, recorded_by: str, recorded_at: str) -> list[dict]
reproduce(evaluation: Mapping, records: Iterable[Mapping]) -> Reproduction
```

The packaged `binary-brier-v1` rule, version `1`, computes `d = p - outcome; loss = d * d` in binary64 after its eligibility checks. For a declared probability of `0.7` and resolved outcome `1`, the raw loss is `0.09000000000000002`, displayed as `0.09` to two decimal places. This is arithmetic for one eligible forecast, not evidence of predictive skill.

```python
records = RecordStore("evidence-store").records()
report = comuvia.evaluate(records)
for result in report.results:
    print(result.status, result.loss, result.reasons)
```

`evaluate` selects latest forecast and narrative candidate revisions while resolving pinned question references against all supplied records. Assertion and interpretation records remain narrative exclusions; extraction confidence never becomes an event probability. An unknown rule or version passed to `load_rule` raises `KeyError`.

| Return type | Fields and interpretation |
|---|---|
| `ScoringRule` | `id`, `version`, `digest`, `descriptor`; identifies the packaged policy. |
| `EvaluationReport` | `scoring_rule`, `counts`, `results`; `as_json()` returns serializable data with `aggregate: null`. |
| `CandidateResult` | `record`, `kind`, `status`, `reasons`, `outcome`, `loss`; status is `scored`, `unresolved` or `excluded`. |
| `Reproduction` | `result`, `stored`, `recomputed`, `reason`; result is `match`, `mismatch` or `unavailable`. |

Counts distinguish candidates, forecasts, independent questions, forecast families, eligible/scored results, unresolved results and exclusions. Named reasons include `conditional_scenario`, `incomplete_question`, `missing_required_time`, `no_declared_probability` and `question_mismatch`. A missing or unresolved outcome can produce status `unresolved`; non-scores are not automatically errors.

`evaluation_records` returns evaluation bodies for forecasts paired with outcomes; it does not append them. Supply explicit evaluation and recording times and append the returned bodies with a writer. Each evaluation pins its forecast, outcome and rule. `reproduce(evaluation, records)` re-derives that result without overwriting it, including after newer revisions arrive. There is no aggregate skill scorer, interval scorer or automatic conversion of statements into probabilities.

## Constants and supported record types

| Export | Value |
|---|---|
| `comuvia.__version__` | `"0.1.0"` |
| `comuvia.SCHEMA_VERSION` | `"comuvia/0.1"` |
| `comuvia.RECORD_TYPES` | `resource_snapshot`, `source_assertion`, `interpretation`, `assessment`, `question`, `forecast`, `outcome`, `evaluation`, `artifact_descriptor` |

Package, schema, store and scoring-policy versions have separate meanings. Unknown/withheld markers and nullability are defined per field by the schemas; do not replace every absent value with `null`.

## Command line

An installed core provides `comuvia`; `python -m comuvia` runs the same interface. These examples assume the named JSON/source files already exist and `evidence-store` is initially new. Append reference targets before dependants.

```sh
comuvia --version
comuvia validate question.json forecast.json snapshot.json outcome.json
comuvia digest snapshot.json
comuvia digest retained-source.csv --content
comuvia store init evidence-store
comuvia store append evidence-store question.json forecast.json snapshot.json outcome.json
comuvia --json store show evidence-store --latest --bodies
comuvia verify snapshot.json --content retained-source.csv --store evidence-store
comuvia --json evaluate --store evidence-store
comuvia --json store verify evidence-store
comuvia --json reproduce --store evidence-store
```

The last command reproduces stored evaluation records; a read-only `evaluate` call creates none. To append an evaluation for a synthetic tutorial run, use explicit fictional timestamps:

```sh
comuvia evaluate --store evidence-store --append --evaluated-at 2027-02-17T00:00:00Z --recorded-by evaluator:tutorial --recorded-at 2027-02-17T00:00:00Z
comuvia reproduce --store evidence-store
```

Exit codes: **0** success, including named exclusions; **1** refusal or mismatch; **2** usage error. A successful identity check or CLI exit does not imply every reference was checked, every candidate was scored, or a forecast is accurate: inspect the structured results.
