Comuvia SDK / Documentation

Developer reference · 0.1.0 alpha

Evidence & data quality

Trace a claim or number to its source; understand what the SDK can check today.

Public alpha reference. Version 0.1.0 was published on PyPI on 2026-10-01; pin the version. Release status and verification

Comuvia helps an application keep a checkable record of where information came from, how someone used it, and what changed later. This gives readers, researchers and organizations a practical way to ask: “Show me the evidence behind this statement or number.”

The benefit is accountability through inspectable evidence. A publisher can show its work, acknowledge uncertainty and correct mistakes without quietly replacing its earlier account. A reader can challenge a particular claim rather than having to trust or reject an entire organization.

What is a dependency?

A dependency is something a result relies on. An article may rely on an interview. A chart may rely on a table, a calculation and a choice of dates. A forecast may rely on observations, model assumptions and a definition of the event being predicted.

Linking to a website is useful, but the page may change. Comuvia records can reference a specific version of another record, identified by its digest: a fingerprint of its contents. Where permitted source bytes are retained and checked against a previously trusted fingerprint, a reviewer can detect a later change.

This is an example an application can assemble from records. The SDK does not automatically discover sources, extract every claim or build a complete dependency graph. The application chooses what to record and retains the underlying files.

Example 1: an article, speech or video

Suppose a minister says: “We expect the policy to reduce fuel prices by 10% within a year.” This is a fictional teaching example.

The application can keep:

  1. The source: the permitted transcript or recording, its location, capture date and limitations.
  2. The statement: who said those words, the exact quotation and its location in the retained source.
  3. The interpretation: which policy, prices, starting date and conditions the editor thinks the speaker meant. Unknown details remain unknown.
  4. The published result: the exact article, video or caption file, with an explanation of what it used.
  5. The review: whether the quotation supports the article's wording, what was disputed, and any later correction.

“A minister predicted a reduction” is different from “prices will fall” or “the policy caused prices to fall.” An official recording can support the attribution; it does not prove the prediction or its causal explanation.

Core 0.1 can check supported text quotations against supplied UTF-8 text. A recording's digest identifies the recording; it does not check that a clip, spoken quotation or caption has the right context or timing. Applications still need those checks.

If the editor dropped the words “we expect,” the correction should preserve the original edition and explain the changed interpretation. It should not rewrite the speaker's source record to conceal the editorial error.

Example 2: a statistical table and a chart

Suppose a fictional quarterly output index is 100 in Q1 and 96 in Q2. An analyst reports a 4% decline: (96 − 100) / 100 × 100.

A useful evidence trail retains the table version, the two selected cells, units, frequency, calculation and exact chart. If the provider later revises Q2 to 98, a new calculation gives a 2% decline. Both editions matter: the earlier result explains what the analyst could know then; the revision explains what is known now.

Comuvia 0.1 can identify source and result files with records and artifact descriptors. It does not parse arbitrary tables, check the calculation above, track every cell, or automatically find affected charts. Dataset-specific metadata, calculation checks and impact reports belong to the application; reusable versions are proposed work.

A quarterly series is not expected to contain one observation for every month. An absent value is not zero. A year-to-date total is not a monthly value. A file can pass an identity check while all these meanings are wrong.

Quality is several questions, not one trust score

Question How records help today What still needs checking
Is this the same file or record? Digests and exact references identify retained versions The expected digest must be retained independently of the possibly changed input
Where did the claim come from? Snapshots, assertions and supported selectors retain attribution Whether all material sources and context were captured
Does the evidence support the conclusion? Separate interpretations and assessments preserve a reviewer's findings Meaning, counterevidence and causal reasoning; no automatic truth verdict
Is the number comparable? Records can describe units and retain source/result artifacts Periods, frequency, revisions, missingness, formulas and population coverage
Was a forecast accurate? Eligible binary forecasts can be scored and the evaluation reproduced A declared event probability, matching question/outcome and required dates/rules
What changed, and who reviewed it? Successor revisions preserve recorded corrections and assessments Authenticated identity, completeness of disclosure and independent publication-time evidence

The assessment record can preserve a verdict, evidence, rubric and coverage. A core assessment reviews a recorded source assertion; a dataset-wide quality report can be retained as an artifact. The SDK validates supported structure and coverage arithmetic; it does not decide whether the verdict is justified. “Confidence in extracting this sentence” is not “probability this event will happen.”

What the current library provides

Building block Plain-language role
resource_snapshot and source_assertion Record a source version and what it says
interpretation and assessment Keep the user's reading and the review separate from the source
artifact_descriptor Describe and identify a retained dataset, article, chart or other file
Exact record references and verification Check supported references against supplied records and retained content
RecordStore revisions Append a correction through the SDK without replacing the previous record
Questions, forecasts, outcomes and evaluations Reproduce supported binary forecast evaluation, or explain why input is excluded

References placed in application-defined extensions are covered by record identity, but core 0.1 does not traverse or validate their dependency meanings. An artifact's declared units, content schema or rights reference are not a unit converter, dataset validator or license decision.

How this can help the common good

  • Readers and citizens: inspect evidence behind a consequential claim and understand what remains uncertain.
  • Responsible media and organizations: demonstrate corrections, preserve context and make review work reusable.
  • Researchers and data users: compare versions and locate assumptions before relying on a result.
  • Reviewers: challenge specific evidence or methods with a reproducible record instead of an unexplained reputation label.

These are benefits to test with real users, not measured outcomes already proved by the SDK. A transparent trail can reveal weak reasoning; it cannot force an organization to participate or make an unsupported claim true.

Share enough to inspect, protect what must remain private

An application may share selected records, source links, methods at an appropriate level and correction history while keeping licensed documents, private data and proprietary model internals restricted. Metadata-only disclosure can be useful, but it must say which checks an outside reader cannot reproduce. Metadata and excerpts also need their own rights review.

The SDK is a local library, not a central registry or public rating service. It neither uploads records nor authenticates every named organization. Its append-only API does not prevent someone with filesystem control replacing a whole store; independent retained digests, copies or checkpoints are needed for stronger assurance. Hashes are fingerprints, not encryption or proof of truth.

Start with source-linked evidence and record-store limits. Broader dataset validation and publication review are application responsibilities, not implemented core features.