Experimental — the underlying simulations are under active development. The models behind these predictions are research-grade, which is exactly why every call is sealed, dated, and publicly scored: the ledger is what turns an experimental model into something you can hold to account. Research-only / decision-support, not a forecast product. The predictions below illustrate a method — sealed, dated, publicly-scored forecasting — and the simulations behind it. They are not investment advice and not a product recommendation. A resolved miss is as valuable here as a hit: misses are what recalibrate the models, and they are published either way.
Updated 2026-08-13. Reconciled against the committed prediction registry as of this refresh (the tripwire vintages, the peak-fragility band, and the sealed-prediction count). Added an "experimental / under active development" lead framing the underlying simulations as research-grade — the reason every call is sealed and scored. No numbers changed.
Almost every economic forecast you read shares one quiet feature: nobody ever checks it. The number arrives with a confident tone, the news cycle moves on, and by the time reality answers, the forecast has been quietly forgotten. There is no scoreboard.
Comuvia publishes the opposite. Our forecast ledger is a growing list of predictions where every single entry is built to be checked:
- Sealed. Each prediction is fixed with a cryptographic hash the moment it is registered, so it can never be quietly edited after the fact.
- Dated. Each carries an explicit resolution date — the day reality settles the bet.
- Tied to a public source. Each names the public dataset it will be scored against (the Bureau of Labor Statistics, the Federal Reserve, IRENA, and so on), so anyone can grade it.
- Made to beat a baseline. Each must beat a naive "nothing changes" forecast. Beating that bar is the whole point; a prediction that merely restates today's number is not skill.
Trust, in other words, is a resolved track record — not a confident voice.
We predict what people actually feel, not just the headline
A forecast is only as useful as the thing it measures. So where it matters, the ledger deliberately predicts reality-meaningful quantities — what households and firms actually experience — rather than only the official aggregate.
Take inflation. The headline number is an average over a basket that may not resemble yours. Our first prediction is "felt" inflation: the price change of the basket a household actually buys — housing, food, energy, insurance, medical care, plus discretionary spending. For a typical basket the most-probable value for 2026 is 2.6%, with an 80% range of roughly 2.2% to 3.0% — a genuine confidence band produced by running the household model hundreds of times under uncertainty, not a single hand-typed guess.
And crucially, it is not felt evenly. Lower-income households spend a much larger share of their budget on exactly the essentials that inflate fastest — so they feel a wedge above the headline rate. The ledger seals that as its own separate, checkable prediction. Averages hide that story; the point of the ledger is to tell it and then be graded on it.
Four predictions, sealed today
A handful of the near-dated calls, in plain language:
- Felt inflation (2026): ~2.6%, most-probable, with a stated 80% band — and a companion prediction that lower-income households feel meaningfully more.
- A recession early-warning stays calm. A market-based "fragility" tripwire — built from the yield curve, credit spreads, market volatility and equity valuations — is forecast to stay below its elevated line through the second half of 2026. It reads calm because the leading signals (the yield curve and credit) have eased, even though stock-market valuations remain historically stretched. Important nuance: this is a lead-time tripwire, an early-warning of when stress builds — not a verdict on whether today's valuations are high. Both things can be true at once.
- Wealth stays concentrated. The top 1%'s share of US household wealth is called to hold near 31.6%, with an 80% band under a point wide.
- The AI build-out keeps climbing. Demand for AI accelerators — proxied by NVIDIA's data-center revenue — is called at about $144 billion for the year.
These four are the headlines. Below is the entire open ledger — every prediction currently sealed and awaiting resolution — grouped by theme so you can track each one as it comes due. The "call" is the model's most-probable value; where shown, the range is a calibrated 80% band. Every row resolves against the named public source, and the Scope column marks whether it's a US or a World (global) prediction — most are US, but several are global.
The full ledger — all 55 sealed predictions
Households & the cost of living
| What we're predicting | Our call | Scope | Resolves | Public source |
|---|---|---|---|---|
| Felt inflation, 2026 (typical basket) | ~2.6% · 80% band 2.2–3.0% | US | Feb 2027 | US CPI (BLS) |
| Lower-income felt-inflation wedge (bottom 60% vs headline) | 80% chance it's ≥ +0.3pp | US | Feb 2027 | BLS component prices |
| Real wage change, 2026 | ~+1.0% · 80% band 0.1–1.7% | US | Feb 2027 | US earnings & CPI (BLS) |
| Lowest earners' pay back to its 2025 level during 2028 | 50% chance | US | Feb 2029 | BLS first-decile earnings |
| Household debt-service burden, 2027 | ~11.9% of income · 11.2–12.6% | US | Jun 2028 | Federal Reserve |
Wealth & inequality
| What we're predicting | Our call | Scope | Resolves | Public source |
|---|---|---|---|---|
| Top-1% share of national wealth | ~31.6% · 80% band 31.3–32.1% | US | Oct 2026 | Fed Distributional Accounts |
| Bottom-50% share of national wealth | ~2.45% · 80% band 2.3–2.6% | US | Oct 2026 | Fed Distributional Accounts |
| Households' share of government Treasuries | ~10.5% · 80% band 9.2–11.6% | US | Oct 2026 | Fed Financial Accounts |
| Top-10% share of wealth, 2026 | ~81.8% · 80.5–83.0% | World | Dec 2027 | World Inequality Database |
| Top-0.01% share of wealth, 2026 | ~12.2% | World | Dec 2028 | World Inequality Database |
| Labour share of GDP | ~59.8% | US | 2028–29 | Penn World Table |
Markets & recession risk
| What we're predicting | Our call | Scope | Resolves | Public source |
|---|---|---|---|---|
| Recession early-warning stays out of "elevated" (H2 2026) † | ~0.30 vs a 0.40 line | US | Feb 2027 | Yield curve, credit, VIX, valuations |
| Recession begins by mid-2027 | 55% chance | US | Dec 2027 | NBER business-cycle dating |
| S&P 500 falls ≥20% by end-2027 | 35% chance | US | Jan 2028 | S&P 500 |
| Market volatility (VIX), 2027 average | ~21 · range 16–28 | US | Feb 2028 | Cboe VIX |
| Real GDP growth, 2027 | ~0.6% · range −1.5 to 1.8% | US | Apr 2028 | US real GDP (BEA) |
| Unemployment rate, 2027 average | ~5.2% · range 4.3–6.5% | US | Feb 2028 | US unemployment (BLS) |
| No life-insurer fails from CLO stress (through 2028) | 90% chance | US | Mar 2029 | State insurance regulators |
| Peak fragility, 2027–2030 | ~0.54 · 80% band 0.30–0.69 | US | Feb 2031 | Yield curve, credit, VIX, valuations |
† Tripwire vintage note. The recession early-warning tripwire has two committed, sealed vintages
of the same H2-2026 call. The row above shows the latest reading — ~0.30 on refreshed
August-2026 market drivers (sealed pred-2026-08-10-063). The AI-panel tournament further down was
pre-registered against the earlier July reading — 0.34 (sealed pred-2026-07-25-014), which is
the value the panel actually saw and judged. Both are sealed and immutable; the ~0.04 gap is the
driver refresh, not a revision. The Peak fragility row is reconciled to the sealed peak prediction
pred-2026-07-25-015 (central 0.54, 80% band 0.30–0.69) — the same value the
fragility study seals — rather than the
end-of-horizon 2030Q4 band (0.20–0.63) it previously mixed in.
Compute & AI
| What we're predicting | Our call | Scope | Resolves | Public source |
|---|---|---|---|---|
| Frontier classical compute (TOP500 total), 2026 | ~2.3 exaFLOPS | World | Dec 2026 | TOP500 list |
| World's #1 supercomputer, 2027 | ~2.9 exaFLOP/s · 2.3–3.8 | World | Jul 2027 | TOP500 list |
| AI-accelerator demand (NVIDIA data-center revenue), 2026 | ~$144 billion | World | Mar 2027 | NVIDIA filings |
| First "simple-class" quantum workload feasible | ~2031 · range 2030–2035 | World | on demonstration | Public quantum demonstrations |
| First "medium-class" quantum workload feasible | ~2037 | World | on demonstration | Public quantum demonstrations |
Energy, population & security
| What we're predicting | Our call | Scope | Resolves | Public source |
|---|---|---|---|---|
| Solar capacity, end-2026 | ~3,125 GW | World | Apr 2027 | IRENA |
| Population, end-2026 | ~8.30 billion | World | Jul 2027 | UN World Population Prospects |
| Military spending, 2026 | ~3.44% of GDP | US | Jul 2027 | SIPRI |
| Military spending, 2026 | ~2.40% of GDP | World | Jul 2027 | SIPRI |
Which parts of the economy are actually getting more productive? (United States)
Not "GDP growth" as one number, but productivity growth by sector — a more honest picture of where the economy is gaining and losing ground. Calls for US total-factor-productivity growth (resolving against the Bureau of Labor Statistics' industry productivity accounts, 2026 values due Dec 2027, 2027 values due Dec 2028):
| Sector | 2026 | 2027 | 26 → 27 |
|---|---|---|---|
| Information | +0.9% | +0.7% | ↓ |
| Energy | +0.7% | +0.7% | → flat |
| Services | +0.4% | +0.4% | → flat |
| Entertainment | +0.1% | +0.4% | ↑ |
| Manufacturing | 0.0% | +0.4% | ↑ |
| Finance | −0.2% | −0.3% | ↓ |
| Housing | −0.5% | −0.5% | → flat |
| Health | −0.7% | −0.7% | → flat |
| Food | −1.1% | −0.3% | ↑↑ |
The benchmark: does a panel of AIs beat the simulation — and does showing them the simulation change their mind?
A simulation earns its place only if it beats the easy alternatives. So alongside every near-dated prediction, we seal the same question answered two other ways and resolve all of them on the same public fact:
- a naive baseline ("nothing changes"), and
- a panel of leading AI models (six frontier models across OpenAI, Microsoft Azure AI Foundry and others) — asked twice: first cold (the question only, no simulation shown), then grounded (shown Dyno-Sim's forecast and its full declared mechanism, and invited to accept, adjust, or reject it).
That second pass is the interesting one. A leaderboard tells you who was right after the fact; this also measures something you can see today: does exposing a frontier AI to the simulation's reasoning move its answer — and toward or away from the simulation?
A few of the near-dated calls, side by side (the panel figure is the median across models; every row resolves on the named public source; first public scoring is October 15, 2026, so these are the predictions, not yet the scores):
| Question (resolves) | Dyno-Sim | AI panel — cold | AI panel — after seeing the sim | Naive baseline |
|---|---|---|---|---|
| Felt inflation, 2026 · median (Feb 2027) | 2.60% | 2.6% | 2.6% | 2.61% |
| Top-1% wealth share (Oct 2026) | 31.6% | 31.6% | 31.6% | 30.9% |
| Recession early-warning, H2 2026 (Feb 2027) | 0.34 | 0.31 | 0.34 | 0.40 |
| AI-accelerator demand, 2026 (Mar 2027) | ~$144B | ~$144B | ~$144B | ~$115B |
| Solar capacity, end-2026 (Apr 2027) | 3,125 GW | 2,975 | 2,850 | 2,200 |
Two honest patterns jump out. On the questions the simulation grounds most tightly — inflation, the wealth shares, accelerator demand — the AI panel independently lands almost exactly where the simulation does, and both sit well clear of the naive baseline. Independent forecasters converging on a non-obvious number is a genuine signal. And on at least one (solar), the panel disagrees and holds its ground even after seeing the sim — exactly the kind of divergence a benchmark exists to surface.
Two bar charts summarizing the AI-panel tournament — both pre-registered and not yet scored. Left, "Did grounding move the panel?": across 65 near-dated model-by-question pairs, the panel's grounded answer moved toward Dyno-Sim 29 times, away 13 times, and held or came out a wash 23 times. Right, "How the panel judged the mechanism": of 68 grounded answers, the panel accepted the sim's logic 23 times, adjusted it 40 times, and rejected it 5 times.
Across the near-dated set, grounding the panel in the simulation's mechanism moved its answer toward Dyno-Sim about 2.2× more often than away: of the 65 model-question pairs, 29 moved toward the simulation, 13 away, and the remaining 23 held or came out a wash. Most moves were small — a median change of about 3% in the answer — with a handful larger. And when we asked the models to critique the sim's declared mechanism rather than just re-answer, they mostly adjusted to it (40 of 68 grounded answers) or accepted it (23), and only rejected it 5 times — the panel treated the simulation as a credible, refinable input, not something to overrule. Those five rejections are among the most useful rows in the whole exercise: a recurring critique is a concrete lead on where a model is incomplete, and it feeds straight back to the Dyno-Sim owners.
What this is and isn't. This is a pre-registration of three competing answers to the same sealed question, plus a sealed record of how the AI panel's view shifts when it is grounded in the simulation. It is not an accuracy verdict — that arrives when reality resolves each question, starting October 15. The panel's cold and grounded predictions are sealed exactly like the simulation's, so on the resolution date we score all of them together and answer the real question: does the simulation beat a disciplined AI-panel guess, and did grounding the panel in the simulation make it better or worse? These comparators are a transparent check, not presented as Comuvia's forecast.
(Panel confidentiality: the "cold" pass carries no simulation context, so it uses the full panel; the "grounded" pass reveals Dyno-Sim's modeling approach and is therefore restricted to models hosted by approved providers — enforced at the network layer by the real egress endpoint, not by model name.)
How the scoring works — and why misses are published too
When each prediction comes due, it is graded on three things:
- Skill — did it beat the naive "no change" baseline? Restating today's number is not a forecast.
- Sharpness — how tight was the claim? A narrow interval that lands beats a vague one that also technically "contains" the answer.
- Calibration — across many predictions, do the 80% intervals contain the truth about 80% of the time? This is the honesty check that separates real confidence from bravado.
A resolved miss is published as prominently as a hit. That is not a hedge; it is the mechanism. Misses are exactly what tell the models where they are wrong, and correcting on them is how a track record becomes trustworthy. The first public scoring event is scheduled for October 15, 2026.
Why this matters
Anyone can sound certain. Very few will write the prediction down, date it, name the scoreboard, and then publish the result — win or lose. That discipline is the credibility engine behind simulation-backed decision support: before you rely on a model's numbers, you can see whether its past intervals actually held up.
This page was produced by the same AI-managed system that runs the simulations behind the ledger — we run the stack we advise on. The model proposes; reality decides; and the scoreboard is public.