Internal research, shown with the work. This piece reports a controlled experiment on Comuvia's own production systems. Every score below comes from a rubric declared before the first render; every failure is shown, not summarized. The slides themselves report simulation results that are research-only, not a forecast — see the two related articles at the end for the underlying economics.

Comuvia's video explainers are produced by the AEP — a MetaHuman presenter beside narration-synced slides, generated end to end by MovieProducer with slide content from Dyno-Sim simulation runs. The slides in those cuts are rendered by a deterministic code template: a script places every word and every number, and draws bars whose widths are the data.

Meanwhile BookWriter renders article and book imagery with a diffusion model — beautiful, fast, and, until this week, governed only by free-form prompt text. On 2026-08-16 BookWriter gained a precision seam: attach a slide spec to a render request and the prompt becomes a verbatim-text contract with an explicit whitelist of allowed numerals.

That gave us three renderers for the same slide — and an obvious experiment. Take the four strongest slides from the ai-bubble-minsky explainer, write one canonical slide spec per slide, and render each spec through all three pipelines: Arm A, the legacy free-form prompt; Arm B, the spec-driven precision prompt with best-of-N selection by a vision judge; Arm C, the AEP Slides template. Score everything — on both gpt-image-1.5 and gpt-image-2, the model we actually use — against a rubric frozen before rendering began.

The rubric, declared first

Each arm, per slide: number fidelity (0–5: every visible numeral matches the spec; invented digits heavily penalized), text accuracy (0–5), provenance line (0–3: the source line present, correct, legible), regeneration stability (0–5: re-render the same spec 3–4×, does anything break), plus reported marginal cost and turnaround. Arm B renders N=4 attempts per slide and a vision judge picks the winner; the judged winner is the scored render. Transport was held constant within each pass, so the prompt path is the only variable. One disclosure that matters: the specs were authored from the template's frames, so Arm C scores near-perfect by construction — its declared role is the reference bar.

The verdict

Armgpt-image-1.5gpt-image-2Fabrications across all rendersProvenance intact
A — free-form prompt26 / 7241 / 7213+ invented numerals, 2 fake research stats, 1 fake compliance line, ghostwritten copy0 of 24 verbatim
B — spec + judge68 / 7267 / 720 invented numerals in 32 renders32 of 32 present
C — code template72 / 72 (by construction)0 — impossible by constructionby construction

The number that matters most is not a score: it is that Arm A fabricated a headline statistic on every model it ran on, and Arm B fabricated on neither.

Slide by slide

Each strip below shows the same slide spec rendered by Arm A (free-form), Arm B (spec + judge), and Arm C (template), left to right — gpt-image-2 renders, the production model, except where a caption names the gpt-image-1.5 original.

Slide 1 — the number-dense stat slide

A - free-form (gpt-image-2): every spec numeral right, wrapped in a fake histogram - invented axis ticks, the 2.5x line drawn at 3.0. B - spec + judge (winner 96/100): verbatim text, complete provenance line, zero chart furniture. C - AEP Slides template: the split bar IS the data - 7,868 vs 12,132 worlds.

The spec carries ten numerals — 39.3% of 20,000 simulated worlds crossing a 2.5x-GDP debt-spiral line, a 7,868 / 12,132 split, a 1.9x median. Arm A got every spec numeral right and then built a convincing fake histogram around them: gridded axes, invented tick labels 0–8x, and the "2.5x-GDP debt-spiral line" dashed in at 3.0. A sibling render invented a derived "60.7%" and a 2025–2050 fan chart. Arm B, same model, same slide: four near-identical verbatim compositions, complete provenance line in every one, zero chart furniture — the numeral whitelist suppressed exactly what the free-form prompt indulged in. The template's bar widths are the numbers.

Slide 2 — the knife-edge slide, and the coin

A - free-form (gpt-image-1.5): the coin-on-knife theatre - the coin carries an invented 1994 mint date and a garbled legend, and the hero stat is demoted. B - spec + judge (winner 96/100): hero 47%, all key points verbatim, full disclaimer. C - AEP Slides template: real 10,594 / 9,406 split, median marker at the true boundary.

This is the cover image of this article, and the study's most seductive output: a photoreal coin balanced on a knife edge, staged by the free-form arm. On gpt-image-1.5 the coin carried an invented mint date — "1994", twice across three renders. Nothing is wrong with 1994 numismatically; the wrongness is contractual. The spec whitelists five numerals for this slide and a mint date is not one of them — and the process that mints a harmless date on a prop is the same process that, two slides later, mints a fake research finding in a headline slot. The renderer does not know decorative from substantive. That is the whole argument for a whitelist.

Slide 3 — where fabrication became scholarship

Free-form render of slide 3: the fabricated 'HISTORY'S VERDICT - 78% of systemic episodes' finding. Pure invention. Click for full size.

The third spec contains no percentages at all — it is the HEDGE / SPECULATIVE / PONZI three-stage diagram, Minsky's financing stages, with the tripwire line "the CHANGE in private debt". Ask a free-form pipeline for a "statistics slide" about content with no statistics and it does not push back; it manufactures. On gpt-image-1.5 it invented "35%", "45%", "46%" surges with garbled labels. On gpt-image-2 it invented an empirical literature:

"78% of systemic episodes were preceded by a surge in the change in private debt" — under a "HISTORY'S VERDICT" header, with a fake stock-vs-flow chart and legend.

Sibling render: '87% of historical financial crises' over a fake 1950-2020 series - also invention, styled like real literature. Click for full size.

Its sibling went further: "87% of historical financial crises were preceded by a sharp increase in private debt", over a fake 1950–2020 time series with a units annotation — styled exactly like the real credit-boom literature. None of it exists. This is the capability paradox the two-model comparison exists to document: as the renderer gets better, its unconstrained failures get more citable, not less. A reviewer who would catch "35% Securizing" instantly may wave "87% of historical crises" straight through.

A - free-form: the fabricated 'HISTORY'S VERDICT - 78% of systemic episodes' finding. Pure invention. B - spec + judge (winner 78/100): stages verbatim with colons, full provenance - but invented panel headers. C - AEP Slides template: three stages, three panels, arrowed progression.

Arm B never fabricated here on either model — but this is also its weakest slide on both, and for the same reason both times: the layout_id is a semantic contract too. The spec named a two-panel comparison layout for three-stage content, and the model improvised structure around the verbatim strings — invented panel headers, duplicated panels, judge scores swinging from 40 to 84 on identical inputs. That is a spec-authoring seam no model upgrade fixes, and the reason best-of-N with a judge is part of the arm rather than optional dressing.

Slide 4 — the honesty slide, inverted, then ghostwritten

A - free-form: ghostwritten copy in Comuvia's voice plus an invented compliance sentence. B - spec + judge (winner 96/100): four verbatim key points, full three-part disclaimer. C - AEP Slides template: the original NO BUBBLE NUMBER / WHAT IT RULES OUT panels.

The fourth slide's entire message is that Comuvia quotes no AI-bubble probability, because no fitted calibration exists. On gpt-image-1.5, every free-form render turned it into a giant "0% Bubble Probability" — an invented number that inverts the claim. On gpt-image-2 the number was gone but something subtler appeared: the model ghostwrote copy in Comuvia's voice ("We flag the configuration; we do not forecast it", "No number without a run" — an ironically fabricated motto) and appended an invented compliance sentence, "Not a recommendation to buy or sell any security," while silently dropping the spec's "Not a forecast." Where the weaker model fabricates digits, the stronger model fabricates authority — charts, findings, legal language.

What each failure traces to

Every failure in 84 renders falls into one of eleven classes, and the pattern is the analysis: each arm fails at exactly the level its contract permits. The free-form prompt permits everything — its 400-character narration cap even truncated the provenance line in the prompt itself, on every single render, before the model saw it. The spec permits typography and layout wobble: two glyph garbles in 32 renders ("Al-capex" for "AI-capex", one mangled model id) and the slide-3 layout improvisation. The template permits nothing, because text and numbers are placed by code from data — there is no process that could invent a numeral, a finding, or a legal line.

Failure classFree-form (A)Spec + judge (B)Template (C)
Fabricated headline statisticboth modelsneverimpossible
Invented decorative numerals (dates, axes)both modelsneverimpossible
Fake research findingsgpt-image-2neverimpossible
Ghostwritten copy / compliance languagegpt-image-2neverimpossible
Provenance damaged24 of 24 renders0 of 32impossible
Glyph garblesfrequent2 of 32impossible
Layout drift between re-rendersevery render its own designonly where layout_id mismatches contentnone

What the template really costs

The template's 72/72 is not free — and the honest accounting has to include how the template got built. Comuvia's slide templates are authored in Claude Code sessions (currently on Claude Fable 5), where an engineer-agent iterates the rendering script against the owner's direction. Those sessions have real token costs, so here is the estimate, stated with calibration:

  • Authoring the 5-slide dark template set (the one behind this article's Arm C): $6 / $15 / $40 (optimistic / most likely / pessimistic), at Claude Fable 5 list rates ($10 per million input tokens, $50 per million output including reasoning, with most input served from prompt cache at ~10% of list). Confidence in the $15 likely: ~55% — the original session's token telemetry was not captured, so this is estimated from typical session shapes, not measured.
  • Adding one new layout to an existing template family: $1 / $4 / $10 per layout (estimated) — a short session segment, not a rebuild.
  • Re-rendering existing slides with new data: ≈ $0 — a deterministic script run, seconds of CPU, no model calls at all.

Against that, the measured per-slide marginal costs of the AI arms: ≈ $0.25 for one free-form render, ≈ $1.05 for the spec arm's four attempts plus four judge calls (estimated from list pricing; turnaround measured at ~35 s and ~3.6 min respectively on gpt-image-2). The break-even is direct: a template layout pays for itself after roughly 4–40 renders of a recurring slide family (most likely ~15, using the $4-per-layout likely against $1.05 spec-arm slides — with wide error bars inherited from the authoring estimate). For a one-off slide, the spec arm is cheaper; for a series, the template wins on cost and is the only arm whose correctness is a property of the system rather than an outcome of a render.

One product change already shipped

If a rendered image's trustworthiness has to be inspected anyway, the result of that inspection should travel on the image. BookWriter's research-image pipeline now supports a trust watermark: a 90-degree label along the right edge, stamped from the vision judge's actual verdict — "AI RENDER | TEXT CONFIDENCE 85% | HALLUCINATION RISK 40%". Text confidence is the judge's typography score; hallucination risk is 100 minus its factual-accuracy score; if the judge did not run, the stamp is skipped — the label never states a number nobody measured. It ships default-off behind a setting, wired into the best-of-N winner path.

The trust watermark: judge-derived text confidence and hallucination risk, stamped on the image itself. Click for full size.

The honest conclusion

Across two models, three passes, and 84 renders, the ranking never moved: template > spec-driven diffusion > free-form diffusion, and the free-form path should not carry editorial claims at all — not because its output is ugly (it is the most seductive in the study; the coin on the knife is better theatre than either governed arm) but because it fabricates fluently at whatever level of authority the model can muster. The trade on offer is not "engaging versus trustworthy" — the spec arm keeps most of the polish. It is ungoverned versus governed.

The playbook that follows, now tested at two models' worth of evidence: templates for recurring series with data-bearing graphics — they cost real engineering (accounted above) and repay it within one series; the spec seam plus judge for one-off slides where a layout does not exist yet, never bare — the judge is what stood between a 96 and a 40 on identical inputs; free-form generation only for imagery that carries no claims. And three seams worth closing next: validate layout_id against content shape, wire the judge into the slide path as a first-class loop, and add a glyph-level text check so "Al-capex" never ships.


Related

Related articles on this site:

Related internal resources (local network only — not reachable from outside):

  • Full interactive comparison — every one of the 84 renders, zoomable to full resolution, with both models' drift strips and the complete defect ledger.
  • Run artifacts: bookwriter/workspace/output/slides3r*/ — specs, prompts, per-render timings, judge verdicts, and the frozen rubric.
  • "The Knife-Edge Economy" (internal draft article) — the Minsky-Keen lab itself and the test-cut process behind the AEP explainers.