How Do You Know?

Five words that beat three blind rounds

The Frame

You can prove a specialist works — or just make it look like it works

A team built an AI Storyteller specialist and set out to validate it. Validation is the whole point — and the hard part is telling real proof from a convincing appearance.

So which did they end up with?

Run right, it validated cleanly — confirmed, not projected

10 narrative traps were planted. 10 were excluded — confirmed, not projected. Claims came from findings only, and every statement traced.

But getting there was anything but clean.

10/10
Narrative traps excluded
15/15
Claims from findings only
24/24
Statements traced

The newest of six specialists, put through a frozen blind tournament

Canvas-specialists is a bundle of 6 AI specialist agents — a Researcher, a Formatter, a Data Analyzer, a Competitive Analyst, a Writer, and the newest addition: the Storyteller.

To test it: one frozen 8-document evidence pack, randomized labels, two independent blind judges per run.

Three rounds, three different winners — that's noise, not signal

The blind A/B/C tournament ran the same on-device SLM strategy memo three times. Every variant won exactly once.

If more rounds only muddied the answer, what was the testing actually telling them?

RunABC
Run 135.534.046.5
Run 242.533.041.5
Run 344.546.543.5

Every variant won once. That's noise, not signal.

Eighteen scorecards, exactly one real signal

Cross-run aggregation of 18 scorecards found a single significant result: Narrative Memorability at p=.011 — Storyteller 4.83 vs Writer-only 3.67 (of 5), Cohen's d > 2.0.

One signal in a mountain of testing. So why did they keep testing?

p=.011
Narrative Memorability — the only criterion to clear significance (omnibus p=0.352 overall)

Run 1 had already named the answer

Run 1 told them: the Storyteller adds narrative lift but bleeds precision when its output feeds the Writer. Runs 2 and 3 generated noise because n=6 is underpowered.

Three rounds to characterize a problem one round already named. More testing wasn't the path.

Act 4 — The Turning Point

Five words replaced projection with proof

“How do you know they are all 3 star? Is this confirmed through testing or theoretical?”

There was only one way to answer: stop projecting, run the chains, and look at the output.

An instruction fix that looks right is not the same as one that works

Live smoke tests split the result. The fragile part was the old most-used chain, not the new specialist — and the real bug was one missing input type: the Writer had no story-output case.

Solved with one new input type — because someone asked how do you know, and got an honest answer.

Researcher → Writer ✓ Genuinely fixed
Formatter auto-inserted. 9 findings traced to 9 S-numbered claims. 347 words, all structural blocks present.
Comp Analysis → Writer ✗ Still broken
Markdown tables and prose instead of the pipe-delimited matrix. Zero source URLs.
Sources

Sources & Research Methodology

Data as of: March 6, 2026  ·  Status: Narrative artifact / validation case study in the amplifier-stories bundle (live).

Primary source: docs/story-how-do-you-know.html in the checked-out repo ramparte/amplifier-stories, with siblings docs/story-building-the-storyteller.html and case-studies/canvas-becomes-a-specialist-platform.html.

Research performed:

Methodology (as stated by the deck): 3 blind runs × 2 independent judges × 3 variants = 18 scorecards; frozen 8-document evidence pack (Doc A–H); randomized X/Y/Z labels; 5 rubric criteria scored 1–5; same scenario (on-device SLM integration strategy memo); Kruskal-Wallis omnibus p=0.352, p=.011 Narrative Memorability; power analysis ~26 scorecards/variant.

Gaps: The canvas-specialists source (agents, blind A/B/C harness, cited PR #16 / commits 9bbb3ba and 4fe5214) is not checked out under this tree and those commits do not resolve via git — treated as deck-asserted references, not source-verified. All tournament statistics are asserted by the deck; the raw scorecards/dataset were not independently re-derivable. Deck's own noted open items: Competitive Analysis format compliance not yet re-addressed; single-domain test scenario; chain smoke tests not independently re-verified.

Primary contributor: Chris Park / cpark4x (product direction, all decisions) with Amplifier AI (analysis, implementation, testing); deck promoted by sadlilas and regenerated spine-first by Sam Schillace.

More Amplifier Stories