How we test — and how you can check us.
We make one claim about quality: our stories are more factual and better-sourced than the raw models you’d otherwise use. This is how we prove it — and how you can audit it yourself. Nothing here is a number we picked after the fact.
The question, written down first
Same story, every engine
What we measure
Two kinds of evidence — one that no one can argue with, and one that removes us from the room.
What we’ll claim — decided in advance
We wrote the thresholds down before the run, so the honest result is the only result we can report:
Equal prose, verifiable receipts.We don’t claim to out-write a frontier model — on overall reader preference the gap is close, and we say so. We claim to out-source it: far more of our numbers trace to a named source, and far fewer are unsupported. Reported over 30 stories, with 95% confidence intervals, on the benchmark page.
Check us yourself
Preregistered · same source packet for every engine · blind, position-swapped judging by independent models · deterministic metrics shared with the production checker · all drafts and judge verdicts retained for audit.