The wire
TickerFed holds at 4.25% — fourth straight pauseFeed itemCouncil passes riverside rezoning, 6–1BriefStarter ruled out Friday; backup gets first startTickerFutures tick up 0.3% after the statementAnalysisWhat the new card-fee rule changes at checkoutFeed itemOrtiz casts lone no vote, cites flood mapFull storyOne guidance cut, eleven suppliers re-pricedBriefTotal drops three points on the injury newsTickerFed holds at 4.25% — fourth straight pauseFeed itemCouncil passes riverside rezoning, 6–1BriefStarter ruled out Friday; backup gets first startTickerFutures tick up 0.3% after the statementAnalysisWhat the new card-fee rule changes at checkoutFeed itemOrtiz casts lone no vote, cites flood mapFull storyOne guidance cut, eleven suppliers re-pricedBriefTotal drops three points on the injury news
Back to the benchmark
Methodology · Preregistered

How we test — and how you can check us.

We make one claim about quality: our stories are more factual and better-sourced than the raw models you’d otherwise use. This is how we prove it — and how you can audit it yourself. Nothing here is a number we picked after the fact.

01

The question, written down first

Does N7produce more factual, better-sourced stories than the same class of frontier model used raw? We committed the metrics, the judge rubric, and the thresholds for what we’d claim beforerunning anything. The commit history is the timestamp — we can’t move the goalposts after seeing the score.
02

Same story, every engine

Every engine is handed the identical source packet and asked for the same article. Then three ways to write it: a raw model with no tools, a raw model with web search— the skeptic’s “just use GPT or Claude” — and N7, our engine, which adds a sourcing, verification, and editorial layer on top. We benchmark against the real alternatives: GPT-5.6 and Claude Opus 4.8, run exactly the way a team would use them.
03

What we measure

Two kinds of evidence — one that no one can argue with, and one that removes us from the room.

Deterministic
Receipts, counted by machine
The share of numbers in a story that trace to a source, and the count that trace to nothing. No judgment — it uses the exact same checker our live product runs. The eval and the newsroom share one ruler.
Blind preference
Judged by other labs, blind
Two independent judge models, blind to which system wrote which, each judging both orderings — a win only counts if it survives the swap. A model never scores its own writing.
04

What we’ll claim — decided in advance

We wrote the thresholds down before the run, so the honest result is the only result we can report:

Equal prose, verifiable receipts.

We don’t claim to out-write a frontier model — on overall reader preference the gap is close, and we say so. We claim to out-source it: far more of our numbers trace to a named source, and far fewer are unsupported. Reported over 30 stories, with 95% confidence intervals, on the benchmark page.

05

Check us yourself

Every story we publish ships with its receipts — the sources it used and the verdict of the check — so you never have to take the score on faith. The benchmark harness saves every draft and every raw judge verdict verbatim, and we’re open-sourcing it: the same test, run against your own newsrooms, so the number is yours, not ours.

Preregistered · same source packet for every engine · blind, position-swapped judging by independent models · deterministic metrics shared with the production checker · all drafts and judge verdicts retained for audit.