Provenance
How these articles are checked
These articles are written by language models. The rewrite process starts with a researched fact sheet; an independent model audit checks the resulting prose against it. Coverage differs by article. A research record does not mean an article has passed an audit: each record shows the stages completed.
This page is the accounting, and the accounting includes what has not landed. Every number here is generated from the pipeline's own output files, so the page cannot drift into flattery. Last rebuilt 6 Sep 2026.
Where the corpus stands
The timeline carries 1,973 events with an article behind them. Research, rewriting, auditing and publication are counted separately. Some published rewrites are still awaiting an independent audit.
- Researched 1,973 / 1,973
A fact sheet exists, built from the open web.
- Rewritten 1,969 / 1,973
The article was written from that fact sheet, and nothing else.
- Audited 1,011 / 1,973
A second model checked the article against the sheet, line by line.
- Live on the site 1,969 / 1,973
The rewritten text is the text this site serves. Only these carry a provenance line.
Coverage was deliberately ordered by importance — the events most people actually click were done first.
| Importance | Events | Researched | Rewritten | Audited | Live |
|---|---|---|---|---|---|
| 90+ | 232 | 232 | 232 | 232 | 232 |
| 80–89 | 771 | 771 | 771 | 771 | 771 |
| 70–79 | 970 | 970 | 966 | 8 | 966 |
What happens to one article
-
Research
OpenAI Codex, web search on
A researcher model searches the open web and fills a fixed JSON schema: what happened on the day, who was there, verified numbers, the immediate cause and consequence, what the sources leave uncertain, and the common traps. Every fact carries the URL it came from. Numbers without a stating source are not allowed in.
The same pass audits the article that was already on the site and lists its factual errors. That list is why this exists.
-
Write
claude-fable-5 · claude-opus-5
A different model writes the article from the fact sheet alone. It does not see the old text and does not search the web, so it has nothing to embellish from — anything it asserts either came off the sheet or is an invention the next stage is built to catch.
-
Audit
OpenAI Codex, adversarial posture
A third pass, by a different model family from the writer, reads the sheet and the article side by side and is told to find problems — a pass has to be earned. Any critical issue makes the verdict FIX. Every flagged claim comes with the correction the sheet supports.
-
Correct, then audit again
same auditor, higher reasoning effort
Flagged issues become a fixlist, the fixlist is applied, and the whole cohort goes back through the audit — the last pass at higher reasoning effort than the first. The defect rates below are what that repetition bought.
The audit is not a vibe check. It scores against a fixed rubric, and any of these failing is enough to send the article back:
-
H1-hookDoes the dek + first sentence make a curious stranger stop scrolling? Concrete, surprising, or human — not 'X was an important event.' -
E1-event-specificIs the body about THIS event on THIS date — what happened, who was there, what changed by nightfall — rather than the parent topic's general history? Rough bar: >=70% of sentences would not fit under a sibling event's title. -
Y1-anchoredEvent year (and day, when known) appears in the first two paragraphs and frames the piece. -
F1-groundedEvery named number, person, place, and causal claim traces to the fact sheet / cited source. No invented specifics. Uncertainty stated as uncertainty ('possibly', 'tradition holds'). -
D1-dedupNo sentence appears verbatim in any sibling article; shared topic background lives in the topic backgrounder, not repeated per event. -
K1-kickerThe closing sentence of the article (and of each section) must state a sheet-grounded fact or an explicitly attributed claim — never a rhetorical synthesis that exceeds the sheet ('city of millions', 'burned nothing', 'was one page', 'the moment it became public'). If the kicker reaches for resonance, quote-check every noun and verb against the fact sheet before keeping it.
2 further criteria (V1-voice, P1-pacing) cover voice and pacing. They are recorded but never block publication — a dull sentence is not a factual error.
What the audits found
The same 227 articles — the top importance band — went through the audit three times, with the errors from each pass corrected before the next. Critical issues raised per article:
- Pass 1 · medium effort 0.60
136 critical issues across 227 articles
- Pass 2 · medium effort 0.24
54 critical issues across 227 articles
- Pass 3 · high effort 0.13
30 critical issues across 227 articles
A 78% drop, and the last pass ran at higher reasoning effort than the first — so the final number was measured by a stricter reader than the one that produced the first. The curve is flattening, not finished: the last pass still found 30 critical issues. Some articles on this site are still wrong.
Dates worth not trusting
A timeline is a machine for stating dates with more confidence than the evidence carries. The research grades how firm each date actually is. Of the 1,973 events researched so far:
- 1,290 rest on a date the sources fix directly.
- 158 carry a traditional date — the conventional one, not one a contemporary record pins down.
- 525 are disputed: the sources disagree with each other.
Those events keep their conventional position on the timeline, because that is where a reader will look for them. What they also carry is a flag on the provenance line — traditional date, date disputed — and, where the research landed somewhere else entirely, the date the sources actually point to.
As a check on the researcher itself, 236 events were researched a second time by a different system with no sight of the first answer. The two agreed on the date for 184 of them and disagreed on 52. That disagreement rate is itself a useful number: roughly 22% of historical dates are unstable enough that two careful researchers land in different places.
What is not verified
Research records distinguish completed audits from articles not yet audited. Articles without a matching researched rewrite have no record.
- 4 Researched, not yet rewritten. A fact sheet exists; the article has not been rewritten from it.
- 958 Rewritten, not yet audited. Written from a fact sheet, but no second model has checked it line by line.
The honest version
Nothing here was checked by a historian. The articles are written by a model and audited by a model, and both of those models can be confidently wrong. An adversarial reviewer that shares a blind spot with the writer will miss the error they share. The fact sheets are only as good as the pages the researcher found, and it prefers primary and scholarly sources but does not always get them. The illustrations are synthetic and are labelled as such; on articles about atrocities they are suppressed entirely, because a generated image of a real massacre is a lie no caption fixes.
What this is better than is the thing it replaced. The previous text was a model writing history from memory, with a Wikipedia link stapled underneath and no check of any kind — the failure mode of most AI-written reference sites, which simply do not tell you this is what they are. The research pass was also pointed at the text already on the site, and asked to list its factual errors. It found at least one in 1,923 of the 1,973 articles it read — 9,388 errors in total, an average of 4.8 per article. Those articles read exactly as fluently as the corrected ones do. Fluency was never the signal.
So the claim on this site is narrow and specific: for each article that carries a provenance line, a named set of sources was consulted on a named date, the prose matches the researched rewrite, and its audit history is shown explicitly, including no audit and unresolved findings. Where the date is uncertain, the record says so. The sources are linked so readers can inspect the evidence.