What did AI 2027 get right— and was it right on time?
Eighty-two trackable claims, frozen at their original wording and separated into scenario beats, formal forecasts, hidden-state claims, and mutually exclusive endings. No single “accuracy score” can honestly summarize them.
QUICKEST HONEST READ08·21·26
OpenAI has documented a real two-week pause in deployment-focused frontier training, with its largest planned frontier RL run still on hold and monitoring consuming roughly 20% of inference compute. That is concrete development friction from cyber risk—not evidence that AI 2027’s broader slowdown branch has occurred.
01
Signal: capability risk is now imposing measurable cost and delay inside a frontier lab.
02
Unresolved: Astra’s Critical cyber threshold remains preliminary, without the frozen horizon test.
03
Hinge: the pause is temporary and lab-led, not government oversight or an industry-wide slowdown.
82records extracted
14conservatively judged here
70–90%authors’ own pace estimates*
~0.70×independent speed estimate*
*Different methods. Neither is an accuracy percentage. Why this matters
External comparator · 53 claims · 27 Jul 2026
Another tracker sees a mixed picture.
We preserve this as a comparison layer, not ground truth. Its categories are broader and more permissive than the resolution rules used in our core ledger.
The rail is ordered by the date in the original material. “Today” is fixed to the evidence cutoff, not your device clock.
30 visible2 met10 partial2 missed
ORIGINAL WINDOWPREDICTIONOBSERVED READ
21 AUG 2026TODAY / EVIDENCE CUTOFF
Forecast drift
The headline date was never the whole distribution.
Scenario mode, personal median, model output, and conditional takeoff date are different objects. Reporting them as one “AGI prediction” creates fake certainty.
APR 20252027
Scenario / modal year
Published story date; not every author’s median.
APR 20252028–33
Formal SC medians
Individual and aggregate all-things-considered forecasts varied widely.
NOV 2025~2030
Daniel’s stated median
He said progress looked somewhat slower than the scenario.
APR 2026JAN 2029
Daniel SC p50
Current all-things-considered forecast; p10 Apr 2027, p90 Jun 2036.
AUG 2026NOV 2027
Daniel AC p50
New uplift/revenue anchors; Eli p50 Jan 2030 and Brendan p50 Jan 2029.
REVISION WATCH
The August update says reality is moving at roughly 70–90% of the original pace, but it removes older grading items, adds uplift and revenue anchors, and conditions the forecast on no policy slowdown. Those are useful revisions—not a comparable accuracy series.
A vivid narrative creates many opportunities to claim a loose resemblance as a hit. The ledger separates wording, interpretation, evidence, timing, and forecasting credit.
●Substantially met
The necessary substance is supported by public evidence.
◐Partial / analogous
A material part matched, but the complete claim did not.
◆Disputed
Credible evidence or reasonable operationalizations conflict.
×Missed / behind
The target window or magnitude is materially off.
○Pending
The window is open or the claim has not yet become judgeable.
?Not publicly verifiable
The claim concerns private, classified, or hidden state.
↳Conditional branch
This only applies after a branch condition and is not scored now.
Five non-negotiables
01
Freeze the claim. Original wording and date stay visible beside later edits.
02
Split compound claims. A matching compute share does not confirm nationalization, chip allocation, and espionage.
03
Separate event from credit. A true statement already observable at publication earns little predictive credit.
04
Do not upgrade announcements. Planned GW is not active GW; a contract ceiling is not spend.
05
Keep uncertainty legible. “No public evidence” is not the same as false.
X / NEWS INTAKEDiscovery first. Evidence second.
Public X posts from labs, evaluators, officials, researchers, and event organizers are monitored as leads and contemporaneous statements. No capability, attendance, spending, or operational-capacity claim resolves from one post. Material claims need a durable primary source or corroboration.
Source register
Read the evidence, not just the verdict.
First-party sources establish what was claimed. Independent evaluators, official data, and adversarial critiques carry more weight when judging whether it happened.