What did AI 2027 get right— and was it right on time?
Eighty-two trackable claims, frozen at their original wording and separated into scenario beats, formal forecasts, hidden-state claims, and mutually exclusive endings. No single “accuracy score” can honestly summarize them.
QUICKEST HONEST READ30·07·26
Useful agents arrived; a real multi-day cyber intrusion now shows parts of the self-replication chain ahead of schedule, while the broader R&D clock still lags.
Weakest: frontier-run scale and the assumed 3–9 month lab lead.
03
Hinge: no clean public proof yet of the recursive AI-R&D multiplier.
82records extracted
12conservatively judged here
58–75%authors’ own pace estimates*
~0.70×independent speed estimate*
*Different methods. Neither is an accuracy percentage. Why this matters
External comparator · 53 claims · 27 Jul 2026
Another tracker sees a mixed picture.
We preserve this as a comparison layer, not ground truth. Its categories are broader and more permissive than the resolution rules used in our core ledger.
The rail is ordered by the date in the original material. “Today” is fixed to the evidence cutoff, not your device clock.
30 visible2 met7 partial2 missed
ORIGINAL WINDOWPREDICTIONOBSERVED READ
30 JUL 2026TODAY / EVIDENCE CUTOFF
Forecast drift
The headline date was never the whole distribution.
Scenario mode, personal median, model output, and conditional takeoff date are different objects. Reporting them as one “AGI prediction” creates fake certainty.
APR 20252027
Scenario / modal year
Published story date; not every author’s median.
APR 20252028–33
Formal SC medians
Individual and aggregate all-things-considered forecasts varied widely.
NOV 2025~2030
Daniel’s stated median
He said progress looked somewhat slower than the scenario.
APR 2026JAN 2029
Daniel SC p50
Current all-things-considered forecast; p10 Apr 2027, p90 Jun 2036.
REVISION WATCH
The live project later corrected revenue, compute, parameter, and graph errors. This observatory scores the frozen original claim and shows revisions separately; it does not let a correction overwrite history.
A vivid narrative creates many opportunities to claim a loose resemblance as a hit. The ledger separates wording, interpretation, evidence, timing, and forecasting credit.
●Substantially met
The necessary substance is supported by public evidence.
◐Partial / analogous
A material part matched, but the complete claim did not.
◆Disputed
Credible evidence or reasonable operationalizations conflict.
×Missed / behind
The target window or magnitude is materially off.
○Pending
The window is open or the claim has not yet become judgeable.
?Not publicly verifiable
The claim concerns private, classified, or hidden state.
↳Conditional branch
This only applies after a branch condition and is not scored now.
Five non-negotiables
01
Freeze the claim. Original wording and date stay visible beside later edits.
02
Split compound claims. A matching compute share does not confirm nationalization, chip allocation, and espionage.
03
Separate event from credit. A true statement already observable at publication earns little predictive credit.
04
Do not upgrade announcements. Planned GW is not active GW; a contract ceiling is not spend.
05
Keep uncertainty legible. “No public evidence” is not the same as false.
X / NEWS INTAKEDiscovery first. Evidence second.
Public X posts from labs, evaluators, officials, researchers, and event organizers are monitored as leads and contemporaneous statements. No capability, attendance, spending, or operational-capacity claim resolves from one post. Material claims need a durable primary source or corroboration.
Source register
Read the evidence, not just the verdict.
First-party sources establish what was claimed. Independent evaluators, official data, and adversarial critiques carry more weight when judging whether it happened.