FORECAST OBSERVATORY / 01Evidence cutoff 30 JUL 2026MONITORING ACTIVE

A living evidence ledger

What did AI 2027 get right—
and was it right on time?

Eighty-two trackable claims, frozen at their original wording and separated into scenario beats, formal forecasts, hidden-state claims, and mutually exclusive endings. No single “accuracy score” can honestly summarize them.

QUICKEST HONEST READ30·07·26

Useful agents arrived; a real multi-day cyber intrusion now shows parts of the self-replication chain ahead of schedule, while the broader R&D clock still lags.

  1. 01

    Strongest: coding agents, infrastructure, defense—and cyber autonomy.

  2. 02

    Weakest: frontier-run scale and the assumed 3–9 month lab lead.

  3. 03

    Hinge: no clean public proof yet of the recursive AI-R&D multiplier.

82records extracted
12conservatively judged here
58–75%authors’ own pace estimates*
~0.70×independent speed estimate*

*Different methods. Neither is an accuracy percentage. Why this matters

External comparator · 53 claims · 27 Jul 2026

Another tracker sees a mixed picture.

We preserve this as a comparison layer, not ground truth. Its categories are broader and more permissive than the resolution rules used in our core ledger.

16Confirmed
3Ahead
8On track
4Behind
13Emerging
9Not testable
Inspect the independent methodology

Prediction tape

Scroll the claim. Open the evidence.

The rail is ordered by the date in the original material. “Today” is fixed to the evidence cutoff, not your device clock.

30 visible2 met7 partial2 missed
30 JUL 2026TODAY / EVIDENCE CUTOFF

Forecast drift

The headline date was never the whole distribution.

Scenario mode, personal median, model output, and conditional takeoff date are different objects. Reporting them as one “AGI prediction” creates fake certainty.

APR 20252027

Scenario / modal year

Published story date; not every author’s median.

APR 20252028–33

Formal SC medians

Individual and aggregate all-things-considered forecasts varied widely.

NOV 2025~2030

Daniel’s stated median

He said progress looked somewhat slower than the scenario.

APR 2026JAN 2029

Daniel SC p50

Current all-things-considered forecast; p10 Apr 2027, p90 Jun 2036.

REVISION WATCH

The live project later corrected revenue, compute, parameter, and graph errors. This observatory scores the frozen original claim and shows revisions separately; it does not let a correction overwrite history.

Open the project changelog ↗

Forecaster audit

There is no honest “team accuracy” number.

The five bylined contributors have different roles, methods, domains, and amounts of public scoring evidence. Pooling them would misattribute skill.

01

Eli Lifland

Core forecaster · timelines

8,418Metaculus forecast updates

Strong conventional forecaster; mixed AI-specific record

  • GJOpen Brier score 0.23 versus a reported 0.301 median
  • #1 all-time on RAND/INFER; 1,275 scored Metaculus questions
  • Early AI tournaments were less dominant; he documented major MATH/MMLU misses

Caveat: Short-horizon geopolitical forecasting skill does not automatically transfer to AGI timing or takeoff.

02

Daniel Kokotajlo

Lead author · scenario

19 / 35fully correct in one post-hoc audit

Prescient narrative forecaster; limited proper-score evidence

  • A 2021 scenario anticipated multimodal chatbots, reasoning traces, agents, export controls, and very large training runs
  • The same audit found 8 partial and 6 wrong claims
  • Ranked 41/413 on ten questions in the AI 2025 survey

Caveat: The 19/35 result is a subjective retrospective classification, not a pre-registered accuracy rate.

03

Thomas Larsen

Core forecaster · goals & policy

16 / 413AI 2025 survey rank

Promising near-term result; small public sample

  • Top-4% placement on ten scored near-term AI questions
  • Especially strong on Cybench and public salience
  • Was too bullish on RE-Bench

Caveat: Ten correlated questions are too few to estimate a stable long-horizon forecasting rate.

04

Romeo Dean

Core forecaster · compute & security

N/Dpublic scored record

Relevant domain experience; no public scorecard found

  • Produced AI-chip, security, and government-intervention forecasts at Constellation
  • Authored the AI 2027 compute and security supplements
  • No confidently attributable public tournament or retrospective score was found

Caveat: This means not documented, not that no private track record exists.

05

Scott Alexander

Editorial collaborator

Editorrole in AI 2027

Experienced calibration writer; not the project forecaster

  • Has published and graded annual probabilistic predictions for years
  • Contributed prose and publicity to AI 2027
  • Explicitly says he cannot take credit for the forecast itself

Caveat: His personal calibration work should not be pooled into the AI 2027 team’s accuracy.

Resolution protocol

Designed to resist hindsight.

A vivid narrative creates many opportunities to claim a loose resemblance as a hit. The ledger separates wording, interpretation, evidence, timing, and forecasting credit.

Substantially met

The necessary substance is supported by public evidence.

Partial / analogous

A material part matched, but the complete claim did not.

Disputed

Credible evidence or reasonable operationalizations conflict.

×Missed / behind

The target window or magnitude is materially off.

Pending

The window is open or the claim has not yet become judgeable.

?Not publicly verifiable

The claim concerns private, classified, or hidden state.

Conditional branch

This only applies after a branch condition and is not scored now.

Five non-negotiables

  1. 01

    Freeze the claim. Original wording and date stay visible beside later edits.

  2. 02

    Split compound claims. A matching compute share does not confirm nationalization, chip allocation, and espionage.

  3. 03

    Separate event from credit. A true statement already observable at publication earns little predictive credit.

  4. 04

    Do not upgrade announcements. Planned GW is not active GW; a contract ceiling is not spend.

  5. 05

    Keep uncertainty legible. “No public evidence” is not the same as false.

X / NEWS INTAKEDiscovery first. Evidence second.

Public X posts from labs, evaluators, officials, researchers, and event organizers are monitored as leads and contemporaneous statements. No capability, attendance, spending, or operational-capacity claim resolves from one post. Material claims need a durable primary source or corroboration.

Source register

Read the evidence, not just the verdict.

First-party sources establish what was claimed. Independent evaluators, official data, and adversarial critiques carry more weight when judging whether it happened.