The Turbine Appears Once
In one line: I scored every claim in Leopold Aschenbrenner’s Situational Awareness whose window has closed. On forward calls he is 13 right, 3 partial, 7 wrong. He was right about the machines and wrong about what they would cost, and the word “turbine” appears once in 165 pages, in the subordinate clause where both of his power misses live.
1. What I did, and the rules I set before I started
In June 2024 Leopold Aschenbrenner published Situational Awareness: The Decade Ahead, 165 pages arguing that AGI arrives around 2027 and that the constraint is industrial rather than intellectual. In July 2026 the hedge fund he built on that thesis was liquidated after margin calls. That sequence invites a cheap article and I am not writing it.
Here are the rules I fixed before reading, because a scorecard whose rules are set after the fact is not a scorecard.
Only closed windows are scored. Roughly sixty of his claims are dated 2027, 2028, 2030 or "the end of the decade," including his headline thesis. None of them is due. Judging a forecast before its resolution date is the exact error my own forecast wing exists to avoid, so those sixty carry no verdict here, not even a hint of one.
Retrospective claims are scored as sourcing accuracy, not prediction skill, and labelled that way. A man who correctly reports what MMLU scores were in 2024 has demonstrated care, not foresight.
Where the evidence is thin I say the evidence is thin rather than forcing a verdict to fill a cell.
That leaves 27 genuinely forward-looking claims whose windows have closed. Those are the prediction test.
2. The scorecard
| Claims | Right | Partial | Wrong | Undetermined | |
|---|---|---|---|---|---|
| Forward claims, window closed | 27 | 13 | 3 | 7 | 4 |
| Retrospective sourcing accuracy | 43 | 33 | 4 | 1 | 5 |
| Arithmetic and internal consistency | 9 | 6 | 0 | 2 | 1 |
On the only part that tests prediction, he is 13 right, 3 partial, 7 wrong, 4 undetermined. For a two-year-old document making dated quantitative claims about an industry in flux, that is a respectable record and better than most of what was written about AI in 2024.
3. Where he was right, including where he was too cautious
The hits are not vague. They are numbers with dates.
He wrote that Nvidia would do over $200 billion of revenue in calendar 2025 against a sell-side consensus he named as $120 to $130 billion. Nvidia's fiscal 2026 came in at $215.9 billion, up 65%, with data centre at $193.7 billion. He named the consensus, named his own number, and beat both. That is the cleanest falsifiable market call in the document.
He tabulated a 2024 cluster at roughly 100,000 H100-equivalents and about 100 megawatts. Three months after publication, xAI's Colossus came online in Memphis with 100,000 H100s and over 100 megawatts approved. He tabulated a 2026 cluster at roughly a million H100-equivalents, tens of billions of dollars, and about a gigawatt. In January 2026, Colossus 2 came online at about a gigawatt. That is a two-year-out forecast of a physical object, and the object exists.
And on aggregate AI investment he was wrong in the direction nobody expects. He projected roughly $500 billion for 2026. Actual hyperscaler AI infrastructure spend is running near $725 billion, above a trillion if you count Stargate. The most-mocked number in the document was too small.
4. Where he was wrong, and the shape of it
The misses cluster as tightly as the hits, and in a different place.
His demand-side revenue timing ran about a year late. He put OpenAI at a $10 billion run rate by late 2024 or early 2025; it got there in June 2025. He projected a big-tech company at $100 billion of AI revenue by mid-2026; Microsoft, the only one of the three he named that discloses a comparable figure, reported $37 billion. He also said we would have $10 trillion companies by now. The largest company on earth is around $4.7 trillion.
There is a fairness note that cuts toward him on the second one. His assumed doubling rate was, if anything, too slow. Anthropic went from roughly $1 billion to $47 billion of run-rate revenue in seventeen months. He asked the right question of the wrong firms.
But the weakest part of the document is the power build, and it is weak in a specific and instructive way.
5. The turbine appears once
Aschenbrenner identified the binding constraint correctly and early. "Probably the single biggest constraint on the supply-side will be power," he wrote, and he was right, and much of the subsequent industry conversation traces to him saying it.
Then he priced it. Gas plant capex "under $1000 per kW." Combined cycle plants buildable "in about two years."
Neither survived contact. Gas turbine prices are reported up roughly 300% in three years. Current analyst estimates put GE Vernova's heavy-duty turbine equipment alone near $790 per kW and its aeroderivative units near $1,800, before you have built anything around them, and 2026 orders are pricing ten to twenty points above late 2025.
I searched the full text. The word "turbine" appears once in 165 pages, in a subordinate clause: "The harder part would be building enough generators / turbines; this wouldn't be trivial, but it seems doable."
That single clause is where both of his power misses live. He saw the constraint and waved at its supply chain in eleven words.
I want to be careful about what that does and does not prove. It is not evidence that he is careless; the essay is about capability trajectories and it is entitled to compress the industrial detail. It is evidence about which kind of claim survives two years: his claim about the physics held, and his claim about the price did not, because the price lives in a supply chain he did not model.
I have written 5,000 words on that supply chain from the other direction, and I reached the same place by a different road. My own paper on the nuclear buildout found roughly 14.7 gigawatts of announced hyperscaler nuclear capacity against 50 megawatts of new construction actually underway, with turbine and interconnection lead times as the binding term. Neither of us is contradicting the other. He identified the constraint; the constraint turned out to be more expensive and slower than one clause allows.
6. Where his claims and my priced slate touch the same object
On 1 August 2026 I submitted fourteen pre-registered forecasts with binding resolution criteria to an external forecasting competition, and the slate froze permanently at submission. Three of them price things his essay also discusses. This is the comparison worth making, and it is about form rather than merit.
Power. He wrote that power is the biggest constraint, undated, with no threshold, closing on three modal verbs: it "can, must, and will be solved." I priced one auction: PJM's 2029/30 base residual auction clears at or above $325.00 per megawatt-day, 72%, resolving on PJM's own published results. Same underlying claim. One of the two can be graded by a stranger who does not care about either of us.
For what it is worth to the underlying question, PJM has now cleared at its regulatory collar three consecutive times, and the 2028/29 auction cleared 6,831 megawatts short of its reliability requirement, which PJM says is among the first times in its history the entire RTO fell short. Uncollared, PJM estimates every zone would have cleared near $555.
Hyperscaler capex. This is the tightest like-for-like in the exercise: the same four companies, the same line item. He gave 2024 dollar thresholds per company and went three for four. His single miss was Meta at $39.23 billion against his "$40 billion plus," a 2% shortfall that becomes 7% if you exclude finance-lease principal, and which definition applies was never specified. My version of that forecast specifies the exclusions in writing, before the fact, because I read his and saw where the ambiguity was.
Note the direction of aggression, which does not flatter me. His implied growth rate was 100% a year. Mine is 20%, and I still only price it at 62%.
Chips and China. Here the interesting finding is an absence. Across 165 pages: "rare earth" appears zero times. "Chokepoint" appears zero times. "ASML" appears zero times. His supply-side risk model runs in one direction, from American capability outward. Two of my fourteen forecasts price the chokepoints running the other way.
7. What this does not show
My slate has resolved zero of fourteen. No accuracy comparison between us is computable, and I am not offering one. He has a two-year record with closed windows. I have a book that has not started scoring.
Being resolvable is not the same as being right. A forecast with a binding criterion and a date can be graded, which is a different virtue from being correct, and the whole point of publishing mine is that I might be graded badly. If my fourteen come back at ten wrong, the pre-registration will have worked exactly as designed and I will have been wrong in public on schedule.
A directional essayist and a pre-registered forecaster are doing different jobs. He was writing to change what people believe about the decade, and by that standard the essay worked: the trillion-dollar cluster is now a normal phrase, and the industry conversation about power partly descends from him. Pre-registration would have made it a worse essay and a better scorecard. Those are not the same document and I am not pretending he failed at a task he did not set himself.
And the fund is not evidence about the essay. A liquidation is a statement about leverage, correlation and margin, which I have written about separately. It says nothing about whether AGI arrives in 2027. The sixty undue claims stay undue.
8. The one thing I would take from this
His hardware calls landed and his price calls did not, and the difference is that the hardware was a trend he could extrapolate while the price lived in a supply chain of turbines, transformers, interconnection queues and skilled labour that no trendline reaches.
That is the transferable lesson, and it is the same one my own power research produced independently: when a forecast has a physical bottleneck in it, the forecast is only as good as your model of the bottleneck, and the bottleneck is almost never in the part of the document you enjoyed writing.
Method. Every claim was extracted from the June 2024 PDF with page citations and classified as checkable-now or not-yet-due before any scoring. Verification used primary sources where they exist: SEC filings and company results for revenue and capex, PJM's published auction releases for capacity prices, EIA series for electricity. Trade press is flagged inline as secondary wherever it was the only source available. The full extraction and scorecard, including the roughly sixty undue claims left unscored, is available on request.
Educational research only. Not investment advice. This note concerns published claims and their resolution, and makes no assertion about any individual's conduct or ability.