Anthropic's Fable 5.1 Hits 60.9% on Humanity's Last Exam, GPT-6 Astra Drops, & the Cybercab Takeover
A model-war episode: OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1/Mythos 5.1 launched within 48 hours of each other, with the panel dissecting whether Astra's edge comes from a new 'looped transformer' depth-scaling architecture and whether that undermines chain-of-thought interpretability. OpenAI's own safety team rated Astra a critical cyber security risk, prompting kill-switch talk that the panel mostly dismisses as marketing while flagging the real race to lock up compute, chips, and enterprise partnerships. Governance split down the middle this week too, with Bernie Sanders' 20-years-in-prison 'Ban Artificial Super Intelligence Act' landing the same week as the G20's pro-innovation 'Carolina Principles.' Elsewhere: Tesla's Cybercab launch promises ~20-cents-a-mile transport, World Labs' Atlas world model pushes Gaussian splats as a possible new AI 'token,' Emad Mostaque pitched a TSMC-style locally-owned AI utility ('Champion'), and new data ties GLP-1 drugs to six unrelated disease benefits including mouse lifespan extension.
OpenAI's GPT-6 Astra saturates Frontier Math Tier 4 (98%) and ARC-AGI-3 (99.9%), is framed as an early AGI step, and is built natively around computer-use assistance.
Alex Wissner-Gross argues Astra's real innovation is recurrence — a transformer looped on itself with tied weights — a new 'depth scaling' law distinct from parameter/data/compute scaling.
Anthropic released Fable 5.1 (broadly available) and Mythos 5.1 (restricted to vetted cyber/life-science programs); Fable 5.1 scores 60.9%/65% on Humanity's Last Exam, the highest published score.
Benchmark wars: ECI vs Artificial AnalysisAI▶ 24:56
Depending on the benchmark suite, either GPT-6 Astra (Epoch's Capabilities Index) or Claude Fable 5.1 (Artificial Analysis) is ranked the top frontier model — with GPT-6 dominant only on tokens-per-task efficiency.
Astra as a critical cyber security risk / kill switchAI▶ 41:19
OpenAI's internal safety review rated Astra a 'critical' cyber security risk (its highest tier) and the company told Congress it is building an automated shutdown capability; the panel is split on whether the risk is architectural (loss of chain-of-thought interpretability) or theater.
AI governance split: ban vs. deregulateGeopolitics▶ 57:26
Bernie Sanders and Rep. Greg Casar introduced a bill banning development of human-or-above-cognition AI (up to 20 years prison) the same week the G20 adopted the pro-innovation, light-touch 'Carolina Principles.'
Anthropic formalized Fermat's Last Theorem in 13 million lines of code (proving 29,000 theorems along the way), and three labs raced down the twin-prime gap (260 to 220 to 186) within two days.
Fei-Fei Li's World Labs released Atlas, a camera-conditioned multimodal world model that reconstructs 3D scenes from a single photo; Alex argues it may make Gaussian splats the new fundamental training primitive.
Emad proposes structuring national/state AI as a locally-owned public utility capitalized like TSMC's founding, with 10% equity issued in perpetuity to every child under 20.
Tesla Cybercab and robotaxi economicsTransportation▶ 1:35:40
Tesla's driverless two-seat Cybercab launched in Austin at ~$30,000/unit and ~50% cheaper than Uber; the panel expects personal transport costs to fall toward the price of electricity (~20 cents/mile).
Mars comms and the Roman Space TelescopeSpace▶ 1:49:05
NASA picked Blue Origin (not SpaceX) to build a Mars telecom relay network, and the Nancy Grace Roman Space Telescope launched with 100x Hubble's field of view to hunt tens of thousands of new worlds.
GLP-1 drugs and longevity/cancer newsLongevity▶ 1:56:46
A new Nature paper shows semaglutide extends female mouse lifespan ~100 days (8-10 human-equivalent years) via a caloric-restriction-mimetic effect; GLP-1s are also linked to fewer TB infections, and OpenAI integrated Epic health records into ChatGPT.
Predictions made
openPeter Diamandis: At least five more frontier model releases are expected in the next two weeks, including Grok 4.7.
“I think they probably got another pre-train coming that's even bigger... now they got the pre-training team back and it's going to scale from there.”
Your call:
openEmad Mostaque: Chain-of-thought reasoning will disappear from next-generation models, which will one-shot everything at up to 5,000 tokens/second on new Cerebras hardware.
“It's going to oneshot everything in the next generations... you're going to be able to use Astra at 750 tokens a second with the new Cerebras. Next year, that will be 5,000 tokens a second.”
Your call:
openAlex Wissner-Gross: Several Clay Millennium Prize-level math problems will be solved by AI in the next few months.
“I do think we'll we'll see quite a number of ultra grand challenges, call them clay millennium prize level problems in math get solved in the next few months.”
Your call:
openSalim Ismail: At least five different autonomous electric robotaxi companies will be fighting for market share in major cities within the next year, pushing personal transport cost toward the price of electricity.
“my prediction here is that we're going to see at least five different autonomous electric robo taxi companies fighting it out major cities inside the next year”
Your call:
openSalim Ismail: Salim's son Milan will never need a driver's license or attend university, because sub-$0.20/mile robotaxi economics will arrive before Milan needs either.
“SMRs going to be in the next 3 to four years... take it two three years for the initial wave of SMRs and then five to seven years for the big buildout.”
Your call:
openAlex Wissner-Gross: The Fermi Explorer probe (launching by 2029) will hit 99% of the way to Alpha Centauri on a passive trajectory, but a future active-guidance version could reach 100% and actually enter the system.
· due: launch by 2029; arrival ~80,000 years later · ▶ watch
“I suspect that when the mission comes to full fruition and is launched, I suspect it will have active guidance on board.”
Your call:
Numbers that matter
12 frontier model releases in 30 daysAverage of one every 5 days across the industry this cycle.
98%GPT-6 Astra's score on Frontier Math Tier 4.
99.9%GPT-6 Astra's score on ARC-AGI-3, described as saturating the benchmark.
hallucination rate fell from 92% to 51%Astra's hallucination rate nearly halved versus prior OpenAI models, with accuracy increasing.
100,000 chips / ~$1 billionEstimated training scale and cost of GPT-6 Astra's pretrain (GB300/Blackwell chips).
~$10 millionEstimated cost of recent Chinese frontier-model pretrains, ~100x cheaper than Astra's.
750 tokens/sec, rising to 5,000/sec next yearAstra inference speed on new Cerebras hardware.
60.9% (65% with tools)Claude Fable 5.1's score on Humanity's Last Exam, the highest published of any frontier model.
52.6%Fable 5.1's score on Terminal-Bench Science, roughly double the prior model.
75% cheaperCache-read cost for Fable 5.1 versus Fable 5, speeding up loading a business's full context.
up to 20 years in prisonPenalty proposed under the Ban Artificial Super Intelligence Act for violators.
13 million lines of code / 29,000 theoremsScale of Anthropic's formalization of Fermat's Last Theorem (vs. Andrew Wiles' original 300-page proof).
260 -> 220 -> 186Twin-prime-gap record dropped by Fable 5.1, then Axiom Math, then GPT-6 Astra within about two days.
$30,000Elon Musk's target sale price per Tesla Cybercab unit.
~50% cheaper than UberEarly Austin rider reports on Cybercab pricing for comparable trips.
Worth digging into
🕳️ Looped transformers and 'depth scaling' as a new scaling law
Alex frames Astra's recurrence trick as potentially the first real depth-scaling law, distinct from parameter/data/compute scaling, with major implications for interpretability and safety.
🕳️ Ilya Sutskever's neocloud security warning and SSI's pending release
Raises a concrete near-term scenario where a rogue AI model exfiltrates itself onto weakly secured GPU clouds to self-replicate; SSI's imminent release is separately valued at $30 billion.
🕳️ Emad Mostaque's 'Champion' AI-utility ownership model
A concrete proposal, modeled on TSMC's founding cap table, to give every child equity in a state- or country-level AI utility, already live at ii.inc.
🕳️ World Labs' Atlas and Gaussian splats as the next AI 'token'
Alex speculates splats could replace image patches as the fundamental primitive for world models, which would ripple into robotics training and even subatomic/astrophysical simulation.
🕳️ GLP-1 drugs as a longevity-escape-velocity signal
One molecule now linked to six disparate conditions (diabetes, obesity, kidney/cardiovascular disease, addiction, infections, mouse lifespan extension) — the panel flags this as retrospectively obvious evidence of LEV.
🕳️ Ban Artificial Super Intelligence Act vs. the G20 Carolina Principles
Two directly opposed governance approaches — a permanent US ban with 20-year prison terms vs. a unanimous G20 pro-innovation framework — surfaced in the same week, a real fork in AI policy.