Which Industries Survive AI, The New AI Benchmarks, and the 2026 Recursive Learning Timeline | #218
Matt Fitzpatrick, CEO of Invisible Technologies and former global head of McKinsey's QuantumBlack Labs, joins the Moonshots crew to argue that AI's impact will hit industries unevenly -- media, legal services, and BPOs face structural disruption while oil & gas and real estate change less. He walks through why most enterprise AI pilots fail (dirty data, no clear KPI-owning operator, 'let a thousand flowers bloom' science projects), why Klarna's fully-agentic contact center rollback happened, and why he expects a proliferation of thousands of hyper-narrow, task-specific benchmarks to replace broad public ones. Alex Wissner-Gross pushes back hard on whether human-in-the-loop labor marketplaces like Invisible's Meridial survive as reinforcement fine-tuning gets more data-efficient and recursive self-improvement approaches; Fitzpatrick argues human feedback becomes more, not less, necessary as models specialize. The episode closes with concrete enterprise case studies (Charlotte Hornets draft scouting, Lifespan MD, US Navy underwater drones, Swiss Gear inventory forecasting) and Fitzpatrick's 2026 predictions: multi-agent orchestration, a multimodal leap, and RL gyms/'mirror world' simulation environments.
Uneven AI disruption across industriesEconomy▶ 4:56
Fitzpatrick argues AI will not impact all industries equally -- media, legal services, and business process outsourcing face major structural change, while oil & gas and real estate largely retain their existing function and decision-making.
Discussion of whether companies should build in-house AI expertise, hire a chief AI officer, or rent/partner externally (e.g. with Invisible); most companies lack in-house skills, especially smaller firms without a CTO.
Klarna announced a fully agentic customer-service operation handling 2.3 million calls a month, then rolled it back to human agents 8-12 months later; the panel debates why, with Fitzpatrick arguing a properly designed system should always keep humans in the loop for complex, non-first-line issues.
Fitzpatrick argues enterprises need thousands of task- and vertical-specific evals (e.g. per industry, per document type) rather than relying on broad public coding/cognition benchmarks, since real deployment risk hinges on task-specific accuracy.
RLHF vs. reinforcement fine-tuning debateAI▶ 40:44
Alex Wissner-Gross challenges whether Invisible's human-labor marketplace (Meridial) for RLHF has a long-term future as reinforcement fine-tuning becomes more data-efficient and AI researchers approach human-level ML research capability; Fitzpatrick counters that human feedback becomes more essential as models specialize into low-precedent-data domains.
Fitzpatrick cites an MIT finding that only about 5% of enterprise AI initiatives reach production, attributing failure to dirty/fragmented data, lack of focus, and the 'let a thousand flowers bloom' pattern where no clear operational KPI owner exists.
Salim Ismail argues true AI transformation means redesigning the entire functional flow of a business (e.g. a printer company) around automated functions rather than simply automating existing human job roles; Diamandis frames this as the innovation-at-the-edge vs. legacy-core dynamic.
Fitzpatrick details Invisible client work: Charlotte Hornets draft-prep computer vision on player movement patterns, Lifespan MD's HIPAA-compliant multi-tenant data platform (Neuron), US Navy/SAIC underwater drone swarm decisioning, and Swiss Gear inventory forecasting across 750 combined data tables.
Proprietary data protection vs. frontier LLM APIsAI▶ 53:36
Dave Blundin raises the risk of feeding proprietary enterprise data (banks, insurers, hospitals) into third-party LLM APIs; Fitzpatrick notes not all company data is equally sensitive and expects continued growth of on-premise/small-language-model approaches for the truly proprietary slice.
Future of human expertise and last jobs standingEconomy▶ 1:09:34
The panel debates which job categories survive AI longest -- Fitzpatrick points to oil & gas field expertise, real estate judgment, and physical/high-touch trades; Wissner-Gross offers competing hypotheses (politicians, top scientists, or high-authenticity/high-touch roles).
AI in government and regulatory processesGeopolitics▶ 1:14:57
Salim Ismail highlights AI's potential to cut permitting and public-sector process timelines dramatically, citing studies on energy/data-center permitting and OECD findings on licensing and compliance cycle times.
Fitzpatrick previews his 2026 predictions: proliferation of orchestrated multi-agent task-specific systems, a multimodal (audio/video) leap in how people interact with models, and growing use of 'mirror world'/RL gym simulated environments to test agents before real-world deployment.
Predictions made
openAlex Wissner-Gross: A form of recursive self-improvement -- an AI researcher that is as good as or better than human ML researchers at building models -- will be achieved.
EP #? · · due: 2-3 years (outer bound stated as 10-15 years) · ▶ watch
“My timelines are approximately two to three years for a sub-element of recursive self-improvement where we get our AI researcher that's as good, if not stronger, than the human researchers for building ML models as a conservative outer bound.”
Your call:
openDave Blundin: Self-improving massive foundation models will reach superhuman IQ.
“It's looking more and more likely that these self-improving massive foundation models are going to get to, you know, superhuman IQ this year. This year being 2026.”
Your call:
openMatt Fitzpatrick: Multi-agent teams -- task-specific agents orchestrated by an LLM rather than one decisioning agent -- will become a dominant enterprise AI deployment architecture.
“You'll train task-specific agents for individual tasks, usually orchestrated by an LLM.”
Your call:
openMatt Fitzpatrick: AI interaction will take a 'multimodal leap' with video, image, and especially audio becoming a much bigger part of how people engage with models, less text-based than historically.
“I don't think that will all be text-based like it has been historically.”
Your call:
openMatt Fitzpatrick: Adoption of 'mirror world'/RL gym simulated digital-twin environments for testing AI agents and tasks before real-world rollout will grow among both model builders and enterprises.
“They were saying that job profile I think will two, three, four X over the next couple years.”
Your call:
Numbers that matter
Frontier models have shown 50-100% improvement on most benchmark dimensions over the last 3 yearsFitzpatrick citing broad public benchmark trends as evidence models are improving rapidly overall.
Klarna's AI reportedly replaced 700 full-time agents, handled 2.3 million calls/month, projected $40M/year savingsCited as the widely publicized (and later reversed) agentic customer-service success story.
Standard VC term sheet legal fees are capped at $50,000 but consistently bill out to just under that capDave Blundin's example of highly templated, automatable legal work in venture financings.
US spends ~$13,000-$14,000 per patient per capita on health care vs. $2,500-$3,000 in Germany/Canada; 30-40% is admin costUsed to argue AI's biggest near-term healthcare win is administrative burden, not clinical decisioning.
A major global bank runs 300 separate, siloed customer databasesIllustrates enterprise data fragmentation as a core blocker to enterprise AI deployment.
Only about 5% of enterprise AI models/pilots reach production (MIT report)Central statistic framing why enterprise AI adoption has lagged despite model capability gains.
Invisible combined 750 data tables for Swiss Gear; inventory coverage grew ~30% and reliably forecasted SKUs roughly doubledConcrete outcome metrics from Invisible's inventory-forecasting engagement, delivered within a couple of months.
~25% of each US high school graduating class enters a field that did not exist when they were in high schoolFitzpatrick's evidence that labor markets historically absorb technological disruption via new job categories.
~20% of US employment is in 'digital ecosystem' jobs; ~9% of US citizens are full-time social media influencers (WSJ)Cited as evidence of how much the economy has already shifted toward digitally native work.
AI-assisted permitting could cut energy/data-center project timelines by 50%; OECD: AI could shrink public-sector licensing/compliance cycle times by 70%Cited by Salim Ismail as the biggest near-term positive societal use of AI in government.
Worth digging into
🕳️ Klarna's agentic contact-center rollback
A widely publicized real-world case of an AI deployment being announced as a massive success (700-agent replacement, $40M projected savings) and then reversed within a year -- a direct counterpoint to enterprise AI hype.
🕳️ Invisible's Meridial marketplace vs. AI-driven recursive self-improvement
Alex Wissner-Gross directly challenges whether human-labor marketplaces for RLHF/fine-tuning survive as AI researchers approach human-level ML research capability -- a load-bearing tension for Invisible's business model.
🕳️ MIT's '5% of enterprise AI pilots reach production' finding
This is the episode's central statistic explaining why enterprise AI adoption has underperformed model capability gains, but the underlying study and methodology are not named.
🕳️ BloombergGPT as a cautionary precedent
Wissner-Gross cites BloombergGPT's proprietary-data approach being leapfrogged within months by generalist frontier models as a parable for whether proprietary fine-tuning strategies can survive.
🕳️ Apple's ~18 secret internal disruption teams
Dave Blundin proposes Apple's skunkworks model (small stealth teams sent to disrupt adjacent industries, e.g. producing the Apple Watch) as the organizational template other large companies should adopt for AI transformation.
🕳️ OECD and permitting-AI timeline studies
Salim Ismail cites specific, high-stakes figures (50% cut in energy/data-center permitting time, 70% cut in public-sector cycle times) with major implications for housing and infrastructure policy.