CLAIMS, CHECKED — RECEIPTS INCLUDED

THE HYPE SCOREBOARD

Every big local-AI headline, run down to the source. Red means the claim is busted, green means it checks out, amber means it's more complicated than the headline let on. Each verdict opens to its receipts, and every one links to the video where we walk through it. Read the sources yourself — that's the point.

$ scoreboard → 5 claims checked · 2 busted · 3 mixed

THE SCOREBOARD

5 claims

“AI finally beat the unbeatable AGI test.”

BUSTED

Real progress, wildly oversold. GPT-5.6 "Sol" set a genuine ARC-AGI-3 record and was the first model to fully win ONE game — but that's 7.8% vs an untrained human's 100%.

Receipts4 sources
  • GPT-5.6 Sol: 7.78% semi-private / 13.33% public, new SOTA, first to fully win a game (FT09, 152 actions vs 208 human): arcprize.org/results/openai-gpt-5-6-sol + @arcprize
  • Human baseline 100%, median 7.4 min untrained: ARC-AGI-3 Technical Report (arcprize.org)
  • March 2026 best was Gemini 3.1 Pro at 0.37%: officechai.com
  • ~$20,000 per max-reasoning eval: thealgorithmicbridge.com
Watch the check →AI "Beat" the Impossible Test: 7.8% · 1:24

“The AI bubble is bursting — it all crashes next week.”

MIXED

The froth in AI valuations and capex is real — but "it all crashes next week" is an unprovable guess, and the part that survives the frenzy is the local, open AI you run yourself.

Receipts5 sources
  • BIS Annual Report (~Jun 29 2026; >$1T hyperscaler capex, canal/railway/dot-com manias): bis.org/publ/arpdf/ar2026e1.htm
  • Draft US Treasury report (Jul 6, via NOTUS; AI leaders more mature/profitable than dot-com, no crash date): notus.org
  • MIT NANDA "GenAI Divide" 95% of enterprise pilots no P&L: fortune.com
  • Circular Nvidia/OpenAI/Oracle financing ~$800B: gfmag.com
  • Palantir −30% YTD vs +85% revenue: stockanalysis.com/stocks/pltr
Watch the check →Is the AI Bubble About to Pop? · 1:39

“It tops SWE-Bench, so it's the best AI coder money can buy.”

BUSTED

The leaderboard is real, but OpenAI's own audit found roughly a third of SWE-Bench Pro's tasks are broken — so a headline coding score can't reliably prove which model is actually the best engineer.

Receipts6 sources
  • OpenAI "Separating signal from noise in coding evaluations" (Jul 8 2026; ~30% of SWE-Bench Pro broken, recommendation retracted): openai.com/index/separating-signal-from-noise-coding-evaluations
  • 731 tasks, automated flagged 200 (27.4%): alphasignal.ai
  • Five engineers flagged 249 (34.1%), methods agreed ~74%: officechai.com
  • Flaw categories (overly strict/underspecified/low-coverage tests): startuphub.ai
  • Top score 80.3%: benchlm.ai/benchmarks/swePro
  • Precedent — OpenAI dropped SWE-Bench Verified after ~59% of 138 sampled failed tasks proved flawed: openai.com/index/why-we-no-longer-evaluate-swe-bench-verified
Watch the check →The "Best AI Coder" Benchmark Broken? · 1:05

“Hand over your health data to Samsung's AI, or it wipes everything.”

MIXED

The scary dialog was real; the "hands over your data or we wipe everything" mass-deletion was NOT.

Receipts6 sources
  • Dialog + consent wording: 9to5Google (Jul 13)
  • Samsung's clarification: 9to5Google (Jul 15)
  • Four data categories: Android Authority
  • "Decline doesn't delete your data": Android Police / SamMobile
  • Export/erase how-to: Samsung support ANS10001379
  • Unpacked Jul 22: SamMobile
Watch the check →Did Samsung REALLY Threaten to Delete Your Health Data? · 1:21

“A free open model beats GPT — just download it and run frontier AI at home.”

MIXED

The open weights really are free — but GLM-5.2 needs 200GB+ to even load, so for almost everyone it's a rented API, not something your gaming PC can run.

Receipts5 sources
  • 744B params + benchmark wins (SWE-bench Pro 62.1 vs 58.6): VentureBeat, apidog
  • Memory to load: 2-bit ~245GB (avenchat), INT4 372GB (Spheron)
  • RTX 5090 = 32GB current flagship: nvidia.com
  • API pricing (~$1.40/M in, $4.40/M out): apidog/OpenRouter
  • Weights: Hugging Face (MIT license)
Watch the check →GLM 5.2 & Kimi: Free Giants You Can't Run · 5:35

HYPE, CHECKED — WEEKLY

One email a week: what actually shipped in local AI, what was hype, receipts included. No spam, unsubscribe any time.