GPT-6 SOL AND LUNA ARE OUT!!!
Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V
Read Video · Transkript & Insights
Diese Folge hat ein vollständiges Transkript + KI-Insights
Kostenloses Konto · keine Karte nötig · 150 Credits nach Registrierung, genug für diese Episode
- 📄 Vollständiges Transkript mit Zeitstempeln
- ✨ KI-Zusammenfassung, Keywords & Mindmap
- 💡 Kernaussagen & Zitate
- Sprecher 1
Episoden-Zeitlinie
Introduction and overview of Claude Opus 5.5 release
- Anthropic released Claude Opus 5.5, the first model since Dario's essay calling for pacing the frontier, and it performs at Fable 5.1 level for most tasks at 40% lower cost than Opus 5.
- Opus 5.5 is faster, cheaper, and on many benchmarks actually takes the frontier position from Fable 5.1, which the host calls crazy to see.
- The host notes this is the first major model release since Anthropic began pacing the frontier, and external evaluators Meter and Frontier Design tested it before release.
Benchmark deep dive: Terminal Bench, Frontier Code, Cursor Bench, GDPval
- Opus 5.5 dominates Terminal Bench 4.0 with 66.4% versus Astra's 57.9% and Fable 5.1's 55.8%, a 10-point jump over Fable.
- On GDPval 2.1, OpenAI's benchmark for real-world knowledge work, Opus 5.5 scores 1846 ELO versus Fable 5.1's 1735, a 300+ point jump over GPT6 Astra's 1542.
- On Frontier Code V1.1 Opus 5.5 hits 54.4% versus Fable's 50% and Astra's 53.3%, and on Cursor Bench it scores 57.8% versus Fable 5.1's 51%.
- Opus 5.5 did not take first place on Automation Bench (Astra 41.4% vs 40%) or Terminal Bench Science, where GPT6 Astra still leads.
Pricing, speed, and cost-per-task efficiency analysis
- Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, a 20% price decrease from Opus 5, with cache reads at 20 cents versus 50 cents.
- The host argues cost per task completed matters more than cost per token, since a cheap model that uses 10x tokens is still expensive overall.
- On Automation Bench and Frontier Code, Opus 5.5 sits in the top-left quadrant of quality versus cost, with medium effort beating max effort on cost efficiency.
- Opus 5.5 generates output more than 30% faster than Opus 5 and requires less compute to serve.
Noch keine Inhalte generiert.
Noch keine Inhalte generiert.
Noch keine Inhalte generiert.
Transkript und Insights werden KI-generiert und können Fehler enthalten. Die Genauigkeit hängt von der Audioqualität und der Deutlichkeit der Sprecher ab — bei Unklarheiten ist das Originalaudio die maßgebliche Quelle.
Episoden und Videos zum Lesen
Podcast-Episoden

How I turn joy into art | Yinka Ilori
TED Talks Daily
5. Aug. 202611:20EN
Essentials: Control Your Brain Chemistry for Focus, Motivation & Well-Being
Huberman Lab
6. Aug. 202634:37EN
(Preview) Astra (and AGI?) Arrives, Meta’s Muse and the Agent Opportunity, Anthropic and the Revival of (P)Doom Angst
Sharp Tech with Ben Thompson
10. Sept. 202624:58EN
#6: Offseason vs. Preseason mit Niklas Jauch
We talking about practice
21. März 20211:00:10DE
EP678 | 🎮
Gooaye 股癌
11. Juli 202652:46ZH-Hant
163.- “Mujeres, dinero y el miedo a incomodar” con Maca Riva
LA MAGIA DEL CAOS con Aislinn Derbez
26. Mai 20261:27:54ES
Videos

What is Loop Engineering?
KodeKloud
14. Juli 20266:47EN
Brad Gerstner: No AI Bubble, Semis Eat the Nasdaq & AI's Take Off Problem
All-In Podcast
17. Sept. 202618:10EN
Shortened: GoogleAC_English(GB)_AYF-Choice_9x16_20s
Video ad upload channel for 208-574-4978
7. März 20260:09EN
#32 Die Schlacht von Worringen - Der Freiheitskampf der Kölner
99 mal Geschichte
12. Nov. 202553:29DE
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
4. Aug. 202654:10SK