GPT-6 SOL AND LUNA ARE OUT!!!
Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V
Leggi video · Trascrizione e analisi
Questo episodio ha una trascrizione completa e analisi AI
Account gratuito · nessuna carta necessaria · 150 crediti all'iscrizione, sufficienti per sbloccare questo episodio
- 📄 Trascrizione completa con timestamp
- ✨ Riepilogo AI, parole chiave e mappa mentale
- 💡 Conclusioni e citazioni chiave
- Parlante 1
Cronologia dell'episodio
Introduction and overview of Claude Opus 5.5 release
- Anthropic released Claude Opus 5.5, the first model since Dario's essay calling for pacing the frontier, and it performs at Fable 5.1 level for most tasks at 40% lower cost than Opus 5.
- Opus 5.5 is faster, cheaper, and on many benchmarks actually takes the frontier position from Fable 5.1, which the host calls crazy to see.
- The host notes this is the first major model release since Anthropic began pacing the frontier, and external evaluators Meter and Frontier Design tested it before release.
Benchmark deep dive: Terminal Bench, Frontier Code, Cursor Bench, GDPval
- Opus 5.5 dominates Terminal Bench 4.0 with 66.4% versus Astra's 57.9% and Fable 5.1's 55.8%, a 10-point jump over Fable.
- On GDPval 2.1, OpenAI's benchmark for real-world knowledge work, Opus 5.5 scores 1846 ELO versus Fable 5.1's 1735, a 300+ point jump over GPT6 Astra's 1542.
- On Frontier Code V1.1 Opus 5.5 hits 54.4% versus Fable's 50% and Astra's 53.3%, and on Cursor Bench it scores 57.8% versus Fable 5.1's 51%.
- Opus 5.5 did not take first place on Automation Bench (Astra 41.4% vs 40%) or Terminal Bench Science, where GPT6 Astra still leads.
Pricing, speed, and cost-per-task efficiency analysis
- Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, a 20% price decrease from Opus 5, with cache reads at 20 cents versus 50 cents.
- The host argues cost per task completed matters more than cost per token, since a cheap model that uses 10x tokens is still expensive overall.
- On Automation Bench and Frontier Code, Opus 5.5 sits in the top-left quadrant of quality versus cost, with medium effort beating max effort on cost efficiency.
- Opus 5.5 generates output more than 30% faster than Opus 5 and requires less compute to serve.
Nessun contenuto ancora generato.
Nessun contenuto ancora generato.
Nessun contenuto ancora generato.
Trascrizione e analisi sono generate dall'AI e possono contenere errori. L'accuratezza dipende dalla qualità dell'audio e dalla chiarezza dei parlanti: in caso di dubbi, l'audio originale resta la fonte attendibile.
Episodi e video pronti da leggere
Episodi podcast

Cook Samin Nosrat (‘Salt Fat Acid Heat’) Returns with ‘Good Things’
Talk Easy with Sam Fragoso
7 set 20251:12:44EN
As Trump Purges Immigration Judges, One Speaks Out
The Daily
23 giu 202635:41EN
How to make learning as addictive as social media | Luis von Ahn
TED Talks Daily
7 set 202614:23EN
EP688 | 🥽
Gooaye 股癌
15 ago 202651:29ZH-Hant
Rotten to the Core: Apple attackiert OpenAI-Kultur | ohne Belege: 200 Ökonomen warnen vor KI-Jobverlust | Nadella warnt vor KI-Daten-Risiko #579
Doppelgänger
14 lug 20261:13:47DE
163.- “Mujeres, dinero y el miedo a incomodar” con Maca Riva
LA MAGIA DEL CAOS con Aislinn Derbez
26 mag 20261:27:54ES
Video

Anthropic's CEO: ‘We Don’t Know if the Models Are Conscious’ | Interesting Times with Ross Douthat
Interesting Times
12 feb 20261:02:30EN
Staphylococcus aureus
Osmosis from Elsevier
14 ott 202014:32EN
Give me 11 Minutes and I'll Make you Dangerously Persuasive
Daniel Pink
23 dic 202510:39EN
10 habitudes qui m’ont VRAIMENT fait perdre du poids
leawellnesss
15 apr 202614:34FR
#35 Das Attentat von Anagni
99 mal Geschichte
5 dic 202541:08DE