GPT-6 SOL AND LUNA ARE OUT!!!
Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V
Read Video · Transkripsi & Wawasan
Episode ini punya transkripsi lengkap + AI insights
Akun gratis · tanpa kartu · 150 kredit saat daftar, cukup untuk episode ini
- 📄 Transkripsi lengkap dengan stempel waktu
- ✨ Ringkasan AI, kata kunci & peta pikiran
- 💡 Poin utama & kutipan
- Pembicara 1
Garis waktu episode
Introduction and overview of Claude Opus 5.5 release
- Anthropic released Claude Opus 5.5, the first model since Dario's essay calling for pacing the frontier, and it performs at Fable 5.1 level for most tasks at 40% lower cost than Opus 5.
- Opus 5.5 is faster, cheaper, and on many benchmarks actually takes the frontier position from Fable 5.1, which the host calls crazy to see.
- The host notes this is the first major model release since Anthropic began pacing the frontier, and external evaluators Meter and Frontier Design tested it before release.
Benchmark deep dive: Terminal Bench, Frontier Code, Cursor Bench, GDPval
- Opus 5.5 dominates Terminal Bench 4.0 with 66.4% versus Astra's 57.9% and Fable 5.1's 55.8%, a 10-point jump over Fable.
- On GDPval 2.1, OpenAI's benchmark for real-world knowledge work, Opus 5.5 scores 1846 ELO versus Fable 5.1's 1735, a 300+ point jump over GPT6 Astra's 1542.
- On Frontier Code V1.1 Opus 5.5 hits 54.4% versus Fable's 50% and Astra's 53.3%, and on Cursor Bench it scores 57.8% versus Fable 5.1's 51%.
- Opus 5.5 did not take first place on Automation Bench (Astra 41.4% vs 40%) or Terminal Bench Science, where GPT6 Astra still leads.
Pricing, speed, and cost-per-task efficiency analysis
- Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, a 20% price decrease from Opus 5, with cache reads at 20 cents versus 50 cents.
- The host argues cost per task completed matters more than cost per token, since a cheap model that uses 10x tokens is still expensive overall.
- On Automation Bench and Frontier Code, Opus 5.5 sits in the top-left quadrant of quality versus cost, with medium effort beating max effort on cost efficiency.
- Opus 5.5 generates output more than 30% faster than Opus 5 and requires less compute to serve.
Belum ada konten.
Belum ada konten.
Belum ada konten.
Transkrip dan wawasan dihasilkan oleh AI dan mungkin mengandung kesalahan. Akurasi bergantung pada kualitas audio dan kejelasan pembicara — jika ada yang tidak tepat, audio asli selalu menjadi sumber kebenaran.
Episode dan video siap dibaca
Episode podcast

The shape-shifting sounds of the accordion | Maria Telesheva
TED Talks Daily
27 Agu 202613:22EN
Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
26 Mei 202657:25EN
Scaling a $300K Moving Company in 60 Minutes
The Game with Alex Hormozi
14 Jul 202634:15EN
Ideasi & Pengelolaan SDM (part 1)
Creatalks
21 Jun 201951:13ID
#71 Die Deutschen im Amerikanischen Unabhängigkeitskrieg
Wer wir sind und warum das nicht klappte ...
19 Agu 202644:55DE特番|从蜂窝网络到手机革命:杨旸谈移动通信浪潮三十年
忽左忽右
16 Sep 20261:19:05ZH-Hans
Video

Give Me 12 Minutes and I’ll Give You 30 Years of Productivity Advice
Daniel Pink
14 Sep 202511:59EN
GPUs, TPUs, & The Economics of AI Explained | Gavin Baker Interview
Invest Like The Best
9 Des 20251:28:22EN
Neuroscientist: You Will NEVER Feel Stressed Again | Andrew Huberman
RESPIRE
27 Feb 202311:05EN
Ngaji Al Muhadzab Syirozy 1 Bagian 84
Miftahul Huda
4 Jun 202138:28ID
COMO SER GENTIL ESTÁ ACABANDO COM A SUA AUTORIDADE? | Fabiana Bertotti #152
Como Você Fez Isso?
30 Jul 20261:19:47PT