GPT-6 SOL AND LUNA ARE OUT!!!
Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V
Read Video · Transkript & İçgörüler
Bu bölümün tam transkripti + AI içgörüleri var
Ücretsiz hesap · kart gerekmez · kayıtta 150 kredi, bu bölümü açmaya yeterli
- 📄 Zaman damgalı tam transkript
- ✨ AI özeti, anahtar kelimeler ve zihin haritası
- 💡 Ana çıkarımlar ve alıntılar
- Konuşmacı 1
Bölüm zaman çizelgesi
Introduction and overview of Claude Opus 5.5 release
- Anthropic released Claude Opus 5.5, the first model since Dario's essay calling for pacing the frontier, and it performs at Fable 5.1 level for most tasks at 40% lower cost than Opus 5.
- Opus 5.5 is faster, cheaper, and on many benchmarks actually takes the frontier position from Fable 5.1, which the host calls crazy to see.
- The host notes this is the first major model release since Anthropic began pacing the frontier, and external evaluators Meter and Frontier Design tested it before release.
Benchmark deep dive: Terminal Bench, Frontier Code, Cursor Bench, GDPval
- Opus 5.5 dominates Terminal Bench 4.0 with 66.4% versus Astra's 57.9% and Fable 5.1's 55.8%, a 10-point jump over Fable.
- On GDPval 2.1, OpenAI's benchmark for real-world knowledge work, Opus 5.5 scores 1846 ELO versus Fable 5.1's 1735, a 300+ point jump over GPT6 Astra's 1542.
- On Frontier Code V1.1 Opus 5.5 hits 54.4% versus Fable's 50% and Astra's 53.3%, and on Cursor Bench it scores 57.8% versus Fable 5.1's 51%.
- Opus 5.5 did not take first place on Automation Bench (Astra 41.4% vs 40%) or Terminal Bench Science, where GPT6 Astra still leads.
Pricing, speed, and cost-per-task efficiency analysis
- Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, a 20% price decrease from Opus 5, with cache reads at 20 cents versus 50 cents.
- The host argues cost per task completed matters more than cost per token, since a cheap model that uses 10x tokens is still expensive overall.
- On Automation Bench and Frontier Code, Opus 5.5 sits in the top-left quadrant of quality versus cost, with medium effort beating max effort on cost efficiency.
- Opus 5.5 generates output more than 30% faster than Opus 5 and requires less compute to serve.
Henüz içerik yok.
Henüz içerik yok.
Henüz içerik yok.
Transkript ve içgörüler yapay zeka tarafından oluşturulur ve hatalar içerebilir. Doğruluk, ses kalitesine ve konuşmacıların netliliğine bağlıdır — bir şey yanlış görünse orijinal ses her zaman doğru kaynaktır.
Bölümler ve videolar okunmaya hazır
Podcast bölümleri
INTJ Personality Type Advice - 0088
Personality Hacker Podcast
19 Eki 20151:05:25EN
A journalist's trick for talking to people you can't stand | Joshua Johnson
TED Talks Daily
3 Ağu 20269:31EN
Humanity's First Star Probe, Architect Labs Beats NVIDIA 3.4x, Musk Wants Satellites to Cool Earth | EP #285
Moonshots with Peter Diamandis
2 Eyl 20261:57:22EN
Kadınlar ve Erkekler Neden Arkadaş Olamaz?
Kendine İyi Davran
22 Eyl 202617:55TR
Vorbei am Gatekeeper: Mit der richtigen Leadquelle direkt zum Entscheider
B2B Sales on Air
16 Eki 202533:15DE
SÉRIE: RELIGIÃO TÓXICA - A GRAÇA NÃO É O QUE VOCÊ PENSA| PR.YAN AUGUSTO
Minha Igreja Na Cidade
27 Tem 202655:27PT
Videolar

Ajeya Cotra – "This might be the clearest warning shot we ever get"
Dwarkesh Patel
1 Eyl 20262:20:33EN
🟣 The Collison Brothers LIVE on TBPN
TBPN
30 Nis 20261:28:56EN
Take Upwork jobs, let AI do them (Astra etc). It's absurd.
Greg Isenberg
23 Eyl 202648:05EN
Hannah Arendt e a Banalidade do Mal
A Filosofia Explica
8 Haz 202213:50PT
Un pays qui s'embrase, une caste qui s'embrasse... Avec Alexis Poulin
Idriss J. Aberkane, Ph.D x3
18 Ağu 20262:12:55FR