Anthropic went CRAZY (Opus 5.5)
Check out all our Opus 5.5 tests: https://forwardfuture.com/opus-5-5-review Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com Thanks to Thariq for joining! https://x.com/trq212 My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://www.anthropic.com/claude-opus-5-5
Read Video · 文字稿與深度分析
本集已有完整文字稿 + AI 深度分析
免費註冊 · 無需信用卡 · 註冊即獲 150 積分,足夠解鎖本集
- 📄 完整文字稿含時間戳
- ✨ AI 摘要、關鍵詞與心智圖
- 💡 核心要點與精彩引言
- 說話人 1
節目時間軸
Anthropic releases Claude Opus 5.5, the first frontier model since calling for pacing the frontier, delivering top benchmark scores at lower cost and higher speed.
- Opus 5.5 is the first major Anthropic release since Dario's essay calling for pacing the frontier, yet it still lands as the absolute frontier model, beating Fable 5.1 on most benchmarks.
- It performs at Fable 5.1 level for most tasks while costing 40% less to run than Opus 5 and generating output over 30% faster.
- The last .5 release, Opus 4.5, was the inflection point where coding models became capable of handling large multi-hour autonomous tasks, so expectations for 5.5 are high.
Benchmark results show Opus 5.5 dominating coding, knowledge work, and reasoning evaluations.
- On Terminal Bench 4.0, Opus 5.5 scores 66.4% versus 57.9% for Astra and 55.8% for Fable 5.1, a 10-point jump over the previous frontier.
- On GDPval 2.1, OpenAI's real-world knowledge work benchmark, it hits an 1846 ELO score versus 1735 for Fable 5.1, a 300+ point jump.
- It also leads on Frontier Code V1.1 (54.4), Cursorbench (57.8), and Humanity's Last Exam (67%), though it places second to Astra on Automation Bench and Terminal Bench Science.
Pricing drops modestly per token, but cost-per-task efficiency is the real story.
- Opus 5.5 costs $4 per million input tokens versus $5 for Opus 5, and $20 versus $25 per million output tokens, with cache reads at 20 cents versus 50 cents.
- The 40% total cost reduction on typical workloads comes from combining lower token prices with fewer tokens needed to complete a task.
- On quality-versus-cost-per-task charts for Automation Bench, Frontier Code, GDPval, and Terminal Bench, Opus 5.5 consistently occupies the ideal upper-left quadrant.
關鍵概念
- Opus 5.5— Anthropic's new frontier model that dominates most benchmarks while being cheaper and faster than Opus 5.
- pacing the frontier— Anthropic's stated policy of slowing the release of its most capable models while making frontier intelligence broadly accessible.
- Terminal Bench 4.0— The key agentic coding benchmark where Opus 5.5 scored 66.4%, a 10-point jump over Fable 5.1.
精選金句
It performs at the level of Fable 5.1 for most tasks, but it costs 40% less to run than Opus 5.
🤯— A frontier model that is simultaneously cheaper and more capable overturns the assumption that top intelligence must cost more.
That is over a 300 point ELO jump.
🤯— A 300-point ELO gain on GDPval is an enormous leap that reveals how fast frontier capability is compounding.
可執行的洞察
📊AI Model Evaluation
Cost per task completed matters far more than cost per million tokens when comparing models.
This week, run the same real task through two models at different effort settings and log the total cost and quality of each.
Medium thinking effort often beats max effort on both cost and quality for coding tasks.
Test your most common coding task at medium and max effort this week and compare the results before defaulting to max.
🤖Agentic Workflow
Concise model explanations are critical when managing many parallel agents and context switching.
This week, add an instruction to your agent prompts requiring a three-bullet summary of completed work before any detail.
Tool calling and computer use are improving fast enough to automate real desktop workflows.
Pick one repetitive desktop task this week and try automating it with a computer-use agent.
轉錄文字與 AI 洞察均由模型自動生成,可能存在少量誤差。辨識效果與音訊品質、語速及發音清晰度相關——若有內容看起來有誤,以原始音訊為準。
Podcast 與影片,已可閱讀
音訊 Podcast

The Mystery of Sea Creatures (2/5): A giant Jurassic sea dragon, unearthed | Dean R. Lomax
TED Talks Daily
2026年7月18日15:33EN
Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
2026年5月26日57:25EN
Scaling a $300K Moving Company in 60 Minutes
The Game with Alex Hormozi
2026年7月14日34:15EN
EP197 | 減重名醫:每一種飲食方法都會失敗 ft. 初日診所院長 宋晏仁醫師
博音
2025年10月20日1:13:44ZH-Hant
♯23 ホテル八木のMVVを今一度整理してみる
旅館経営と観光の現場から - ホテル八木のリアルトーク
2024年10月26日12:21JA
#67 Wie stark war August der Starke?
Wer wir sind und warum das nicht klappte ...
2026年7月22日1:24:57DE
影片

Did Elon catch up? (Grok 4.7 is here)
Matthew Berman
2026年9月22日16:57EN
#1 Mindset Expert: Simple Mindset Shifts That Transform Your Body, Energy, & Life
Mel Robbins
2025年12月20日1:20:36EN
The Confidence Trick | Ian Robertson
CenterforBrainHealth
2020年8月24日59:41EN
2026/08/24(一) 輝達伺服器傳漲價15%:AI成本暴增,成本誰吸收?
財女珍妮
2026年8月24日30:06ZH-Hant
#31 Der Kölner Dom - Wer war Meister Gerhard?
99 mal Geschichte
2025年11月5日54:34DE