Did Elon catch up? (Grok 4.7 is here)
Check out Zapier! https://bit.ly/3TLc4tQ Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://x.ai/news/grok-4-7 https://x.com/GavinSBaker/status/2091542026072338623 https://artificialanalysis.ai/
Read Video · 文字稿与深度分析
本集已有完整文字稿 + AI 深度分析
免费注册 · 无需信用卡 · 注册即获 150 积分,足够解锁本集
- 📄 完整文字稿含时间戳
- ✨ AI 摘要、关键词与思维导图
- 💡 核心要点与精彩引言
节目时间轴
Elon Musk's prediction that Grok 4.7 would match Opus 5.0 and the model's release at the same price and speed as Grok 4.6
- Elon tweeted about a week before release that Grok 4.7 should be roughly on par with Opus 5.0, not 5.1, setting a specific benchmark expectation.
- Grok 4.7 launched as a notable improvement over Grok 4.6 at the same price and speed, with xAI claiming it works longer on difficult tasks and checks its own work more carefully.
- The blog post describes Grok 4.7 as xAI's most capable model for coding and knowledge work with their best calibrated safeguards to date.
Cursorbench 4.0 results show Grok 4.7 is competitive with Opus 5 but not beating it, at roughly half the cost
- On Cursorbench 4.0, Grok 4.7 sits just behind Opus 5 on max thinking settings but costs about half as much to run the benchmark.
- There is a major performance delta between low thinking effort at 33% and extra-high thinking effort at 46.3%, one of the biggest effort curves outside of GPT 5.6 Soul.
- The host notes xAI now owns Cursor, so the benchmark should be viewed with that ownership relationship in mind.
Token efficiency and steps-per-task comparisons show Fable 5.1 leading on quality while Grok 4.7 wins on cost-effectiveness
- On average output tokens per task, Grok 4.7 is very comparable to Opus 5, and at lowest thinking effort it uses few tokens but scores poorly.
- Fable 5.1 is the quality winner but is multiple times more expensive than Grok 4.7, so its cost per completed task is much higher despite similar token usage.
- GPT 5.6 Soul appears most efficient on steps taken per task, which matches the host's personal experience of it having the most direct shot to task completion.
关键概念
- Grok 4.7— The newly released xAI model that is the central subject of the episode.
- cost per task completed— The key metric the host argues matters most when evaluating models for real workloads.
- CursorBench— The coding benchmark used to compare Grok 4.7 against Opus 5 and other frontier models.
精选金句
Grok 4.7 should be roughly on par with Opus 50, not 5.1.
🔥— Elon's own prediction sets a high bar that the host then tests, revealing whether the claim holds up.
It is not beating it. It's just behind it on the max thinking setting for both models but it is much less expensive about half the cost to run that benchmark.
🤯— Shows that near-frontier performance at half the cost is the real value proposition, not outright superiority.
可执行的洞察
📊AI Model Evaluation
Cost per task completed is often more important than raw benchmark scores for real-world use.
This week, calculate the cost per task for your most common AI workflow using your current model and compare it to Grok 4.7's pricing.
Benchmark tables can be cherry-picked; always look for missing models or metrics.
Next time you see a model comparison, check if all relevant competitors are included and look for independent evaluations like Artificial Analysis.
💼Business Strategy
Open-weight models are gaining enterprise traction due to cost, control, and privacy.
Evaluate whether an open-weight model like Muse Spark 1.3 could replace a proprietary model in your pipeline this quarter.
Near-frontier models offer a sweet spot for industries that don't need the absolute best answer.
Identify tasks in your organization where a 5-10% performance drop is acceptable in exchange for 50% cost savings.
转录文字与 AI 洞察均由模型自动生成,可能存在少量误差。识别效果与音频质量、语速和发音清晰度相关——如有内容看起来不对,以原始音频为准。
播客与视频,已可阅读
音频播客

How schools can nurture every student's genius | Trish Millines Dziko
TED Talks Daily
2026年9月7日17:50EN
The companies with the biggest gender pay gaps
The Daily Aus
2024年2月27日18:06EN
Andrew Private Conversations - The War Room
EMERGENCY MEETING TATE SPEECH
2025年1月26日1:24:19EN
E250|mRNA的第二战场:对话英博,拆解Moderna人类首个肿瘤疫苗三期突破
硅谷101
2026年8月28日1:16:50ZH-Hans
#18 「太宰のこと治って呼ぼうと思う」
朝井リョウ・加藤千恵 信頼できない語り手
2026年6月5日52:27JA
#406 | O ALERTA QUE O MERCADO ESTÁ DANDO SOBRE O CRÉDITO NO BRASIL
Market Makers
2026年8月30日2:05:02PT
视频

Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history
Dwarkesh Patel
2024年6月4日4:32:07EN
If you want 2026 to be the best year of your life, please watch this video…
Daniel Pink
2025年12月29日25:59EN
Deepseek did it again...
Matthew Berman
2026年9月11日16:39EN
从「上瘾模型」到「专注力训练」,如何在被算法理解的世界里重新找回主动?| 英文访谈 S9E33
声动活泼
2025年10月16日50:48ZH-Hans
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
2026年8月4日54:10SK