Did Elon catch up? (Grok 4.7 is here)
Check out Zapier! https://bit.ly/3TLc4tQ Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://x.ai/news/grok-4-7 https://x.com/GavinSBaker/status/2091542026072338623 https://artificialanalysis.ai/
Read Video · 文字起こしと深掘り分析
このエピソードには全文文字起こし + AI インサイトがあります
無料アカウント · カード不要 · 登録で150クレジット獲得、このエピソードのアンロックに十分
- 📄 タイムスタンプ付き全文文字起こし
- ✨ AI 要約・キーワード・マインドマップ
- 💡 重要ポイントと名言
エピソードのタイムライン
Elon Musk's prediction that Grok 4.7 would match Opus 5.0 and the model's release at the same price and speed as Grok 4.6
- Elon tweeted about a week before release that Grok 4.7 should be roughly on par with Opus 5.0, not 5.1, setting a specific benchmark expectation.
- Grok 4.7 launched as a notable improvement over Grok 4.6 at the same price and speed, with xAI claiming it works longer on difficult tasks and checks its own work more carefully.
- The blog post describes Grok 4.7 as xAI's most capable model for coding and knowledge work with their best calibrated safeguards to date.
Cursorbench 4.0 results show Grok 4.7 is competitive with Opus 5 but not beating it, at roughly half the cost
- On Cursorbench 4.0, Grok 4.7 sits just behind Opus 5 on max thinking settings but costs about half as much to run the benchmark.
- There is a major performance delta between low thinking effort at 33% and extra-high thinking effort at 46.3%, one of the biggest effort curves outside of GPT 5.6 Soul.
- The host notes xAI now owns Cursor, so the benchmark should be viewed with that ownership relationship in mind.
Token efficiency and steps-per-task comparisons show Fable 5.1 leading on quality while Grok 4.7 wins on cost-effectiveness
- On average output tokens per task, Grok 4.7 is very comparable to Opus 5, and at lowest thinking effort it uses few tokens but scores poorly.
- Fable 5.1 is the quality winner but is multiple times more expensive than Grok 4.7, so its cost per completed task is much higher despite similar token usage.
- GPT 5.6 Soul appears most efficient on steps taken per task, which matches the host's personal experience of it having the most direct shot to task completion.
主要な概念
- Grok 4.7— The newly released xAI model that is the central subject of the episode.
- cost per task completed— The key metric the host argues matters most when evaluating models for real workloads.
- CursorBench— The coding benchmark used to compare Grok 4.7 against Opus 5 and other frontier models.
注目の名言
Grok 4.7 should be roughly on par with Opus 50, not 5.1.
🔥— Elon's own prediction sets a high bar that the host then tests, revealing whether the claim holds up.
It is not beating it. It's just behind it on the max thinking setting for both models but it is much less expensive about half the cost to run that benchmark.
🤯— Shows that near-frontier performance at half the cost is the real value proposition, not outright superiority.
実行可能なテイクアウェイ
📊AI Model Evaluation
Cost per task completed is often more important than raw benchmark scores for real-world use.
This week, calculate the cost per task for your most common AI workflow using your current model and compare it to Grok 4.7's pricing.
Benchmark tables can be cherry-picked; always look for missing models or metrics.
Next time you see a model comparison, check if all relevant competitors are included and look for independent evaluations like Artificial Analysis.
💼Business Strategy
Open-weight models are gaining enterprise traction due to cost, control, and privacy.
Evaluate whether an open-weight model like Muse Spark 1.3 could replace a proprietary model in your pipeline this quarter.
Near-frontier models offer a sweet spot for industries that don't need the absolute best answer.
Identify tasks in your organization where a 5-10% performance drop is acceptable in exchange for 50% cost savings.
文字起こしとAIインサイトは自動生成されたものであり、誤りが含まれる場合があります。認識精度は音質や話者の発話の明瞭さに左右されます——内容に不就がある場合は、元の音声をご確認ください。
すぐに読めるエピソードと動画
ポッドキャスト

Why you need the growth mindset of a weed | Stephen Nobles
TED Talks Daily
2026年8月20日18:24EN
ChatGPT – The Super Assistant Era | BG2 Guest Interview
BG2Pod with Brad Gerstner and Bill Gurley
2026年3月15日1:03:40EN
Episode #218 ... Dostoevsky - Notes From Underground
Philosophize This!
2024年12月17日33:37EN
#27 「加藤さんって実生活が物足りないんでしょ?」
朝井リョウ・加藤千恵 信頼できない語り手
2026年8月7日1:01:05JA
No.221 雷鸟 CEO:新技术越来越多,我们为什么还需要一副智能眼镜?
三五环
2026年5月27日1:27:43ZH
Folge 1 - Ein verschwundenes Land
Zeitreise DDR
2025年11月9日22:49DE
動画

#1 Mindset Expert: Simple Mindset Shifts That Transform Your Body, Energy, & Life
Mel Robbins
2025年12月20日1:20:36EN
The Confidence Trick | Ian Robertson
CenterforBrainHealth
2020年8月24日59:41EN
Sleep and Relaxation: NSDR Yoga Nidra with Kelly Boys
The Alembic
2022年12月11日47:54EN
Un pays qui s'embrase, une caste qui s'embrasse... Avec Alexis Poulin
Idriss J. Aberkane, Ph.D x3
2026年8月18日2:12:55FR
#31 Der Kölner Dom - Wer war Meister Gerhard?
99 mal Geschichte
2025年11月5日54:34DE