Did Elon catch up? (Grok 4.7 is here)
Check out Zapier! https://bit.ly/3TLc4tQ Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://x.ai/news/grok-4-7 https://x.com/GavinSBaker/status/2091542026072338623 https://artificialanalysis.ai/
Read Video · Transkripsi & Wawasan
Episode ini punya transkripsi lengkap + AI insights
Akun gratis · tanpa kartu · 150 kredit saat daftar, cukup untuk episode ini
- 📄 Transkripsi lengkap dengan stempel waktu
- ✨ Ringkasan AI, kata kunci & peta pikiran
- 💡 Poin utama & kutipan
Garis waktu episode
Elon Musk's prediction that Grok 4.7 would match Opus 5.0 and the model's release at the same price and speed as Grok 4.6
- Elon tweeted about a week before release that Grok 4.7 should be roughly on par with Opus 5.0, not 5.1, setting a specific benchmark expectation.
- Grok 4.7 launched as a notable improvement over Grok 4.6 at the same price and speed, with xAI claiming it works longer on difficult tasks and checks its own work more carefully.
- The blog post describes Grok 4.7 as xAI's most capable model for coding and knowledge work with their best calibrated safeguards to date.
Cursorbench 4.0 results show Grok 4.7 is competitive with Opus 5 but not beating it, at roughly half the cost
- On Cursorbench 4.0, Grok 4.7 sits just behind Opus 5 on max thinking settings but costs about half as much to run the benchmark.
- There is a major performance delta between low thinking effort at 33% and extra-high thinking effort at 46.3%, one of the biggest effort curves outside of GPT 5.6 Soul.
- The host notes xAI now owns Cursor, so the benchmark should be viewed with that ownership relationship in mind.
Token efficiency and steps-per-task comparisons show Fable 5.1 leading on quality while Grok 4.7 wins on cost-effectiveness
- On average output tokens per task, Grok 4.7 is very comparable to Opus 5, and at lowest thinking effort it uses few tokens but scores poorly.
- Fable 5.1 is the quality winner but is multiple times more expensive than Grok 4.7, so its cost per completed task is much higher despite similar token usage.
- GPT 5.6 Soul appears most efficient on steps taken per task, which matches the host's personal experience of it having the most direct shot to task completion.
Konsep utama
- Grok 4.7— The newly released xAI model that is the central subject of the episode.
- cost per task completed— The key metric the host argues matters most when evaluating models for real workloads.
- CursorBench— The coding benchmark used to compare Grok 4.7 against Opus 5 and other frontier models.
Kutipan penting
Grok 4.7 should be roughly on par with Opus 50, not 5.1.
🔥— Elon's own prediction sets a high bar that the host then tests, revealing whether the claim holds up.
It is not beating it. It's just behind it on the max thinking setting for both models but it is much less expensive about half the cost to run that benchmark.
🤯— Shows that near-frontier performance at half the cost is the real value proposition, not outright superiority.
Tindakan nyata
📊AI Model Evaluation
Cost per task completed is often more important than raw benchmark scores for real-world use.
This week, calculate the cost per task for your most common AI workflow using your current model and compare it to Grok 4.7's pricing.
Benchmark tables can be cherry-picked; always look for missing models or metrics.
Next time you see a model comparison, check if all relevant competitors are included and look for independent evaluations like Artificial Analysis.
💼Business Strategy
Open-weight models are gaining enterprise traction due to cost, control, and privacy.
Evaluate whether an open-weight model like Muse Spark 1.3 could replace a proprietary model in your pipeline this quarter.
Near-frontier models offer a sweet spot for industries that don't need the absolute best answer.
Identify tasks in your organization where a 5-10% performance drop is acceptable in exchange for 50% cost savings.
Transkrip dan wawasan dihasilkan oleh AI dan mungkin mengandung kesalahan. Akurasi bergantung pada kualitas audio dan kejelasan pembicara — jika ada yang tidak tepat, audio asli selalu menjadi sumber kebenaran.
Episode dan video siap dibaca
Episode podcast

Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
26 Mei 202657:25EN
Scaling a $300K Moving Company in 60 Minutes
The Game with Alex Hormozi
14 Jul 202634:15EN
The Let Them Theory by Mel Robbins & Sawyer Robbins
Deep Dive Reads: Self-Help Book Reviews & Literary Insights for Growths
26 Mar 202622:08EN
Ideasi & Pengelolaan SDM (part 1)
Creatalks
21 Jun 201951:13ID
EP78《贪婪的多巴胺》:如何像沉迷游戏一样沉迷学习?
纵横四海
15 Mar 20264:10:35ZH
Episode-105:日経新春杯の予想と春の展望について
LOVE競馬!!
16 Jan 202118:39JA
Video

Google is SO back...
Wes Roth
17 Sep 202614:31EN
We Finally Got a Robot on the Show | EP 161
Hard Fork
7 Nov 20251:09:50EN
Screensharing top takes in AI/startups
Greg Isenberg
9 Jul 20261:24:53EN
Ngaji Al Muhadzab Syirozy 1 Bagian 84
Miftahul Huda
4 Jun 202138:28ID
«Современный урок по ФГОС: требования, этапы, цифровые решения»
ЯКласс
11 Apr 20231:39:24RU