Did Elon catch up? (Grok 4.7 is here)
Check out Zapier! https://bit.ly/3TLc4tQ Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://x.ai/news/grok-4-7 https://x.com/GavinSBaker/status/2091542026072338623 https://artificialanalysis.ai/
동영상 읽기 · 대본 및 인사이트
이 에피소드에는 전체 대본과 AI 인사이트가 있습니다
무료 계정 · 카드 불필요 · 가입 시 크레딧 150개, 이 에피소드를 잠금 해제하기에 충분합니다
- 📄 타임스탬프가 포함된 전체 대본
- ✨ AI 요약, 키워드와 마인드맵
- 💡 주요 결론과 인용문
에피소드 타임라인
Elon Musk's prediction that Grok 4.7 would match Opus 5.0 and the model's release at the same price and speed as Grok 4.6
- Elon tweeted about a week before release that Grok 4.7 should be roughly on par with Opus 5.0, not 5.1, setting a specific benchmark expectation.
- Grok 4.7 launched as a notable improvement over Grok 4.6 at the same price and speed, with xAI claiming it works longer on difficult tasks and checks its own work more carefully.
- The blog post describes Grok 4.7 as xAI's most capable model for coding and knowledge work with their best calibrated safeguards to date.
Cursorbench 4.0 results show Grok 4.7 is competitive with Opus 5 but not beating it, at roughly half the cost
- On Cursorbench 4.0, Grok 4.7 sits just behind Opus 5 on max thinking settings but costs about half as much to run the benchmark.
- There is a major performance delta between low thinking effort at 33% and extra-high thinking effort at 46.3%, one of the biggest effort curves outside of GPT 5.6 Soul.
- The host notes xAI now owns Cursor, so the benchmark should be viewed with that ownership relationship in mind.
Token efficiency and steps-per-task comparisons show Fable 5.1 leading on quality while Grok 4.7 wins on cost-effectiveness
- On average output tokens per task, Grok 4.7 is very comparable to Opus 5, and at lowest thinking effort it uses few tokens but scores poorly.
- Fable 5.1 is the quality winner but is multiple times more expensive than Grok 4.7, so its cost per completed task is much higher despite similar token usage.
- GPT 5.6 Soul appears most efficient on steps taken per task, which matches the host's personal experience of it having the most direct shot to task completion.
핵심 개념
- Grok 4.7— The newly released xAI model that is the central subject of the episode.
- cost per task completed— The key metric the host argues matters most when evaluating models for real workloads.
- CursorBench— The coding benchmark used to compare Grok 4.7 against Opus 5 and other frontier models.
주목할 인용문
Grok 4.7 should be roughly on par with Opus 50, not 5.1.
🔥— Elon's own prediction sets a high bar that the host then tests, revealing whether the claim holds up.
It is not beating it. It's just behind it on the max thinking setting for both models but it is much less expensive about half the cost to run that benchmark.
🤯— Shows that near-frontier performance at half the cost is the real value proposition, not outright superiority.
실행 가능한 주요 결론
📊AI Model Evaluation
Cost per task completed is often more important than raw benchmark scores for real-world use.
This week, calculate the cost per task for your most common AI workflow using your current model and compare it to Grok 4.7's pricing.
Benchmark tables can be cherry-picked; always look for missing models or metrics.
Next time you see a model comparison, check if all relevant competitors are included and look for independent evaluations like Artificial Analysis.
💼Business Strategy
Open-weight models are gaining enterprise traction due to cost, control, and privacy.
Evaluate whether an open-weight model like Muse Spark 1.3 could replace a proprietary model in your pipeline this quarter.
Near-frontier models offer a sweet spot for industries that don't need the absolute best answer.
Identify tasks in your organization where a 5-10% performance drop is acceptable in exchange for 50% cost savings.
대본과 인사이트는 AI가 생성하므로 오류가 있을 수 있습니다. 정확도는 오디오 품질과 화자의 명료도에 따라 달라지며, 이상한 부분이 있다면 원본 오디오가 항상 기준입니다.
바로 읽을 수 있는 에피소드와 동영상
팟캐스트 에피소드

He Couldn't Walk Away || How Jetha Devapura Built Sri Lanka's Biggest Crisis Line
The Giving Habit
2026년 7월 15일55:53EN
The shape-shifting sounds of the accordion | Maria Telesheva
TED Talks Daily
2026년 8월 27일13:22EN
Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
2026년 5월 26일57:25EN
#74 Deutschland: Die Idee einer Nation
Wer wir sind und warum das nicht klappte ...
2026년 9월 8일1:06:37DE
Café com Deus Pai | 21 de maio
Café Com Deus Pai | Podcast oficial
2026년 5월 21일4:53PT
商业小样39 | 汽车轮胎涨价两轮,为何无人知晓?
商业就是这样
2026년 5월 10일9:26ZH
동영상

Gemini Robotics – AI for the Physical World, with Keerthana and Ted of Google DeepMind
Cognitive Revolution "How AI Changes Everything"
2025년 5월 17일1:48:11EN
Jev: ChatGPT Co-Creator’s Answer to RLHF
AI Council
2026년 6월 19일36:39EN
Staphylococcus: Aureus, Epidermidis, Saprophyticus
Ninja Nerd
2021년 10월 21일1:01:18EN
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
2026년 8월 4일54:10SK
2026/08/24(一) 輝達伺服器傳漲價15%:AI成本暴增,成本誰吸收?
財女珍妮
2026년 8월 24일30:06ZH-Hant