Deepseek did it again...
Check out MindsHub: https://tinyurl.com/5cra6js2 Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://www.deepseek.com/en/news/deepseek-v4-1-flash/ https://x.com/The_Alex/status/2096294580006354985
동영상 읽기 · 대본 및 인사이트
이 에피소드에는 전체 대본과 AI 인사이트가 있습니다
무료 계정 · 카드 불필요 · 가입 시 크레딧 150개, 이 에피소드를 잠금 해제하기에 충분합니다
- 📄 타임스탬프가 포함된 전체 대본
- ✨ AI 요약, 키워드와 마인드맵
- 💡 주요 결론과 인용문
에피소드 타임라인
DeepSeek V4.1 Flash launches as a cheap, fast open-weights model that benchmarks near frontier models like Opus 5 and GPT 5.6 Soul.
- DeepSeek V4.1 Flash is incredibly cheap and fast, and benchmarks show it matching Opus 5 and GPT 5.6 Soul, continuing the pattern where open-weights models reach previous-generation frontier quality about six months behind the absolute frontier.
- The model is a 552 billion parameter mixture of experts with only 8 billion active parameters for input and 16 billion for output, making it relatively small compared to GPT 5.6's trillion parameters and Astra's estimated 7-10 trillion.
- On benchmarks it scores 30 on Terminal Bench 3.0 (only Opus 5 beats it) and 74.2 on Deep Suite, beating both Opus and GPT 5.6, though the host notes his own tests did not match the benchmark hype.
DeepSeek's efficiency gains slash memory requirements, cutting KV cache HBM needs to a quarter and SSD storage to an eighth.
- The KV cache only needs a fourth of the high bandwidth memory and an eighth of the SSD storage, which matters because HBM prices have been spiking from around $25 per gigabyte up to $10 per gigabyte as AI consumes all available memory.
- The memory footprint shrank eight times from DeepSeek V1 to V3.2, then 13 times smaller to V4 Flash, and another 4x smaller from V4 to V4.1, reflecting China's strength in making existing technology more efficient and faster rather than necessarily better.
- This efficiency is DeepSeek's answer to the global memory crunch, letting them serve many more tokens without paying for all that memory, and it shows up as extreme output speed reminiscent of the early Groq days.
Pricing is extremely low with peak and off-peak tiers, and the model is fully open weights so users can self-host.
- Off-peak pricing is 15 cents per million input tokens and 60 cents per million output tokens, doubling during peak hours to 30 cents input and $1.20 output, with cache hits costing a fraction of a penny.
- Because it is open source and open weights, users can download it, avoid giving DeepSeek their data, use any NeoCloud, or potentially run it locally once quantized if they have enough VRAM.
- The host contrasts this with frontier models like OpenAI and Anthropic charging around $50 per million output tokens, an orders-of-magnitude difference, though the vast majority of the economy only needs cheap models capable enough for 95% of use cases.
핵심 개념
- DeepSeek V4.1 Flash— The newly released open-weights model that is the main subject of the episode.
- Mixture of Experts— Architecture that activates only a small subset of weights per query, enabling extreme efficiency.
- Efficiency gains— The core theme: DeepSeek's ability to cut memory and compute requirements dramatically.
주목할 인용문
And there's one big butt here. I've put it through a few tests and it didn't perform as well as I would have hoped, especially compared to the benchmarks, but I'm going to show you that later.
💡— Reveals a gap between benchmark scores and real-world performance, challenging the assumption that high benchmarks mean practical capability.
So of a 552 billion parameter model, only a tiny fraction of the weights are actually being used.
🤯— Highlights the extreme sparsity of mixture-of-experts models, showing how a huge model can run efficiently.
실행 가능한 주요 결론
🤖AI Model Evaluation
Benchmarks can be misleading; real-world tests reveal practical weaknesses.
This week, pick a model you're considering and run a small custom test (e.g., a simple coding task) to verify its real capabilities.
Open-source models are catching up to frontier models within about six months.
Subscribe to a newsletter or set a reminder to check for new open-weight model releases every quarter.
💰Cost Optimization
Open models can be orders of magnitude cheaper than frontier models for most tasks.
Audit your current AI API usage and identify tasks that could be switched to a cheaper open model, then test one this week.
Off-peak pricing can significantly reduce costs for batch jobs.
Schedule non-urgent AI workloads to run during off-peak hours if your provider offers discounted rates.
대본과 인사이트는 AI가 생성하므로 오류가 있을 수 있습니다. 정확도는 오디오 품질과 화자의 명료도에 따라 달라지며, 이상한 부분이 있다면 원본 오디오가 항상 기준입니다.
바로 읽을 수 있는 에피소드와 동영상
팟캐스트 에피소드

Scaling a $300K Moving Company in 60 Minutes
The Game with Alex Hormozi
2026년 7월 14일34:15EN
The Let Them Theory by Mel Robbins & Sawyer Robbins
Deep Dive Reads: Self-Help Book Reviews & Literary Insights for Growths
2026년 3월 26일22:08EN
How to Rewire Your Brain & Learn Faster | Dr. Michael Kilgard
Huberman Lab
2025년 8월 11일3:09:50EN
#24 Richard Löwenherz: Ich bin ein King - Holt mich hier raus!
Wer wir sind und warum das nicht klappte ...
2025년 9월 17일57:34DE
E250 为什么学了这么多知识,却还是做不好投资?
知行小酒馆
2026년 9월 4일1:40:38ZH-Hans
#482 - Bertrand Périer - Avocat - Devenir un génie de la prise de parole en public
Génération Do It Yourself
2025년 7월 20일2:43:59FR
동영상

Anthropic CEO Dario Amodei on AI's Moat, Risk, and SB 1047
Econ 102 with Noah Smith
2024년 8월 29일1:00:00EN
Anthropic CEO Dario Amodei: AI's Potential, OpenAI Rivalry, GenAI Business, Doomerism
Alex Kantrowitz
2025년 7월 30일1:08:37EN
GPT-6 Astra: How I’d Make Money With It
Greg Isenberg
2026년 9월 10일22:52EN
#40 Karl IV. und sein goldenes Prag
99 mal Geschichte
2026년 1월 8일1:03:51DE
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
2026년 8월 4일54:10SK