Deepseek did it again...
Check out MindsHub: https://tinyurl.com/5cra6js2 Join My Newsletter for Regular AI Updates 👇🏼 https://forwardfuture.com My Links 🔗 👉🏻 X: https://x.com/matthewberman 👉🏻 Forward Future X: https://x.com/forwardfuture 👉🏻 Instagram: https://www.instagram.com/matthewberman_ai 👉🏻 Discord: https://discord.gg/u7wTTGWhuJ 👉🏻 Spotify: https://open.spotify.com/show/6dBxDwxtHl1hpqHhfoXmy8 Media/Sponsorship Inquiries ✅ https://bit.ly/44TC45V Links: https://www.deepseek.com/en/news/deepseek-v4-1-flash/ https://x.com/The_Alex/status/2096294580006354985
Read Video · Bản Ghi & Phân Tích
Tập này có bản ghi đầy đủ + AI insights
Tài khoản miễn phí · không cần thẻ · 150 tín dụng khi đăng ký, đủ để mở khóa tập này
- 📄 Bản ghi đầy đủ có dấu thời gian
- ✨ Tóm tắt AI, từ khóa & bản đồ tư duy
- 💡 Điểm chính và trích dẫn nổi bật
Dòng thời gian tập
DeepSeek V4.1 Flash launches as a cheap, fast open-weights model that benchmarks near frontier models like Opus 5 and GPT 5.6 Soul.
- DeepSeek V4.1 Flash is incredibly cheap and fast, and benchmarks show it matching Opus 5 and GPT 5.6 Soul, continuing the pattern where open-weights models reach previous-generation frontier quality about six months behind the absolute frontier.
- The model is a 552 billion parameter mixture of experts with only 8 billion active parameters for input and 16 billion for output, making it relatively small compared to GPT 5.6's trillion parameters and Astra's estimated 7-10 trillion.
- On benchmarks it scores 30 on Terminal Bench 3.0 (only Opus 5 beats it) and 74.2 on Deep Suite, beating both Opus and GPT 5.6, though the host notes his own tests did not match the benchmark hype.
DeepSeek's efficiency gains slash memory requirements, cutting KV cache HBM needs to a quarter and SSD storage to an eighth.
- The KV cache only needs a fourth of the high bandwidth memory and an eighth of the SSD storage, which matters because HBM prices have been spiking from around $25 per gigabyte up to $10 per gigabyte as AI consumes all available memory.
- The memory footprint shrank eight times from DeepSeek V1 to V3.2, then 13 times smaller to V4 Flash, and another 4x smaller from V4 to V4.1, reflecting China's strength in making existing technology more efficient and faster rather than necessarily better.
- This efficiency is DeepSeek's answer to the global memory crunch, letting them serve many more tokens without paying for all that memory, and it shows up as extreme output speed reminiscent of the early Groq days.
Pricing is extremely low with peak and off-peak tiers, and the model is fully open weights so users can self-host.
- Off-peak pricing is 15 cents per million input tokens and 60 cents per million output tokens, doubling during peak hours to 30 cents input and $1.20 output, with cache hits costing a fraction of a penny.
- Because it is open source and open weights, users can download it, avoid giving DeepSeek their data, use any NeoCloud, or potentially run it locally once quantized if they have enough VRAM.
- The host contrasts this with frontier models like OpenAI and Anthropic charging around $50 per million output tokens, an orders-of-magnitude difference, though the vast majority of the economy only needs cheap models capable enough for 95% of use cases.
Khái niệm chính
- DeepSeek V4.1 Flash— The newly released open-weights model that is the main subject of the episode.
- Mixture of Experts— Architecture that activates only a small subset of weights per query, enabling extreme efficiency.
- Efficiency gains— The core theme: DeepSeek's ability to cut memory and compute requirements dramatically.
Trích dẫn nổi bật
And there's one big butt here. I've put it through a few tests and it didn't perform as well as I would have hoped, especially compared to the benchmarks, but I'm going to show you that later.
💡— Reveals a gap between benchmark scores and real-world performance, challenging the assumption that high benchmarks mean practical capability.
So of a 552 billion parameter model, only a tiny fraction of the weights are actually being used.
🤯— Highlights the extreme sparsity of mixture-of-experts models, showing how a huge model can run efficiently.
Hành động cụ thể
🤖AI Model Evaluation
Benchmarks can be misleading; real-world tests reveal practical weaknesses.
This week, pick a model you're considering and run a small custom test (e.g., a simple coding task) to verify its real capabilities.
Open-source models are catching up to frontier models within about six months.
Subscribe to a newsletter or set a reminder to check for new open-weight model releases every quarter.
💰Cost Optimization
Open models can be orders of magnitude cheaper than frontier models for most tasks.
Audit your current AI API usage and identify tasks that could be switched to a cheaper open model, then test one this week.
Off-peak pricing can significantly reduce costs for batch jobs.
Schedule non-urgent AI workloads to run during off-peak hours if your provider offers discounted rates.
Phiên âm và thông tin chi tiết được tạo bởi AI và có thể chứa lỗi. Độ chính xác phụ thuộc vào chất lượng âm thanh và sự rõ ràng của người nói — nếu có gì sai, âm thanh gốc luôn là nguồn chính xác nhất.
Tập podcast và video sẵn sàng để đọc
Tập podcast

The missing half of music history | Gabriella Di Laccio
TED Talks Daily
3 thg 9, 202610:29EN
(Preview) Microsoft’s Plan for Platform Survival, Meta and the Market’s Permission, A Lack of Situational Awareness
Sharp Tech with Ben Thompson
6 thg 8, 202628:47EN
The Multidisciplinary Approach to Thinking | Peter D. Kaufman [Outliers]
The Knowledge Project
13 thg 1, 202626:04EN
#12 「くぼちん電報ありがと〜!」
朝井リョウ・加藤千恵 信頼できない語り手
24 thg 4, 202652:32JA
#8 Die Völkerwanderung - Völker...? Wanderung...?
Wer wir sind und warum das nicht klappte ...
28 thg 5, 202539:45DE
Capitalismo
História em Meia Hora
28 thg 8, 202132:28PT
Video

Give Me 12 Minutes and I’ll Give You 30 Years of Productivity Advice
Daniel Pink
14 thg 9, 202511:59EN
GPUs, TPUs, & The Economics of AI Explained | Gavin Baker Interview
Invest Like The Best
9 thg 12, 20251:28:22EN
Neuroscientist: You Will NEVER Feel Stressed Again | Andrew Huberman
RESPIRE
27 thg 2, 202311:05EN
INDIA’S GOT LATENT S2 EP4 ft. Karan Aujla, Tanmay Bhat, Gurleen Pannu, Rahul Dua
Samay Raina
2 thg 8, 202653:54HI
#32 Die Schlacht von Worringen - Der Freiheitskampf der Kölner
99 mal Geschichte
12 thg 11, 202553:29DE