The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
From pushing inference beyond 4,000 tokens per second to working with OpenAI on a new generation of ultra-fast AI infrastructure, Cerebras is betting that speed doesn’t just make models faster, it makes entirely new kinds of AI possible. In this episode, Cerebras co-founder and CTO Sean Lie joins swyx and Vibhu fresh off Hot Chips to unpack CS4, preview CS5, and explain why 100–200 tokens per second may soon feel like “batch mode.” We go deep on Cerebras’ wafer-scale architecture, its partnership with OpenAI, and the emerging hardware stack for frontier inference. Sean breaks down OpenAI’s Jalapeño chip, why AI-first chip design could transform the semiconductor industry, what NVIDIA, Groq, AMD, Etched, and other challengers are getting right and wrong, and why the future of AI infrastructure will increasingly be heterogeneous and disaggregated. We also discuss model-hardware co-design, the path toward 10,000-token-per-second inference, new approaches to memory and 3D packaging, and the growing strategic competition between the US and China across both models and semiconductors. We discuss: • Why Cerebras believes 100–200 tokens per second is becoming the new “batch mode” • CS4 and how Cerebras is pushing inference beyond 4,000 tokens per second • CS5 and the path toward 10,000 TPS on medium-sized models • Running frontier models at up to 5,000 tokens per second • Why Cerebras is effectively sold out of its current capacity • How OpenAI is using ultra-fast inference internally for incident response and critical research • Why faster inference can enable more reasoning loops and more capable agents • OpenAI’s Jalapeño chip and why Sean sees AI-first chip design as the future • How Cerebras and OpenAI could combine CS5 and Jalapeño into a new inference stack • Sean’s critique of Groq and the limitations of SRAM architectures for frontier-scale models • Why inference is breaking apart into specialized workloads like prefill, decode, attention, and expert routing • Why future data centers may increasingly be designed like one giant computer • How models designed around NVIDIA GPUs leave performance gains on the table for alternative architectures • Why hardware-model co-design could unlock massive additional gains • NVIDIA, AMD, TPU, and Trainium — and where traditional accelerator architectures are heading • Sean’s take on Etched and what next-generation AI hardware startups should actually innovate on • Why memory bandwidth, 3D integration, power, and cooling are becoming central bottlenecks • The rise of Chinese open models and China’s increasingly independent AI hardware ecosystem — Sean Lie • Cerebras: https://www.cerebras.ai/ • X: https://x.com/seanlie Timestamps 00:00:00 Hook 00:01:19 Introduction 00:03:40 CS4 and 4,400 Tokens Per Second 00:07:07 From “Impossible” Wafer Scale to OpenAI 00:10:13 CS5 and the Road to 10,000 TPS 00:12:05 Cerebras Is Sold Out — and OpenAI Wants the Capacity 00:14:53 Why OpenAI Is Sharing Ultra-Fast Inference 00:15:53 OpenAI Jalapeño and AI-Designed Chips 00:20:11 Performance, Power, and the Groq Debate 00:21:44 Groq’s 31B Benchmark and SRAM Limitations 00:24:44 The Future of Heterogeneous AI Inference 00:26:35 Breaking Inference Into Specialized Hardware 00:29:41 Model-Hardware Co-Design 00:32:00 NVIDIA, AMD, TPU, and Trainium 00:33:44 Sean’s Take on Etched 00:35:30 Why the Next Breakthrough Goes Beyond the Chip 00:38:31 Memory Bandwidth and 3D Integration 00:39:19 Power, Cooling, and the Hard Parts of Wafer Scale 00:40:42 China, Open Models, and the Semiconductor Race 00:43:04 Cerebras’ IPO and Closing Thoughts
동영상 읽기 · 대본 및 인사이트
이 에피소드에는 전체 대본과 AI 인사이트가 있습니다
무료 계정 · 카드 불필요 · 가입 시 크레딧 150개, 이 에피소드를 잠금 해제하기에 충분합니다
- 📄 타임스탬프가 포함된 전체 대본
- ✨ AI 요약, 키워드와 마인드맵
- 💡 주요 결론과 인용문
- 화자 1
- 화자 2
- 화자 1
- 화자 2
에피소드 타임라인
Sean Lie discusses the evolution of Hot Chips and the golden age of hardware innovation.
- Hot Chips has transformed from a niche gathering of computer architecture enthusiasts to a venue for revealing industry-changing hardware.
- Sean Lie believes the semiconductor industry is experiencing unprecedented innovation across chip design, interconnect, system design, software, and optics.
- The buzz at Hot Chips reflects a unique time where AI is driving rapid hardware development.
Cerebras CS4 launch details: new Nexus platform, 2x performance, and record inference speeds.
- CS4 is built on a new modular platform called Nexus, providing twice the power and interconnect bandwidth with half the latency compared to previous generation.
- The CS4 system demonstrated GPT-OSS running at over 4,400 tokens per second, which Sean Lie describes as 'mind-blowing'.
- This speed enables new applications and more capable agents by allowing more reasoning loops in real-time.
The shift from technology demonstration to solving real problems with wafer-scale integration.
- The first-generation chip was primarily a technology demonstration to prove wafer-scale integration was possible.
- Now Cerebras is focused on scaling capacity and running models that users care about, transitioning to solving real problems.
- The partnership with OpenAI exemplifies this shift, running frontier models at 14 times faster than GPU speeds.
핵심 개념
- wafer-scale integration— Core technology enabling Cerebras to aggregate massive SRAM for large models.
- ultra-fast inference— The key value proposition of Cerebras, enabling new applications and agentic loops.
- CS4— Cerebras' latest system with 2x performance improvement, running GPT-OSS at 4,400 TPS.
주목할 인용문
what used to be considered fast at like 1,000 uh 100 or or 200 tokens per second is quickly becoming the new batch mode.
🤯— Reveals how the definition of 'fast' is shifting dramatically, making previous speeds seem obsolete.
we're showing uh GPTOSS uh running at over 4,400 TPS which is just mind-blowing. It's like it it it it almost feels like it's fake, right?
🤯— The sheer speed of 4,400 tokens per second is so high it seems unreal, highlighting the leap in performance.
실행 가능한 주요 결론
💻Hardware Innovation
Wafer-scale integration enables massive SRAM aggregation, solving memory bandwidth bottlenecks for large models.
Research wafer-scale and chiplet architectures to understand their trade-offs for AI workloads.
Power and cooling are often the real challenges in chip design, not just logic integration.
When evaluating new hardware, consider its power and cooling requirements as primary constraints.
⚡AI Inference Trends
Ultra-fast inference (thousands of tokens per second) enables new real-time applications and more capable agents.
Prototype an agentic workflow that relies on sub-second response times to see the difference.
Prefill-decode disaggregation and heterogeneous architectures are emerging to optimize different phases of inference.
Study how disaggregated inference can improve cost and latency for your deployment.
대본과 인사이트는 AI가 생성하므로 오류가 있을 수 있습니다. 정확도는 오디오 품질과 화자의 명료도에 따라 달라지며, 이상한 부분이 있다면 원본 오디오가 항상 기준입니다.
바로 읽을 수 있는 에피소드와 동영상
팟캐스트 에피소드

PA Replay: GOING ALL IN....What You Need To Know
The Pure Athlete Podcast
2026년 7월 14일52:14EN
3 steps to turn everyday get-togethers into transformative gatherings | Priya Parker
TED Talks Daily
2026년 8월 15일12:38ENPulling Back the Curtain on Sportsbooks & VIP Programs w/ Dillon Borgida | Ep 63
The Risk Takers Podcast
2024년 3월 14일1:38:10EN
Año nuevo en septiembre parte 5
Ático Primera con Laia Castel
2026년 9월 2일42:03ES
Dr. Gerald Hüther - Warum sind wir alle so unruhig?
Hotel Matze
2026년 5월 13일2:51:38DE
Le mystère de la chambre secrète
Passages, le podcast d'histoires vraies de Louie Media
2026년 5월 20일41:51FR
동영상

Doomberg Sees Hundreds of Billions Flowing Into Venezuela’s Oil Sector
In it to Win it
2026년 9월 2일29:27EN
Doomberg: Energy, AI, and the Calls Nobody Else Made
AdvisorAnalyst
2026년 8월 11일1:25:04EN
The AI Tsunami is Here & Society Isn't Ready | Dario Amodei x Nikhil Kamath | People by WTF
Nikhil Kamath
2026년 2월 24일1:08:34EN
#27 Die Nibelungen - Wer war Siegfried?
99 mal Geschichte
2025년 10월 8일1:01:13DE
RESUMO: REFORMA PROTESTANTE (Luteranismo, Calvinismo, Anglicanismo e Contrarreforma) Débora Aladim
Débora Aladim
2023년 2월 13일48:51PT