Stanford CS329A Self-Improving AI Agents | Part 3 | Robust Verification
Want to dive deeper? This curriculum is covered in the following online courses: - Agentic AI professional education course: https://stanford.io/4zOyjPN - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This lecture recording from Stanford's CS329A, Self-Improving AI Agents, taught by Azalia Mirhoseini on September 29, 2025, traces the evolution of verification methods for large language model outputs across four research papers. It covers OpenAI's "Training Verifiers to Solve Math Word Problems," which introduced the GSM8K dataset and outcome-based reward models, and "Let's Verify Step by Step," which compares outcome-supervised and process-supervised reward models using the PRM800K dataset of human-labeled reasoning steps. The lecture also covers Math-Shepherd, which automates step-level annotation without human labels, and Weaver, a Stanford paper that combines ensembles of weak verifiers, including reward models and LLM judges, to close the generation-verification gap. Topics include majority voting and self-consistency baselines, credit assignment in process versus outcome supervision, reward hacking, and using trained verifiers as reward signals for reinforcement learning fine-tuning. Speaker Bio: Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.
Read Video · Transkript & Insights
Transkript & KI-Insights generieren — Kostenlos testen
Kostenloses Konto · keine Karte nötig · 150 Credits nach Registrierung, genug für diese Episode
- 📄 Vollständiges Transkript mit Zeitstempeln
- ✨ KI-Zusammenfassung, Keywords & Mindmap
- 💡 Kernaussagen & Zitate
Episoden und Videos zum Lesen
Podcast-Episoden

The missing half of music history | Gabriella Di Laccio
TED Talks Daily
3. Sept. 202610:29EN
(Preview) Microsoft’s Plan for Platform Survival, Meta and the Market’s Permission, A Lack of Situational Awareness
Sharp Tech with Ben Thompson
6. Aug. 202628:47EN
The Multidisciplinary Approach to Thinking | Peter D. Kaufman [Outliers]
The Knowledge Project
13. Jan. 202626:04EN
#72 Die Deutschen und die Französische Revolution
Wer wir sind und warum das nicht klappte ...
26. Aug. 202658:49DE
Capitalismo
História em Meia Hora
28. Aug. 202132:28PT
EP09.《远见》职业生涯45年,你该如何规划?
纵横四海
16. Feb. 20231:54:34ZH
Videos

Give Me 12 Minutes and I’ll Give You 30 Years of Productivity Advice
Daniel Pink
14. Sept. 202511:59EN
GPUs, TPUs, & The Economics of AI Explained | Gavin Baker Interview
Invest Like The Best
9. Dez. 20251:28:22EN
Neuroscientist: You Will NEVER Feel Stressed Again | Andrew Huberman
RESPIRE
27. Feb. 202311:05EN
#32 Die Schlacht von Worringen - Der Freiheitskampf der Kölner
99 mal Geschichte
12. Nov. 202553:29DE
COMO SER GENTIL ESTÁ ACABANDO COM A SUA AUTORIDADE? | Fabiana Bertotti #152
Como Você Fez Isso?
30. Juli 20261:19:47PT