Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning
Want to dive deeper? This curriculum is covered in the following online courses: - Agentic AI professional education course: https://stanford.io/4zOyjPN - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Azalia Mirhoseini on October 6, 2025, covers three papers on planning and multi-step reasoning in language model agents. LATS, or Language Agent Tree Search, combines reasoning, acting, and search using Monte Carlo Tree Search with LLM-judge and self-consistency scoring, tested on HotpotQA and WebShop. SPRINT fine-tunes reasoning models such as DeepSeek-R1 to generate independent plans for parallel execution, reducing sequential token count while improving accuracy on math, Countdown, and GPQA Diamond benchmarks. SWiRL generates offline synthetic multi-step tool-use trajectories scored by an LLM judge and trains models through multi-step reinforcement learning without executing tools during training, showing generalization across HotpotQA and GSM8K. The lecture addresses trade-offs including inference cost, irreversible actions, and the comparative effect of process-filtered versus outcome-filtered training data. Speaker Bio: Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.
Read Video · Transkript & Insights
Transkript & KI-Insights generieren — Kostenlos testen
Kostenloses Konto · keine Karte nötig · 150 Credits nach Registrierung, genug für diese Episode
- 📄 Vollständiges Transkript mit Zeitstempeln
- ✨ KI-Zusammenfassung, Keywords & Mindmap
- 💡 Kernaussagen & Zitate
Episoden und Videos zum Lesen
Podcast-Episoden

(Preview) Doom Debates Go Mainstream, AI Religion and the Economic Future, Several Vectors of the China Question
Sharp Tech with Ben Thompson
18. Sept. 202633:19EN
How video games can level up the way you learn | Kris Alexander
TED Talks Daily
7. Sept. 202614:25EN
Why a New Class of AI “Judgment Models” Could Have Big Business Implications
The AI Daily Brief: Artificial Intelligence News and Analysis
16. Sept. 202625:19EN
Folge 2 - Fahnenappell und Mathe am Samstag: Schule in der DDR
Zeitreise DDR
9. Nov. 202518:43DE
FRANCK LOPVET : "Ça me gonfle le bien-être !"
Entre vous et moi - Dominique Lagrou Sempère
8. Juni 20261:50:01FR
商业小样50 | 都在讨论家务机器人,不如关心到底有什么家务
商业就是这样
20. Sept. 202614:49ZH-Hans
Videos

Screensharing top takes in AI/startups
Greg Isenberg
9. Juli 20261:24:53EN
The Inference Frontier: from 100 to 10,000 tokens per second — Sean Lie, Cerebras CTO
Latent Space
2. Sept. 202644:03EN
The rise of taste, human authenticity and judgment in an AI world | Adam Mosseri (Head of IG)
Lenny's Podcast
9. Juli 20261:08:29EN
#29 Der Sachsenspiegel
99 mal Geschichte
22. Okt. 202550:26DE
«Современный урок по ФГОС: требования, этапы, цифровые решения»
ЯКласс
11. Apr. 20231:39:24RU