Readpodcast AI

Stanford CS329A Self-Improving AI Agents | Part 5 | Planning and Multi-Step Reasoning

3.8.20261:14:5617.120 AufrufeAuf YouTube ansehen

Want to dive deeper? This curriculum is covered in the following online courses: - Agentic AI professional education course: https://stanford.io/4zOyjPN - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video Summary: This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Azalia Mirhoseini on October 6, 2025, covers three papers on planning and multi-step reasoning in language model agents. LATS, or Language Agent Tree Search, combines reasoning, acting, and search using Monte Carlo Tree Search with LLM-judge and self-consistency scoring, tested on HotpotQA and WebShop. SPRINT fine-tunes reasoning models such as DeepSeek-R1 to generate independent plans for parallel execution, reducing sequential token count while improving accuracy on math, Countdown, and GPQA Diamond benchmarks. SWiRL generates offline synthetic multi-step tool-use trajectories scored by an LLM judge and trains models through multi-step reinforcement learning without executing tools during training, showing generalization across HotpotQA and GSM8K. The lecture addresses trade-offs including inference cost, irreversible actions, and the comparative effect of process-filtered versus outcome-filtered training data. Speaker Bio: Azalia Mirhoseini Assistant Professor of Computer Science, Stanford University Azalia Mirhoseini is a co-founder of Ricursive Intelligence, a frontier lab dedicated to recursive self-improvement through AI that designs the chips that fuel it. She is also an Assistant Professor of Computer Science at Stanford University where she directs Scaling Intelligence, a lab focused on developing scalable and self-improving AI systems and methodologies toward the goal of artificial general intelligence. Previously, she spent several years in industry AI labs, including Google Brain, Anthropic, and Google DeepMind, working on the development of Claude and Gemini. Her past work includes Mixture-of-Experts (MoE) neural architectures, now predominantly used in leading generative AI models; AlphaChip, a pioneering work on deep reinforcement learning for layout optimization used in the design of advanced chips like Google AI accelerators (TPUs) and data center CPUs; as well as pioneering research on LLM Test-Time Scaling. Her work has been recognized through the Okawa Research Grant, the Google ML and Systems Junior Faculty Award, MIT Technology Review's 35 Under 35 Award, the Best ECE Thesis Award at Rice University, publications in flagship venues such as Nature, and coverage by various media outlets, including WSJ, NYT, Forbes, MIT Technology Review, IEEE Spectrum, WIRED, and TechCrunch.

Read Video · Transkript & Insights

Transkript & KI-Insights generieren — Kostenlos testen

Kostenloses Konto · keine Karte nötig · 150 Credits nach Registrierung, genug für diese Episode

  • 📄 Vollständiges Transkript mit Zeitstempeln
  • ✨ KI-Zusammenfassung, Keywords & Mindmap
  • 💡 Kernaussagen & Zitate

Episoden und Videos zum Lesen