Readpodcast AI

Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents

3.8.20261:12:2710.382 AufrufeAuf YouTube ansehen

Want to dive deeper? This curriculum is covered in the following online courses: - Agentic AI professional education course: https://stanford.io/4zOyjPN - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video summary: This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on October 17, 2025, examines self-improvement through search. AlphaCode pretrains a masked language model on GitHub and CodeContests data and generates large numbers of samples before clustering and selecting a final answer, while AlphaCode2 fine-tunes Gemini Pro with a learned scoring model, reaching an 85th percentile ranking on competitive programming contests. The lecture explains how solve rate scales with sample budget and where selection and clustering become bottlenecks even after large-scale sample generation. It then introduces Search-O1, a method that triggers search queries when a reasoning model expresses uncertainty and reasons over retrieved documents, which outperforms standard and agentic retrieval-augmented generation on GPQA and multi-hop question-answering benchmarks including HotpotQA and Bamboogle. The session closes by comparing Search-O1's prompting-based approach to Search-R1's reinforcement-learning-based approach for teaching models when to search. Speaker Bio: Aakanksha Chowdhery Adjunct Professor of Computer Science, Stanford University Dr. Aakanksha Chowdhery is pushing the frontier of agentic LLMs, focusing on recursive self-improvement and long-horizon agents that learn and deploy in the real world. She is one of the few researchers globally who has led frontier model training end-to-end, across both dense and mixture-of-experts (MoE) architectures. At Google, she led the 540B PaLM model, the largest densely trained language model in the world at the time. She subsequently drove pre-training and scaling of Gemini's MoE models across multiple generations, and contributed key components to PaLM-E, Med-PaLM, and the Pathways infrastructure underpinning Google's large-model efforts. She went on to build and lead pretraining teams for open intelligence efforts at Reflection and Meta. Earlier, she held research roles at Microsoft Research and Princeton. At Stanford, where she earned her PhD, she teaches CS329A (Self-Improving AI Agents) and serves as Program Chair for MLSys 2026.

Read Video · Transkript & Insights

Transkript & KI-Insights generieren — Kostenlos testen

Kostenloses Konto · keine Karte nötig · 150 Credits nach Registrierung, genug für diese Episode

  • 📄 Vollständiges Transkript mit Zeitstempeln
  • ✨ KI-Zusammenfassung, Keywords & Mindmap
  • 💡 Kernaussagen & Zitate

Episoden und Videos zum Lesen