Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents
Want to dive deeper? This curriculum is covered in the following online courses: - Agentic AI professional education course: https://stanford.io/4zOyjPN - XCS329 graduate course: https://online.stanford.edu/courses/cs329a-self-improving-ai-agents A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents A similar curriculum is covered in XCS329z https://online.stanford.edu/courses/cs329z-engineering-ai-agents Follow along with the course schedule and syllabus: https://cs329a.stanford.edu/ View the course playlist: https://www.youtube.com/playlist?list=PLangBM27OtEA Video summary: This lecture video from Stanford's CS329A, Self-Improving AI Agents, taught by Aakanksha Chowdhery on October 17, 2025, examines self-improvement through search. AlphaCode pretrains a masked language model on GitHub and CodeContests data and generates large numbers of samples before clustering and selecting a final answer, while AlphaCode2 fine-tunes Gemini Pro with a learned scoring model, reaching an 85th percentile ranking on competitive programming contests. The lecture explains how solve rate scales with sample budget and where selection and clustering become bottlenecks even after large-scale sample generation. It then introduces Search-O1, a method that triggers search queries when a reasoning model expresses uncertainty and reasons over retrieved documents, which outperforms standard and agentic retrieval-augmented generation on GPQA and multi-hop question-answering benchmarks including HotpotQA and Bamboogle. The session closes by comparing Search-O1's prompting-based approach to Search-R1's reinforcement-learning-based approach for teaching models when to search. Speaker Bio: Aakanksha Chowdhery Adjunct Professor of Computer Science, Stanford University Dr. Aakanksha Chowdhery is pushing the frontier of agentic LLMs, focusing on recursive self-improvement and long-horizon agents that learn and deploy in the real world. She is one of the few researchers globally who has led frontier model training end-to-end, across both dense and mixture-of-experts (MoE) architectures. At Google, she led the 540B PaLM model, the largest densely trained language model in the world at the time. She subsequently drove pre-training and scaling of Gemini's MoE models across multiple generations, and contributed key components to PaLM-E, Med-PaLM, and the Pathways infrastructure underpinning Google's large-model efforts. She went on to build and lead pretraining teams for open intelligence efforts at Reflection and Meta. Earlier, she held research roles at Microsoft Research and Princeton. At Stanford, where she earned her PhD, she teaches CS329A (Self-Improving AI Agents) and serves as Program Chair for MLSys 2026.
Read Video · Transkript & Insights
Transkript & KI-Insights generieren — Kostenlos testen
Kostenloses Konto · keine Karte nötig · 150 Credits nach Registrierung, genug für diese Episode
- 📄 Vollständiges Transkript mit Zeitstempeln
- ✨ KI-Zusammenfassung, Keywords & Mindmap
- 💡 Kernaussagen & Zitate
Episoden und Videos zum Lesen
Podcast-Episoden

Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
26. Mai 202657:25EN
Scaling a $300K Moving Company in 60 Minutes
The Game with Alex Hormozi
14. Juli 202634:15EN
The Let Them Theory by Mel Robbins & Sawyer Robbins
Deep Dive Reads: Self-Help Book Reviews & Literary Insights for Growths
26. März 202622:08EN
#67 Wie stark war August der Starke?
Wer wir sind und warum das nicht klappte ...
22. Juli 20261:24:57DE
EP78《贪婪的多巴胺》:如何像沉迷游戏一样沉迷学习?
纵横四海
15. März 20264:10:35ZH
Episode-105:日経新春杯の予想と春の展望について
LOVE競馬!!
16. Jan. 202118:39JA
Videos

Marketing Engineer: The $1M Job with AI Agents
Greg Isenberg
31. Aug. 202635:19EN
A Cheeky Pint with Anthropic CEO Dario Amodei
Stripe
6. Aug. 20251:02:48EN
How Confidence Affects Your Health with Ian Robertson, PhD
CenterforBrainHealth
13. Dez. 20211:04:57EN
#36 Ludwig der Bayer - der dem Papst trotzt
99 mal Geschichte
11. Dez. 202559:57DE
从「上瘾模型」到「专注力训练」,如何在被算法理解的世界里重新找回主动?| 英文访谈 S9E33
声动活泼
16. Okt. 202550:48ZH-Hans