Readpodcast AI

What's Next After RLHF? — Diogo Almeida, TypeSafe AI

可讀2026/7/3118:041,806 次觀看在 YouTube 觀看

RLHF made models that are extraordinary at pleasing the human in the loop, and Diogo Almeida, a GPT-4 co author, argues that is exactly the problem. Optimizing for human preference optimizes for engagement and for overpromising, the same pressure that makes a model confidently agree that a fart audio file is a symphony. That produces two camps: one where models act as assistants with a human catching mistakes, where RLHF shines, and one where they operate autonomously with real stakes, where the same instinct to please quietly becomes a liability. So what comes next is not the Claude Code era but a shift in what you optimize. Almeida frames it through Sutton's bitter lesson: the task matters more than the data, and reinforcement learning with verifiable rewards points the model at real automation instead of human approval. He is careful that pre trained models are already incredibly capable and that the trap is bolting preference optimization on top, which teaches confidence and drops modes. The through line is that assistance and automation pull in different directions in optimization space, and the field is only starting to say plainly which one it is building. Speaker info: - https://x.com/CompleteSkeptic - https://www.linkedin.com/in/diogomda/ - https://typesafe.ai/ Timestamps: 0:00 - Not the Claude Code era 1:40 - The state of the field 3:14 - Two camps: assistance and autonomy 4:31 - Why models please the human in the loop 6:37 - How RLHF actually works 7:31 - Preference versus what's true 8:10 - When the consequences get real 8:47 - So what's next 9:35 - Assistance is not automation 14:31 - Is pre-training the problem? 15:43 - RLVR and Sutton's bitter lesson

Read Video · 文字稿與深度分析

本集已有完整文字稿 + AI 深度分析

免費註冊 · 無需信用卡 · 註冊即獲 150 積分,足夠解鎖本集

  • 📄 完整文字稿含時間戳
  • ✨ AI 摘要、關鍵詞與心智圖
  • 💡 核心要點與精彩引言
  • 說話人 1

轉錄文字與 AI 洞察均由模型自動生成,可能存在少量誤差。辨識效果與音訊品質、語速及發音清晰度相關——若有內容看起來有誤,以原始音訊為準。

需要協助或想分享意見?support@readpodcast.ai

Podcast 與影片,已可閱讀