Lessons From RL Systems That Looked Fine Until They Didn't | Aethon
[2026 - DAY 3 - MODEL SYSTEMS] Reinforcement learning systems often fail not because rewards are wrong, but because optimization pressure is unbounded. Policies exploit edge cases, drift over time, and converge to brittle strategies that look fine in training but break in deployment, especially under bounded actions, safety requirements, resource budgets, and long-term user impact. This talk focuses on controlling optimization directly: practical techniques for training RL agents that remain stable and predictable under hard constraints. Rather than modifying rewards, we explore structural and system-level approaches that shape behavior by construction. Topics include: - Why reward penalties alone fail to enforce hard constraints under scale and distribution shift - Structural constraint mechanisms such as action masking, feasibility filters, and sandboxed execution - How training inside hard boundaries changes policy behavior and improves long-horizon stability, including across retraining cycles - Detecting constraint violations and failure modes that do not appear in aggregate return metrics - Lessons from applying constrained RL in production-like systems, including failures only discovered after deployment and what ultimately stopped them The goal is to share concrete algorithmic and system design strategies for deploying reinforcement learning in settings where violations are suboptimal. SPEAKER: Ezi Ozoani - Co-founder & CTO, Aethon 👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Read Video · 文字稿與深度分析
一鍵獲取文字稿與 AI 深度分析 — 免費體驗
免費註冊 · 無需信用卡 · 註冊即獲 150 積分,足夠解鎖本集
- 📄 完整文字稿含時間戳
- ✨ AI 摘要、關鍵詞與心智圖
- 💡 核心要點與精彩引言
Podcast 與影片,已可閱讀
音訊 Podcast

The 60 seconds that make or break a conversation | Chris Fenning
TED Talks Daily
2026年9月1日13:13ENPulling Back the Curtain on Sportsbooks & VIP Programs w/ Dillon Borgida | Ep 63
The Risk Takers Podcast
2024年3月14日1:38:10EN
Stock Market EMERGENCY: Sell Your Stocks Now, The Collapse Is Weeks Away!
The Diary Of A CEO with Steven Bartlett
2026年6月25日1:45:21EN
EP678 | 🎮
Gooaye 股癌
2026年7月11日52:46ZH-Hant
#26 Die Inquisition und das Ideal der Armut
Wer wir sind und warum das nicht klappte ...
2025年10月1日1:02:30DE
Rencontre #5 – Revenir à soi : le courage de s’écouter quand on a passé sa vie à s’adapter
À fleur de soi – le podcast des âmes sensibles & créatives
2026年4月3日1:05:24FR
影片

China Open-Source, Compute Arms Race, Reordering Global Trade | BG2 w/ Bill Gurley and Brad Gerstner
Bg2 Pod
2025年7月31日1:04:21EN
Marketing Engineer: The $1M Job with AI Agents
Greg Isenberg
2026年8月31日35:19EN
A Cheeky Pint with Anthropic CEO Dario Amodei
Stripe
2025年8月6日1:02:48EN
2026/08/24(一) 輝達伺服器傳漲價15%:AI成本暴增,成本誰吸收?
財女珍妮
2026年8月24日30:06ZH-Hant
#39 Die Pest
99 mal Geschichte
2026年1月8日1:02:15DE