Lessons From RL Systems That Looked Fine Until They Didn't | Aethon
[2026 - DAY 3 - MODEL SYSTEMS] Reinforcement learning systems often fail not because rewards are wrong, but because optimization pressure is unbounded. Policies exploit edge cases, drift over time, and converge to brittle strategies that look fine in training but break in deployment, especially under bounded actions, safety requirements, resource budgets, and long-term user impact. This talk focuses on controlling optimization directly: practical techniques for training RL agents that remain stable and predictable under hard constraints. Rather than modifying rewards, we explore structural and system-level approaches that shape behavior by construction. Topics include: - Why reward penalties alone fail to enforce hard constraints under scale and distribution shift - Structural constraint mechanisms such as action masking, feasibility filters, and sandboxed execution - How training inside hard boundaries changes policy behavior and improves long-horizon stability, including across retraining cycles - Detecting constraint violations and failure modes that do not appear in aggregate return metrics - Lessons from applying constrained RL in production-like systems, including failures only discovered after deployment and what ultimately stopped them The goal is to share concrete algorithmic and system design strategies for deploying reinforcement learning in settings where violations are suboptimal. SPEAKER: Ezi Ozoani - Co-founder & CTO, Aethon 👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Leggi video · Trascrizione e analisi
Ottieni trascrizione e analisi AI per questo episodio — Inizia gratis
Account gratuito · nessuna carta necessaria · 150 crediti all'iscrizione, sufficienti per sbloccare questo episodio
- 📄 Trascrizione completa con timestamp
- ✨ Riepilogo AI, parole chiave e mappa mentale
- 💡 Conclusioni e citazioni chiave
Episodi e video pronti da leggere
Episodi podcast

Why Jensen Huang Believes We’ve Reached AGI and Inside OpenAI’s German Website Hijack | #287
Moonshots with Peter Diamandis
9 set 20262:23:43EN
Episode 68: How Timothy Baxter Built Baxter Research Into the Gold Standard of Criminal Research
Behind the Screens: Conversations with Background Screening Pros hosted by Les Rosen
20 lug 202659:28EN
A guerrilla gardener in South Central LA | Ron Finley
TED Talks Daily
15 ago 202613:08EN
我們不是被 AI 取代,是被自己的懶惰取代《未來1000天》|槓桿說書EP2
思維槓桿
10 lug 202621:36ZH
Cold Calling Mastery: Vom Opener zum Termin
B2B Sales on Air
30 ott 202538:43DE
#17 「準備の鬼、パワポの悪魔」
朝井リョウ・加藤千恵 信頼できない語り手
29 mag 202652:25JA
Video

How to Build Things with Jev & OpenJevs
Sam Witteveen
21 set 202619:19EN
Tom Holland and Jon Bernthal Team Up While Eating Spicy Wings | Hot Ones
First We Feast
23 lug 202629:21EN
Ryan Greenblatt – What happens once AI can automate AI research?
Dwarkesh Patel
11 ago 20262:12:32EN
#32 Die Schlacht von Worringen - Der Freiheitskampf der Kölner
99 mal Geschichte
12 nov 202553:29DE
10 habitudes qui m’ont VRAIMENT fait perdre du poids
leawellnesss
15 apr 202614:34FR