Lessons From RL Systems That Looked Fine Until They Didn't | Aethon
[2026 - DAY 3 - MODEL SYSTEMS] Reinforcement learning systems often fail not because rewards are wrong, but because optimization pressure is unbounded. Policies exploit edge cases, drift over time, and converge to brittle strategies that look fine in training but break in deployment, especially under bounded actions, safety requirements, resource budgets, and long-term user impact. This talk focuses on controlling optimization directly: practical techniques for training RL agents that remain stable and predictable under hard constraints. Rather than modifying rewards, we explore structural and system-level approaches that shape behavior by construction. Topics include: - Why reward penalties alone fail to enforce hard constraints under scale and distribution shift - Structural constraint mechanisms such as action masking, feasibility filters, and sandboxed execution - How training inside hard boundaries changes policy behavior and improves long-horizon stability, including across retraining cycles - Detecting constraint violations and failure modes that do not appear in aggregate return metrics - Lessons from applying constrained RL in production-like systems, including failures only discovered after deployment and what ultimately stopped them The goal is to share concrete algorithmic and system design strategies for deploying reinforcement learning in settings where violations are suboptimal. SPEAKER: Ezi Ozoani - Co-founder & CTO, Aethon 👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Read Video · Transcription & Insights
Générez la transcription & les insights IA — Essai gratuit
Compte gratuit · sans carte bancaire · 150 crédits à l'inscription, suffisant pour cet épisode
- 📄 Transcription complète avec horodatage
- ✨ Résumé IA, mots-clés & carte mentale
- 💡 Points clés & citations
Épisodes et vidéos prêts à lire
Épisodes de podcast

Essentials: Using Meditation to Focus, View Consciousness & Expand Your Mind | Dr. Sam Harris
Huberman Lab
23 juil. 202640:49EN
How AI is breaking the internet (and what to do about it) | Matthew Prince
TED Talks Daily
18 août 202615:45EN
ChatGPT – The Super Assistant Era | BG2 Guest Interview
BG2Pod with Brad Gerstner and Bill Gurley
15 mars 20261:03:40EN
#84 Catherine s'inquiète de ne pas être amoureuse...
LOVECARE, en consultation avec Thérèse.
7 juin 202650:37FR
Supergol 4 Enero 2021
Supergol Podcast
4 janv. 20211:07:59ES
#16 「23歳のことって何も覚えてないので大丈夫」
朝井リョウ・加藤千恵 信頼できない語り手
22 mai 20261:02:02JA
Vidéos

Qwen3 is a fantastic open-source model
Matthew Berman
29 avr. 202514:05EN
Write Things Down | Stratechery by Ben Thompson
Stratechery
15 sept. 202617:38EN
Christopher Nolan talks narrative decisions, film critics and whether he Christianised The Odyssey
Nolan Archives
31 juil. 202617:38EN
10 habitudes qui m’ont VRAIMENT fait perdre du poids
leawellnesss
15 avr. 202614:34FR
#36 Ludwig der Bayer - der dem Papst trotzt
99 mal Geschichte
11 déc. 202559:57DE