From Spans to Trajectories: Observability for Long-Running Agents | HoneyHive
[2026 - DAY 2 - AI ENGINEERING] Agents have evolved. We've moved from orchestration frameworks — where agents operate through definite steps and turns you defined upfront — to harnesses, where the LLM uses skills and tools to chart its own trajectory. Modern agents run for hours or days, producing hundreds to thousands of steps in a single session. This calls for a fundamentally different methodology for monitoring and evaluating them in production. This talk shares what we've learned building observability infrastructure for agent harnesses at HoneyHive. We'll start with why the harness — not the model — has become the hardest engineering problem in production AI, and why traditional APM breaks down when traces are 10,000 spans deep and failures happen four tool calls deep. We'll walk through a live trajectory view to see what long-running agent traces actually look like at scale, and the specific challenges they create: context rot, semantic failure modes, and the needle-in-a-haystack problem of finding the moment that mattered. Then we'll dig into skills as the new unit of behavior and the dual role of clustering in agent development: unsupervised clustering for discovering emergent patterns and identifying where guardrails are needed, and supervised classifiers for production evaluation at scale. We'll close on what comes next — swarm observability for multi-agent systems. SPEAKER: Sunny Bakhda - Founding Engineer, HoneyHive 👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Read Video · 文字稿与深度分析
一键获取文字稿与 AI 深度分析 — 免费体验
免费注册 · 无需信用卡 · 注册即获 150 积分,足够解锁本集
- 📄 完整文字稿含时间戳
- ✨ AI 摘要、关键词与思维导图
- 💡 核心要点与精彩引言
播客与视频,已可阅读
音频播客

The 60 seconds that make or break a conversation | Chris Fenning
TED Talks Daily
2026年9月1日13:13ENPulling Back the Curtain on Sportsbooks & VIP Programs w/ Dillon Borgida | Ep 63
The Risk Takers Podcast
2024年3月14日1:38:10EN
Stock Market EMERGENCY: Sell Your Stocks Now, The Collapse Is Weeks Away!
The Diary Of A CEO with Steven Bartlett
2026年6月25日1:45:21EN
vol.37 “我只是做自己 却冒犯到了一些人”该怎么办
天真不天真
2025年9月30日1:07:03ZH-Hans
Rencontre #5 – Revenir à soi : le courage de s’écouter quand on a passé sa vie à s’adapter
À fleur de soi – le podcast des âmes sensibles & créatives
2026年4月3日1:05:24FR
#66 Die Türken vor Wien - Wer war der Türkenlouis?
Wer wir sind und warum das nicht klappte ...
2026年7月15日1:05:55DE
视频

Give Me 12 Minutes and I’ll Give You 30 Years of Productivity Advice
Daniel Pink
2025年9月14日11:59EN
GPUs, TPUs, & The Economics of AI Explained | Gavin Baker Interview
Invest Like The Best
2025年12月9日1:28:22EN
Neuroscientist: You Will NEVER Feel Stressed Again | Andrew Huberman
RESPIRE
2023年2月27日11:05EN
从「上瘾模型」到「专注力训练」,如何在被算法理解的世界里重新找回主动?| 英文访谈 S9E33
声动活泼
2025年10月16日50:48ZH-Hans
INDIA’S GOT LATENT S2 EP4 ft. Karan Aujla, Tanmay Bhat, Gurleen Pannu, Rahul Dua
Samay Raina
2026年8月2日53:54HI