Build Trustworthy LLM Apps Powered by Agentic Evals | Meta
[2026 - DAY 3 - LIGHTNING TALK] Agentic AI has certainly secured mindspace of industry. As more & more people continue to play with LLMs, they begin developing an intuition aka vibes around whether LLM is performing the task at hand well or not. While usecases like math problem, classification are easier to verify; generative content creation etc are not. Adding agents to the mix makes things lot more complex. This talk will highlight the particular challenges one may face as they continue their journey towards productionizing their Agentic application: outputs are non deterministic, ground truth is hard to find, and what used to work before no longer does. Yet """"LGTM"""" isn't a deployment strategy. While typical LLM chatbot can be seen as a more sophisticated Google search, we're beginning to expect more from Agents: Don't just give an answer but actually perform the task. Security, bias, privacy etc suddenly become non-negotiable. Only after we handle these complexities, we can unblock real usecases in healthcare, finance, legal etc. This talk tackles why agent evaluation is fundamentally harder than traditional ML testing: multi-step reasoning chains, tool use side effects & more. How to build evaluation datasets that actually reflect production scenarios, not just cherry-picked examples. We'll cover automated evaluation pipelines using LLM-as-judge patterns, and when you can not avoid human in the loop. The session addresses detecting regressions before users do: setting up continuous evaluation that catches model degradation. Tricky cases when agent aces public evals but fails in production, and how to build evaluations that predict real-world performance. SPEAKER: Rushabh Mehta - Tech Lead, Meta 👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf
Leggi video · Trascrizione e analisi
Ottieni trascrizione e analisi AI per questo episodio — Inizia gratis
Account gratuito · nessuna carta necessaria · 150 crediti all'iscrizione, sufficienti per sbloccare questo episodio
- 📄 Trascrizione completa con timestamp
- ✨ Riepilogo AI, parole chiave e mappa mentale
- 💡 Conclusioni e citazioni chiave
Episodi e video pronti da leggere
Episodi podcast

What a harness is and how to build one with Claude Agent SDK
How I AI
8 lug 202624:35EN
Miele’s Andreas Wieser and Yulia Kalner at Kantar on reinventing German Brands Through Emotion
DMEXCO Podcast by Verena Gründel
29 lug 202631:55EN
Why AI is an even bigger deal than you think | Reed Hastings
TED Talks Daily
20 lug 202619:08EN
EP81《深度关系》:如何“给情绪价值”?包教包会
纵横四海
8 mag 20264:29:11ZH
8月28日(金)NVIDIA最高益 AI需要創出へ全方位、科学論文 安保上「懸念ある」組織との共著倍増
ながら日経
27 ago 20269:47JA
Café com Deus Pai | 25 de maio
Café Com Deus Pai | Podcast oficial
25 mag 20264:30PT
Video

The Bubble Most Will Get Wrong | Aswath Damodaran on How He Is Investing in a World of AI
Excess Returns
16 gen 20261:02:11EN
Three Lab Warnings in Five Days, Researcher Flags “Gambling with Our Lives,” and Labs Race
Peter H. Diamandis
11 set 20262:43:14EN
Why AI is going vertical (again) | Dianne Penn (Anthropic)
Lenny's Podcast
26 lug 20261:33:51EN
#35 Das Attentat von Anagni
99 mal Geschichte
5 dic 202541:08DE
«Современный урок по ФГОС: требования, этапы, цифровые решения»
ЯКласс
11 apr 20231:39:24RU