Readpodcast AI

Welcome to Our Data Benchmark, Where Everything's Made Up and the Points Don't Matter

17/06/202633:58746 visualizzazioniGuarda su YouTube

[2026 - DAY 2 - ANALYTICS & DATA SCI] SPIDER, DSBench, and every other analytics benchmark treat data work like a pub quiz: here's a question, here's the answer, did you match? But real analytics is arguing about whether "revenue" means bookings or collections, discovering that Stripe amounts are in cents while your platform stores dollars, and figuring out why the numbers don't tie to last quarter's deck. Current benchmarks can't express any of that—they just check if you got 47.3. Worse: smarter models keep making the same mistakes. Opus is clearly more intelligent than Sonnet, but it falls into the same traps—path of least resistance, accepts the first answer, doesn't ask clarifying questions. We'll show specific examples where industry standard benchmarks fail (including our own) and share some ideas for evals that test what analysts actually do: learn a messy warehouse over time, not answer a frozen question on day zero. SPEAKER: Izzy Miller - AI Engineer & AI Research Lead, Hex 👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf

Leggi video · Trascrizione e analisi

Ottieni trascrizione e analisi AI per questo episodio — Inizia gratis

Account gratuito · nessuna carta necessaria · 150 crediti all'iscrizione, sufficienti per sbloccare questo episodio

  • 📄 Trascrizione completa con timestamp
  • ✨ Riepilogo AI, parole chiave e mappa mentale
  • 💡 Conclusioni e citazioni chiave

Episodi e video pronti da leggere