Readpodcast AI

No Dropped Frames: designing a VLM around a latency budget | Moondream

21/06/202622:40176 visualizzazioniGuarda su YouTube

[2026 - Day 1 - INFERENCE SYSTEMS] Moondream is a vision language model that runs in real time on video streams. This talk covers the model-side work behind it. I'll start with architecture: upcycling from dense to MoE, and the tradeoffs when you're optimizing for latency rather than just parameter count. Then tokenization: why we built a custom SuperBPE tokenizer and what it bought us. The goal throughout was to avoid modeling decisions that would hurt us at inference time. I'll also cover training infrastructure. We wrote custom training engines and RL systems because existing open source projects were pushing us toward design decisions that didn't fit. I'll talk about where we diverged and what we got out of it. Finally, inference. Real-time VLM isn't just a serving problem or a modeling problem. We built a custom inference engine alongside the model, and I'll cover how the two informed each other. SPEAKER: Vik Korrapati - CTO, Moondream 👉 Sign up for our "No BS" Newsletter to get the latest technical data & AI content: https://aicouncil.com/newsletter ABOUT AI COUNCIL: AI Council brings together the brightest minds in data to share industry knowledge, technical architectures and best practices in building cutting edge data & AI systems and tools. FIND US: Website: https://aicouncil.com/ LinkedIn: https://www.linkedin.com/company/aicouncilconf/ X: https://x.com/aicouncilconf

Leggi video · Trascrizione e analisi

Ottieni trascrizione e analisi AI per questo episodio — Inizia gratis

Account gratuito · nessuna carta necessaria · 150 crediti all'iscrizione, sufficienti per sbloccare questo episodio

  • 📄 Trascrizione completa con timestamp
  • ✨ Riepilogo AI, parole chiave e mappa mentale
  • 💡 Conclusioni e citazioni chiave

Episodi e video pronti da leggere