Readpodcast AI

Qwen3.8-27B & How to Serve it Fast

可读2026/8/1818:00228,370 次观看在 YouTube 观看

In this video, I look at the long awaited Qwen3.8-27B model. Both what it can do and how to serve it at the maximum tokens per second Thanks to Dell for Sponsoring the Compute #DellProPrecision #DellProMax #DellTech #NVIDIA 📖 Website: https://qwen.ai/ 🤗 HF: https://huggingface.co/collections/Qwen/qwen38 SGLang: https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B Twitter: https://x.com/Sam_Witteveen 🕵️ Interested in building LLM Agents? Fill out the form below Building LLM Agents Form: https://drp.li/dIMes 👨‍💻Github: https://github.com/samwit/llm-tutorials ⏱️Time Stamps: 00:00 Intro 00:50 ThinkingCap 01:25 Qwen3.8 - 27B 01:59 Different Versions on Hugging Face 02:12 Benchmarks 03:14 Artificial Analysis Benchmark 04:10 Qwen3.8-27B on Hugging Face 06:40 Demo 14:17 SGLang

Read Video · 文字稿与深度分析

本集已有完整文字稿 + AI 深度分析

免费注册 · 无需信用卡 · 注册即获 150 积分,足够解锁本集

  • 📄 完整文字稿含时间戳
  • ✨ AI 摘要、关键词与思维导图
  • 💡 核心要点与精彩引言

转录文字与 AI 洞察均由模型自动生成,可能存在少量误差。识别效果与音频质量、语速和发音清晰度相关——如有内容看起来不对,以原始音频为准。

需要帮助或想分享反馈?support@readpodcast.ai

播客与视频,已可阅读