Qwen3.8-27B & How to Serve it Fast
In this video, I look at the long awaited Qwen3.8-27B model. Both what it can do and how to serve it at the maximum tokens per second Thanks to Dell for Sponsoring the Compute #DellProPrecision #DellProMax #DellTech #NVIDIA 📖 Website: https://qwen.ai/ 🤗 HF: https://huggingface.co/collections/Qwen/qwen38 SGLang: https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B Twitter: https://x.com/Sam_Witteveen 🕵️ Interested in building LLM Agents? Fill out the form below Building LLM Agents Form: https://drp.li/dIMes 👨💻Github: https://github.com/samwit/llm-tutorials ⏱️Time Stamps: 00:00 Intro 00:50 ThinkingCap 01:25 Qwen3.8 - 27B 01:59 Different Versions on Hugging Face 02:12 Benchmarks 03:14 Artificial Analysis Benchmark 04:10 Qwen3.8-27B on Hugging Face 06:40 Demo 14:17 SGLang
Read Video · 文字稿与深度分析
本集已有完整文字稿 + AI 深度分析
免费注册 · 无需信用卡 · 注册即获 150 积分,足够解锁本集
- 📄 完整文字稿含时间戳
- ✨ AI 摘要、关键词与思维导图
- 💡 核心要点与精彩引言
节目时间轴
Qwen releases the 27B model after the massive 2.4T-parameter Qwen 3.8 Max, and its benchmarks surpass Meta's Muse Glimmer and the older Qwen 3.6 27B.
- Qwen 3.8 Max has 2.4 trillion parameters, making local deployment impractical for most users, so the 27B variant released on Friday is the model everyone was actually waiting for.
- Qwen's own benchmarks show the 3.8 27B substantially outperforming Meta's Muse Glimmer and the Qwen 3.6 27B across nearly every category, including vision and computer-use tasks where it even beats Opus 4.6 Max.
- Artificial Analysis scored the model 52 on its intelligence index, just behind GLM 5.2 at 53 and DeepSeek V4 Pro, and far ahead of any other open model runnable without serious hardware investment.
- On the Agentic Index the model beats GLM 5.2 and some GPT 5.6 variants, which is remarkable for something that can run locally at decent token speeds.
The HuggingFace ecosystem already offers many versions of Qwen 3.8 27B, from full BF16 down to 4-bit quants, fine-tunes, abliterated variants, and MLX builds for Mac.
- Qwen publishes both a BF16 full-resolution version and an FP8 version, while Unsloth has already released 4-bit NVFP4 quantizations that not all GPUs can actually run.
- Multiple fine-tunes and abliterated (uncensored) versions exist, with the Black Frost AI team among the first to release one, and the MLX community has shipped BF16, 8-bit, and 4-bit builds for Mac users without AMD or Nvidia GPUs.
- Testing was done on a Dell T2 Pro Max with an RTX Pro 6000 GPU providing 96GB of VRAM, enough to load even the full 16-bit BF16 model without difficulty.
Reasoning level dramatically changes token consumption and output quality, with X-high thinking burning tens of thousands of tokens for marginal gains.
- In the HTML website test, no-thinking mode produced a solid site with zero reasoning tokens, low thinking used about 512 tokens, and medium thinking unexpectedly used fewer tokens than low on this particular task.
- X-high thinking consumed 17,500 tokens and never finished within the 32k output limit, with runs reaching as high as 22,000 thinking tokens, meaning users need a very large context window and patience.
- In the pelican SVG test, X-high used 11,000 thinking tokens, medium dropped to under 1,000 with a still-good result, and no-thinking produced a noticeably ugly pelican, making medium the sweet spot.
- A red dragon variant confirmed the pattern: no-thinking got the bicycle right but the dragon looked childish, medium regressed slightly on the dragon, and X-high used 35,000 total tokens (21,000 for thinking) for a result that was better but still not great.
关键概念
- Qwen 3.8 27B— The newly released 27-billion-parameter open-weight model that is the main subject of the episode.
- local AI inference— The core theme of running large models on your own hardware rather than via cloud APIs.
- reasoning token budget— How many thinking tokens the model uses at different reasoning levels, which dramatically changes speed and quality.
精选金句
while I've limited it to 32k max tokens out, I've actually run out because I've spent 17,500 tokens on thinking.
🤯— Reveals that the highest reasoning level can consume more than half the output budget on thinking alone, making it impractical for many tasks.
I've had the thinking tokens be as high as 22,000 thinking tokens.
🤯— Shows the extreme token cost of X high reasoning, which most users would not expect from a 27B model.
可执行的洞察
🧠Model Selection
Qwen 3.8 27B substantially outperforms its predecessor and Meta's Muse Glimmer on most benchmarks, making it the new default for local AI.
Download the FP8 version from HuggingFace this week and run your standard coding or agent task to compare against your current model.
Different quantizations (FP8, NVFP4, MLX) offer different trade-offs between speed, quality, and hardware compatibility.
Benchmark at least two quantizations on your own hardware using a fixed prompt and record tokens per second and output quality.
⚙️Serving Configuration
Speculative decoding with MTP set to 3 provides a significant speed boost with minimal quality loss.
Enable MTP=3 in your vLLM or SGLang config and measure the tokens-per-second improvement on your hardware.
SGLang with NVFP4 weights and Docker offloading achieved the fastest speeds, up to 200+ tokens per second on an RTX Pro 6000.
If you have a Blackwell GPU, try the SGLang cookbook Docker image with their NVFP4 weights this week.
转录文字与 AI 洞察均由模型自动生成,可能存在少量误差。识别效果与音频质量、语速和发音清晰度相关——如有内容看起来不对,以原始音频为准。
播客与视频,已可阅读
音频播客

As Trump Purges Immigration Judges, One Speaks Out
The Daily
2026年6月23日35:41EN
How to make learning as addictive as social media | Luis von Ahn
TED Talks Daily
2026年9月7日14:23EN
How to Eliminate Self-Doubt Forever & Build Unshakeable Confidence
The Mel Robbins Podcast
2026年5月11日1:19:02EN
345-如何让你小孩愿意跟你讲话?
独树不成林
2026年6月11日27:44ZH-Hans
Rotten to the Core: Apple attackiert OpenAI-Kultur | ohne Belege: 200 Ökonomen warnen vor KI-Jobverlust | Nadella warnt vor KI-Daten-Risiko #579
Doppelgänger
2026年7月14日1:13:47DE
163.- “Mujeres, dinero y el miedo a incomodar” con Maca Riva
LA MAGIA DEL CAOS con Aislinn Derbez
2026年5月26日1:27:54ES
视频

Anthropic's CEO: ‘We Don’t Know if the Models Are Conscious’ | Interesting Times with Ross Douthat
Interesting Times
2026年2月12日1:02:30EN
Staphylococcus aureus
Osmosis from Elsevier
2020年10月14日14:32EN
Give me 11 Minutes and I'll Make you Dangerously Persuasive
Daniel Pink
2025年12月23日10:39EN
从「上瘾模型」到「专注力训练」,如何在被算法理解的世界里重新找回主动?| 英文访谈 S9E33
声动活泼
2025年10月16日50:48ZH-Hans
10 habitudes qui m’ont VRAIMENT fait perdre du poids
leawellnesss
2026年4月15日14:34FR