Readpodcast AI

Qwen3.8-27B & How to Serve it Fast

Sam Witteveen
Sẵn sàng18/8/202618:00230.556 lượt xemXem trên YouTube

In this video, I look at the long awaited Qwen3.8-27B model. Both what it can do and how to serve it at the maximum tokens per second Thanks to Dell for Sponsoring the Compute #DellProPrecision #DellProMax #DellTech #NVIDIA 📖 Website: https://qwen.ai/ 🤗 HF: https://huggingface.co/collections/Qwen/qwen38 SGLang: https://lmsysorg.mintlify.app/cookbook/autoregressive/Qwen/Qwen3.8-27B Twitter: https://x.com/Sam_Witteveen 🕵️ Interested in building LLM Agents? Fill out the form below Building LLM Agents Form: https://drp.li/dIMes 👨‍💻Github: https://github.com/samwit/llm-tutorials ⏱️Time Stamps: 00:00 Intro 00:50 ThinkingCap 01:25 Qwen3.8 - 27B 01:59 Different Versions on Hugging Face 02:12 Benchmarks 03:14 Artificial Analysis Benchmark 04:10 Qwen3.8-27B on Hugging Face 06:40 Demo 14:17 SGLang

Read Video · Bản Ghi & Phân Tích

Tập này có bản ghi đầy đủ + AI insights

Tài khoản miễn phí · không cần thẻ · 150 tín dụng khi đăng ký, đủ để mở khóa tập này

  • 📄 Bản ghi đầy đủ có dấu thời gian
  • ✨ Tóm tắt AI, từ khóa & bản đồ tư duy
  • 💡 Điểm chính và trích dẫn nổi bật

Phiên âm và thông tin chi tiết được tạo bởi AI và có thể chứa lỗi. Độ chính xác phụ thuộc vào chất lượng âm thanh và sự rõ ràng của người nói — nếu có gì sai, âm thanh gốc luôn là nguồn chính xác nhất.

Cần trợ giúp hoặc muốn góp ý?support@readpodcast.ai

Tập podcast và video sẵn sàng để đọc