Moonshot AI - Kimi K3 tested
Kimi API (Extra 15% bonus for new users' first purchase, until Nov 15) https://platform.kimi.ai?track_id=track-f23ab1e6f2dd45ff84369c9ce034108c In this video I am going to be looking at the latest open-weight model from Moonshot AI - Kimi K3, an affordable frontier alternative to ChatGPT and Claude. I run through a few tests which are: 1. Performance 2. Memory 3. Agency 4. OpenAI Human Eval 5. Kanban 6. Sand Physics 7. Dungeon Crawler 8. Blender 9. Godot If you're interested in local LLMs, AI and homelabs - feel free to subscribe! Model: https://www.kimi.ai/blog/kimi-k3 My configs: https://github.com/lukesdevlab/youtube HumanEval: https://github.com/openai/human-eval Patreon: https://www.patreon.com/cw/LukesDevLab #localllm #localai #homelab #llamacpp #homelab #openai #moonshotai #kimi #k3 Chapters: 0:00 Coming Up 0:20 Intro 0:29 Model 1:39 Tests Overview 2:11 Performance 3:33 Memory 4:43 Reasoning 5:15 OpenAI HumanEval 5:55 Kanban 7:22 Sand Physics 8:20 Dungeon Crawler 9:24 Blender 10:54 Godot Platforms 13:43 Godot Dungeon Crawler 15:18 Conclusion
Read Video · Transkript & Insights
Diese Folge hat ein vollständiges Transkript + KI-Insights
Kostenloses Konto · keine Karte nötig · 150 Credits nach Registrierung, genug für diese Episode
- 📄 Vollständiges Transkript mit Zeitstempeln
- ✨ KI-Zusammenfassung, Keywords & Mindmap
- 💡 Kernaussagen & Zitate
Episoden-Zeitlinie
Introduction to Kimi K3 by Moonshot AI and its pricing and benchmark setup
- Kimi K3 is a frontier-level open-weight model from Moonshot AI that rivals OpenAI and Anthropic offerings, though most users will access it via API rather than running it locally.
- API pricing is 30 cents per million tokens for cached input, $3 per million for uncached input, and $15 per million for output, making it significantly cheaper than Claude.
- The video is sponsored by Moonshot AI, but the host emphasizes the model will be tested with the exact same benchmark suite as always, with no special treatment.
- The test suite includes performance metrics, a memory test, reasoning benchmarks, OpenAI HumanEval with 164 Python challenges, a Kanban front-end test, sand physics, dungeon crawler, and MCP challenges with Blender and Godot.
Performance metrics: prefill and decode speeds on the Kimi API
- Prefill speed starts around 160 uncached and 240 cached tokens per second and reaches nearly 10,000 cached and 11,800 non-cached tokens per second at large prefill, which is unusual since uncached normally lags behind.
- The host suspects the benchmark tool is not handling the quantization information correctly, so the cached and uncached numbers may actually reflect two similar runs rather than a true distinction.
- Decode speed hovers around 43 tokens per second, which is impressive for a 3-trillion-parameter model with 1 million context.
Memory test at 256K context and reasoning benchmark results
- The memory test inserts data at 0, 25, 50, 75, or 100% context depth and asks the model to retrieve it three times per depth, totaling 15 runs, and the model found the data without issue on every run.
- The host could not run the memory test at the full 1 million context because the API kept timing out, so 256K was used to stay comparable with other models.
- On the reasoning benchmark the model scored 42 out of 48 for an 88% pass rate, but surprisingly failed twice on medium difficulty while only failing once on hard and three times on expert.
- OpenAI HumanEval results showed an 86% pass rate overall, a 94% pass rate on answered questions, and a 91% answer rate, indicating the model does not overthink within its 8,000-token limit.
Schlüsselkonzepte
- Kimi K3— The open-weight frontier model from Moonshot AI that is the central subject of the entire review.
- Moonshot AI— The Chinese company behind Kimi K3 that also sponsored the video and provided API credits.
- open-weight model— The key differentiator that lets users run Kimi K3 on their own hardware, unlike closed models from OpenAI and Anthropic.
Bemerkenswerte Zitate
So remember this is a 3 trillion parameter model with 1 million context. So the speed is pretty great considering.
🤯— It is genuinely surprising that a model of this enormous scale can maintain 43 tokens per second, revealing how far inference optimization has come.
So it's got a 42 out of 48, and that gives us an 88% pass rate. So, a very respectable pass rate overall. But yeah, you can see we're not getting 12 out of 12 on medium. It's fallen down on two reasoning on medium, but then only one on hard and only three on expert.
💡— It overturns the assumption that models fail on harder questions first, showing Kimi K3 paradoxically struggles more on medium difficulty than hard or expert.
Konkrete Handlungen
🧪AI Model Evaluation
Frontier models like Kimi K3 may find standard benchmark suites too easy, requiring harder tests to differentiate capabilities.
This week, design a custom benchmark task that pushes a frontier model beyond its comfort zone, such as a complex 3D game with multiple interacting systems.
Benchmark contamination is a real concern, as HumanEval may be in training data, so results should be interpreted with caution.
Create or find a private, unpublished coding challenge to test models on, ensuring the task is not in any public dataset.
💰Cost Optimization
Using max reasoning mode is often overkill and significantly increases cost without proportional benefit for many tasks.
Experiment with lowering the reasoning setting to medium or high on your next API call and measure the quality difference versus cost savings.
Kimi K3's API pricing is substantially cheaper than alternatives like Claude, making it a strong contender for budget-conscious developers.
Calculate the cost of running your typical workload on Kimi K3 versus your current model and consider switching if the savings are significant.
Transkript und Insights werden KI-generiert und können Fehler enthalten. Die Genauigkeit hängt von der Audioqualität und der Deutlichkeit der Sprecher ab — bei Unklarheiten ist das Originalaudio die maßgebliche Quelle.
Episoden und Videos zum Lesen
Podcast-Episoden

Ep #82 Polar Explorers
Case by Case
9. Mai 202438:29EN
(Preview) Doom Debates Go Mainstream, AI Religion and the Economic Future, Several Vectors of the China Question
Sharp Tech with Ben Thompson
18. Sept. 202633:19EN
How video games can level up the way you learn | Kris Alexander
TED Talks Daily
7. Sept. 202614:25EN
#12 Die Wikinger kommen
Wer wir sind und warum das nicht klappte ...
25. Juni 202536:34DE
EP17.《分手心理学》:新伤口和旧伤痛的治愈指南
纵横四海
11. Apr. 20232:46:00ZH
#5「夢を叶えても独り」
朝井リョウ・加藤千恵 信頼できない語り手
6. März 20261:11:04JA
Videos

Why the AI’s honeymoon is ending (and tech workers are feeling it) | Noam Segal
Lenny's Podcast
12. Juli 20261:36:29EN
Leopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of history
Dwarkesh Patel
4. Juni 20244:32:07EN
If you want 2026 to be the best year of your life, please watch this video…
Daniel Pink
29. Dez. 202525:59EN
#32 Die Schlacht von Worringen - Der Freiheitskampf der Kölner
99 mal Geschichte
12. Nov. 202553:29DE
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
4. Aug. 202654:10SK