Moonshot AI - Kimi K3 tested
Kimi API (Extra 15% bonus for new users' first purchase, until Nov 15) https://platform.kimi.ai?track_id=track-f23ab1e6f2dd45ff84369c9ce034108c In this video I am going to be looking at the latest open-weight model from Moonshot AI - Kimi K3, an affordable frontier alternative to ChatGPT and Claude. I run through a few tests which are: 1. Performance 2. Memory 3. Agency 4. OpenAI Human Eval 5. Kanban 6. Sand Physics 7. Dungeon Crawler 8. Blender 9. Godot If you're interested in local LLMs, AI and homelabs - feel free to subscribe! Model: https://www.kimi.ai/blog/kimi-k3 My configs: https://github.com/lukesdevlab/youtube HumanEval: https://github.com/openai/human-eval Patreon: https://www.patreon.com/cw/LukesDevLab #localllm #localai #homelab #llamacpp #homelab #openai #moonshotai #kimi #k3 Chapters: 0:00 Coming Up 0:20 Intro 0:29 Model 1:39 Tests Overview 2:11 Performance 3:33 Memory 4:43 Reasoning 5:15 OpenAI HumanEval 5:55 Kanban 7:22 Sand Physics 8:20 Dungeon Crawler 9:24 Blender 10:54 Godot Platforms 13:43 Godot Dungeon Crawler 15:18 Conclusion
Read Video · Transkript & İçgörüler
Bu bölümün tam transkripti + AI içgörüleri var
Ücretsiz hesap · kart gerekmez · kayıtta 150 kredi, bu bölümü açmaya yeterli
- 📄 Zaman damgalı tam transkript
- ✨ AI özeti, anahtar kelimeler ve zihin haritası
- 💡 Ana çıkarımlar ve alıntılar
Bölüm zaman çizelgesi
Introduction to Kimi K3 by Moonshot AI and its pricing and benchmark setup
- Kimi K3 is a frontier-level open-weight model from Moonshot AI that rivals OpenAI and Anthropic offerings, though most users will access it via API rather than running it locally.
- API pricing is 30 cents per million tokens for cached input, $3 per million for uncached input, and $15 per million for output, making it significantly cheaper than Claude.
- The video is sponsored by Moonshot AI, but the host emphasizes the model will be tested with the exact same benchmark suite as always, with no special treatment.
- The test suite includes performance metrics, a memory test, reasoning benchmarks, OpenAI HumanEval with 164 Python challenges, a Kanban front-end test, sand physics, dungeon crawler, and MCP challenges with Blender and Godot.
Performance metrics: prefill and decode speeds on the Kimi API
- Prefill speed starts around 160 uncached and 240 cached tokens per second and reaches nearly 10,000 cached and 11,800 non-cached tokens per second at large prefill, which is unusual since uncached normally lags behind.
- The host suspects the benchmark tool is not handling the quantization information correctly, so the cached and uncached numbers may actually reflect two similar runs rather than a true distinction.
- Decode speed hovers around 43 tokens per second, which is impressive for a 3-trillion-parameter model with 1 million context.
Memory test at 256K context and reasoning benchmark results
- The memory test inserts data at 0, 25, 50, 75, or 100% context depth and asks the model to retrieve it three times per depth, totaling 15 runs, and the model found the data without issue on every run.
- The host could not run the memory test at the full 1 million context because the API kept timing out, so 256K was used to stay comparable with other models.
- On the reasoning benchmark the model scored 42 out of 48 for an 88% pass rate, but surprisingly failed twice on medium difficulty while only failing once on hard and three times on expert.
- OpenAI HumanEval results showed an 86% pass rate overall, a 94% pass rate on answered questions, and a 91% answer rate, indicating the model does not overthink within its 8,000-token limit.
Temel kavramlar
- Kimi K3— The open-weight frontier model from Moonshot AI that is the central subject of the entire review.
- Moonshot AI— The Chinese company behind Kimi K3 that also sponsored the video and provided API credits.
- open-weight model— The key differentiator that lets users run Kimi K3 on their own hardware, unlike closed models from OpenAI and Anthropic.
Önemli alıntılar
So remember this is a 3 trillion parameter model with 1 million context. So the speed is pretty great considering.
🤯— It is genuinely surprising that a model of this enormous scale can maintain 43 tokens per second, revealing how far inference optimization has come.
So it's got a 42 out of 48, and that gives us an 88% pass rate. So, a very respectable pass rate overall. But yeah, you can see we're not getting 12 out of 12 on medium. It's fallen down on two reasoning on medium, but then only one on hard and only three on expert.
💡— It overturns the assumption that models fail on harder questions first, showing Kimi K3 paradoxically struggles more on medium difficulty than hard or expert.
Uygulanabilir çıkarımlar
🧪AI Model Evaluation
Frontier models like Kimi K3 may find standard benchmark suites too easy, requiring harder tests to differentiate capabilities.
This week, design a custom benchmark task that pushes a frontier model beyond its comfort zone, such as a complex 3D game with multiple interacting systems.
Benchmark contamination is a real concern, as HumanEval may be in training data, so results should be interpreted with caution.
Create or find a private, unpublished coding challenge to test models on, ensuring the task is not in any public dataset.
💰Cost Optimization
Using max reasoning mode is often overkill and significantly increases cost without proportional benefit for many tasks.
Experiment with lowering the reasoning setting to medium or high on your next API call and measure the quality difference versus cost savings.
Kimi K3's API pricing is substantially cheaper than alternatives like Claude, making it a strong contender for budget-conscious developers.
Calculate the cost of running your typical workload on Kimi K3 versus your current model and consider switching if the savings are significant.
Transkript ve içgörüler yapay zeka tarafından oluşturulur ve hatalar içerebilir. Doğruluk, ses kalitesine ve konuşmacıların netliliğine bağlıdır — bir şey yanlış görünse orijinal ses her zaman doğru kaynaktır.
Bölümler ve videolar okunmaya hazır
Podcast bölümleri

He Couldn't Walk Away || How Jetha Devapura Built Sri Lanka's Biggest Crisis Line
The Giving Habit
15 Tem 202655:53EN
The shape-shifting sounds of the accordion | Maria Telesheva
TED Talks Daily
27 Ağu 202613:22EN
Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
26 May 202657:25EN
Kadınlar ve Erkekler Neden Arkadaş Olamaz?
Kendine İyi Davran
22 Eyl 202617:55TR
商业小样43 | AI时代,谁在给服务器“降温”
商业就是这样
21 Haz 202612:16ZH
#71 Die Deutschen im Amerikanischen Unabhängigkeitskrieg
Wer wir sind und warum das nicht klappte ...
19 Ağu 202644:55DE
Videolar

Building a Software Factory that actually works (Full Course)
Greg Isenberg
14 Eyl 202631:30EN
Loop Engineering from First Principles — Kyle Mistele, HumanLayer
AI Engineer
25 Tem 202617:40EN
Why the AI’s honeymoon is ending (and tech workers are feeling it) | Noam Segal
Lenny's Podcast
12 Tem 20261:36:29EN
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
4 Ağu 202654:10SK
2026/08/24(一) 輝達伺服器傳漲價15%:AI成本暴增,成本誰吸收?
財女珍妮
24 Ağu 202630:06ZH-Hant