Moonshot AI - Kimi K3 tested
Kimi API (Extra 15% bonus for new users' first purchase, until Nov 15) https://platform.kimi.ai?track_id=track-f23ab1e6f2dd45ff84369c9ce034108c In this video I am going to be looking at the latest open-weight model from Moonshot AI - Kimi K3, an affordable frontier alternative to ChatGPT and Claude. I run through a few tests which are: 1. Performance 2. Memory 3. Agency 4. OpenAI Human Eval 5. Kanban 6. Sand Physics 7. Dungeon Crawler 8. Blender 9. Godot If you're interested in local LLMs, AI and homelabs - feel free to subscribe! Model: https://www.kimi.ai/blog/kimi-k3 My configs: https://github.com/lukesdevlab/youtube HumanEval: https://github.com/openai/human-eval Patreon: https://www.patreon.com/cw/LukesDevLab #localllm #localai #homelab #llamacpp #homelab #openai #moonshotai #kimi #k3 Chapters: 0:00 Coming Up 0:20 Intro 0:29 Model 1:39 Tests Overview 2:11 Performance 3:33 Memory 4:43 Reasoning 5:15 OpenAI HumanEval 5:55 Kanban 7:22 Sand Physics 8:20 Dungeon Crawler 9:24 Blender 10:54 Godot Platforms 13:43 Godot Dungeon Crawler 15:18 Conclusion
Read Video · Transcript & Insight
This episode has a full transcript + AI insights
Free account · no card needed · 150 credits on signup, enough to unlock this episode
- 📄 Full transcript with timestamps
- ✨ AI summary, keywords & mind map
- 💡 Key takeaways & quotes
Episode Timeline
Introduction to Kimi K3 by Moonshot AI and its pricing and benchmark setup
- Kimi K3 is a frontier-level open-weight model from Moonshot AI that rivals OpenAI and Anthropic offerings, though most users will access it via API rather than running it locally.
- API pricing is 30 cents per million tokens for cached input, $3 per million for uncached input, and $15 per million for output, making it significantly cheaper than Claude.
- The video is sponsored by Moonshot AI, but the host emphasizes the model will be tested with the exact same benchmark suite as always, with no special treatment.
- The test suite includes performance metrics, a memory test, reasoning benchmarks, OpenAI HumanEval with 164 Python challenges, a Kanban front-end test, sand physics, dungeon crawler, and MCP challenges with Blender and Godot.
Performance metrics: prefill and decode speeds on the Kimi API
- Prefill speed starts around 160 uncached and 240 cached tokens per second and reaches nearly 10,000 cached and 11,800 non-cached tokens per second at large prefill, which is unusual since uncached normally lags behind.
- The host suspects the benchmark tool is not handling the quantization information correctly, so the cached and uncached numbers may actually reflect two similar runs rather than a true distinction.
- Decode speed hovers around 43 tokens per second, which is impressive for a 3-trillion-parameter model with 1 million context.
Memory test at 256K context and reasoning benchmark results
- The memory test inserts data at 0, 25, 50, 75, or 100% context depth and asks the model to retrieve it three times per depth, totaling 15 runs, and the model found the data without issue on every run.
- The host could not run the memory test at the full 1 million context because the API kept timing out, so 256K was used to stay comparable with other models.
- On the reasoning benchmark the model scored 42 out of 48 for an 88% pass rate, but surprisingly failed twice on medium difficulty while only failing once on hard and three times on expert.
- OpenAI HumanEval results showed an 86% pass rate overall, a 94% pass rate on answered questions, and a 91% answer rate, indicating the model does not overthink within its 8,000-token limit.
Key Concepts
- Kimi K3— The open-weight frontier model from Moonshot AI that is the central subject of the entire review.
- Moonshot AI— The Chinese company behind Kimi K3 that also sponsored the video and provided API credits.
- open-weight model— The key differentiator that lets users run Kimi K3 on their own hardware, unlike closed models from OpenAI and Anthropic.
Notable Quotes
So remember this is a 3 trillion parameter model with 1 million context. So the speed is pretty great considering.
🤯— It is genuinely surprising that a model of this enormous scale can maintain 43 tokens per second, revealing how far inference optimization has come.
So it's got a 42 out of 48, and that gives us an 88% pass rate. So, a very respectable pass rate overall. But yeah, you can see we're not getting 12 out of 12 on medium. It's fallen down on two reasoning on medium, but then only one on hard and only three on expert.
💡— It overturns the assumption that models fail on harder questions first, showing Kimi K3 paradoxically struggles more on medium difficulty than hard or expert.
Actionable Takeaways
🧪AI Model Evaluation
Frontier models like Kimi K3 may find standard benchmark suites too easy, requiring harder tests to differentiate capabilities.
This week, design a custom benchmark task that pushes a frontier model beyond its comfort zone, such as a complex 3D game with multiple interacting systems.
Benchmark contamination is a real concern, as HumanEval may be in training data, so results should be interpreted with caution.
Create or find a private, unpublished coding challenge to test models on, ensuring the task is not in any public dataset.
💰Cost Optimization
Using max reasoning mode is often overkill and significantly increases cost without proportional benefit for many tasks.
Experiment with lowering the reasoning setting to medium or high on your next API call and measure the quality difference versus cost savings.
Kimi K3's API pricing is substantially cheaper than alternatives like Claude, making it a strong contender for budget-conscious developers.
Calculate the cost of running your typical workload on Kimi K3 versus your current model and consider switching if the savings are significant.
Transcript and insights are AI-generated and may contain errors. Accuracy depends on audio quality and speaker clarity — if something looks off, the original audio is always the source of truth.
Episodes & Videos Ready to Read
Podcast episodes

He Couldn't Walk Away || How Jetha Devapura Built Sri Lanka's Biggest Crisis Line
The Giving Habit
Jul 15, 202655:53EN
The shape-shifting sounds of the accordion | Maria Telesheva
TED Talks Daily
Aug 27, 202613:22EN
Bad Maps and Good Intentions; Sophie Radice on the trials and tribulations of life beyond the comfort zone S5 E11
How to have Extraordinary Relationships
May 26, 202657:25EN
商业小样43 | AI时代,谁在给服务器“降温”
商业就是这样
Jun 21, 202612:16ZH
#71 Die Deutschen im Amerikanischen Unabhängigkeitskrieg
Wer wir sind und warum das nicht klappte ...
Aug 19, 202644:55DE
SÉRIE: UMA VEZ SALVO, SALVO PARA SEMPRE - A CERTEZA DA SALVAÇÃO: PARTE 2| PR.PEDRO ESTRELLA
Minha Igreja Na Cidade
Aug 31, 202659:49PT
Videos

Anthropic CEO Dario Amodei on AI's Moat, Risk, and SB 1047
Econ 102 with Noah Smith
Aug 29, 20241:00:00EN
Anthropic CEO Dario Amodei: AI's Potential, OpenAI Rivalry, GenAI Business, Doomerism
Alex Kantrowitz
Jul 30, 20251:08:37EN
GPT-6 Astra: How I’d Make Money With It
Greg Isenberg
Sep 10, 202622:52EN
#40 Karl IV. und sein goldenes Prag
99 mal Geschichte
Jan 8, 20261:03:51DE
Firma bez šéfov: Funguje to? - Money Talk 113 s Ferom Baníkom
Milan Dubec
Aug 4, 202654:10SK