
Kimi K3, a 2.8T parameter MoE model, is positioned as a leading open-weight AI model, surpassing previous open models in performance and capabilities. It ranks highly in competitive benchmarks and is part of a broader trend of Chinese AI labs adopting open-source models to enhance global diffusion and economic competitiveness. The model's efficiency is attributed to architectural innovations like Kimi Delta Attention, which enable better scaling and intelligence conversion despite limited compute power.
Reflections on AI's writing ability and how AI models get more capable.

After a few long years of finding time to document my lessons from training open models, my post-training book is done!

Musings on model alignment, what determines safety, and where we go from here.
