
The evaluation of open models like Gemma 4, DeepSeek V4, and GLM-5.1 revealed significant performance gaps compared to closed models, with open models lagging in benchmarks and real-world tasks. The assessment highlighted the lack of standardized benchmarks, such as those used in Claude Code or OpenCode, which exacerbates the disparity in capabilities. Open models, while showing promise in certain areas, still face challenges in performance and scalability, emphasizing the need for more comprehensive and fair benchmarking frameworks.
Reflections on AI's writing ability and how AI models get more capable.

After a few long years of finding time to document my lessons from training open models, my post-training book is done!

Musings on model alignment, what determines safety, and where we go from here.
