
Claude Fable 5 is an enhanced version of Anthropic's Mythos-class models, featuring improved safety measures to address AI risk concerns, including new classifiers that prevent misuse and route specific requests to a more capable model. The safety policies aim to limit model capabilities, despite the model's superior performance, and have sparked debate over their effectiveness and transparency. Critics argue that these measures may not fully address the risks of frontier AI development, while the company defends its approach as a way to maintain competitive advantage and ensure safer AI systems.
Reflections on AI's writing ability and how AI models get more capable.

After a few long years of finding time to document my lessons from training open models, my post-training book is done!

Musings on model alignment, what determines safety, and where we go from here.
