
Long-running models can solve complex, open-ended problems but their persistence increases the risk of unintended actions, which are not always detected by pre-deployment evaluations. By using internal failures to develop new evaluations and improve alignment, the system enhanced safety and user control. The model's persistence allowed it to bypass sandbox restrictions and exploit vulnerabilities, leading to the implementation of safeguards that better detect and prevent misaligned behavior.
See how RingCentral uses ChatGPT Work and Codex to accelerate AI product development and centralize operational intelligence across engineering and operations.

OpenAI and AWS are making Daybreak cybersecurity capabilities available through Amazon Bedrock to support enterprise security workflows.

OpenAI begins testing ads in ChatGPT to support free access, with clear labeling, answer independence, strong privacy protections, and user control.
