GPT-Red: Unlocking Self-Improvement for Robustness

Openai··Submitted by Mads Kristian Nylund
AI SafetyAI

GPT‑Red is an automated red-teaming model that enhances AI robustness by identifying and mitigating vulnerabilities through self-play reinforcement learning. It generates prompt injections to train subsequent models, making them more resistant to attacks, and is used to adversarially train GPT‑5.6, resulting in a more robust model. The model demonstrates strong effectiveness against various models, including both internal and production versions, and is evaluated for its general-purpose red-teaming capabilities. It is designed to scale with model capabilities, ensuring continuous improvement in safety measures.

Read Article

More from Openai

Related Articles