Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

Technologyreview··Submitted by Mads Kristian Nylund
AI SecurityAI DevelopmentAI Evaluation

OpenAI developed GPT-Red, an LLM super-hacker designed to enhance the security of its models by simulating cyberattacks. The latest version of GPT-5.6 was trained against GPT-Red, resulting in the most secure release yet. GPT-Red automates red-teaming by identifying vulnerabilities through prompt injection attacks, but it struggles with back-and-forth conversations and image-based attacks. OpenAI claims GPT-Red is more robust than any copycat model but is not perfect and misses some attack types. The company will continue to refine its security processes despite not releasing GPT-Red.

Read Article

More from Technologyreview

Related Articles