Back to timeline

Fri, July 1702:48ResearchModel releasesAI safetyModel releases guide

OpenAI Releases GPT-Red Automated Red-Teaming Model, Beats Human Red Team 84% vs 13%

Decision Brief

What changedOpenAI trained an internal automated red-teaming model, GPT-Red, which defeated human red team 84% to 13% in prompt injection tests.
Why it mattersGPT-Red uses self-play reinforcement learning to automatically discover attacks, reducing GPT-5.6 Sol's prompt injection failure rate by 6x, greatly improving automated red-teaming for security engineers.
Who should careAll AI builders
Affected stackOpenAI
Source confidenceMedium · Reliable media or first-hand reporting

OpenAI trained GPT-Red, an internal automated red-teaming attack model, using self-play reinforcement learning against a set of defensive LLMs. In a replicated indirect prompt injection arena, GPT-Red defeated human red team members 84% to 13% and discovered a new attack category "false chain-of-thought." On OpenAI's hardest direct prompt injection benchmark, GPT-Red reduced GPT-5.6 Sol's failure rate by 6x. However, OpenAI acknowledged the model still struggles with multi-turn and image-based attacks.

Summary basis: official / RSS sourceCompiled from the source scope noted above; the original remains authoritative.

Sources

  • MarkTechPost

    Fast research-paper and ML tooling summaries, useful for infra and agent updates.

  • MarkTechPost

Related intel

留言

登入后即可留言,和其他 builder 交换实测心得。

还没有留言,抢头香。