Monday, 17 August 2026 Login

Virtual Tech. Real Impact.

BREAKING
Remote Workflows

OpenAI unveils AI that hunts its own flaws

OpenAI unveils AI that hunts its own flaws - ai security
OpenAI unveils AI that hunts its own flaws

OpenAI developed an internal artificial intelligence system named GPT-Red to identify security flaws in its own models before release.

The tool automates red teaming, a process that tests software for vulnerabilities. While human security teams traditionally handle this work, GPT-Red performs attacks at scale, running far more tests than manual efforts could.

How GPT-Red operates

GPT-Red sends prompts to a target model, examines responses, and refines its approach with each attempt. Failed attacks are discarded, while successful ones advance for further review. The system learned through self-play reinforcement, acting as an attacker against defender models in varied scenarios.

Attackers receive rewards for successful exploits, while defenders are rewarded for resistance. As defenses strengthen, the AI adapts by developing more advanced attacks, creating a cycle of continuous improvement. The method outperforms human red-teamers, succeeding in 84% of test scenarios compared to 13% for human testers.

A key success involved “fake chain-of-thought” attacks, which previously worked over 95% of the time against GPT-5.1. Against GPT-5.6, these attacks now succeed less than 10% of the time. The tool also uncovered vulnerabilities in autonomous agents, including a vending machine agent that could be manipulated to change prices or cancel orders.

Related: Ex-Cisco team launches AI security startup Tenet

Limitations and internal deployment

GPT-Red remains an internal tool and will not be released publicly. OpenAI maintains strict separation from its deployed models to prevent misuse of its attack methods. Findings from the system are integrated into training, with earlier versions used since GPT-5.3.

Some weaknesses persist. The system struggles with multi-turn conversational attacks that unfold over several exchanges and has limited effectiveness against image-based prompt injection. Human testers will continue addressing these gaps.

The risks increase as AI models gain autonomy. Nikhil Kandpal, a co-creator of the tool, stated that the risk surface and potential impact expand alongside model capabilities. Dylan Hunn, another co-creator, noted that the model excels at identifying effective exploits compared to human testers.

Jessica Ji, a senior research analyst at Georgetown University’s Center for Security and Emerging Technology, described the results as promising but stressed that human expertise remains essential.

OpenAI unveiled GPT-Red alongside the release of GPT-5.6, positioning it against rivals like Anthropic’s Claude. Prompt injection remains one of the most persistent unsolved challenges in AI security, and tools like GPT-Red help address it before models reach users.

Tags:

Leave a Reply

Your email address will not be published. Required fields are marked *