The Wire
01:47Beauty Pop-Ups Are Redefining How Brands Connect With Customers01:46Three century-old Korean temple toilets set for heritage status01:46Yash Raj Films Enters Music With Raah Records, Debuts Aman’s ‘Jaadugari’01:46Hybe breaks 1 trillion won quarterly sales on BTS comeback momentum01:46Tom Waits Returns with Spoken-Word Track ‘The Fly’, a Family Collaboration01:47Beauty Pop-Ups Are Redefining How Brands Connect With Customers01:46Three century-old Korean temple toilets set for heritage status01:46Yash Raj Films Enters Music With Raah Records, Debuts Aman’s ‘Jaadugari’01:46Hybe breaks 1 trillion won quarterly sales on BTS comeback momentum01:46Tom Waits Returns with Spoken-Word Track ‘The Fly’, a Family Collaboration
SPOTLIGHT NO. 412 · SINGAPORE · THU 6 AUG 2026 · 18:13 +00:00 Sign in Subscribe
Uncategorized

OpenAI Built an AI Super-Hacker Called GPT-Red to Attack Its Own Models

OpenAI built GPT-Red, an AI system designed to attack its own models. It says the tool cut successful attacks against GPT-5.6 to under 23%, down from over 90%.

OpenAI Built an AI Super-Hacker Called GPT-Red to Attack Its Own Models

OpenAI has built an internal AI system named GPT-Red whose only job is to break other OpenAI models, and the company says pitting its newest release against it produced its most attack-resistant model to date, according to MIT Technology Review.

The firm describes GPT-Red as an automated red-teamer, a stand-in for the human testers who normally probe software for weaknesses before release. It released the latest version of its flagship model, GPT-5.6, last week, and credits the sparring sessions with GPT-Red for hardening it against a range of cyberattacks.

Why OpenAI wants a machine attacker

Red-teaming has traditionally been slow, manual work. A group of people tries to find as many ways as possible to hijack a system, then engineers patch the holes. That model strains as LLMs move into agent territory, where they read files, browse websites, run third-party code, and coordinate with other agents.

"The risk surface grows and the blast radius also grows," said Nikhil Kandpal, an OpenAI research scientist who co-created GPT-Red. The company frames the tool as a way to keep pace with attack methods that human teams cannot fully anticipate.

Dylan Hunn, another research scientist and co-creator, said the intent is to stay ahead of future models rather than react to them. The team says GPT-Red has already surfaced attack types that had not been documented before.

The self-play dojo

To build it, OpenAI took a model that had no hacking training and placed it in a self-play loop against several other models. GPT-Red's objective was to attack; the others' objective was to defend. Across repeated rounds, the attacker sharpened its methods while the defenders learned to resist them.

The training environment, which OpenAI calls a dojo, simulated real deployment scenarios: browsing the web, reading emails and calendars, and editing code. When GPT-Red discovered a viable attack, it tested variations to isolate the most effective one for a given setting.

"Compared to a human red-teamer, the model is very, very good at finding exactly what will work," Hunn said. The team describes it as persistent to the point of drilling relentlessly into any weakness it locates.

Much of the work targeted prompt injection, where hidden instructions embedded in text, code, or a webpage push a model to act against its developers' or users' intent, such as leaking confidential data or sabotaging a codebase. One method GPT-Red produced was new to the researchers: a fake chain of thought. A chain of thought is the running set of notes a model keeps as it works through a problem. GPT-Red learned to insert a forged entry into another model's notes, tricking it into treating false information as verified.

"It's like if I told you that 1+1=3 and that you have verified this already," said research scientist Chris Choquette-Choo, describing how the target model then accepts and repeats the bad result.

The numbers OpenAI is citing

OpenAI rerun a 2025 experiment in which human testers hunted for flaws in an earlier version of GPT-5. Given the same task, GPT-Red found effective attacks more often than the humans did. In a separate test, it hacked Vendy, a vending-machine agent from Andon Labs, changing item prices and canceling a customer order.

The headline claim: more than 90% of GPT-Red's strongest attacks worked against GPT-5, released in August of last year, while fewer than 23% worked against GPT-5.6. Those figures come from OpenAI's own testing and have not been independently verified.

Jessica Ji, a senior research analyst on AI security at Georgetown University's Center for Security and Emerging Technology, called the self-play approach sound. "The results look very promising," she said, while adding that human testing still matters and that the harder problem is knowing where it is most needed.

Limits and gatekeeping

GPT-Red is not comprehensive. It struggles with multi-turn attacks that unfold over a back-and-forth conversation, and it is weak at using images to smuggle text into a model, both of which human attackers handle more easily. OpenAI positions it as a supplement to human red-teamers, sometimes handing it a human-devised attack and asking it to generate every variant.

The company will not release GPT-Red, and argues that a copycat would be hard to reproduce. Choquette-Choo said the team spent more than a year on the model with the compute resources of one of the world's richest companies behind it, framing that as a barrier to anyone hoping to train a comparable attacker.

That argument cuts both ways. A safety tool this capable, kept private and defended as too expensive to replicate, concentrates both the offensive capability and the judgment about how to use it inside a single vendor, with outside verification limited to what OpenAI chooses to publish.

The Brief · Every weekday

The people and forces shaping Asia.

One email, every morning, in five minutes — in the language the world reads.

FREE · UNSUBSCRIBE ANYTIME