TL;DR
OpenAI built GPT-Red to attack its own models before shipping them, then published the results, including where it fell short. The headline: automated attacks broke models far more often than human testers did. The takeaway for AI vendors: automated, continuous red teaming, done openly, is becoming the standard for proving safety, not claiming it.
OpenAI’s GPT-Red doesn’t fix the problems in your agent built on GPT. But it does point to something much more important that’s happening in this space.
OpenAI just revealed GPT-Red: an internal system built to attack its own models before anyone else can, and to prove it worked, they published the results.
For anyone who builds, sells, or embeds AI, this is the clearest signal yet of the bar that “trustworthy AI” is about to be measured against.
The vendors who move first will define what “trusted” looks like for everyone else.
What GPT-Red Is (and Isn't)
GPT-Red is not a product. It's not a tool you can buy, download, or plug in.
It's an internal system OpenAI built to find and patch weaknesses in its own models before they ship, and OpenAI has said it keeps GPT-Red strictly internal, walled off from its deployed models.
This containment is deliberate. It keeps the advanced capabilities of the system out of the hands of adversaries, while still hardening the models that actually reach the users at the same time.
What OpenAI made public is the method that it went about. GPT-Red generates adversarial attacks, especially prompt injections, and feeds every successful break back into training to harden the next model.
Sharing that method openly sets a new operational baseline: automated, pre-deployment adversarial testing as the standard the rest of the market will be measured against.
The Evidence
The numbers published are what makes this publication hard to ignore.
On unseen scenarios in a replicated prompt-injection benchmark, GPT-Red found successful attacks in 84% of cases, compared with 13% for human red teamers probing the same model.
Automated stress-testing did not just match manual testing. It outpaced it by a wide margin.
It also held up against real systems. Pointed at a live autonomous agent running an office vending machine, GPT-Red achieved all of its objectives: cutting the price of an expensive in-stock item to the minimum allowed, listing a new $100+ item at that same floor price, and cancelling another customer's order.
Why This Matters: Transparency Is the Real Signal
The most important thing OpenAI did was choosing to be open.
OpenAI published the method, the win rates, and the failures, including how vulnerable its own earlier models were. It disclosed the live-system vulnerabilities it found and said fixes are being tested, and it pointed to a forthcoming pre-print with the details.
Being upfront about the disparities between model generations, where they were weak, and by how much, is the opposite of the usual instinct to publish only the wins.
As AI becomes more regulated and more scrutinized, buyers, regulators, and end users are increasingly asking vendors to show their work rather than assert that their systems are safe.
Trust, safety, and security are moving to the center of how AI products get evaluated, and transparency about your own weaknesses, not just your strengths, is what earns that trust.
This splits the market in two:
- Vendors who can demonstrate automated, adversarial resilience, and do it openly, become the trusted partners of choice.
- Vendors who cannot demonstrate this automated resilience, face escalating scrutiny, lost deals, and reputational risk.
You Don't Need to Build Your Own GPT-Red
Nobody should be spending millions to train a frontier-grade attacker. That isn't feasible for the vast majority of AI teams, and it isn't the point.
Our recommendation is straightforward: source a proven, automated red-teaming platform.
The real shift isn't owning the biggest attacker; it's moving from manual, ad hoc red teaming to a continuous practice that pairs human expertise with the scale and depth of automation. Sourcing a mature capability is both the sensible path and the faster one.
This is where Alice comes in. Alice is an AI security platform that delivers automated red teaming. Its stress-testing is built on a database of billions of real-world adversarial scenarios and covers both out-of-the-box and fully custom policies, so it reflects how your agents actually get attacked rather than a generic checklist.
Guardrails Have to Evolve Too
Red teaming is only half the story, and it’s now clear that static guardrails won't survive contact with RL-trained adversaries. The same techniques behind a system like GPT-Red are exactly what break a fixed rule set.
The fix is to make guardrails dynamic and adaptive, and to feed them red teaming insights directly. Guardrails informed by real adversarial testing are calibrated to how systems actually get broken.
With Alice, the guardrails used to secure agents are built on top of the same red teaming assessment that finds your weaknesses. They learn from real attack data, which keeps them accurate, low-latency, and tuned to hold false positives down. The same system that finds the holes is the system that closes them.
Where to Start
The path runs in three stages, and with Alice they operate as one continuous loop rather than three separate projects:
- Stress-test. Audit your exposure to advanced prompt injection and adversarial attacks, and run automated red-teaming pilots.
With Alice, this is available now, so you can pilot quickly and see where your agents are vulnerable across both out-of-the-box and custom policies.
- Guardrail. Take your red teaming findings and convert them to guardrails trained to withstand RL-trained attackers, documenting and sharing your testing as you go.
With Alice, the exact results taken from the red teaming are used to create solid guardrails designed for protecting your agent.
- Continuously assess. Move from random, occasional stress-testing to continuous automated adversarial testing across the AI lifecycle.
Alice provides constant red teaming even after the guardrails are activated to continuously learn about new threats and vulnerabilities in the system, and to optimize the guardrails accordingly.
The Opportunity
OpenAI has shown what responsible AI looks like in this era: test your own systems adversarially, at scale, and automatically, and be transparent about what you find, including where you fall short.
Every vendor navigating this world will be expected to put trust, safety, and security at the core of what they ship, and to prove it rather than promise it.
The vendors who embrace that openness become the ones that the public trusts. The ones who wait will be explaining, later, why they didn't.
Want to know how your agents hold up? See Alice in action →
What’s New from Alice
AI in Healthcare: Protecting Patient Data Without Falling Behind
Your doctor knows things about you that almost nobody else does. So what happens when AI gets access to all of it? Sandy Dunn has spent much of her career worrying about exactly that. She's a healthcare CISO, and her answer is calmer than you'd think: the things that can go wrong aren't new, it's how fast they happen and how far the damage spreads. In this episode, she and Mo get into why HIPAA has become paperwork that protects almost nobody, why the safest data is the data you never collected, and what happens to trust when AI is in the exam room.
It Takes AI to Break AI: The Case for AI Red Teaming
As AI systems gain autonomy, organizations need security approaches built specifically for AI behavior. Learn why AI-driven red teaming is becoming a critical defense layer.
Demystifying AI Red Teaming
Your AI passed every check. That doesn't mean it's safe. Learn how to red team AI systems before adversaries find the gaps you missed.

