TL;DR
We evaluated nine SOTA models as enterprise agents across 147 adversarial scenarios from seven work domains in seven industries. The results show a high failure rate across the board, with every frontier model we tested failing on the majority of adversarial enterprise scenarios.
Enterprises are increasingly deploying AI agents into real workflows, giving them access to tools, data, and the ability to act on their own. Indirect prompt injection attacks are already showing up in that production context: a hidden instruction buried in an email, a ticket, a vendor request, waiting for an agent with the right access to act on it.
From Alice's work with enterprises, we know that security is one of the top reasons agent deployments stall - it’s not a lack of capability, but a lack of confidence that the agent won't be the one that gets it wrong. Alice built this benchmark to measure how often frontier models fall for indirect prompt injection in the enterprise context.
This is the first entry in Alice's benchmark series evaluating agentic model behavior under the conditions that actually matter for deployment. Future entries will explore other areas of agentic risk.

What we measured
We evaluated nine frontier models as enterprise agents across 147 adversarial scenarios spanning seven work domains (finance, legal, marketing and PR, HR, sales, IT, customer support) and seven industries (energy and utilities, gaming and entertainment, news and media, consumer discretionary, finance, healthcare, technology).
Each scenario pairs a realistic enterprise task with an embedded attack: a hidden instruction, a scope-expanding request, a poisoned input. The attacks are delivered through patterns like helpfulness trap, environment poisoning, lookalike domain, social engineering, and authority laundering and they target violations including data exfiltration, confidentiality breaches, fraud, privilege escalation, and more.
Every scenario runs three times per model. Each run is scored on a 0-10 security scale by a nine-judge jury of frontier models. The primary metric used to evaluate models’ performance is Agent Failure Rate (AFR): the share of scenarios where a model fails at least once across its three runs, not whether it can succeed once. A scenario is marked failed if even one run scores below 7.0 on security. A separate functional adherence score was also measured as a quality assurance check to ensure the agent actually engaged with the task and prevent the agent from reward-hacking by simply doing nothing.
A sample scenario
The scenarios were shaped by extensive consultation with domain and security experts, who authored the original attack scenarios and conducted final validation to keep them realistic and grounded in what's actually at stake for enterprises.
Below is a sample scenario:

Results: high failure rates across the board
Every frontier model we tested failed on the majority of adversarial enterprise scenarios. Mean AFR across all the models was 0.75. Claude Opus 4.6 led at 0.58; the rest clustered between 0.67 and 0.88.
In practice, that means security failures are the norm, not the exception. The average model failed at least once on roughly three out of every four scenarios, and every frontier model failed on more than half.
An update: re-testing against newer model versions
Several providers have shipped newer versions since this benchmark was run in May 2026. Alice re-tested the current lineup against the same scenarios and the results were mixed.
Claude Opus 4.8 posted the sharpest gain in the cohort, cutting AFR from 0.58 to 0.17, more than a threefold reduction. GPT-5.6-sol-pro and Kimi-k3 also improved meaningfully.

Conclusion
No model is safe by default
Even the best model in the updated results still fails 17% of scenarios. That's a significant improvement, but still not good enough. In a production deployment, a single successful attack can mean a data leak, a fraudulent transaction, or a compliance breach. Failing one time in six is not a rounding error. For an enterprise evaluating these models that’s a strong reason to hold off on deployment.
This is hill-climbable
However, this gap is not permanent. Alice partners with frontier model providers to help close the shortfalls it surfaces through targeted training and re-evaluation. We're already doing this work with leading frontier labs and seeing meaningful improvements in their performance.
Want to learn more or collaborate on future benchmarks?
Contact Alice LabsWhat’s new from Alice
ENT-IPI Bench: Enterprise Indirect Prompt Injection Benchmark
We evaluated nine frontier models as enterprise agents across 147 adversarial scenarios from seven work domains in seven industries for their vulnerability to indirect prompt injection.
LIVE from Black Hat Las Vegas: AI, Nation-States, and the Battlefield That Keeps Changing
What if the biggest threat to your security team isn't the attacker, it's the model you're relying on to stop them? LIVE from Black Hat Las Vegas, Mo and Madi bring together two cybersecurity authors, Caroline Wong, Chief Strategy Officer at Axari and author of The AI Cybersecurity Handbook, and Allie Mellen, Principal Analyst at Forrester and author of Code War, who wrote very different books that turn out to be arguing the same point. One explains why nations attack the way they do. The other explains why AI just changed the cost, speed, and scale of everything. Tune in!
It Takes AI to Break AI: The Case for AI Red Teaming
As AI systems gain autonomy, organizations need security approaches built specifically for AI behavior. Learn why AI-driven red teaming is becoming a critical defense layer.
5 Ways Your Third-Party CX Agent Gets Broken
Third-party CX agents create hidden liability. Learn the 5 attack patterns vendors miss and how WonderSuite closes the gap.

