TL;DR
This benchmark evaluates 10 frontier models and agents on a selection of 12 cyber capability tasks across the full security lifecycle. Each task runs on real software with real vulnerabilities and objective scoring at every step. Qwen 3.8 Max leads with a 0.72 mean score and 7 of 12 tasks solved, while Mistral Medium 3.5 trails at 0.36.

Overview
If you ask a frontier model to perform a cybersecurity task it will usually produce something. The question is whether the output is actually right. Most cyber benchmarks don't answer that well: they grade pass/fail, become contaminated once used in training, ignore false positives entirely, and don't capture how far a model actually got or how accurate its output is.
We built this benchmark on real code we wrote, with vulnerabilities modeled on real threat actor techniques and planted decoys to test both completion and accuracy, not just volume. The results show that no model tested can reliably cover the range of tasks it would take to replace core parts of a security analyst's job.
Note: v1.0 covers an initial selection of 12 tasks across four gyms. Future versions will expand coverage of number of tasks and across the full RL security lifecycle.
What Was Measured
- 10 frontier models: Opus 5.5, Qwen 3.8 Max, Kimi K3, DeepSeek V4.1 Flash, GLM-5.3, Gemini 3.1 Pro, Muse Spark 1.3, GPT-6 Luna, Grok 4.7, Mistral Medium 3.5.
- 12 tasks across four gyms: malware analysis, vulnerability detection, vulnerability patching, and blackbox penetration testing (3 tasks per gym)
- 3 runs per task per model, graded on a 0-1 scale (0 is no progress and 1 is full completion), with partial credit awarded based on the verifier built for each task in advance.
- Tasks blocked by a model's guardrails are scored 0, the same as any other non-completion
A Sample Task
Each task plants ground-truth objectives (vulnerabilities, IOCs, or exploit chains), several of them with deliberate decoys.

The Results




Observations
Note: The following are observations based on performance across the benchmark's current selection of 12 tasks (3 runs each), and are likely to evolve as task coverage expands.
- Some models fell for planted false positives. Of the 43 false positives planted across the vulnerability detection and malware analysis tasks, Mistral Medium 3.5 fell for 12, GPT-6 Luna for 8, and Grok 4.7 for 6.
- Speed doesn't indicate performance. GPT-6 Luna is the fastest model on every pentest task but scores just 0.00-0.25 there.
- Consistency matters more than peak performance. On one of the tasks, three models each scored 1.00 at least once, but none did it reliably across all three runs.
- Some labs block offensive security work by design. Opus 5.5's guardrails refused every pentest task at the API layer. GPT-6 Sol did the same (tested separately, not in this report). A deliberate choice, but a real constraint for legitimate offensive security work.
Looking Forward
Future versions of this benchmark will expand task coverage across new and existing gyms, add efficiency scoring for time and token usage, and introduce more sophisticated decoys.
Want to learn more or collaborate on future benchmarks? Contact us
Want to collaborate on future benchmarks?
Talk to our team →What’s new from Alice
Cyber RL Benchmark v1.0: Measuring Frontier Cybersecurity Capability
Alice's Cyber RL Benchmark v1.0 tested 10 frontier models on 12 real cybersecurity tasks, with fresh code, real vulnerabilities, planted decoys, and objective scoring. The finding: no model can reliably cover the range of a security analyst's job.
Making Sense of AI: Trust, Scale, and the Human Role
Curiosity might be our most important security tool. In the first episode of Curiouser & Curiouser, Mo Sadek sits down with longtime security leader Julie Tsai to explore AI, security, and the human judgment that still matters most. Together, they cut through hype and fear to talk about what’s actually changing, what isn’t, and how we build systems we can truly trust.
Virtual Fireside Chat: The TAKE IT DOWN Act, Six Months In - What's Changed on Deepfakes and NCII
Six months after the TAKE IT DOWN Act took effect, the NCII landscape looks different and even more complicated. Join Alice for a live fireside chat with Google's Nidhi Lahoti on what's actually changed, what hasn't, and where T&S fit into it all - Join us live on October 9th, 2pm EST.
5 Ways Your Third-Party CX Agent Gets Broken
Third-party CX agents create hidden liability. Learn the 5 attack patterns vendors miss and how WonderSuite closes the gap.
