Alice raises $140M for frontier AI safety and security [Read the news]
Back
Blog

Cyber RL Benchmark v1.0: Measuring Frontier Cybersecurity Capability

Alice Labs
-
Oct 6, 2026
See the full results, scores, and methodology.
Read the full breakdown →

TL;DR

This benchmark evaluates 10 frontier models and agents on a selection of 12 cyber capability tasks across the full security lifecycle. Each task runs on real software with real vulnerabilities and objective scoring at every step. Qwen 3.8 Max leads with a 0.72 mean score and 7 of 12 tasks solved, while Mistral Medium 3.5 trails at 0.36.

‍

Overview

If you ask a frontier model to perform a cybersecurity task it will usually produce something. The question is whether the output is actually right. Most cyber benchmarks don't answer that well: they grade pass/fail, become contaminated once used in training, ignore false positives entirely, and don't capture how far a model actually got or how accurate its output is. 

We built this benchmark on real code we wrote, with vulnerabilities modeled on real threat actor techniques and planted decoys to test both completion and accuracy, not just volume. The results show that no model tested can reliably cover the range of tasks it would take to replace core parts of a security analyst's job.

Note: v1.0 covers an initial selection of 12 tasks across four gyms. Future versions will expand coverage of number of tasks and across the full RL security lifecycle.

‍

What Was Measured

  • 10 frontier models: Opus 5.5, Qwen 3.8 Max, Kimi K3, DeepSeek V4.1 Flash, GLM-5.3, Gemini 3.1 Pro, Muse Spark 1.3, GPT-6 Luna, Grok 4.7, Mistral Medium 3.5.
  • 12 tasks across four gyms: malware analysis, vulnerability detection, vulnerability patching, and blackbox penetration testing (3 tasks per gym)
  • 3 runs per task per model, graded on a 0-1 scale (0 is no progress and 1 is full completion), with partial credit awarded based on the verifier built for each task in advance.
  • Tasks blocked by a model's guardrails are scored 0, the same as any other non-completion

‍

A Sample Task

Each task plants ground-truth objectives (vulnerabilities, IOCs, or exploit chains), several of them with deliberate decoys.

A sample vulnerability-detection task: a working Flask app with one exploitable flaw and seven calibrated red herrings, testing whether a model finds the real issue and rules out the decoys.
Flask app audit task: one real vulnerability, seven decoys, and a five-step attack chain

The Results

Malware analysis: three static-analysis tasks on a staged ELF loader, a malicious npm package, and a Windows persistence loader. Opus 5.5 leads at 78%, with most models clustered in the high 60s to low 70s.
Malware analysis mean scores for 10 models, led by Opus 5.5 at 78%
Vulnerability detection: three whitebox audits that reward finding real vulnerabilities while ruling out decoys. Qwen 3.8 Max leads at 89%, and the scores spread widely, the clearest sign that handling false positives is what separates the models.
Vulnerability detection mean scores for 10 models, led by Qwen 3.8 Max at 89%

‍

Vulnerability patching: the benchmark's highest-scoring gym, where each fix has to hold against proof-of-concept variants and regression checks. Six models reached a perfect 100%, with the rest close behind.
Vulnerability patching mean scores, the benchmark's highest gym, with six models tied at 100%
Blackbox pentesting: the hardest gym, where the agent must chain exploits from a single target host up to remote code execution. No model clears 42%. Opus 5.5 scored zero because its API-layer guardrails blocked offensive work by design, a deliberate choice rather than a capability gap.
Blackbox pentesting mean scores, the benchmark's hardest gym, topping out at 42%

‍

Observations

Note: The following are observations based on performance across the benchmark's current selection of 12 tasks (3 runs each), and are likely to evolve as task coverage expands.

  • Some models fell for planted false positives. Of the 43 false positives planted across the vulnerability detection and malware analysis tasks, Mistral Medium 3.5 fell for 12, GPT-6 Luna for 8, and Grok 4.7 for 6.
  • Speed doesn't indicate performance. GPT-6 Luna is the fastest model on every pentest task but scores just 0.00-0.25 there.
  • Consistency matters more than peak performance. On one of the tasks, three models each scored 1.00 at least once, but none did it reliably across all three runs.
  • Some labs block offensive security work by design. Opus 5.5's guardrails refused every pentest task at the API layer. GPT-6 Sol did the same (tested separately, not in this report). A deliberate choice, but a real constraint for legitimate offensive security work.

‍

Looking Forward

Future versions of this benchmark will expand task coverage across new and existing gyms, add efficiency scoring for time and token usage, and introduce more sophisticated decoys.

Want to learn more or collaborate on future benchmarks? Contact us

Want to collaborate on future benchmarks?

Talk to our team →
Share

What’s new from Alice

5 Ways Your Third-Party CX Agent Gets Broken

whitepaper
Jul 31, 2026
,
 
Jul 31, 2026
 -
This is some text inside of a div block.
 min read
Jul 31, 2026
 -
This is some text inside of a div block.
 min watch
July 31, 2026

Third-party CX agents create hidden liability. Learn the 5 attack patterns vendors miss and how WonderSuite closes the gap.

Learn More
Red-Team Lab