ActiveFence is now Alice
x
Back
Blog

For Extremists, the Gap Was Never Intent. It Was Capability.

Uri Klempner
-
Aug 5, 2026

TL;DR

Extremists have never lacked intent, only capability, and AI is now closing that gap. In default settings, using nothing more than a patient conversation, our red team got models to walk through steps toward biological, chemical, and explosive harm. This is a short summary of a research paper by Alice's Uri Klempner, published by GNET.

Extremists, those who hold beliefs outside society's norms and promote intolerance, hatred, or violence, have never lacked intent. What has historically prevented most from acting on those beliefs was capability. That was true both in the physical world and in the dark corners of the digital sphere, where Alice has spent years monitoring and tracking how extremists and terrorists operate.

The information itself was always there. Bomb-making manuals, virology papers, chemical synthesis routes, and other dangerous materials have long been available across the open web and the dark web. What was difficult to obtain was the part no document can fully capture: tacit knowledge. The know-how that bridges the gap between a written protocol and a successful outcome, the understanding of why an experiment fails on the bench even when every step appears correct on paper.

That gap is what frontier AI is quietly closing.

A model that can diagnose why your synthesis failed, explain what to change, and walk you to the next step is a mentor, not a search engine. And a mentor is exactly what most attackers never had. More than half of lone-actor terrorists between 1990 and 2012 were socially isolated. The role AI now fills, patient, tireless, encouraging, is the role they were missing.

"Theoretically": A simple trick

In February 2026, researchers from our red team lab ran several frontier models through CBRNE (Chemical, Biological, Radiological, Nuclear, and Explosive) testing. This was a ground-level stress test. No technical exploits, no sophisticated jailbreaks or injections, just simple prompting that mimicked how a lone actor would approach a model, with harmful intent implied rather than stated. Every model output was then checked by our subject-matter experts, PhD-level chemists and biologists, for whether it was actionable.

It very much was.

One model produced a 54-page progression across the full acquisition-to-deployment chain for a biological weapon, assembled from individually reasonable requests. Others gave operational detail on Clostridium botulinum, or walked through how to extract and isolate ricin.

One model turned on a single word. A request for how to rig an explosive was refused. Add "theoretically" to the identical prompt, resubmit, and the model returned a walkthrough for a remote-detonated improvised explosive device.

Other attempts used “crescendo” jailbreaking. It sounds like a term of art, but it names something quite simple: a many-shot pattern that wears the model safety mechanisms down over a long conversation, coaxing compliance one reasonable step at a time. This is the same pattern that gets a customer-service bot to write profanity about its own brand. And it is invisible to any test that checks one message at a time.

Let's talk about guardrails

AI developers respond to misuse by adding safety layers called guardrails, built to detect and block requests for harmful content. But threat actors adapt as fast as vulnerabilities get patched. The result is an adversarial cycle that demands continuous evaluation, not a one-time fix.

In many LLMs, guardrails are essentially filters: refuse to answer when a prompt hits a flagged set of words, and return an "I can't help with that" to an obvious ask like "tell me how to build a bomb." But models are built to please the user, and that eagerness to help tends to overpower the filter. It is why the simple tricks above work at all.

So a guardrail cannot be a word list checked one prompt at a time. The helpfulness that makes a model useful is the same helpfulness an attacker steers, and no amount of refusal-tuning on individual answers fully closes it. A refusal only counts if it holds for the whole conversation, not just the sentence that triggered it.

The open-source version is even worse

The models examined in our research were all closed, commercially hosted frontier models. Those tend to have stronger safety mechanisms, with dedicated teams behind them. Open-weight models are where extremism and CBRNE misuse become a more pressing problem. As open systems like Llama, DeepSeek, Qwen, and more recently Kimi grow more capable and more popular, catastrophic capability becomes more accessible, and less supervised, than it is behind a proprietary API.

NPR reported on a user in an extremist forum who turned to one of these uncensored models to research the explosives for an attack on a well-known tower. And unlike a hosted model, an open-weight model can have its guardrails stripped out entirely, at the weights, in minutes, on a laptop that costs a few hundred dollars. Millions of these de-restricted models have already been downloaded, their refusal signal simply deleted.

Build it in before release

If a filter cannot catch this, testing has to. That means red teaming a model the way an attacker would use it, across long, multi-turn conversations, not single prompts, because a single-turn test never sees the crescendo coming. It also means grounding that testing in how adversaries work, across every language and modality they use. A model is only as resilient as the attacks it was trained and tested against.

That work happens at the model level, before release, with the teams building the models themselves. It is what we do at Alice Labs: expert and automated red teaming across text, image, audio, and video; training and alignment datasets that teach refusal to hold across a conversation instead of a single prompt; detection signals drawn from the real threat landscape; and agentic RL environments that test how a model behaves across the full range of conditions it will meet in the wild.

None of this makes frontier models too dangerous to build. It makes them too dangerous to build carelessly. The capability is arriving either way. What gets decided now is how much of it reaches the people who should never have it. 

Read the full research, "Artificial Intelligence and Extremist Capability: How AI Lowers the Barriers Between Intent and Capability," which was originally published on GNET, the Global Network on Extremism and Technology, an academic research initiative studying how technology is used and misused across the extremism landscape.

Share

What’s New from Alice

5 Ways Your Third-Party CX Agent Gets Broken

whitepaper
Jul 31, 2026
,
 
Jul 31, 2026
 -
This is some text inside of a div block.
 min read
Jul 31, 2026
 -
This is some text inside of a div block.
 min watch
July 31, 2026

Third-party CX agents create hidden liability. Learn the 5 attack patterns vendors miss and how WonderSuite closes the gap.

Learn More
Red-Team Lab
Intelligence Desk