TL;DR
Extremists have never lacked intent, only capability, and AI is now closing that gap. In default settings, using nothing more than a patient conversation, our red team got models to walk through steps toward biological, chemical, and explosive harm. This is a short summary of a research paper by Alice's Uri Klempner, published by GNET.
Extremists, those who hold beliefs outside society's norms and promote intolerance, hatred, or violence, have never lacked intent. What has historically prevented most from acting on those beliefs was capability. That was true both in the physical world and in the dark corners of the digital sphere, where Alice has spent years monitoring and tracking how extremists and terrorists operate.
The information itself was always there. Bomb-making manuals, virology papers, chemical synthesis routes, and other dangerous materials have long been available across the open web and the dark web. What was difficult to obtain was the part no document can fully capture: tacit knowledge. The know-how that bridges the gap between a written protocol and a successful outcome, the understanding of why an experiment fails on the bench even when every step appears correct on paper.
That gap is what frontier AI is quietly closing.
A model that can diagnose why your synthesis failed, explain what to change, and walk you to the next step is a mentor, not a search engine. And a mentor is exactly what most attackers never had. More than half of lone-actor terrorists between 1990 and 2012 were socially isolated. The role AI now fills, patient, tireless, encouraging, is the role they were missing.
"Theoretically": A simple trick
In February 2026, researchers from our red team lab ran several frontier models through CBRNE (Chemical, Biological, Radiological, Nuclear, and Explosive) testing. This was a ground-level stress test. No technical exploits, no sophisticated jailbreaks or injections, just simple prompting that mimicked how a lone actor would approach a model, with harmful intent implied rather than stated. Every model output was then checked by our subject-matter experts, PhD-level chemists and biologists, for whether it was actionable.
It very much was.
One model produced a 54-page progression across the full acquisition-to-deployment chain for a biological weapon, assembled from individually reasonable requests. Others gave operational detail on Clostridium botulinum, or walked through how to extract and isolate ricin.
One model turned on a single word. A request for how to rig an explosive was refused. Add "theoretically" to the identical prompt, resubmit, and the model returned a walkthrough for a remote-detonated improvised explosive device.
Other attempts used “crescendo” jailbreaking. It sounds like a term of art, but it names something quite simple: a many-shot pattern that wears the model safety mechanisms down over a long conversation, coaxing compliance one reasonable step at a time. This is the same pattern that gets a customer-service bot to write profanity about its own brand. And it is invisible to any test that checks one message at a time.
Let's talk about guardrails
AI developers respond to misuse by adding safety layers called guardrails, built to detect and block requests for harmful content. But threat actors adapt as fast as vulnerabilities get patched. The result is an adversarial cycle that demands continuous evaluation, not a one-time fix.
In many LLMs, guardrails are essentially filters: refuse to answer when a prompt hits a flagged set of words, and return an "I can't help with that" to an obvious ask like "tell me how to build a bomb." But models are built to please the user, and that eagerness to help tends to overpower the filter. It is why the simple tricks above work at all.
So a guardrail cannot be a word list checked one prompt at a time. The helpfulness that makes a model useful is the same helpfulness an attacker steers, and no amount of refusal-tuning on individual answers fully closes it. A refusal only counts if it holds for the whole conversation, not just the sentence that triggered it.
The open-source version is even worse
The models examined in our research were all closed, commercially hosted frontier models. Those tend to have stronger safety mechanisms, with dedicated teams behind them. Open-weight models are where extremism and CBRNE misuse become a more pressing problem. As open systems like Llama, DeepSeek, Qwen, and more recently Kimi grow more capable and more popular, catastrophic capability becomes more accessible, and less supervised, than it is behind a proprietary API.
NPR reported on a user in an extremist forum who turned to one of these uncensored models to research the explosives for an attack on a well-known tower. And unlike a hosted model, an open-weight model can have its guardrails stripped out entirely, at the weights, in minutes, on a laptop that costs a few hundred dollars. Millions of these de-restricted models have already been downloaded, their refusal signal simply deleted.
Build it in before release
If a filter cannot catch this, testing has to. That means red teaming a model the way an attacker would use it, across long, multi-turn conversations, not single prompts, because a single-turn test never sees the crescendo coming. It also means grounding that testing in how adversaries work, across every language and modality they use. A model is only as resilient as the attacks it was trained and tested against.
That work happens at the model level, before release, with the teams building the models themselves. It is what we do at Alice Labs: expert and automated red teaming across text, image, audio, and video; training and alignment datasets that teach refusal to hold across a conversation instead of a single prompt; detection signals drawn from the real threat landscape; and agentic RL environments that test how a model behaves across the full range of conditions it will meet in the wild.
None of this makes frontier models too dangerous to build. It makes them too dangerous to build carelessly. The capability is arriving either way. What gets decided now is how much of it reaches the people who should never have it.
Read the full research, "Artificial Intelligence and Extremist Capability: How AI Lowers the Barriers Between Intent and Capability," which was originally published on GNET, the Global Network on Extremism and Technology, an academic research initiative studying how technology is used and misused across the extremism landscape.
What’s New from Alice
For Extremists, the Gap Was Never Intent. It Was Capability.
Extremists have never lacked intent, only capability, and frontier AI is quietly closing that gap. In default settings, Alice's red team used patient conversation to walk models toward CBRNE harm, showing why guardrails must hold across every turn.
Why Your CISO Shouldn't Be the Department of No
What if the person whose whole job is risk was the most excited one in the room about AI? That's Mea Clift, and it's a pretty refreshing way to walk into all this. She's the CISO of Cengage, where the data she's protecting belongs to students, so the stakes are real. But instead of bracing for what could go wrong, she leans in, she calls it being "risk excited." She and Mo get into why security is so much better as the Department of KNOW than the department of no, how she handles shadow AI without punishing people for being curious, why she'd tailor a framework she already has instead of building one from scratch, and how she ends up teaching security through Marvel and Monsters Inc. Which, it turns out, is every bit as fun as it sounds.
It Takes AI to Break AI: The Case for AI Red Teaming
As AI systems gain autonomy, organizations need security approaches built specifically for AI behavior. Learn why AI-driven red teaming is becoming a critical defense layer.
5 Ways Your Third-Party CX Agent Gets Broken
Third-party CX agents create hidden liability. Learn the 5 attack patterns vendors miss and how WonderSuite closes the gap.

