Your policy,
enforced properly in 24 hours
Off-the-shelf classifiers enforce someone else's policy. LLMs behind an API are accurate but too slow and costly to run on every message. Alice builds custom guardrails trained on your exact policy, live in 24 hours, running inline on 100% of your traffic, and improving every day after.
Get a DemoFrontier accuracy, small-model speed
A large model reads your policy and labels examples, recording its reasoning. A compact model learns to reproduce those judgments directly, with no reasoning at inference. You pay for the deliberation once, at training, and run something small and fast in production.
The result: the judgment of a frontier model, scored inline in under 120ms at P95, at a fraction of the per-call cost of an LLM, on every message, not a sample.
Red team, guardrail, repeat
1. Red team before launch
2. Turn the findings into guardrails
3. Optimize on a continuous loop
Each round mines your real traffic for what the model gets wrong, so the classifier keeps matching your policy as it actually appears in production. The whole round runs in 24 hours, unprompted.
How the classifier is built
Book a demoIt starts with a teacher and a student
Alice uses a distillation pipeline. A frontier model (the teacher) labels examples against your policy and records its reasoning. A compact model (the student) reproduces those judgments with no reasoning at inference. The deliberation is compiled once, at training, so what runs in production stays small and fast. Your policy itself arrives however it exists, a PDF, a compliance doc, or rules grown over years, and Alice reads it end to end, returning a validation pack so you confirm it was understood before any model is trained.
Then it's toughened against real-world input
Real traffic isn't clean prose. It arrives as HTML fragments, emoji standing in for words, letters s p a c e d out, lookalike characters, and deliberate evasion by people who know a classifier is watching. Alice normalizes every input to canonical form before scoring, using a preprocessing stack built from years of real abuse traffic, so evasion doesn't win just by reaching the model intact.
Finally, it hunts its own mistakes
Alice doesn't wait for a user to complain or an analyst to spot a pattern. The system scores your real traffic, attacks itself, and treats every error it finds as a lead. Each one becomes new training data, and a new version only ships if it fixes those errors without breaking anything that already worked.
Custom guardrails built for any scenario
8
10B+
<120ms
24/7
Trusted by security and product teams in the world's most regulated industries
Alice brings years of adversarial intelligence expertise to AI security. We give enterprise teams the coverage that generic guardrails and one-time audits can't match.
Get a DemoAlice Data Advantage
Alice is the world’s largest collector and manager of adversarial intelligence data. Our data is the cornerstone for protecting platform, tech, and users online.
Explore Rabbit Hole intelligenceWhat’s new from Alice
Cyber RL Benchmark v1.0: Measuring Frontier Cybersecurity Capability
Alice's Cyber RL Benchmark v1.0 tested 10 frontier models on 12 real cybersecurity tasks, with fresh code, real vulnerabilities, planted decoys, and objective scoring. The finding: no model can reliably cover the range of a security analyst's job.
LIVE from Black Hat Las Vegas: AI, Nation-States, and the Battlefield That Keeps Changing
Recorded live at Black Hat, Caroline Wong and Allie Mellen unpack why AI made nation-state attacks cheaper, why safety guardrails can work against you, and what the OpenAI-Hugging Face incident really exposed.
Virtual Fireside Chat: The TAKE IT DOWN Act, Six Months In - What's Changed on Deepfakes and NCII
Six months after the TAKE IT DOWN Act took effect, the NCII landscape looks different and even more complicated. Join Alice for a live fireside chat with Google's Nidhi Lahoti on what's actually changed, what hasn't, and where T&S fit into it all - Join us live on October 9th, 2pm EST.
5 Ways Your Third-Party CX Agent Gets Broken
Third-party CX agents create hidden liability. Learn the 5 attack patterns vendors miss and how WonderSuite closes the gap.