
As GenAI tools and the LLMs behind them impact the daily lives of billions, this report examines whether these technologies can be trusted to keep users safe.
What you’ll learn:
- How LLMs respond to risky prompts from bad actors and vulnerable users
- Where current models show safety strengths and weaknesses
- Actionable steps to improve LLM safety and reduce harmful outcomes
Overview
In this first independent benchmarking report on the LLM safety landscape, ActiveFence’s subject-matter experts put leading models to the test. More than 20,000 prompts were used to analyze how six LLMs respond across seven major languages and four high-risk abuse areas: child exploitation, hate speech, self-harm, and misinformation. The report provides comparative insight into each model’s relative safety strengths and weaknesses, helping teams understand where gaps exist and where additional resources may be required.
What’s new from Alice
LIVE from Black Hat Las Vegas: AI, Nation-States, and the Battlefield That Keeps Changing
What if the biggest threat to your security team isn't the attacker, it's the model you're relying on to stop them? LIVE from Black Hat Las Vegas, Mo and Madi bring together two cybersecurity authors, Caroline Wong, Chief Strategy Officer at Axari and author of The AI Cybersecurity Handbook, and Allie Mellen, Principal Analyst at Forrester and author of Code War, who wrote very different books that turn out to be arguing the same point. One explains why nations attack the way they do. The other explains why AI just changed the cost, speed, and scale of everything. Tune in!
It Takes AI to Break AI: The Case for AI Red Teaming
As AI systems gain autonomy, organizations need security approaches built specifically for AI behavior. Learn why AI-driven red teaming is becoming a critical defense layer.
5 Ways Your Third-Party CX Agent Gets Broken
Third-party CX agents create hidden liability. Learn the 5 attack patterns vendors miss and how WonderSuite closes the gap.