AI Red Team – GenAI Adversarial Evaluation
About the Position
We are looking for an AI Red Team Gym Creator to design and scale adversarial Evaluation for LLM and AI systems.
The role will focus on creating and reviewing indirect prompt injection attacks, jailbreaks, agent manipulation, tool-use attacks, and other adversarial techniques in controlled and authorized environments.
Responsibilities
- Create novel adversarial attacks.
- Help build scalable systems to generate and test attacks at large scale.
- Review attack quality, effectiveness, novelty, and reproducibility.
- Identify new attack patterns and model vulnerabilities.
- Develop adversarial datasets, benchmarks, and regression tests.
- Share findings, techniques, and structured attack data with the Data Science, Security, and Engineering teams.
- Help improve model robustness and defensive capabilities.
Requirements
- 3+ years of professional experience in cybersecurity, security research, red teaming, or a related security field.
- Bachelor’s degree in Data Science, Computer Engineering, Computer Science, or a closely related technical field.
- Strong understanding of LLMs, prompt injection, and AI security.
- Experience with red teaming, security research, or adversarial testing.
- Strong Python and automation skills.
- Creative problem-solving and ability to develop new attack techniques.
- Ability to analyze results and communicate findings clearly.
Preferred Skills That Set You Apart
- Certifications in offensive cybersecurity (e.g., OSWA, OSWE, OSCE3, SEC542, SEC522)
About Alice
THE CHALLENGES ALONG THE WAY
1. Being Both Strategist and Executioner
One of the hardest parts of this role is that you’re both the visionary and the builder; the one drawing the map and paving the road.
That means switching between high-level strategy and hands-on experimentation daily, and doing it while bringing others along with you. There’s no playbook for this kind of work. You’re paving an unpaved road, one small experiment at a time.
2. Balancing Security and Innovation
ActiveFence is the leading provider of security and safety solutions for online experiences, safeguarding more than 3 billion users, top foundation models, and the world’s largest enterprises and tech platforms every day.
As a trusted ally to major technology firms and Fortune 500 brands that build user-generated and GenAI products, ActiveFence empowers security, AI, and policy teams with low-latency Real-Time Guardrails and a continuous Red Teaming program that pressure-tests systems with adversarial prompts and emerging threat techniques. Powered by deep threat intelligence, unmatched harmful-content detection, and coverage of 117+ languages, ActiveFence enables organizations to deliver engaging and trustworthy experiences at global scale while operating safely and responsibly across all threat landscapes.