Senior GenAI Safety Researcher
About the Position
Alice is seeking a driven, detail-focused Senior Generative AI Researcher to take on a leading role within our US team. In this position, you will operate at the cutting edge of AI Safety and Trust & Safety, analyzing potential vulnerabilities and content safety risks across the newest wave of Generative AI tools.
As a Senior Researcher, you won't just run tests, you will design robust testing methodologies, act as a core content expert to support the Program Lead, and actively expand the team’s internal knowledge base. You will partner closely with cross-functional teams and external stakeholders to secure models across multiple modalities, including LLMs, Text-to-Image, Text-to-Video, and AI Agents.
Key Responsibilities
Methodology & Strategy
- Architect rigorous, scalable testing methodologies and red-teaming frameworks to evaluate foundational models, multimodal systems, and AI agents.
- Develop sophisticated prompt strategies across diverse risk domains (e.g., Hate Speech, Misinformation, IP & Copyright infringement, Child Safety) to expose complex model vulnerabilities.
- Conduct ongoing research into emerging jailbreak tactics, prompt injection techniques, and novel circumvention strategies used against foundational safety measures.
Subject-Matter Expertise
- Serve as a trusted content and domain expert, providing deep technical and policy insight to support the Program Lead in scoping projects, assessing risks, and driving strategy.
- Lead efforts to continuously document, synthesize, and expand Alice’s internal AI Safety knowledge base, standardizing best practices, taxonomies, and research findings across the team.
- Mentor junior analysts, foster a culture of continual learning, and elevate the team’s analytical standards.
Operational Excellence
- Own engagement lifecycles from initial planning and methodology design through execution, quality assurance (QA), and final delivery.
- Oversee complex, multi-language datasets across multiple areas of abuse, ensuring the highest precision, accuracy, and output quality.
- Partner effectively with engineering, product, policy, and client-facing teams to communicate research findings and inform mitigation strategies.
Requirements
Must-Have
- 5+ years of experience in AI Safety, Responsible AI, Trust & Safety, or aligned research domains.
- Proven expertise in research design and building qualitative or quantitative evaluation methodologies for GenAI.
- Strong domain expertise in content risks (e.g., toxicity, copyright, misinformation, safety policy violations).
- Track record of project ownership, leading deliverables end-to-end with high attention to detail in fast-paced, variable environments.
- Deep familiarity with modern Generative AI architectures, prompt engineering, red-teaming, and AI agents.
- Strong communication skills to act as a core subject-matter contact for program leads, internal teams, and clients.
Nice-to-Have
- Proven track record of published research in academia, industry whitepapers, or a research institute.
- Hands-on experience evaluating multimodal systems (Text-to-Image, Text-to-Video, Audio).
- Experience mentoring, leading, or QAing the work of junior analysts and researchers.
The salary range for this role is $105K - $115K OTE - Range may vary based on experience. Salary at the time of offer will be commensurate with experience.
About Alice
THE CHALLENGES ALONG THE WAY
1. Being Both Strategist and Executioner
One of the hardest parts of this role is that you’re both the visionary and the builder; the one drawing the map and paving the road.
That means switching between high-level strategy and hands-on experimentation daily, and doing it while bringing others along with you. There’s no playbook for this kind of work. You’re paving an unpaved road, one small experiment at a time.
2. Balancing Security and Innovation
ActiveFence is the leading provider of security and safety solutions for online experiences, safeguarding more than 3 billion users, top foundation models, and the world’s largest enterprises and tech platforms every day.
As a trusted ally to major technology firms and Fortune 500 brands that build user-generated and GenAI products, ActiveFence empowers security, AI, and policy teams with low-latency Real-Time Guardrails and a continuous Red Teaming program that pressure-tests systems with adversarial prompts and emerging threat techniques. Powered by deep threat intelligence, unmatched harmful-content detection, and coverage of 117+ languages, ActiveFence enables organizations to deliver engaging and trustworthy experiences at global scale while operating safely and responsibly across all threat landscapes.
