Senior AI Researcher
About the Position
We are seeking a Senior AI Researcher to lead post-training evaluation, red-teaming, and reinforcement learning (RL) gym audits on open-weight models. The ideal candidate will establish rigorous benchmarking methodologies, evaluate large language models (LLMs) against complex threats like Indirect Prompt Injections (IPI), and construct post-training evaluation pipelines that accurately measure realistic frontier-level security capabilities.
Key Responsibilities
- RL Post-Training & Benchmarking: Execute post-training runs (e.g. GRPO) using mainstream open-weight generalist models against security-focused RL environments, targeting threat vectors like Indirect Prompt Injection (IPI). Reward Diagnostics & Trace Analysis - Analyze live loss curves and rollout traces to identify reward hacking, lazy policy convergence, and flawed or over/under-specified verifiers.
- Task & Environment Auditing: Review tasks and multi-turn environments (including tool use, web navigation, and computer use) for realism, threat model accuracy, data distribution, and dataset balance.
- Performance Reporting (Gym Cards): Generate comprehensive evaluation cards detailing hill-climbing performance uplift across checkpoints, failure modes, tokens/turns per rollout, and task-level success rates.
- Integration & Orchestration: Integrate dockerized environments (e.g., Harbor format) into internal training frameworks, optimizing reset/statefulness semantics, concurrency, and throughput ceilings.
Requirements
Required Qualifications
- Technical Background: M.S. or Ph.D. in Data Science, Machine Learning, Computer Science, or equivalent practical experience in deep learning.
- RL & Post-Training Expertise: Strong hands-on experience training large-scale models using RL algorithms (e.g. GRPO, PPO) on open-weight architectures.
- AI Security Expertise: Solid understanding of LLM vulnerabilities, red-teaming methodologies, and defensive alignment against IPI attacks.
- Infrastructure Skills: Proficiency in PyTorch, Docker containerization, and distributed training architectures.
- Diagnostic Skills: Ability to analyze agent rollout traces, craft deterministic rubrics/verifiers, and debug complex reward shaping flaws.
Preferred Qualifications
- Prior experience working with standard RL gym formats, such as Harbor.
- Experience evaluating complex agentic workflows in tool-use or web-browser environments.
- Familiarity with evaluating open-weight models similar to Llama or Mistral against adversarial workloads.
About Alice
THE CHALLENGES ALONG THE WAY
1. Being Both Strategist and Executioner
One of the hardest parts of this role is that you’re both the visionary and the builder; the one drawing the map and paving the road.
That means switching between high-level strategy and hands-on experimentation daily, and doing it while bringing others along with you. There’s no playbook for this kind of work. You’re paving an unpaved road, one small experiment at a time.
2. Balancing Security and Innovation
ActiveFence is the leading provider of security and safety solutions for online experiences, safeguarding more than 3 billion users, top foundation models, and the world’s largest enterprises and tech platforms every day.
As a trusted ally to major technology firms and Fortune 500 brands that build user-generated and GenAI products, ActiveFence empowers security, AI, and policy teams with low-latency Real-Time Guardrails and a continuous Red Teaming program that pressure-tests systems with adversarial prompts and emerging threat techniques. Powered by deep threat intelligence, unmatched harmful-content detection, and coverage of 117+ languages, ActiveFence enables organizations to deliver engaging and trustworthy experiences at global scale while operating safely and responsibly across all threat landscapes.