RL Environments Engineer
| Hours | Full-time |
|---|---|
| Location | San Francisco, CA San Francisco, California open_in_new |
About this job
Job Description
San Francisco, California | Primarily On-site
We are seeking an Agent Evaluation Infrastructure Engineer to build the environments, evaluation systems, and supporting infrastructure used to train and assess long-horizon enterprise AI agents.
The OpportunityYou will work on the engineering and research problems behind realistic agent environments, post-training systems, and reliable evaluation of complex multi-step workflows.
Key Responsibilities- Design evaluation environments for long-horizon enterprise agent workflows.
- Define tasks, state, tools, graders, and reward signals used to evaluate and improve agents.
- Build high-fidelity representations of complex enterprise software environments.
- Develop infrastructure for rollouts, orchestration, trajectory inspection, and grader pipelines.
- Measure both correctness and efficiency across multi-step agent behavior.
- Investigate evaluation failures, reward-quality issues, and agent behavior.
- Build production-quality systems rather than notebook-only research prototypes.
- Hands-on experience with AI environments, evaluations, reinforcement learning infrastructure, or related agent-training systems.
- Strong software engineering fundamentals.
- Demonstrated ability to build and ship technical infrastructure.
- Understanding of evaluation methodology, reward design, graders, and agent trajectories.
- Ability to work across languages and technology stacks based on system requirements.
A PhD is not required. Strong engineering and shipped environment or evaluation systems are more important than academic credentials or publication history.
SeniorityThe opportunity is open to exceptional new graduates, early-career engineers, and experienced senior candidates. Selection is based primarily on engineering strength and relevant technical work.
Work ArrangementThe role is anchored in San Francisco with a strong preference for in-person collaboration. Limited flexibility may be considered case by case for exceptional candidates.