Just postedUrgently hiring Use left and right arrow keys to navigate
Based on similar jobs in your market
Estimated Pay info$17 per hour
Hours Full-time
Location Topeka, Kansas

About this job

  • Build long-horizon RL environments and tasks for agent training, spanning many steps and hours of realistic effort rather than single-shot prompts

  • Shape environments end to end: stateful, resumable systems with snapshotting, checkpointing, and branching rollouts, at multi-node scale where needed

  • Generate and refine ideas for tasks and agentic trajectories across multi-turn, tool-using, and computer-use agent loops with persistent state across turns

  • Design reward structure for sparse-reward settings: milestone and process rewards, subgoal and task decomposition, and credit assignment across long trajectories

  • Validate that the work is correct, hard, and covers the right edge cases, including rubric design for partial credit, contamination avoidance, and defenses against reward hacking

  • Design curricula that ramp task difficulty rather than shipping fixed-difficulty tasks

  • Build automated pipelines that curate high-quality long-horizon RL environments and tasks at scale

  • Write reliable, well-tested Python infrastructure rather than one-off research scripts

  • Work directly with our research team with high ownership, building novel systems rather than maintaining legacy ones


Nearby locations

Posting ID: 1292282274 Posted: 2026-09-01 Job Title: Long Horizon