jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

engineering

Posted May 23

Software Engineer, RL Training Infra

at openai

San Francisco, United StatesRemote

Requirements

  • ABOUT THE TEAM The Post-Training Frontiers team is responsible for training the frontier agents OpenAI ships to the world (GPT-Next).
  • We train the flagship agentic models behind Codex, ChatGPT, and the API through large-scale reinforcement learning.
  • Second, RL scaling: executing the final large-scale reinforcement learning run, making sure GPUs are used efficiently and training stays healthy.
  • experience in some layer of ML infrastructure.
  • Experience supporting large-scale model training, async RL systems, or high-throughput ML infrastructure. -
  • Experience debugging distributed systems across GPUs, networking, orchestration, or inference stacks. - A background in performance optimization, scaling, or production-critical infrastructure. -
  • Experience directly supporting researchers or fast-moving model teams. About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence

Additional details

  • First, execution and science: working with teams across OpenAI to decide what can go into the final model and how, using scientific experiments and evals that are representative of the final pipeline so issues can be recognized early.
  • Third, research: improving horizontal capabilities like instruction following, factuality, memory, and multi-agent behavior, where the team’s broad visibility helps identify cross-cutting improvements across teams and domains.
  • Fourth, engineering: maintaining the infrastructure stack and internal tools to ensure that both the final run and all integrations go as smoothly as possible and that the systems are easy to work with.
  • ABOUT THE ROLE This role focuses on keeping our frontier RL training runs fast, reliable, and unblocked.
  • You will work across engineering and infrastructure problems as they emerge, from scaling and orchestration issues to inference bottlenecks, numerical problems, and hardware failures, as well as supporting large horizontal integrations in the big run, like multi-agent capabilities or memory.
  • This is a role for a strong generalist who quickly learns anything needed for the task, has high attention to detail, debugs deeply, and is motivated by fixing the highest-impact problem in front of the team. IN THIS ROLE,
  • YOU WILL: - Keep large-scale async RL training runs moving by jumping into the most urgent engineering and infrastructure problems.
  • - Debug issues across training systems, inference, orchestration, scaling, and distributed infrastructure.
  • - Improve the reliability and efficiency of RL training runs.
  • - Help researchers who are developing infrastructure-heavy integrations, such as multi-agent capabilities or memory.

Find more real-time jobs on JobLoom.