engineering
Posted Oct 29, 2025Training: ML Framework Engineer
at openai
San Francisco, United StatesHybrid
Requirements
- About the Team Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs.
- About the Role As a Training: ML Framework Engineer, you will work on improving the training throughput for our internal training framework, while enabling researchers to experiment with new ideas.
- This requires good engineering (for example designing, implementing, and optimizing state-of-the-art AI models), writing bug-free machine learning code (surprisingly difficult!), and acquiring deep knowledge of the performance of supercomputers.
- Since our training framework is used for large runs with massive numbers of GPUs, performance improvements here will have a large impact.
- you will: - Apply the latest techniques in our internal training framework to achieve impressive hardware efficiency for our training runs - Profile and optimize our training framework - Work with researchers to enable them to develop the next generation of models You might thrive in this role if you: - Have run small scale ML experiments - Love figuring out how systems work and continuously come up with ideas for how to make them faster while minimizing complexity and maintenance burden - Have strong
Additional details
- With a dual mandate to accelerate researchers and enable frontier scale, we’re building a unified, modular runtime that meets researchers where they are and moves with them up the scaling curve.
- Our work focuses on three pillars: high-performance, asynchronous, zero-copy tensor and optimizer-state-aware data movement; performant, high-uptime, fault-tolerant training frameworks (training loop, state management, resilient checkpointing, deterministic orchestration, and observability); and distributed process management for long-lived, job-specific and user-provided processes.
- We integrate proven large-scale capabilities into a composable, developer-facing runtime so teams can iterate quickly and run reliably at any scale, partnering closely with model-stack, research, and platform teams.
- Success for us is measured by raising both training throughput (how fast models train) and researcher throughput (how fast ideas become experiments and products).
- In all the projects this role pursues, the ultimate goal is to push the field forward.
- We’re looking for people who love optimizing performance, understanding distributed systems, and who cannot stand having bugs in their code.
- We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role,