infrastructure
Posted Sep 18, 2025Software Engineer, Data Infrastructure - Research
at openai
San Francisco, United StatesOn-site
Responsibilities
- - Build proactive testing and scale validation pipelines for dataset loading at GPU scale.
- - Collaborate with teammates to integrate datasets seamlessly into training and inference pipelines, ensuring smooth adoption and a great user experience.
Requirements
- ABOUT THE TEAM The Workload team is responsible for designing and running OpenAI’s LLM training and inference infrastructure that powers frontier models at massive scale.
- Our systems unify how researchers train and serve models, abstracting away the complexity of performance, parallelism, and execution across vast GPU/accelerator fleets.
- By providing this foundation, the Workload team ensures that researchers can focus on advancing model capabilities while we handle the scale, efficiency, and reliability required to bring those models to life.
- You will be responsible for building standardized dataset interfaces, scaling pipelines across thousands of GPUs, and proactively testing performance bottlenecks.
- YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have strong engineering fundamentals with
- experience in distributed systems, data pipelines, or infrastructure. - Have
- experience building APIs, modular code, and scalable abstractions, while recognizing that abstractions ultimately serve the users and UX is an important part of the abstractions design. - Are comfortable debugging bottlenecks across large fleets of machines. - Take pride in building infrastructure that “just works,” and find joy in being the guardian of reliability and scale. - Are collaborative, humble, and excited to own a foundational (if not glamorous) part of the ML stack.
- Bonus points if you: - Have background knowledge in data math, probability, or distributed data theory. - Have worked with GPU-scale distributed systems or dataset scaling for real-time data About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Additional details
- In this role, you will collaborate closely with the multimodal researchers, and other infra groups to ensure datasets are unified, efficient, and easy to consume. IN THIS ROLE,
- YOU WILL: - Design and maintain standardized dataset APIs, including for multimodal (MM) data that cannot fit in memory.
- - Document and maintain dataset interfaces so they are discoverable, consistent, and easy for other teams to adopt.
- - Establish safeguards and validation systems to ensure datasets remain reproducible and unchanged once standardized.
- - Debug and resolve performance bottlenecks in distributed dataset loading (e.g., straggler systems slowing global training).
- - Provide visualization and inspection tools to surface errors, bugs, or bottlenecks in datasets.