infrastructure
Posted 1 weeks agoData Center Compute Infrastructure
at openai
San Francisco, United StatesRemote
Responsibilities
- - Build tools, processes, systems, or infrastructure that improve execution at scale.
Requirements
- About the Team OpenAI’s Compute organization turns ambitious AI research into real-world capability by delivering the compute infrastructure behind our most advanced models.
- As the demand for frontier AI grows, so does the complexity of the systems required to support it.
- Scaling this infrastructure means solving problems that cut across distributed systems, ML infrastructure, GPU fleets, power, cooling, networking, manufacturing, supply chain, and data center delivery.
- Our work is focused on expanding the compute foundation that enables OpenAI to train more capable models, including systems like GPT-5.6, and make frontier AI available to more people, products, and workflows.
- We’re looking for exceptional people across many disciplines to help build the next generation of AI infrastructure at a scale few organizations have attempted.
- About the Role We are hiring across a broad range of roles to help design, build, scale, and operate OpenAI’s compute infrastructure.
- Depending on your background, you may work on large-scale distributed systems, ML infrastructure, hardware systems, manufacturing, supply chain, data center development, or the physical engineering systems required to bring massive compute capacity online.
- This is an opportunity to work on one of the most important infrastructure challenges in AI: building the compute foundation required to train and serve increasingly capable frontier models. Key
- - Want your work to directly support the development and deployment of frontier AI.
- experience with AI infrastructure, high-performance computing, distributed systems, GPU clusters, or cloud-scale platforms.
- experience operating in fast-moving environments where technical depth and execution speed both matter.
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Additional details
- The team works across software, hardware, facilities, operations, and engineering disciplines to make enormous amounts of compute available, reliable, and efficient.
- You’ll work with teams across research, engineering, hardware, operations, and infrastructure to solve high-impact problems at extraordinary scale.
- This may include improving system reliability, accelerating deployment timelines, increasing operational efficiency, designing new infrastructure, or helping bring new compute platforms and facilities from concept to production.
- Responsibilities - Help build, scale, and operate OpenAI’s global compute infrastructure.
- - Solve complex problems across software, hardware, manufacturing supply chain, and data center systems.
- - Improve the reliability, performance, efficiency, and scalability of critical infrastructure.
- - Partner with cross-functional teams to bring new compute capacity online quickly and reliably.
- - Identify bottlenecks across technical, operational, and physical systems, and develop practical solutions.
- - Contribute to the long-term architecture and operational maturity of OpenAI’s compute footprint.
- experience building, scaling, or operating complex technical systems.