engineering
Posted Nov 19, 2025Software Engineer, Fleet Management
at openai
San Francisco, United StatesHybrid
Responsibilities
- - Develop tools that integrate low-level hardware metrics with high-level job scheduling and cluster management algorithms.
- - Automate infrastructure processes, reducing repetitive toil and improving system reliability.
- - Collaborate with hardware, infrastructure, and research teams to ensure seamless integration across the stack.
Requirements
- We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency.
- Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT.
- We prioritize safety, reliability, and responsible AI deployment over unchecked growth.
- You will design and develop solutions that integrate individual nodes and servers into unified clusters, directly contributing to advancing AI research by streamlining the overall research user experience.
- experience in large-scale infrastructure environments.
- - Possess broad knowledge of cluster-level systems (e.g., Kubernetes, CI/CD pipelines, Terraform, cloud providers).
- - Have deep expertise in server-level systems (e.g., systems, containerization, Chef, Linux kernels, firmware management, host routing).
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Additional details
- The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development.
- About the Role The Software Engineer, Operating Systems & Orchestration will focus on building systems to manage hardware, configurations, vendors, and the people interacting with our infrastructure.
- We use a hybrid work model of 3 days in the office per week and offer relocation assistance to new employees. In this role,
- you will: - Design and build systems to manage both cloud and bare-metal fleets at scale.
- - Leverage LLMs to coordinate vendor operations and optimize infrastructure workflows.
- - Continuously improve tools, automation, processes, and documentation to enhance operational efficiency.
- You might thrive in this role if you: - Have strong software engineering skills with
- - Are passionate about optimizing the performance and reliability of large compute fleets.
- - Thrive in dynamic environments and are eager to solve complex infrastructure challenges.
- - Value automation, efficiency, and continuous improvement in everything you build.