infrastructure
Posted Feb 5Software Engineer, GPU Infrastructure - HPC
at openai
San Francisco, United StatesOn-site
Requirements
- We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency.
- Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT.
- We prioritize safety, reliability, and responsible AI deployment over unchecked growth.
- Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change.
- Experience managing large-scale server environments. - A balance of strengths in building and operationalizing. - Proficiency in Python, Go, or similar languages. - Strong Linux, networking, and server hardware knowledge. - Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.
- Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning) - Knowledge of hardware management protocols (e.g., IPMI, Redfish). - High-performance computing (HPC) or distributed systems experience. - Prior
- experience developing, managing, or designing hardware. - Familiarity with monitoring tools (e.g., Prometheus, Grafana).
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Benefits
- Prior hardware expertise is not required for this role. Bonus Skills: -
Additional details
- About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development.
- About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet.
- Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions.
- With increasingly large supercomputers, the stakes continue to rise.
- Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale.
- This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure.
- This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions.
- We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role,