jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

infrastructure

Posted Feb 5

Software Engineer, GPU Infrastructure - HPC

at openai

San Francisco, United StatesOn-site

Requirements

  • We oversee large-scale systems that span data centers, GPUs, networking, and more, ensuring high availability, performance, and efficiency.
  • Our work enables OpenAI’s models to operate seamlessly at scale, supporting both internal research and external products like ChatGPT.
  • We prioritize safety, reliability, and responsible AI deployment over unchecked growth.
  • Our team empowers strong engineers with a high degree of autonomy and ownership, as well as ability to effect change.
  • Experience managing large-scale server environments. - A balance of strengths in building and operationalizing. - Proficiency in Python, Go, or similar languages. - Strong Linux, networking, and server hardware knowledge. - Comfort digging into noisy data with SQL, PromQL, and Pandas or any other tool.
  • Experience with low level details of hardware components, protocols, and associated Linux tooling (e.g., PCIe, Infiniband, networking, power management, kernel perf tuning) - Knowledge of hardware management protocols (e.g., IPMI, Redfish). - High-performance computing (HPC) or distributed systems experience. - Prior
  • experience developing, managing, or designing hardware. - Familiarity with monitoring tools (e.g., Prometheus, Grafana).
  • About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence

Benefits

  • Prior hardware expertise is not required for this role. Bonus Skills: -

Additional details

  • About the team The Fleet team at OpenAI supports the computing environment that powers our cutting-edge research and product development.
  • About the role As a software engineer on the Fleet High Performance Computing (HPC) team, you will be responsible for the reliability and uptime of all of OpenAI’s compute fleet.
  • Minimizing hardware failure is key to research training progress and stable services, as even a single hardware hiccup can cause significant disruptions.
  • With increasingly large supercomputers, the stakes continue to rise.
  • Being at the forefront of technology means that we are often the pioneers in troubleshooting these state-of-the-art systems at scale.
  • This is a unique opportunity to work with cutting-edge technologies and devise innovative solutions to maintain the health and efficiency of our supercomputing infrastructure.
  • This role will require a keen focus on system-level comprehensive investigations and the development of automated solutions.
  • We want people who go deep on problems, investigate as thoroughly as possible, and build automation for detection and remediation at scale. In this role,

Find more real-time jobs on JobLoom.