engineering
Posted Jun 20Hardware Operations Engineer
at openai
San Francisco, United StatesOn-site
Responsibilities
- Responsibilities - Drive technical triage and resolution of complex hardware failures impacting production systems.
- - Lead root cause analysis (RCA) efforts for critical hardware incidents and develop corrective and preventive action plans.
- - Collaborate with Cloud Service Provider operations teams and OEM vendors to coordinate repairs, replacements, upgrades, and hardware lifecycle activities.
- - Support new hardware introductions, validation activities, and production readiness reviews.
- - Develop scalable operational standards and best practices that can be deployed across future Stargate campuses.
- - Mentor on-site technicians and partner teams on advanced troubleshooting methodologies and hardware operational excellence.
Requirements
- About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem.
- Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads.
- We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments.
- As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses.
- About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses.
- You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments.
- The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills.
- experience supporting large-scale datacenter hardware infrastructure, with
- experience in a senior technician, sustaining engineering, or hardware operations leadership role. - Deep expertise with server platforms, GPU systems, storage infrastructure, rack integration, and datacenter hardware architecture. - Strong
- Experience conducting root cause analysis and driving long-term corrective actions.
- - Strong understanding of hardware reliability engineering principles and fleet-health management.
- - Proven ability to partner effectively across engineering, operations, manufacturing, and vendor organizations.
- - Excellent written and verbal communication skills with the ability to influence technical and operational decisions. -
- Experience developing operational processes, maintenance standards, and technical documentation. - Ability to travel occasionally to support new campus deployments and operational readiness activities. PREFERRED QUALIFICATIONS -
- Experience supporting large-scale GPU clusters or AI/ML infrastructure environments. - Familiarity with fleet health systems, telemetry platforms, and hardware monitoring tools. - Ability to identify appropriate data and perform detailed analysis to support all elements of this role, including dashboard development -
- Experience with failure analysis methodologies such as FRACAS, RCCA, 5-Why, Fishbone, or FMEA. - Knowledge of Linux system administration and hardware validation workflows. -
- Experience supporting hyperscale datacenter operations or HPC environments. - Familiarity with server manufacturing, rack integration, and NPI-to-sustaining transitions. - Industry certifications such as CompTIA Server+, OEM hardware certifications, or equivalent experience. -
- Experience applying Environmental Health and Safety (EHS) practices in mission-critical datacenter environments.
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Experience
- Qualifications - 8+ years of
Additional details
- This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability.
- You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems.
- Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts.
- - Partner with Fleet Health Engineering to investigate recurring hardware issues, identify failure patterns, and improve fleet reliability.
- - Establish and continuously improve hardware maintenance procedures, operational runbooks, and troubleshooting standards.
- - Analyze hardware failure trends and operational metrics to identify reliability risks and improvement opportunities.
- - Coordinate spare parts strategy and inventory planning with supply chain and site teams.
- - Partner with Hardware Engineering, Manufacturing, and Infrastructure teams to provide field feedback that improves future platform designs.
- experience diagnosing complex hardware failures and leading repair efforts in production environments. -
- - Comfortable operating independently in high-priority production environments with significant operational responsibility.