jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

engineering

Posted Jun 20

Hardware Operations Engineer

at openai

San Francisco, United StatesOn-site

Responsibilities

  • Responsibilities - Drive technical triage and resolution of complex hardware failures impacting production systems.
  • - Lead root cause analysis (RCA) efforts for critical hardware incidents and develop corrective and preventive action plans.
  • - Collaborate with Cloud Service Provider operations teams and OEM vendors to coordinate repairs, replacements, upgrades, and hardware lifecycle activities.
  • - Support new hardware introductions, validation activities, and production readiness reviews.
  • - Develop scalable operational standards and best practices that can be deployed across future Stargate campuses.
  • - Mentor on-site technicians and partner teams on advanced troubleshooting methodologies and hardware operational excellence.

Requirements

  • About the Team OpenAI, in close collaboration with our capital partners, is building the world’s most advanced AI infrastructure ecosystem.
  • Our Industrial Compute organization develops and deploys large-scale AI campuses designed to support the next generation of frontier model training and inference workloads.
  • We partner closely with Data Center Operations, Fleet Health Engineering, Manufacturing, Network Infrastructure, Capacity Planning, and our infrastructure partners to maintain world-class operational performance across rapidly expanding AI environments.
  • As we scale globally, we are building the operational frameworks, reliability standards, and sustaining engineering practices required to support thousands of GPUs and servers across multiple campuses.
  • About the Role We are seeking a Datacenter Hardware Technician Lead to serve as the senior on-site technical authority for hardware reliability and fleet health at one of OpenAI’s flagship AI campuses.
  • You will help establish hardware maintenance standards, operational procedures, and best practices that scale across future OpenAI infrastructure deployments.
  • The ideal candidate combines deep hands-on datacenter hardware expertise with strong troubleshooting, failure analysis, and cross-functional leadership skills.
  • experience supporting large-scale datacenter hardware infrastructure, with
  • experience in a senior technician, sustaining engineering, or hardware operations leadership role. - Deep expertise with server platforms, GPU systems, storage infrastructure, rack integration, and datacenter hardware architecture. - Strong
  • Experience conducting root cause analysis and driving long-term corrective actions.
  • - Strong understanding of hardware reliability engineering principles and fleet-health management.
  • - Proven ability to partner effectively across engineering, operations, manufacturing, and vendor organizations.
  • - Excellent written and verbal communication skills with the ability to influence technical and operational decisions. -
  • Experience developing operational processes, maintenance standards, and technical documentation. - Ability to travel occasionally to support new campus deployments and operational readiness activities. PREFERRED QUALIFICATIONS -
  • Experience supporting large-scale GPU clusters or AI/ML infrastructure environments. - Familiarity with fleet health systems, telemetry platforms, and hardware monitoring tools. - Ability to identify appropriate data and perform detailed analysis to support all elements of this role, including dashboard development -
  • Experience with failure analysis methodologies such as FRACAS, RCCA, 5-Why, Fishbone, or FMEA. - Knowledge of Linux system administration and hardware validation workflows. -
  • Experience supporting hyperscale datacenter operations or HPC environments. - Familiarity with server manufacturing, rack integration, and NPI-to-sustaining transitions. - Industry certifications such as CompTIA Server+, OEM hardware certifications, or equivalent experience. -
  • Experience applying Environmental Health and Safety (EHS) practices in mission-critical datacenter environments.
  • About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence

Experience

  • Qualifications - 8+ years of

Additional details

  • This role operates at the intersection of hardware operations, sustaining engineering, and fleet reliability.
  • You will partner closely with Cloud Service Provider operations teams, OpenAI fleet-health engineers, hardware engineering teams, and OEM vendors to identify, diagnose, and resolve hardware issues affecting production systems.
  • Beyond day-to-day operational support, you will drive root cause investigations, reliability improvement initiatives, lifecycle management programs, and operational readiness efforts.
  • - Partner with Fleet Health Engineering to investigate recurring hardware issues, identify failure patterns, and improve fleet reliability.
  • - Establish and continuously improve hardware maintenance procedures, operational runbooks, and troubleshooting standards.
  • - Analyze hardware failure trends and operational metrics to identify reliability risks and improvement opportunities.
  • - Coordinate spare parts strategy and inventory planning with supply chain and site teams.
  • - Partner with Hardware Engineering, Manufacturing, and Infrastructure teams to provide field feedback that improves future platform designs.
  • experience diagnosing complex hardware failures and leading repair efforts in production environments. -
  • - Comfortable operating independently in high-priority production environments with significant operational responsibility.

Find more real-time jobs on JobLoom.