engineering
Posted Oct 31, 2025Software Engineer, Hardware
at openai
San Francisco, United StatesHybrid
Responsibilities
- - Develop simulation infrastructure to validate runtime behaviors, test training stack changes, and support early-stage hardware and system development.
Requirements
- ABOUT THE TEAM OpenAI’s Hardware organization develops AI-native silicon and system-level solutions for the unique demands of advanced AI workloads.
- Building on efforts like Jalapeño, the team is developing future generations of AI-native silicon and tightly integrated systems to power the next generation of frontier models.
- By co-designing chips, systems, tools, and methodologies, the team helps deliver faster, more efficient, and production-ready hardware for OpenAI’s supercomputing platform.
- You will work at the intersection of systems programming, ML infrastructure, and high-performance computing, helping to create both ergonomic developer APIs and highly efficient runtime systems.
- YOU WILL: - Design and build APIs and runtime components to orchestrate computation and data movement across heterogeneous ML workloads.
- - Work across a diverse stack, primarily using Rust and Python, with opportunities to influence architecture decisions across the training framework.
- YOU MIGHT THRIVE IN THIS ROLE IF YOU: - Have a deep curiosity for how large-scale systems work and enjoy making them faster, simpler, and more reliable. - Are proficient in systems programming (e.g., Rust, C++) and scripting languages like Python. - Have
- experience in one or more of the following areas: compiler development, kernel authoring, accelerator programming, runtime systems, distributed systems, or high-performance simulation. - Are excited to work in a fast-paced, highly collaborative environment with evolving hardware and ML system demands. - Value engineering excellence, technical leadership, and thoughtful system design.
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Additional details
- ABOUT THE ROLE As a software engineer on the Scaling team, you’ll help build and optimize the low-level stack that orchestrates computation and data movement across OpenAI’s supercomputing clusters.
- Your work will involve designing high-performance runtimes, building custom kernels, contributing to compiler infrastructure, and developing scalable simulation systems to validate and optimize distributed training workloads.
- This means balancing ease of use and introspection with the need for stability and performance on our evolving hardware fleet.
- This role is based in San Francisco, CA, with a hybrid work model (3 days/week in-office).
- - Contribute to compiler infrastructure, including the development of optimizations and compiler passes to support evolving hardware.
- - Profile and optimize system bottlenecks, especially around I/O, memory hierarchy, and interconnects, at both local and distributed scales.
- - Rapidly deploy runtime and compiler updates to new supercomputing builds in close collaboration with hardware and research teams.