infrastructure
Posted 4 days agoSimulation Infrastructure Engineer
at openai
San Francisco, United StatesOn-site
Responsibilities
- - Implement end-to-end automation to run model evaluation in sim (SIL) and orchestrate HIL runs; compute realism and task metrics, generate dashboards and alerts, and ensure evaluation is repeatable and auditable.
- - Create robust APIs and connectors so research, training, and data-collection systems can schedule, seed, and evaluate batches of simulations; support RL rollouts, imitation-data collection, and presubmit model checks.
- - Build scheduling, batching and orchestration for running very large numbers of concurrent rollouts (target tens of thousands of rollouts / large RL workloads), solve engine-level scaling (parallelization, batching multiple runs per engine), and optimize cloud/GPU runtime reliability.
- - Produce metrics and tooling for measuring simulation health, throughput, fidelity regressions, and cost; create presubmit / canary tests that catch sim regressions early.
- - Implement artifact versioning, environment immutability (images / asset versions), experiment provenance, and policies for resource quotas and cost control across the sim farm.
Requirements
- About the Team Our Robotics team is focused on unlocking general-purpose robotics and pushing towards AGI-level intelligence in dynamic, real-world settings.
- We strive to seamlessly blend high-level AI capabilities with the constraints of physical systems to improve peoples’ lives.
- You will collaborate closely with Sim Realism, Sim Environments, Research, and Ops to make simulation an integrated, reproducible, and measurable part of our ML and robotics workflows.
- - Are comfortable with distributed compute and cloud GPU workloads: you know how to get many sims running concurrently (scheduling, batching, GPU orchestration) and optimize throughput/cost.
- - Have built or maintained HIL/SIL workflows or other sim↔hardware integrations and understand the operational challenges of bridging software and hardware testbeds.
- experience with Python/C++/Rust, container orchestration (Kubernetes), distributed task queues, and CI systems; bonus if you’ve worked with RL tooling, task generators, or large-scale data pipelines. - Enjoy collaborating across teams to turn experimental simulation work into dependable production tooling.
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Additional details
- About the Role We are hiring a Sim Infrastructure Engineer to turn simulation systems into reliable, automated, production-quality pipelines that power model training, evaluation, and hardware-in-the-loop validation.
- This role owns the automation, orchestration, and tool integration that apply simulation to concrete robotics tasks: building CI/CD for SIL/HIL, presubmit checks, automatic model evaluation, metric computation and reporting, and the runtime infrastructure to run simulations at scale.
- This role is based in San Francisco, CA, and requires in-person 4 days a week. In this role,
- you will: - Build and maintain presubmit checks, continuous integration and deployment pipelines for simulation code, environments, and tasks so simulation artifacts are testable, versioned, and reproducible.
- - Work closely with Sim Environments, Sim Realism, research, and ops to close the loop—ensuring simulation improvements directly translate into better model evaluation and training results.
- You might thrive in this role if you: - Have deep software engineering & infra
- experience: you’ve built CI/CD at scale, authored reliable pipelines, and shipped production services that coordinate many moving parts.
- - Are strong with automation, observability and metrics: you enjoy instrumenting systems, defining meaningful KPIs, and surfacing regressions early.
- - Can design APIs and developer ergonomics so research and SWE teams can easily submit jobs, reproduce experiments, and interpret results. - Have