data
Posted 2 days agoData Engineer, CPU & Storage
at openai
San Francisco, United StatesOn-site
Responsibilities
- - Build backend services and scheduled workloads that normalize infrastructure data and make it available to engineering and analytics teams.
- - Design data models that create consistent representations of capacity, inventory, utilization, and infrastructure state across systems.
- - Build data-quality and reconciliation mechanisms to identify missing, stale, or inconsistent information across source systems.
- - Develop reusable patterns for onboarding new infrastructure and vendor data sources as OpenAI's footprint grows
Requirements
- About the Team OpenAI's Industrial Compute organization is building the world's most advanced AI infrastructure ecosystem.
- Through a combination of strategic partnerships and self-built campuses, we are scaling the compute, storage, and networking platforms that power frontier AI training and inference.
- As OpenAI's infrastructure footprint grows, CPU and storage data increasingly spans internal platforms, vendor systems, APIs, databases, object storage, capacity management systems, and operational tooling.
- About the Role We are seeking a Data Engineer to build the data systems and integrations that connect OpenAI's CPU, storage, and supporting infrastructure platforms.
- Rather than focusing primarily on traditional analytical pipelines, you will build the software and integrations required to collect, normalize, and make infrastructure data available across a heterogeneous set of systems.
- Success in this role requires strong software engineering fundamentals, comfort working across unfamiliar systems, and the ability to design pragmatic solutions for moving and reconciling data across infrastructure environments.
- Qualifications - Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience. - 4+ years of
- experience in software engineering, data engineering, backend engineering, infrastructure engineering, or a related technical discipline. - Strong Python programming skills and
- experience building production software, services, automation, or data integrations. - Strong SQL skills and
- experience working directly with relational databases and large operational datasets. -
- Experience building reliable batch, scheduled, or asynchronous workloads in production environments. - Strong understanding of software engineering fundamentals, including testing, debugging, version control, observability, and maintainable system design. -
- Experience partnering closely with infrastructure, backend, platform, or systems engineering teams. Preferred Skills -
- Experience working with compute, storage, capacity, fleet management, inventory, or hardware lifecycle data. - Strong backend engineering
- Experience with Airflow or similar scheduling and orchestration systems. -
- Experience working across relational databases, object storage, offline tables, and analytical data systems. -
- Experience building data-quality and reconciliation mechanisms across multiple systems. - Familiarity with CPU platforms, storage systems, distributed systems, or cloud infrastructure. -
- Experience working in hyperscale, cloud, AI infrastructure, or similarly complex distributed environments.
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Additional details
- The Scaling Analytics team builds the data and software systems that help Industrial Compute understand, plan, and operate infrastructure at global scale.
- We work across capacity, hardware, storage, infrastructure software, and operational systems to connect fragmented sources of infrastructure data and make that information reliable and usable for engineering and planning.
- Building reliable connections across these environments is critical to understanding available capacity, utilization, fleet state, and infrastructure growth.
- This role sits at the intersection of data engineering and backend software engineering.
- CPU and storage data may originate from internal infrastructure platforms, vendor APIs, databases, object storage, capacity systems, and operational services.
- You will determine how to reliably connect these systems and where those integrations should live—whether within an existing infrastructure service, an orchestration framework, a scheduled workload, or a purpose-built application.
- You will work closely with Infrastructure Engineering, Capacity Engineering, Storage, Hardware Operations, and Infrastructure Software to create a reliable data foundation for understanding CPU and storage capacity, utilization, inventory, and operational state.
- Responsibilities - Build integrations that collect CPU, storage, capacity, inventory, and operational data from internal systems and external vendors.
- - Connect data across APIs, databases, object storage, infrastructure services, and capacity management platforms.
- - Determine the right architecture for new integrations, whether within existing services, orchestration frameworks, or purpose-built applications.