engineering
Posted May 21, 2025Software Engineer, Inference - Multi Modal
at openai
San Francisco, United StatesOn-site
You are nearing today's limit. Upgrade for unlimited access.
Responsibilities
- - Optimize systems for high-throughput, low-latency delivery of image and audio inputs and outputs.
- - Collaborate closely with researchers, infra teams, and product engineers to deploy state-of-the-art capabilities.
Requirements
- About the Team OpenAI’s Inference team powers the deployment of our most advanced models - including our GPT models, 4o Image Generation, and Whisper - across a variety of platforms.
- experience while pushing the boundaries of what AI can do.
- We’re expanding into multimodal inference, building the infrastructure needed to serve models that handle image, audio, and other non-text modalities.
- - Have worked with GPU-based ML workloads and understand the performance dynamics of large models, especially with complex data like images or audio.
- - Have familiarity with inference tooling like vLLM, TensorRT-LLM, or custom model parallel systems.
- Experience working with image generation or audio synthesis models in production. - Exposure to distributed ML training or system-efficient model design.
- About OpenAI OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence
Additional details
- Our work ensures these models are available, performant, and scalable in production, and we partner closely with Research to bring the next generation of models into the world.
- These workloads are inherently more heterogeneous and experimental, involving diverse model sizes and interactions, more complex input/output formats, and tighter coordination with product and research.
- About the Role We’re looking for a software engineer to help us serve OpenAI’s multimodal models at scale.
- You’ll be part of a small team responsible for building reliable, high-performance infrastructure for serving real-time audio, image, and other MM workloads in production.
- This work is inherently cross-functional: you’ll collaborate directly with researchers training these models and with product teams defining new modalities of interaction.
- You'll build and optimize the systems that let users generate speech, understand images, and interact with models in ways far beyond text. In this role,
- you will: - Design and implement inference infrastructure for large-scale multimodal models.
- - Enable experimental research workflows to transition into reliable production services.
- - Contribute to system-level improvements including GPU utilization, tensor parallelism, and hardware abstraction layers.
- - Enjoy experimental, fast-evolving work and collaborating closely with research.