other
Posted Jun 11Member of Technical Staff
at fireworks
New York, United StatesOn-site
Responsibilities
- RESPONSIBILITIES: - Architect and build scalable, resilient backend infrastructure to support distributed training, inference, and data processing pipelines - Lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems - Design and implement core backend services with a focus on efficiency and low latency - Drive infrastructure optimization initiatives for compute cost, storage lifecycle management, and network performance - Collaborate with
Requirements
- ABOUT US: At Fireworks, we’re building the future of generative AI infrastructure.
- We’ve been independently benchmarked as the leader in LLM inference speed and are driving cutting-edge innovation through projects like our own function calling and multimodal models.
- We’re an ambitious, collaborative team of builders, founded by veterans of Meta PyTorch and Google Vertex AI.
- In the last few months alone we launched Fireworks Training, partnered with Microsoft Azure Foundry, and published research straight from our production systems.
- (blog https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think) - Open source agents with frontier advisors: matching frontier performance through training and harness engineering.
- (blog https://fireworks.ai/blog/open-source-agents-frontier-advisors) - The fine-tuning bottleneck is not the algorithm: integration friction and iteration speed are what actually stall teams; we documented the patterns across dozens of customer engagements.
- (blog) https://fireworks.ai/blog/fine-tuning-bottlenecks THE ROLE: As a Training Infrastructure Engineer, you'll design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for our generative AI platform.
- You'll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems. KEY
- requirements into robust infrastructure solutions - Evaluate and integrate cloud-native and open-source technologies such as Kubernetes, Ray, Kubeflow, and MLFlow to enhance platform reliability - Own end-to-end systems from design to deployment, emphasizing reliability, fault tolerance, and operational excellence MINIMUM
- QUALIFICATIONS: - Bachelor's degree or equivalent in Computer Science or related field plus four (4) years of
- experience in software engineering or related role - 4 years of
- experience designing, building, and optimizing large-scale backend infrastructure and distributed data systems (e.g., PostgreSQL, MySQL, DynamoDB, Apache Spark, Apache Flink, Apache Kafka) in cloud environments (AWS, GCP, Azure, or equivalent), including cloud-native platforms, core infrastructure components, and optimization techniques (caching, indexing, sharding, replication, transactions, ACID) - 4 years of
- experience with major server-side programming languages and frameworks (e.g., Python, C++, Go, TypeScript) - 4 years of
- experience writing technical design documentation, leading cross-functional projects, and collaborating with cross-functional teams to achieve business impact - 3 years of
- experience developing and maintaining data processing and API systems, including client-server communication frameworks (e.g., gRPC, Thrift) - 3 years of
- experience conducting A/B testing and scientific experimentation (e.g., Statsig, Meta Deltoid, Optimizely) to measure software impact - 3 years of
- experience conducting coding interviews and providing systematic feedback for engineering candidates - 2 years of
- experience with cloud-native tools and infrastructure, such as Docker and Kubernetes - 2 years of
- experience defining and implementing data-driven metrics to support company or team goals How to Apply: Submit resume and apply online at http://www.fireworks.ai/careers and search for job by title.
- Fireworks AI is an equal-opportunity employer.
- We celebrate diversity and are committed to creating an inclusive environment for all innovators. WHY FIREWORKS AI?
- - Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
- - Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
- - Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
Additional details
- Our platform delivers the highest-quality models with the fastest and most scalable inference in the industry.
- A few examples of what that looks like in practice: - Frontier RL is cheaper than the mega-cluster narrative suggests: we ran cross-region rollouts using 98% sparse weight deltas and published what we learned.
- We celebrate diversity and are committed to creating an inclusive environment for all innovators.