jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

other

Posted Jun 11

Member of Technical Staff

at fireworks

New York, United StatesOn-site

Responsibilities

  • RESPONSIBILITIES: - Architect and build scalable, resilient backend infrastructure to support distributed training, inference, and data processing pipelines - Lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems - Design and implement core backend services with a focus on efficiency and low latency - Drive infrastructure optimization initiatives for compute cost, storage lifecycle management, and network performance - Collaborate with

Requirements

  • ABOUT US: At Fireworks, we’re building the future of generative AI infrastructure.
  • We’ve been independently benchmarked as the leader in LLM inference speed and are driving cutting-edge innovation through projects like our own function calling and multimodal models.
  • We’re an ambitious, collaborative team of builders, founded by veterans of Meta PyTorch and Google Vertex AI.
  • In the last few months alone we launched Fireworks Training, partnered with Microsoft Azure Foundry, and published research straight from our production systems.
  • (blog https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think) - Open source agents with frontier advisors: matching frontier performance through training and harness engineering.
  • (blog https://fireworks.ai/blog/open-source-agents-frontier-advisors) - The fine-tuning bottleneck is not the algorithm: integration friction and iteration speed are what actually stall teams; we documented the patterns across dozens of customer engagements.
  • (blog) https://fireworks.ai/blog/fine-tuning-bottlenecks THE ROLE: As a Training Infrastructure Engineer, you'll design, develop, and maintain large-scale backend and cloud-native infrastructure to support distributed machine learning training, inference, and data processing pipelines for our generative AI platform.
  • You'll architect scalable, resilient backend infrastructure, lead technical design discussions, mentor engineers, and establish best practices for large-scale machine learning systems. KEY
  • requirements into robust infrastructure solutions - Evaluate and integrate cloud-native and open-source technologies such as Kubernetes, Ray, Kubeflow, and MLFlow to enhance platform reliability - Own end-to-end systems from design to deployment, emphasizing reliability, fault tolerance, and operational excellence MINIMUM
  • QUALIFICATIONS: - Bachelor's degree or equivalent in Computer Science or related field plus four (4) years of
  • experience in software engineering or related role - 4 years of
  • experience designing, building, and optimizing large-scale backend infrastructure and distributed data systems (e.g., PostgreSQL, MySQL, DynamoDB, Apache Spark, Apache Flink, Apache Kafka) in cloud environments (AWS, GCP, Azure, or equivalent), including cloud-native platforms, core infrastructure components, and optimization techniques (caching, indexing, sharding, replication, transactions, ACID) - 4 years of
  • experience with major server-side programming languages and frameworks (e.g., Python, C++, Go, TypeScript) - 4 years of
  • experience writing technical design documentation, leading cross-functional projects, and collaborating with cross-functional teams to achieve business impact - 3 years of
  • experience developing and maintaining data processing and API systems, including client-server communication frameworks (e.g., gRPC, Thrift) - 3 years of
  • experience conducting A/B testing and scientific experimentation (e.g., Statsig, Meta Deltoid, Optimizely) to measure software impact - 3 years of
  • experience conducting coding interviews and providing systematic feedback for engineering candidates - 2 years of
  • experience with cloud-native tools and infrastructure, such as Docker and Kubernetes - 2 years of
  • experience defining and implementing data-driven metrics to support company or team goals How to Apply: Submit resume and apply online at http://www.fireworks.ai/careers and search for job by title.
  • Fireworks AI is an equal-opportunity employer.
  • We celebrate diversity and are committed to creating an inclusive environment for all innovators. WHY FIREWORKS AI?
  • - Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
  • - Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
  • - Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.

Additional details

  • Our platform delivers the highest-quality models with the fastest and most scalable inference in the industry.
  • A few examples of what that looks like in practice: - Frontier RL is cheaper than the mega-cluster narrative suggests: we ran cross-region rollouts using 98% sparse weight deltas and published what we learned.
  • We celebrate diversity and are committed to creating an inclusive environment for all innovators.

Find more real-time jobs on JobLoom.