other
Posted Oct 30, 2025Member of Technical Staff, Evals & Post-Training Product
at fireworks
On-site
Responsibilities
- - Own fine-tuning product experiences: Build and improve user-facing product workflows for post-training, including fine-tuning experiences across SFT, RFT, and related model-improvement capabilities.
- - Build internal eval workflows: Design and scale evaluation tooling used by internal teams to measure model quality, compare model changes, and inform post-training decisions.
Requirements
- ABOUT US: At Fireworks, we’re building the future of generative AI infrastructure.
- We’ve been independently benchmarked as the leader in LLM inference speed and are driving cutting-edge innovation through projects like our own function calling and multimodal models.
- We’re an ambitious, collaborative team of builders, founded by veterans of Meta PyTorch and Google Vertex AI.
- In the last few months alone we launched Fireworks Training, partnered with Microsoft Azure Foundry, and published research straight from our production systems.
- (blog https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think) - Open source agents with frontier advisors: matching frontier performance through training and harness engineering.
- (blog https://fireworks.ai/blog/open-source-agents-frontier-advisors) - The fine-tuning bottleneck is not the algorithm: integration friction and iteration speed are what actually stall teams; we documented the patterns across dozens of customer engagements.
- (blog) https://fireworks.ai/blog/fine-tuning-bottlenecks We are seeking a Member of Technical Staff, Evals & Post-Training Product to help define how developers improve models on Fireworks.
- You will work across the stack—from APIs, SDKs, and backend systems to user-facing product surfaces in the web app—to make it easier for users to author evals, understand results, fine-tune models, and iterate quickly.
- experience with LLM evaluations and/or post-training methods: How to design useful evals and use their results to guide model improvement. - Product Engineering Skills: The ability to work across backend systems and developer-facing product surfaces.
- Comfortable shipping full-stack features when needed. - Understanding of the GenAI Lifecycle: You understand the end-to-end workflow—from prompting a base model to curating a dataset, fine-tuning, and productionizing agents—and how these steps interconnect. - User-Centric Mindset: Willing to talk to users, triage GitHub issues for open-source projects, and build products from scratch to serve emerging needs. PREFERRED
- Experience: Strong familiarity with designing and running evaluations for domain-specific use cases (e.g. medical, legal, coding, or custom internal datasets). - Open Source Contributions: Prior contributions to developer tools or AI/ML repositories. - Inference & Hardware Knowledge: Interest in the hardware side of AI—understanding GPU constraints, inference optimization techniques, and how they relate to model performance. - Startup DNA:
- Experience in fast-paced environments where you own features end-to-end. WHY FIREWORKS AI?
- - Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
- - Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
- - Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
- Fireworks AI is an equal-opportunity employer.
Experience
- QUALIFICATIONS: - 3+ years of software engineering experience. - Domain-Specific Evaluation
Additional details
- Our platform delivers the highest-quality models with the fastest and most scalable inference in the industry.
- A few examples of what that looks like in practice: - Frontier RL is cheaper than the mega-cluster narrative suggests: we ran cross-region rollouts using 98% sparse weight deltas and published what we learned.
- This role sits at the intersection of product engineering, developer experience, and model quality.
- You will build the products and workflows that connect evaluation and post-training into a continuous loop: helping internal teams run evals at scale, enabling external developers through our open-source Eval Protocol SDK, and owning key product experiences for fine-tuning custom models on Fireworks.
- You will also work directly with customers and internal teams to identify friction, support real-world use cases, and turn repeated pain points into reusable product capabilities. KEY
- - Work closely with users: Partner with customers and internal stakeholders to understand evaluation and fine-tuning needs, support high-priority engagements, triage issues, and convert bespoke workflows into productized solutions. MINIMUM
- REQUIREMENTS: - 1 - 7 years of software engineering
- experience (We are hiring at multiple levels for this role). - Hands-on
- We celebrate diversity and are committed to creating an inclusive environment for all innovators.