engineering
Posted 4 days agoAI Field Engineer, Singapore
at fireworks
Singapore, SingaporeOn-site
Responsibilities
- - For customers whose core product is built on GenAI, architect the inference foundations that capability depends on, and size deployments so they can scale in their market without infrastructure becoming the bottleneck.
- Model Strategy and Fine-Tuning - Guide customers on model selection, fine-tuning strategy (SFT, DPO, RFT), and evaluation methodology. - Build and run fine-tuning pipelines directly with customers, navigating trade-offs between model families, compute cost, and quality targets. - Design and implement evaluation frameworks that measure production-quality metrics, not just benchmark scores.
- Build trust and momentum in person, embedding with their teams where the work happens.
Requirements
- ABOUT US: At Fireworks, we’re building the future of generative AI infrastructure.
- We’ve been independently benchmarked as the leader in LLM inference speed and are driving cutting-edge innovation through projects like our own function calling and multimodal models.
- We’re an ambitious, collaborative team of builders, founded by veterans of Meta PyTorch and Google Vertex AI.
- In the last few months alone we launched Fireworks Training, partnered with Microsoft Azure Foundry, and published research straight from our production systems.
- (blog https://fireworks.ai/blog/frontier-rl-is-cheaper-than-you-think) - Open source agents with frontier advisors: matching frontier performance through training and harness engineering.
- (blog https://fireworks.ai/blog/open-source-agents-frontier-advisors) - The fine-tuning bottleneck is not the algorithm: integration friction and iteration speed are what actually stall teams; we documented the patterns across dozens of customer engagements.
- (blog) https://fireworks.ai/blog/fine-tuning-bottlenecks In the last few months alone we launched Fireworks Training, partnered with Microsoft Azure Foundry, and published research straight from our production systems.
- (blog) https://fireworks.ai/blog/fine-tuning-bottlenecks The Role AI Field Engineers at Fireworks are the technical tip of the spear.
- You embed with our most ambitious customers and technology partners to turn complex AI problems into production systems, fast.
- Earn trust with ML engineers and VPs in the same meeting. - Spend time on-site with customers.
- Qualifications - 5+ years in a hands-on, customer-facing technical role: Forward Deployed Engineer, Applied AI Engineer, Solutions Architect, ML Engineer with field exposure, or technical founder. - Demonstrated ability to build production software with customers, not just advise on it.
- You have shipped code running in someone else's production environment. - Strong Python skills.
- Familiarity with Kubernetes and infrastructure engineering. - Working knowledge of the LLM stack: inference trade-offs, model serving, fine-tuning workflows (SFT at minimum; DPO/RFT a strong plus). -
- Experience with cloud infrastructure (AWS, Azure, GCP) and deploying models on GPU infrastructure. - Exceptional communication: able to run a sharp discovery call, present to a VP, and debug a latency issue with an ML engineer in the same afternoon. Preferred
- Experience with inference serving frameworks (vLLM, SGLang, TensorRT-LLM) and tuning deployments for real workloads. -
- Experience operating as a technical authority inside a customer's environment building within their infrastructure, navigating their constraints, and shipping code that runs in their production systems. - Track record taking GenAI POCs from prototype to production-scale deployments. -
- Experience with hyperscaler AI platforms (Azure AI Foundry, AWS Bedrock/SageMaker, GCP Vertex). -
- Experience building or integrating agentic systems, tool-use chains, or AI-native developer toolchains. WHY FIREWORKS AI?
- - Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
- - Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
- - Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
- Fireworks AI is an equal-opportunity employer.
Experience
- Qualifications - 10+ years in technical field or engineering roles. -
Additional details
- Our platform delivers the highest-quality models with the fastest and most scalable inference in the industry.
- A few examples of what that looks like in practice: - Frontier RL is cheaper than the mega-cluster narrative suggests: we ran cross-region rollouts using 98% sparse weight deltas and published what we learned.
- The role sits at the intersection of engineering, product, and customer delivery.
- You ship code, run benchmarks, debug production issues, and architect deployments.
- But you also lead discovery conversations, align stakeholders, and translate customer pain points into product improvements that compress the feedback loop from field to roadmap.
- This is a role for engineers who are comfortable on-site with customers, building the relationships and trust that happen in person, not just over a call.
- These engagements span more stakeholders and longer cycles, so you will manage executive relationships and align teams while staying hands-on in the code.
- The emphasis is on pairing strong technical delivery with the executive presence to earn trust across an org: discovery, solution design, POC execution, and the path to production at enterprise scale.
- What You'll Work On Technical Delivery and Deployment - Build end-to-end POCs and MVPs alongside customer engineering teams, working inside their codebases, infrastructure, and constraints.
- - Run load tests and establish latency, throughput, and cost baselines against realistic customer traffic profiles, and tune deployments to hit those targets - Deploy and validate new model families on inference frameworks (vLLM, SGLang), determining optimal shapes, quantization configs, and serving patterns across workloads.