engineering
Posted 5 days agoSoftware Engineer, ML Inference Platform
at dialpad
ArgentinaOn-site
You are nearing today's limit. Upgrade for unlimited access.
Responsibilities
- Model Server Integration: Work with model-serving frameworks and runtimes such as vLLM, Triton, TGI, or similar systems, adapting them to internal deployment, observability, and release requirements.
- Build and ship agentic AI products that are redefining how companies operate
Requirements
- Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital.
- Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.
- Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve.
- Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.
- At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.
- We are hiring ML Inference Platform Engineers to build the production systems that serve our in-house AI models at scale.
- You will help turn trained models and emerging AI capabilities into reliable, observable, low-latency production services running on NVIDIA GPUs in GCP.
- It is an implementation-heavy systems engineering role focused on the machinery of inference: model serving, runtime optimization, GPU utilization, deployment safety, traffic management, benchmarking, and production reliability.
- The work is practical, deeply technical, and closely tied to the company’s broader AI strategy.
- We are not building one-off demos; we are building the inference platform by which a growing AI organization can repeatedly and safely ship real model-backed products. What you’ll do
- You will design, build, and improve the systems that connect AI capability development to production inference.
- GPU Infrastructure & Utilization: Operate and optimize containerized workloads on Kubernetes/GCP, with a focus on efficient use of NVIDIA GPUs, memory, storage, and networking.
- Efficiency & Scale: Contribute to strategies that improve compute efficiency, GPU utilization, autoscaling behavior, and cost-performance tradeoffs across the inference platform. Skills you’ll bring
- Strong Software Fundamentals: Proficiency in writing maintainable production code in Python, Go, or another backend-oriented language, with strong debugging and systems-thinking skills.
- Experience building, operating, or optimizing high-throughput services, distributed systems, data/ML infrastructure, or runtime platforms where latency, reliability, and resource utilization matter.
- Kubernetes & Linux Fluency: Hands-on
- experience with containers, Kubernetes, Linux environments, CI/CD, deployment automation, and production operations.
- Collaboration: Ability to work closely with model developers, product engineers, infrastructure teams, and technical leadership to turn evolving AI capabilities into reliable production systems. Why Join Dialpad
- Work at the center of the AI transformation in business communications
- Join a team where AI amplifies every employee’s impact
Experience
- Experience: 6+ years of professional software engineering experience, with a track record of shipping backend services, infrastructure systems, or production platforms that matter.
Benefits
- Competitive salary, comprehensive benefits, and real opportunities for growth
Additional details
- Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer
- We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.
- We look for people who are intensely curious and hold themselves to a high bar.
- Our ambition is significant, and achieving it requires a team that operates at the highest level.
- We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role
- This role sits at the intersection of model development, high-performance runtime systems, and cloud infrastructure.
- This is not a research role, and it is not a generic MLOps or support role.
- Our mission is to shorten the path from promising model capability to dependable production impact.
- We build the shared infrastructure, standards, and release pathways that allow models to move from candidate artifacts into scalable, rollback-safe inference services with clear performance, reliability, and cost characteristics.
- This is a new team, so the systems and interfaces are still being shaped.