engineering
Posted 3 days agoAI Evaluation Engineer
at dialpad
Vancouver, CanadaOn-site
Responsibilities
- Build and ship agentic AI products that are redefining how companies operate
Requirements
- Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital.
- Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.
- Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve.
- Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.
- At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.
- As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead.
- This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.
- Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field. 3+ years of
- experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products. •
- Experience designing structured test strategies across manual and automated workflows.
- Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products. •
- Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
- Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.
- Work at the center of the AI transformation in business communications
- Join a team where AI amplifies every employee’s impact
Benefits
- For exceptional talent based in British Columbia, Canada the target base salary range for this position is posted below.
- Our salary ranges are determined by role, level, and location.
- The range displayed on each job posting reflects the target range for new hire salaries for the position.
- Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
- Your recruiter can share more about the specific salary range for your preferred location during the hiring process.
- Please note that the compensation details listed in British Columbia role postings reflect the base salary only, and do not include bonus, equity, or benefits.
- British Columbia, Canada Salary Range $115,500 — $132,750 CAD Why Join Dialpad
- Competitive salary, comprehensive benefits, and real opportunities for growth
Additional details
- Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer
- We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.
- We look for people who are intensely curious and hold themselves to a high bar.
- Our ambition is significant, and achieving it requires a team that operates at the highest level.
- We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role
- A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.
- You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
- You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
- You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
- You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.