jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

engineering

Posted 3 days ago

AI Evaluation Engineer

at dialpad

Vancouver, CanadaOn-site

Responsibilities

  • Build and ship agentic AI products that are redefining how companies operate

Requirements

  • Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital.
  • Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.
  • Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve.
  • Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.
  • At Dialpad, AI isn’t just a feature; it’s how our teams do their best work every day. We put powerful AI tools in every employee’s hands so they can move faster, think bigger, and achieve more.
  • As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead.
  • This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.
  • Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field. 3+ years of
  • experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products. •
  • Experience designing structured test strategies across manual and automated workflows.
  • Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products. •
  • Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
  • Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.
  • Work at the center of the AI transformation in business communications
  • Join a team where AI amplifies every employee’s impact

Benefits

  • For exceptional talent based in British Columbia, Canada the target base salary range for this position is posted below.
  • Our salary ranges are determined by role, level, and location.
  • The range displayed on each job posting reflects the target range for new hire salaries for the position.
  • Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training.
  • Your recruiter can share more about the specific salary range for your preferred location during the hiring process.
  • Please note that the compensation details listed in British Columbia role postings reflect the base salary only, and do not include bonus, equity, or benefits.
  • British Columbia, Canada Salary Range $115,500 — $132,750 CAD Why Join Dialpad
  • Competitive salary, comprehensive benefits, and real opportunities for growth

Additional details

  • Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile. Being a Dialer
  • We believe every conversation matters. And we’ve built the platform that turns those conversations into insight and action, for our customers and ourselves.
  • We look for people who are intensely curious and hold themselves to a high bar.
  • Our ambition is significant, and achieving it requires a team that operates at the highest level.
  • We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role
  • A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.
  • You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
  • You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
  • You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
  • You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.

Find more real-time jobs on JobLoom.