jobloom

JobLoom finds jobs directly from company career sites before many job boards, then routes you into detailed role pages like this one.

engineering

Posted Jun 3

Engineering Manager, Evals

at cursor

San Francisco, United StatesOn-site

Responsibilities

  • - Lead and grow a high-impact team of engineers and researchers building eval datasets and developer-friendly tools to write and run evals.

Requirements

  • - You’ve built and operated evaluation or measurement systems (e.g., AI evals, experimentation platforms, ranking/relevance, search quality, or reliability instrumentation). #LI-DNI

Contact

  • The evaluation systems that this team builds, including CursorBench https://cursor.com/blog/cursorbench, are critical in the development of our coding models and the quality of our Cursor agents https://cursor.com/blog/continually-improving-agent-harness.
  • - Guide the next generation of CursorBench https://cursor.com/blog/cursorbench so it continues to reflect real developer workflows at Cursor, and expand it with new evals that measure other properties developers value.

Additional details

  • The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering.
  • Our organization is very flat, and our team is small and talent dense.
  • We particularly like people who are truth-seeking, passionate, and creative.
  • We enjoy spirited debate, crazy ideas, and shipping code.
  • ABOUT THE ROLE As an Engineering Manager on the Evals team at Cursor, you’ll lead the group responsible for creating high-signal evaluation datasets for coding agents and building the tools engineers use to write and run them.
  • The team also owns online evaluation systems that track agent quality in production, and the close integration between online and offline evaluations.
  • Your impact will compound across every Cursor product and every Cursor model by making quality measurable, comparable, and easy to improve.
  • WHAT YOU’LL DO - Set the eval roadmap end-to-end—what we measure, why it matters, and how signals turn into shipping + training decisions.
  • - Define crisp online quality signals and turn regressions into robust guardrails.
  • - Integrate evals into decision-making cadence for launches, deploys, and model training loops.

Find more real-time jobs on JobLoom.