engineering
Posted Jun 3Engineering Manager, Evals
at cursor
San Francisco, United StatesOn-site
Responsibilities
- - Lead and grow a high-impact team of engineers and researchers building eval datasets and developer-friendly tools to write and run evals.
Requirements
- - You’ve built and operated evaluation or measurement systems (e.g., AI evals, experimentation platforms, ranking/relevance, search quality, or reliability instrumentation). #LI-DNI
Contact
- The evaluation systems that this team builds, including CursorBench https://cursor.com/blog/cursorbench, are critical in the development of our coding models and the quality of our Cursor agents https://cursor.com/blog/continually-improving-agent-harness.
- - Guide the next generation of CursorBench https://cursor.com/blog/cursorbench so it continues to reflect real developer workflows at Cursor, and expand it with new evals that measure other properties developers value.
Additional details
- The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering.
- Our organization is very flat, and our team is small and talent dense.
- We particularly like people who are truth-seeking, passionate, and creative.
- We enjoy spirited debate, crazy ideas, and shipping code.
- ABOUT THE ROLE As an Engineering Manager on the Evals team at Cursor, you’ll lead the group responsible for creating high-signal evaluation datasets for coding agents and building the tools engineers use to write and run them.
- The team also owns online evaluation systems that track agent quality in production, and the close integration between online and offline evaluations.
- Your impact will compound across every Cursor product and every Cursor model by making quality measurable, comparable, and easy to improve.
- WHAT YOU’LL DO - Set the eval roadmap end-to-end—what we measure, why it matters, and how signals turn into shipping + training decisions.
- - Define crisp online quality signals and turn regressions into robust guardrails.
- - Integrate evals into decision-making cadence for launches, deploys, and model training loops.