other
Posted Jun 25, 2025Member of Technical Staff, Pre-Training Data
at Cohere
Toronto, CanadaRemote
Responsibilities
- - Develop robust data modeling techniques to ensure datasets are structured and formatted for optimal training efficiency.
- - Research and implement innovative data curation methods, leveraging Cohere’s infrastructure to drive advancements in natural language processing.
- - Collaborate with cross-functional teams, including researchers and engineers, to ensure data pipelines meet the demands of cutting-edge language models.
Requirements
- We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents.
- We believe that our work is instrumental to the widespread adoption of AI.
- Join us on our mission and shape the future! Why this role? As a Machine Learning Engineer specializing in pretraining data, you will play a pivotal role in developing the data pipeline that underpins Cohere’s advanced language models.
- By combining research and engineering, you will bridge the gap between raw data and cutting-edge AI models, directly contributing to improvements in critical training metrics like throughput and accelerator utilization.
- Your work will be essential to Cohere’s mission of delivering efficient and reliable language understanding and generation capabilities, driving innovation in natural language processing.
- If you are passionate about transforming data into the foundation of AI systems, this role offers a unique opportunity to make a meaningful impact.
- You may be a good fit if you have: - Strong software engineering skills, with proficiency in Python and
- experience building data pipelines. - Familiarity with curriculum learning, data mixing and data attribution. - Familiarity with data processing frameworks such as Apache Spark, Apache Beam, Pandas, or similar tools. -
- Experience working with large-scale datasets, including web data, code data, and multilingual corpora. - Knowledge of data quality assessment techniques and experimentation with data mixtures. - A passion for bridging research and engineering to solve complex data-related challenges in AI model training.
Benefits
- Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP).
- Full-Time Employees at Cohere enjoy these Perks: 🤝 An open and inclusive culture and work environment 🧑💻 Work closely with a team on the cutting edge of AI research 🍽 Weekly lunch stipend, in-office lunches & snacks 🦷 Full health and dental benefits, including a separate budget to take care of your mental health 🐣 100% Parental Leave top-up for up to 6 months 🎨 Personal enrichment
- benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement 🏙 Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend ✈️ 6 weeks of vacation (30 working days!)
Contact
- Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form https://docs.google.com/forms/d/12a6IrLdF3kI2nonKSr4tiFuz18rLQbaeYV-JM9L4o9Q/edit, and we will work together to meet your needs.
Additional details
- Who are we? Our mission is to scale intelligence to serve humanity.
- Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers.
- We like to work hard and move fast to do what’s best for our customers.
- Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft.
- Each person is one of the best in the world at what they do.
- We believe that a diverse range of perspectives is a requirement for building great products.
- In this role, you will conduct data ablations to evaluate data quality and construct pre-training data mixtures to enhance model performance.
- Please Note: We have offices in London, Paris, Toronto, San Francisco and New York but also embrace being remote-friendly! There are no restrictions on where you can be located for this role between EST and EU.
- As a Member of Technical Staff, Pre-Training Data,
- you will: - Conduct data ablations to assess data quality and experiment with data mixtures to enhance model performance.