research
Posted Aug 19, 2025Member of Technical Staff, Integration/RL Team (Research Engineer)
at Cohere
Paris, FranceRemote
Responsibilities
- - Develop new tools to support and accelerate research and LLM training.
- - Coordinate with other engineering teams (Infrastructure, Efficiency, Serving) and the scientific teams (Agent, Multimodal, Multilingual, etc.) to create a strong and integrated post-training ecosystem.
Requirements
- We’re training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents.
- We believe that our work is instrumental to the widespread adoption of AI.
- Join us on our mission and shape the future! The integration team is responsible for developing and scaling machine learning algorithms and infrastructure for LLM post-training, with a focus on large-scale, distributed RL methods.
- We strive for excellence in both engineering and science by meticulously designing experiments and design docs.
- - Proficiency in Python and related ML frameworks such as JAX, Pytorch and/or XLA/MLIR. -
- Experience with distributed training infrastructures (Kubernetes) and associated frameworks (Ray). - [Bonus] Hands-on
- Experience in ML, LLM and RL academic research.
- This role is perfect for you if you: - Have a deep passion for quality work. - Enjoy tuning and optimising large LLM models. - Comfortable working with people with different levels of software engineering skills, from beginner to more advanced. - Comfortable diving into complex ML codebases to identify and resolve issues, ensuring the smooth operation of our systems. - Thrive in a fast-paced, technically challenging environment, where you can contribute your innovative ideas and solutions.
Benefits
- Experience using and debugging large-scale distributed training strategies (memory/speed profiling). - [Bonus]
- experience with the post-training phase of model training, with a strong emphasis on scalability and performance. - [Bonus]
- Full-Time Employees at Cohere enjoy these Perks: 🤝 An open and inclusive culture and work environment 🧑💻 Work closely with a team on the cutting edge of AI research 🍽 Weekly lunch stipend, in-office lunches & snacks 🦷 Full health and dental benefits, including a separate budget to take care of your mental health 🐣 100% Parental Leave top-up for up to 6 months 🎨 Personal enrichment
- benefits towards arts and culture, fitness and well-being, quality time, and workspace improvement 🏙 Remote-flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co-working stipend ✈️ 6 weeks of vacation (30 working days!)
Contact
- Should you require any accommodations during the recruitment process, please submit an Accommodations Request Form https://docs.google.com/forms/d/12a6IrLdF3kI2nonKSr4tiFuz18rLQbaeYV-JM9L4o9Q/edit, and we will work together to meet your needs.
Additional details
- Who are we? Our mission is to scale intelligence to serve humanity.
- Each one of us is responsible for contributing to increasing the capabilities of our models and the value they drive for our customers.
- We like to work hard and move fast to do what’s best for our customers.
- Cohere is a team of researchers, engineers, designers, and more, who are passionate about their craft.
- Each person is one of the best in the world at what they do.
- We believe that a diverse range of perspectives is a requirement for building great products.
- While tasks are assigned according to everyone’s expertise, there is a global team effort to write production code and support the team research efforts, depending on individual interests and organizational needs.
- In particular, this role aims to enhance the global quality of the post-training codebase by implementing new tools to ease and support research, optimizing post-training algorithms, and scaling distributed RL to unprecedented levels.
- Please Note: We have offices in London, Paris, Toronto, San Francisco, New York but we are also remote-friendly! Applicants for this role may work anywhere between UTC−06:00 and UTC+01:00.
- you will: - Design and write high-performing and scalable software for training models.