CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This Machine Learning Engineer role focuses on building and optimizing data pipelines for large language model training, with responsibility for data quality assessment, mixture design, and curation techniques. It suits engineers who want to combine infrastructure work with research to directly improve AI model performance.
Our summary, not Cohere’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Strong software engineering and Python proficiency
- Experience building data pipelines
- Familiarity with curriculum learning, data mixing, and data attribution
- Experience with data processing frameworks like Apache Spark, Apache Beam, or Pandas
- Experience working with large-scale datasets including web data, code data, and multilingual corpora
- Knowledge of data quality assessment techniques and experimentation with data mixtures
Nice to have
- Publication at top-tier venues such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP