CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves building and maintaining the large-scale web data pipelines that feed Cohere's language models, including extraction, deduplication, filtering, and quality analysis of internet-sourced training data. The position suits experienced software engineers who want to bridge data engineering and AI research while working on infrastructure that directly impacts model capabilities.
Our summary, not Cohere’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Strong software engineering and Python proficiency
- Experience building data pipelines
- Familiarity with data processing frameworks like Apache Spark, Apache Beam, or Pandas
- Experience working with large-scale web datasets
- Knowledge of data quality assessment techniques and experimentation with data mixtures
Nice to have
- Publication at top-tier venues (NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, EMNLP)