CA$250,000 – CA$535,000
Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role develops the synthetic data pipelines that power Cohere's language models, combining research methods with engineering to improve model quality and efficiency. You'll work on data curation, ablation studies, and inference optimization across large-scale datasets and GPU clusters.
Our summary, not Cohere’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Strong software engineering and Python skills
- Experience building data pipelines
- Familiarity with data processing frameworks like Spark, Beam, or Pandas
- Experience working with LLMs through projects or experimentation
- Familiarity with LLM inference frameworks such as vLLM and TensorRT
- Experience with large-scale datasets including web data, code data, and multilingual corpora
- Ability to bridge research and engineering for AI model training challenges
Nice to have
- Publication at top-tier ML venues like NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP