Skip to main content
CareerApp

Member of Technical Staff, Pre-Training Data

Cohere

New York, NY · Remote · full time

CA$250,000 – CA$535,000

Listed on Cohere’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This Machine Learning Engineer role focuses on building and optimizing data pipelines for large language model training, with responsibility for data quality assessment, mixture design, and curation techniques. It suits engineers who want to combine infrastructure work with research to directly improve AI model performance.

Our summary, not Cohere’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • Strong software engineering and Python proficiency
  • Experience building data pipelines
  • Familiarity with curriculum learning, data mixing, and data attribution
  • Experience with data processing frameworks like Apache Spark, Apache Beam, or Pandas
  • Experience working with large-scale datasets including web data, code data, and multilingual corpora
  • Knowledge of data quality assessment techniques and experimentation with data mixtures

Nice to have

  • Publication at top-tier venues such as NeurIPS, ICML, ICLR, AIStats, MLSys, JMLR, AAAI, Nature, COLING, ACL, or EMNLP

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.