Skip to main content
CareerApp

Research Scientist, Data

Pika Labs

Palo Alto, CA · Staff

$185,000 – $400,000

Listed on Pika Labs’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Pika is hiring a senior data engineer to design and operate large-scale data pipelines that feed multimodal AI model training, handling datasets across text, image, audio, and video. The role suits experienced engineers who have built production ML data infrastructure and want to own the full lifecycle of data quality and curation for cutting-edge generative models.

Our summary, not Pika Labs’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years building and scaling data pipelines for machine learning
  • Staff or lead-level experience in research or model training environments
  • Data engineering and ML data curation for large-scale multimodal models
  • Distributed data systems (Spark, Hadoop, Ray, or similar)
  • Large-scale dataset processing and ETL workflows
  • Tools for data labeling, filtering, deduplication, and quality assurance
  • Production-grade data infrastructure for ML pipelines
  • Python, SQL, PySpark, or similar programming
  • Cloud data platforms (AWS, GCP, Azure)
  • Privacy, compliance, and ethical data management

Nice to have

  • Knowledge of emerging data engineering and ML data management developments
  • Cross-functional collaboration and communication skills

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.