# Research Scientist, Data

Hiring organization: [Pika Labs](https://career.thegoodapps.co/organizations/pika-labs)

Canonical page: https://career.thegoodapps.co/jobs/123c09b5-fbac-4739-91b5-37f3dd7a4659

Listed on Pika Labs's own careers site. Applications go to them directly.

- Seniority: Staff
- Location: Palo Alto, CA
- Salary: 185000 – 400000 USD per year

## Summary

Pika is hiring a senior data engineer to design and operate large-scale data pipelines that feed multimodal AI model training, handling datasets across text, image, audio, and video. The role suits experienced engineers who have built production ML data infrastructure and want to own the full lifecycle of data quality and curation for cutting-edge generative models.

_Our summary, not Pika Labs's wording._

## Skills named

Amazon Web Services (AWS), Apache Hadoop, Microsoft Azure, PySpark, Python, SPARK, SQL, V-Ray

## Required

- 5+ years building and scaling data pipelines for machine learning
- Staff or lead-level experience in research or model training environments
- Data engineering and ML data curation for large-scale multimodal models
- Distributed data systems (Spark, Hadoop, Ray, or similar)
- Large-scale dataset processing and ETL workflows
- Tools for data labeling, filtering, deduplication, and quality assurance
- Production-grade data infrastructure for ML pipelines
- Python, SQL, PySpark, or similar programming
- Cloud data platforms (AWS, GCP, Azure)
- Privacy, compliance, and ethical data management

## Nice to have

- Knowledge of emerging data engineering and ML data management developments
- Cross-functional collaboration and communication skills

Apply on Pika Labs's site: https://jobs.ashbyhq.com/pika/82835e51-284c-47af-b254-088fff23acd5/application
