$180,000 – $220,000
Listed on Udio’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves building the data infrastructure that powers a generative audio company's research, focusing on ingesting and unifying large datasets from multiple external sources. You'll design systems for entity resolution, deduplication, and data enrichment at scale, working closely with ML researchers to prepare training-ready datasets.
Our summary, not Udio’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Experience with large, heterogeneous datasets from multiple providers or domains
- Strong background in entity resolution, deduplication, or large-scale data integration
- Proficiency in Python for scalable data processing
- Experience with BigQuery, Google Dataflow, or Apache Beam
- Familiarity with data validation, normalization, and reconciliation
- Ability to design matching and decision strategies balancing accuracy and efficiency
- Quick iteration on pragmatic solutions
- Collaboration with ML and research teams
Nice to have
- Google Cloud Platform systems architecture at scale
- Distributed compute frameworks such as Ray, Spark, or Flink
- JAX-based ML pipelines or multihost training setups
- TFRecords or high-volume training data formats
- Ranking, clustering, or statistical similarity modeling
- Go, NextJS, or React Native for full-stack development