Listed on Protege’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role owns the data ingestion and processing layer that transforms raw, high-volume multimodal data into clean, AI-ready datasets for a platform connecting data holders with AI teams. It suits backend engineers experienced in large-scale data systems who thrive in building production infrastructure that handles messy data reliably at massive scale.
Our summary, not Protege’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- 5+ years building and operating production backend or data systems with real data processing at scale experience
- hands-on experience designing and running large-scale data pipelines
- strong Python programming skills
- experience with distributed data processing
- strong AWS proficiency
- comfort with messy, varied, high-volume data and high ambiguity
- attention to detail without losing speed
- bias to action
- excited to work on a data movement and processing product
Nice to have
- experience processing medical imaging (DICOM), text, audio or video at scale
- background with sensitive or regulated data (HIPAA, healthcare compliance, PHI)
- experience with streaming systems or workflow orchestration (Airflow, Dagster)
- GCP and Azure experience
- prior startup experience as founding or early engineer
- familiarity with ML, NLP, LLM systems, embeddings, and fine-tuning