Skip to main content
CareerApp

ML Engineer, Inference & Optimization

Pika Labs

Palo Alto, CA · Senior

$250,000 – $350,000

Listed on Pika Labs’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

A senior or staff-level role optimizing AI model inference performance at a video generation startup. This suits engineers with deep expertise in GPU acceleration, distributed systems, and deploying large language models and video models efficiently at scale.

Our summary, not Pika Labs’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years of engineering experience
  • Expertise in inference optimization and model deployment at scale
  • Proficiency in GPU programming with CUDA and NCCL
  • Knowledge of distributed parallelism strategies (tensor, sequence, pipeline)
  • Familiarity with video generation and large language models
  • Strong cross-functional communication skills

Nice to have

  • Experience with high-throughput video or real-time streaming deployment
  • Knowledge of distributed training and optimization toolkits
  • Open source contributions to AI infrastructure or deep learning compilers
  • Startup or rapid prototyping background
  • Experience improving training efficiency and resource optimization

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.