# ML Engineer, Inference & Optimization

Hiring organization: [Pika Labs](https://career.thegoodapps.co/organizations/pika-labs)

Canonical page: https://career.thegoodapps.co/jobs/f3b97638-eec2-4603-89ba-5fdecaa05274

Listed on Pika Labs's own careers site. Applications go to them directly.

- Seniority: Senior
- Location: Palo Alto, CA
- Salary: 250000 – 350000 USD per year

## Summary

A senior or staff-level role optimizing AI model inference performance at a video generation startup. This suits engineers with deep expertise in GPU acceleration, distributed systems, and deploying large language models and video models efficiently at scale.

_Our summary, not Pika Labs's wording._

## Skills named

CUDA, LLMs (Large Language Models), Model Compression & Quantization

## Required

- 5+ years of engineering experience
- Expertise in inference optimization and model deployment at scale
- Proficiency in GPU programming with CUDA and NCCL
- Knowledge of distributed parallelism strategies (tensor, sequence, pipeline)
- Familiarity with video generation and large language models
- Strong cross-functional communication skills

## Nice to have

- Experience with high-throughput video or real-time streaming deployment
- Knowledge of distributed training and optimization toolkits
- Open source contributions to AI infrastructure or deep learning compilers
- Startup or rapid prototyping background
- Experience improving training efficiency and resource optimization

Apply on Pika Labs's site: https://jobs.ashbyhq.com/pika/fb9e43d7-36b8-46c7-93cd-2b7504e30363/application
