$200,000 – $350,000
Listed on Modal’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
Modal is seeking an engineer to optimize machine learning inference and fine-tuning workloads on their GPU infrastructure platform. This role suits someone with deep experience in ML systems performance who enjoys debugging GPU bottlenecks and shipping optimizations at scale.
Our summary, not Modal’s wording. The full posting is on their site.
What they ask for
Required
- 5+ years writing high-performance code
- Experience with PyTorch and ML inference frameworks
- Knowledge of Nvidia GPU architecture and CUDA
- Demonstrated ML performance engineering work (GPU occupancy tuning, algorithmic optimization, or overhead elimination)
Nice to have
- Linux kernel and operating system internals knowledge
- Container and file system experience