# Member of Technical Staff - Research, Inference

Hiring organization: [Modal](https://career.thegoodapps.co/organizations/modal)

Canonical page: https://career.thegoodapps.co/jobs/543964c5-01a7-471e-b24b-166b232f0207

Listed on Modal's own careers site. Applications go to them directly.

- Seniority: Mid
- Location: New York, NY
- Salary: 150000 – 350000 USD per year

## Summary

Modal is hiring a researcher to lead inference optimizations for their LLM serving platform, focusing on techniques like speculative decoding and quantization that reduce cost and latency. This role suits someone with a background shipping inference systems or research who can independently drive projects from conception through deployment.

_Our summary, not Modal's wording._

## Skills named

Autoscaling, CUDA, Model Compression & Quantization, Python

## Required

- Research or systems background in LLM inference with published or shipped work
- Fluency across the LLM serving stack from kernels to schedulers
- Track record of shipping research or systems that others build upon
- Ability to independently drive research bets from idea through results
- In-person work in NYC or San Francisco office

Apply on Modal's site: https://jobs.ashbyhq.com/modal/73c97bbc-8e27-4c5d-b38b-90b3afdb0d93/application
