Skip to main content
CareerApp

Member of Technical Staff - ML Performance

Modal

New York, NY · Senior

$200,000 – $350,000

Listed on Modal’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Modal is seeking an engineer to optimize machine learning inference and fine-tuning workloads on their GPU infrastructure platform. This role suits someone with deep experience in ML systems performance who enjoys debugging GPU bottlenecks and shipping optimizations at scale.

Our summary, not Modal’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years writing high-performance code
  • Experience with PyTorch and ML inference frameworks
  • Knowledge of Nvidia GPU architecture and CUDA
  • Demonstrated ML performance engineering work (GPU occupancy tuning, algorithmic optimization, or overhead elimination)

Nice to have

  • Linux kernel and operating system internals knowledge
  • Container and file system experience

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.