Skip to main content
CareerApp

Senior Technical Program Manager (Engineering) - AI Tooling & Systems

Deepgram

USA | Remote · Remote · Senior

$152,000 – $208,000

Listed on Deepgram’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

This role leads the design and delivery of ML infrastructure, model serving systems, and AI tooling that enable Deepgram's research and engineering teams to build and deploy voice AI models at scale. You'll coordinate across research, engineering, and product to translate ML requirements into production systems while optimizing for cost, latency, and developer velocity.

Our summary, not Deepgram’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • 5+ years of program management or technical leadership in ML infrastructure, ML platforms, or AI tooling
  • Strong technical acumen in ML systems with hands-on experience as an ML engineer, systems engineer, or ML infrastructure engineer
  • Experience coordinating cross-functional ML programs from training through evaluation, serving, and monitoring
  • Ability to translate ML and research requirements into robust, scalable infrastructure
  • Comfortable navigating complex technical tradeoffs around accuracy, latency, and cost
  • Excellent communication with both technical and non-technical stakeholders
  • Experience in high-growth or startup environments

Nice to have

  • Hands-on experience with model serving frameworks like vLLM, TensorRT, or TorchServe
  • Experience optimizing LLM or speech/audio model inference through quantization, distillation, or KV-cache optimization
  • Familiarity with ML experiment tracking and versioning tools
  • Background with feature stores, vector databases, or real-time ML systems
  • Knowledge of cost optimization for GPU and ML workloads
  • Experience with multi-region model serving or edge deployment
  • Hands-on experience with PyTorch, CUDA, Hugging Face, or cloud ML platforms

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.