Skip to main content
CareerApp
NVIDIA

Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA

Santa Clara, CA · full time · Senior

$224,000 – $431,250

Listed on NVIDIA’s own careers site. You apply with them directly — we never stand between you and the employer.

What this role is

Lead a team building an agentic platform to monitor, debug, and optimize large-scale generative AI models in production. This role suits experienced engineering managers who understand deep learning systems, observability platforms, and distributed inference infrastructure.

Our summary, not NVIDIA’s wording. The full posting is on their site.

Skills this role names

Log in to see which of these are already on your profile.

What they ask for

Required

  • BSc, MS, or PhD in Computer Science, Computer Engineering, or equivalent
  • 8+ years software engineering experience
  • 3+ years engineering management or technical leadership
  • Led teams building large-scale distributed systems, observability, ML infrastructure, or production AI systems
  • Understanding of LLM/VLM inference systems and production deployment patterns
  • Experience with logs, metrics, traces, profiling, alerting, dashboards, or debugging workflows
  • Strong programming and performance analysis skills
  • Cross-organization collaboration and communication

Nice to have

  • GPU performance analysis background
  • Distributed inference or model serving optimization experience
  • Reliability engineering background
  • Built observability or telemetry platforms for AI, ML, cloud, or distributed infrastructure
  • Experience with OpenTelemetry, Prometheus, Grafana, Jaeger, ClickHouse, or Elastic
  • Built agentic systems reasoning over logs, traces, performance data, or operational workflows
  • Hands-on production GenAI serving experience with metrics like TTFT, TPOT, throughput, KV cache pressure, and cost per token

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.