Staff Machine Learning Engineer - LLM Quantization & Deployment
XPengSanta Clara, CA · full time · Staff
$215,280 – $364,320
Listed on XPeng’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves developing and optimizing large language models for deployment on XPeng's custom AI hardware, with a focus on quantization techniques, inference pipelines, and autonomous driving performance validation. It suits experienced machine learning engineers who enjoy bridging research and production systems, particularly those with deep expertise in model optimization and on-device inference.
Our summary, not XPeng’s wording. The full posting is on their site.
What they ask for
Required
- Master's degree in Computer Science, Computer Engineering, or Electrical Engineering, or equivalent experience
- 3-5 years of industry experience
- Strong understanding of Transformer architectures and LLM inference
- Hands-on experience quantizing or deploying deep learning models in production
- Proficiency with PyTorch and at least one inference or compilation stack
- Strong Python programming and software engineering skills
- Ability to work effectively across research, systems, infrastructure, and product teams
Nice to have
- Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization
- Experience with quantization methods such as AWQ, GPTQ, or SmoothQuant
- Strong numerical analysis and systems engineering skills
- Experience with LLM runtimes including TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes
- Experience deploying LLMs on resource-constrained or heterogeneous hardware
- Contributions to model optimization, inference, compiler, or serving projects
- Publications at top-tier ML conferences such as NeurIPS, ICML, ICLR, or ACL