# Staff Machine Learning Engineer - LLM Quantization & Deployment

Hiring organization: [XPeng](https://career.thegoodapps.co/organizations/xpeng)

Canonical page: https://career.thegoodapps.co/jobs/475a00c6-f6d7-4048-a00d-ce4d7b92986e

Listed on XPeng's own careers site. Applications go to them directly.

- Employment type: full time
- Seniority: Staff
- Location: Santa Clara, CA
- Salary: 215280 – 364320 USD per year

## Summary

This role involves developing and optimizing large language models for deployment on XPeng's custom AI hardware, with a focus on quantization techniques, inference pipelines, and autonomous driving performance validation. It suits experienced machine learning engineers who enjoy bridging research and production systems, particularly those with deep expertise in model optimization and on-device inference.

_Our summary, not XPeng's wording._

## Skills named

Python, PyTorch

## Required

- Master's degree in Computer Science, Computer Engineering, or Electrical Engineering, or equivalent experience
- 3-5 years of industry experience
- Strong understanding of Transformer architectures and LLM inference
- Hands-on experience quantizing or deploying deep learning models in production
- Proficiency with PyTorch and at least one inference or compilation stack
- Strong Python programming and software engineering skills
- Ability to work effectively across research, systems, infrastructure, and product teams

## Nice to have

- Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization
- Experience with quantization methods such as AWQ, GPTQ, or SmoothQuant
- Strong numerical analysis and systems engineering skills
- Experience with LLM runtimes including TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes
- Experience deploying LLMs on resource-constrained or heterogeneous hardware
- Contributions to model optimization, inference, compiler, or serving projects
- Publications at top-tier ML conferences such as NeurIPS, ICML, ICLR, or ACL

Apply on XPeng's site: https://job-boards.greenhouse.io/xpengmotors/jobs/8710020002
