Senior Machine Learning Engineer - LLM Quantization & Deployment
XPengSanta Clara, CA · full time · Senior
$174,720 – $295,680
Listed on XPeng’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role focuses on developing and optimizing large language models for deployment on XPENG's custom AI chip, with emphasis on quantization techniques to reduce model size and inference latency. It suits experienced machine learning engineers who have shipped quantized models in production and want to work on real-world autonomous driving applications.
Our summary, not XPeng’s wording. The full posting is on their site.
What they ask for
Required
- Master's degree in CS, CE, EE or equivalent
- 1-3 years of industry experience or new graduate
- Understanding of Transformer architectures and LLM inference
- Hands-on experience quantizing or deploying deep learning models in production
- PyTorch proficiency
- Proficiency with at least one inference or compilation stack
- Strong Python programming and software engineering skills
- Ability to collaborate across research, systems, infrastructure, and product teams
- Excellent communication and problem-solving skills
Nice to have
- Experience with weight-only, activation, KV-cache, dynamic, static, or mixed-precision quantization
- Experience with AWQ, GPTQ, SmoothQuant or related methods
- Strong numerical analysis and systems engineering skills
- Experience with LLM runtimes such as TensorRT-LLM, vLLM, SGLang, llama.cpp, ONNX Runtime, TVM, MLIR, or custom runtimes
- Experience deploying LLMs on resource-constrained or heterogeneous hardware
- Contributions to model optimization, inference, compiler, or serving projects
- Publications at NeurIPS, ICML, ICLR, ACL, or related conferences