$174,720 – $295,680
Listed on XPeng’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves building and optimizing large-scale vision-language-action foundation models that form the core of XPENG's autonomous driving systems. You'll design multi-modal architectures, lead pretraining strategies on massive fleet data, and collaborate across research and infrastructure teams to deploy intelligent models for next-generation vehicles.
Our summary, not XPeng’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- Master's degree or higher in Computer Science, Electrical/Computer Engineering, or related field
- 3+ years of deep learning research or productization experience
- Proficiency in PyTorch and transformer-based model design
- Experience with large-scale pretraining or multi-modal modeling
- Understanding of representation learning, temporal modeling, and self-supervised or reinforcement learning
- Familiarity with distributed training frameworks and large-batch optimization
Nice to have
- PhD in CS/CE/EE or related field with 1+ years of industry experience
- Published work in top-tier AI conferences (CVPR, ICCV, NeurIPS, ICLR, ICML, ECCV)
- Experience building foundation models, end-to-end driving models, or LLM/VLM architectures (ViT, Flamingo, BEVFormer, RT-2, GRPO)
- Familiarity with RLHF, DPO, GRPO, trajectory prediction, or policy learning for control
- Cross-functional collaboration experience with infrastructure, perception, and planning teams