$174,720 – $295,680
Listed on XPeng’s own careers site. You apply with them directly — we never stand between you and the employer.
What this role is
This role involves developing and training predictive world models that simulate how physical environments evolve, using multimodal data from autonomous vehicles and robots to improve driving and robotic policies. It suits machine learning engineers with deep learning expertise who want to work on foundational AI systems for autonomous systems at scale.
Our summary, not XPeng’s wording. The full posting is on their site.
Skills this role names
Log in to see which of these are already on your profile.
What they ask for
Required
- MS or PhD in Engineering, Computer Science, Deep Learning, Computer Vision, Generative Models or equivalent experience
- Strong applied deep learning experience including model architecture design and large-scale training
- 1-3+ years hands-on experience with PyTorch and distributed training frameworks
- Strong Python programming with software design skills
- Understanding of data structures, algorithms, code optimization and large-scale data processing
- Excellent problem-solving and experimental design skills
Nice to have
- Hands-on experience with generative models for video or 3D such as diffusion or flow matching
- Experience with world models or learned simulators for decision making and model-based reinforcement learning
- Experience with multimodal foundation models, video tokenizers or VAEs
- Experience with large-scale training infrastructure and performance optimization techniques