Skill
Vision Transformer (ViT)
Data Science, Analytics and AI/ML
The Vision Transformer is a deep learning architecture, introduced by Google researchers in 2020, that applies the transformer model (originally designed for natural language processing) to image recognition by splitting images into patches treated like sequence tokens. Machine learning engineers and researchers use ViT and its variants for image classification, object detection, and other computer vision tasks, often achieving state-of-the-art results when trained on large datasets. It represents a shift away from convolutional neural networks as the default approach in computer vision.
Open roles requiring Vision Transformer (ViT) (1)
Senior Machine Learning Engineer - Foundation Model
XPeng
Full-time · Santa Clara, CA · $174,720 – $295,680
This role involves building and optimizing large-scale vision-language-action foundation models that form the core of XPENG's autonomous driving systems. You'll design multi-modal architectures, lead pretraining strategies on massive fleet data, and collaborate across research and infrastructure teams to deploy intelligent models for next-generation vehicles.
Listed on XPeng’s careers site · Apply there ↗
Roles that use Vision Transformer (ViT)
Related skills
Curated neighbors in the taxonomy, whether or not employers ask for them together.