Skip to main content
CareerApp

Skill

Vision Transformer (ViT)

Data Science, Analytics and AI/ML

The Vision Transformer is a deep learning architecture, introduced by Google researchers in 2020, that applies the transformer model (originally designed for natural language processing) to image recognition by splitting images into patches treated like sequence tokens. Machine learning engineers and researchers use ViT and its variants for image classification, object detection, and other computer vision tasks, often achieving state-of-the-art results when trained on large datasets. It represents a shift away from convolutional neural networks as the default approach in computer vision.

See who is hiring

Open roles requiring Vision Transformer (ViT) (1)

Senior Machine Learning Engineer - Foundation Model

XPeng

Full-time · Santa Clara, CA · $174,720 – $295,680

This role involves building and optimizing large-scale vision-language-action foundation models that form the core of XPENG's autonomous driving systems. You'll design multi-modal architectures, lead pretraining strategies on massive fleet data, and collaborate across research and infrastructure teams to deploy intelligent models for next-generation vehicles.

Listed on XPeng’s careers site · Apply there ↗

Roles that use Vision Transformer (ViT)

Related skills

Curated neighbors in the taxonomy, whether or not employers ask for them together.

Turn on analytics and we load Google Analytics: Google gets the pages you open and what you do here — searches, jobs you view, jobs you apply to — and sets its own cookies. Leave it off and the only cookies we set are your login, your theme, and this answer. Accept All also records a yes to advertising, which nothing uses yet. Privacy Policy.