# Vision Transformer (ViT)

Skill · Data Science, Analytics and AI/ML

Canonical page: https://career.thegoodapps.co/skills/vision-transformer-vit

The Vision Transformer is a deep learning architecture, introduced by Google researchers in 2020, that applies the transformer model (originally designed for natural language processing) to image recognition by splitting images into patches treated like sequence tokens. Machine learning engineers and researchers use ViT and its variants for image classification, object detection, and other computer vision tasks, often achieving state-of-the-art results when trained on large datasets. It represents a shift away from convolutional neural networks as the default approach in computer vision.

## Roles that use Vision Transformer (ViT)

- [Machine Learning Engineer](https://career.thegoodapps.co/occupations/machine-learning-engineer) — 1

Related skills: [PyTorch](https://career.thegoodapps.co/skills/pytorch), [Deep Learning](https://career.thegoodapps.co/skills/deep-learning), [TensorFlow](https://career.thegoodapps.co/skills/tensorflow), [Computer Vision](https://career.thegoodapps.co/skills/computer-vision)

## Open roles requiring Vision Transformer (ViT) (1)

- [Senior Machine Learning Engineer - Foundation Model](https://career.thegoodapps.co/jobs/ec8b86ef-f9f4-4260-b6cc-4330b8e073bd/senior-machine-learning-engineer-foundation-model-at-xpeng) at [XPeng](https://career.thegoodapps.co/organizations/xpeng) — Full-time, Santa Clara, CA, $174,720 – $295,680
