← World of AI
Deep Learning
Vision Transformers
Transformer models that process images as sequences of visual patches.
A vision transformer divides an image into fixed-size patches and converts each patch into an embedding. The resulting sequence is processed using self-attention in a way similar to tokens in a language model.
Vision transformers can learn long-range relationships between different parts of an image. They are used for classification, detection, segmentation and multimodal understanding.
Also in Deep Learning
JOIN NOW
Begin the first module
It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.
Join any time · Build AI skills at your pace