Skip to content
Learn AI by building
← World of AI

Deep Learning

Vision Transformers

Transformer models that process images as sequences of visual patches.

A vision transformer divides an image into fixed-size patches and converts each patch into an embedding. The resulting sequence is processed using self-attention in a way similar to tokens in a language model.

Vision transformers can learn long-range relationships between different parts of an image. They are used for classification, detection, segmentation and multimodal understanding.

JOIN NOW

Begin the first module

It is free, it is the real curriculum, and if it is not for you, you have lost nothing but an evening.

Join any time · Build AI skills at your pace