Paper Reading: An Image is Worth 16x16 Words (ViT)
This week we're reading **An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale**, the 2020 paper that introduced the Vision Transformer (ViT).
Paper: https://arxiv.org/abs/2010.11929
We'll break down the key ideas and walk through some of the code implementation.
Online event
Ask Maya about this · Get a personalized feed · Continue on WhatsApp