Vision Transformers (ViT) — Roadmap Step & Resources
Applying the transformer architecture to images by treating patches as tokens
Open on CachedInfo