Vision Transformers (ViT) — Roadmap Step & Resources

Applying the transformer architecture to images by treating patches as tokens

Open on CachedInfo