Transformers & Modern Architectures — Deep Learning & Neural Networks…
The self-attention-based architecture behind virtually every recent state-of-the-art model
Steps in Transformers & Modern Architectures
- Self-Attention Mechanism — advanced · How every token attends to every other token to build contextual representations
- The Transformer Architecture (Encoder-Decoder) — advanced · Multi-head attention, positional encoding and the full encoder-decoder stack
- BERT & Encoder-Only Models — advanced · Bidirectional pretraining and encoder-only models built for understanding tasks
- GPT & Decoder-Only Models — advanced · Autoregressive, decoder-only models built for generation tasks
- Vision Transformers (ViT) — advanced · Applying the transformer architecture to images by treating patches as tokens
- Fine-Tuning Pretrained Transformers — advanced · Adapting a large pretrained model to a specific downstream task
- Parameter-Efficient Fine-Tuning (LoRA, Adapters) — advanced · Fine-tuning huge models by updating only a small fraction of their parameters
Part of
- Deep Learning & Neural Networks roadmap — the full learning path