MLOps for Deep Learning — Deep Learning & Neural Networks Roadmap
Managing deep learning experiments and training at scale like real infrastructure
Steps in MLOps for Deep Learning
- Experiment Tracking for Deep Learning — advanced · Logging hyperparameters, metrics and checkpoints across long-running training jobs
- Distributed Training Basics — advanced · Data parallelism and model parallelism for training across multiple GPUs
- Data Pipelines for Large-Scale Training — advanced · Streaming and sharding datasets too large to fit in memory
- Model Versioning for Deep Learning — advanced · Tracking which checkpoint, dataset version and code produced a given model
Part of
- Deep Learning & Neural Networks roadmap — the full learning path