Postmortems for Platform Incidents — Roadmap Step & Resources
Learning from outages that affected multiple downstream teams
- Level: beginner
- Estimated time: 1 days
- Roadmap: Data Platform & Infrastructure Engineering
Before this step
Interview questions
- Design an SLA and on-call structure for a shared data platform serving multiple internal teams
Study resources
- AWS Documentation (Article) — Official documentation for Amazon Web Services, covering compute, storage and data infrastructure. Background reading/viewing for: Postmortems for Platform Incidents.
- Kubernetes Documentation (Article) — Official documentation relevant to reliability engineering for platforms. Relevant to: Postmortems for Platform Incidents.
Part of
- Platform Reliability Engineering — section
- Data Platform & Infrastructure Engineering roadmap — the full learning path