Know something is wrong before your users do, and know where to look when it is.
Everything this course covers, in the order you'll learn it. Open a stage to see why it matters and jump to its lesson.
The three signals and the questions each one answers.
Lesson 1: Metrics, logs and traces →The standard for metrics in cloud-native systems.
Lesson 2: Prometheus · coming soonTurn raw metrics into dashboards people actually read.
Lesson 3: Grafana · coming soonAlerts that wake you up only when something is truly broken.
Lesson 4: Alerting · coming soonSearch every server's logs from one place.
Lesson 5: Logs · coming soonFollow one request through many services.
Lesson 6: Tracing · coming soonWatch the cluster and every pod on it.
Lesson 7: Kubernetes monitoring · coming soonDecide how reliable is reliable enough, and how to respond when you're not.
Lesson 8: SLOs & on-call · coming soonWhat the paid platforms add, and when they're worth it.
Lesson 9: Hosted tools · coming soonStage order follows the community roadmap at roadmap.sh/devops, trimmed to what DevOps work needs.