DevOps module 07

Observe user impact, not dashboard noise.

Monitoring tells you that something is wrong. Observability helps you determine why.

01

Metrics

Numeric trends for utilisation, saturation, errors, latency and business outcomes.

02

Logs

Timestamped events with structured context, correlation identifiers and searchable fields.

03

Traces

End-to-end request flow across services, queues, databases and external dependencies.

04

Events

Deployment, scaling, configuration and platform changes that explain behaviour.

05

SLOs

Explicit reliability targets derived from user-visible service expectations.

06

Alerts

Actionable notifications tied to impact, ownership, runbooks and escalation.

Alert-quality test

01

Is there customer impact?

A technical threshold without impact may belong on a dashboard rather than a pager.

02

Can someone act?

Every alert needs ownership, context and a tested first-response procedure.

03

Can we verify recovery?

Define the signal that proves service health has returned.