DevOps module 08

Troubleshoot with evidence, not panic.

A disciplined incident method reduces recovery time and prevents random changes from making the failure worse.

01

Stabilise

Protect users first: stop harmful automation, freeze risky changes and restore a known-good path.

02

Scope

Identify affected users, regions, services, versions and the precise start time.

03

Correlate

Compare failures with deployments, configuration, infrastructure and dependency events.

04

Hypothesise

Rank likely causes from evidence; do not change multiple variables at once.

05

Test

Run the smallest safe experiment that can confirm or reject one hypothesis.

06

Recover

Rollback, roll forward or fail over with explicit validation criteria.

07

Learn

Capture timeline, contributing factors, control gaps and owned corrective actions.

First five questions

What changed?

Deployment, secret, certificate, network, policy, dependency or infrastructure.

Who is affected?

All users or one path, tenant, region, browser, node or identity group.

What still works?

Healthy paths narrow the fault domain faster than staring at the failure alone.

Can we roll back?

Know the last verified release and the data compatibility implications.