Trustworthy Software Lab
Module 5

Incident Response Basics

A calm, structured way to think when a production system misbehaves.

Common incidents

API is down
Database unavailable
Secret leaked
Unauthorized access attempt
Cost spike
Data accidentally deleted
Deployment broke production

Response questions

  1. What happened?
  2. When did it start?
  3. Who is affected?
  4. What changed recently?
  5. How do we contain it?
  6. How do we recover?
  7. How do we prevent it again?

AWS evidence sources

Incident questionEvidence source
What changed?CloudTrail
What failed?CloudWatch Logs
Which resource was affected?AWS Config history
Was there unusual access?IAM + CloudTrail
Did deployment cause it?CI/CD logs

Practice scenario

A public API suddenly starts returning 500 errors.
Where do you check first? Consider application logs, recent deployments, CloudTrail changes, database health, and API Gateway metrics. There is rarely one answer — trustworthy teams check evidence in parallel.
Mindset
Incident response is a learning moment. Every incident should end with a prevention step, not just a fix.