Module 5
Incident Response Basics
A calm, structured way to think when a production system misbehaves.
Common incidents
API is down
Database unavailable
Secret leaked
Unauthorized access attempt
Cost spike
Data accidentally deleted
Deployment broke production
Response questions
- What happened?
- When did it start?
- Who is affected?
- What changed recently?
- How do we contain it?
- How do we recover?
- How do we prevent it again?
AWS evidence sources
| Incident question | Evidence source |
|---|---|
| What changed? | CloudTrail |
| What failed? | CloudWatch Logs |
| Which resource was affected? | AWS Config history |
| Was there unusual access? | IAM + CloudTrail |
| Did deployment cause it? | CI/CD logs |
Practice scenario
A public API suddenly starts returning 500 errors.
Where do you check first? Consider application logs, recent deployments, CloudTrail changes, database health, and API Gateway metrics. There is rarely one answer — trustworthy teams check evidence in parallel.
Mindset
Incident response is a learning moment. Every incident should end with a prevention step, not just a fix.