Field note
After a capacity incident: a calmer debrief agenda
Begin with a timeline built from monitors and deploy records, not from memory. Mark when customers first felt pain, when alerts fired, when scale actions started, and when service recovered.
Separate contributing factors into workload shape, capacity posture, and human response. Most events involve more than one. Treating them as a single root cause usually leaves the next peak unprotected.
Ask what a monitor would have needed to show thirty minutes earlier. That question produces actionable instrumentation items faster than a generic “improve observability” action.
End with owners and dates for a small number of changes. A long backlog of vague monitoring tasks is how debriefs lose their value before the next busy season arrives.