1. Purpose
A Root Cause Analysis (RCA) is a structured method for determining why a security incident, control failure, vulnerability, or recurring problem occurred.
The objective is not simply to identify what happened, but to determine:
- What happened
- How it happened
- Why it happened
- Which controls failed or were missing
- Which contributing factors enabled the event
- Whether the issue is isolated or systemic
- What corrective actions are required
- Whether the risk has changed
- How to prevent recurrence
The RCA should be evidence-based and focused on improving the security management system rather than assigning blame.
2. When to Perform Root Cause Analysis
RCA should normally be performed for:
- SEV-1 and significant SEV-2 incidents
- Confirmed account compromise
- Cloud compromise
- Data breach
- Ransomware
- Malware incidents
- Privilege escalation
- Significant vulnerability exploitation
- Production security incidents
- Supplier security incidents
- Repeated security incidents
- Major control failures
- Significant audit findings
- Recurring security weaknesses
- Incidents where existing controls failed unexpectedly
For low-risk events, a simplified RCA may be sufficient.
3. Root Cause Analysis Information
| Field | Details |
|---|---|
| RCA ID | |
| Incident ID | |
| Incident Title | |
| Incident Type | |
| Severity | |
| Date of Incident | |
| RCA Start Date | |
| RCA Completion Date | |
| Incident Owner | |
| RCA Owner | |
| Security Lead | |
| Business Owner | |
| Participants | |
| Customer Impact | Yes / No |
| Personal Data Involved | Yes / No / Unknown |
| Supplier Involved | Yes / No |
| Status | Open / In Progress / Completed |
4. Problem Statement
Describe the problem in one clear statement.
The problem statement should describe the observed failure, not the presumed cause.
Example
An unauthorized user obtained access to a developer’s cloud identity and accessed production resources.
Avoid writing:
The incident happened because the developer was careless.
The first statement describes the problem. The second assumes a cause before investigation.
Problem Statement
5. Incident Summary
Document the known facts.
Include:
- What happened
- When it happened
- How it was detected
- Systems involved
- Accounts involved
- Information involved
- Business impact
- Customer impact
- How the incident was contained
Summary
6. Evidence Reviewed
Root cause conclusions should be supported by evidence.
| Evidence ID | Evidence Description | Source | Date/Time | Finding |
|---|---|---|---|---|
Potential evidence includes:
- Authentication logs
- CloudTrail
- Application logs
- Endpoint logs
- Network logs
- Email records
- Vulnerability reports
- Configuration records
- IAM records
- Change records
- Access reviews
- Security alerts
- Incident timeline
- Policies and procedures
- Training records
- Supplier records
- Previous incidents
- Audit findings
7. Establish the Timeline
Build the timeline before determining the root cause.
| Date/Time | Event | Evidence | Confirmed? |
|---|---|---|---|
| Initial activity | |||
| Authentication | |||
| Privilege change | |||
| Resource access | |||
| Detection | |||
| Reporting | |||
| Containment | |||
| Recovery |
The timeline should distinguish between:
- Confirmed facts
- Reasonable conclusions
- Unknown events
- Assumptions
8. What Happened?
Document the technical sequence.
Initial Event
Initial Access
Actions Performed
Privilege Changes
Systems Accessed
Information Accessed
Persistence
Lateral Movement
Exfiltration or Impact
Containment
Recovery
9. Root Cause Categories
Consider the following categories when identifying the cause.
9.1 Technology
Examples:
- Vulnerability
- Misconfiguration
- Weak authentication
- Missing MFA
- Excessive permissions
- Insecure application
- Unpatched software
- Inadequate logging
- Weak network segmentation
- Exposed service
- Insecure API
- Cloud configuration error
9.2 People
Examples:
- Lack of security awareness
- Inadequate training
- Human error
- Insufficient technical knowledge
- Incorrect execution of a procedure
- Failure to report suspicious activity
People-related observations should be supported by evidence and should not automatically be treated as the root cause.
9.3 Process
Examples:
- Missing procedure
- Outdated procedure
- Incomplete access review
- Weak change management
- Inadequate incident escalation
- Incomplete supplier review
- Missing vulnerability management process
9.4 Governance
Examples:
- Unclear ownership
- Inadequate risk assessment
- Risk accepted without appropriate controls
- Missing management oversight
- Security requirement not defined
- Control responsibility not assigned
9.5 Supplier / Third Party
Examples:
- Supplier control failure
- Subprocessor weakness
- Inadequate supplier monitoring
- Contractual security gap
- Third-party vulnerability
- Supplier notification delay
10. Five Whys Analysis
Use the 5 Whys technique to progressively identify the underlying cause.
Problem
A production AWS resource was accessed without authorization.
Why 1 — Why did this happen?
Why 2 — Why did that happen?
Why 3 — Why did that happen?
Why 4 — Why did that happen?
Why 5 — Why did that happen?
Root Cause Identified
The number of “Whys” does not have to be exactly five. Stop when the analysis reaches a cause that can be addressed through a meaningful control, process, technology, governance, or risk-treatment action.
11. Fishbone / Cause Categories
For complex incidents, examine multiple contributing factors.
People
- Training
- Awareness
- Skills
- Staffing
- Human interaction
Process
- Procedures
- Approvals
- Reviews
- Escalation
- Change management
Technology
- Systems
- Applications
- Cloud
- Configuration
- Authentication
- Monitoring
Information
- Data classification
- Data exposure
- Data handling
- Retention
- Access requirements
Supplier
- Third-party service
- Subprocessor
- Contract
- Monitoring
- Dependency
Governance
- Risk decisions
- Ownership
- Policies
- Management oversight
- Security requirements
Environment
- Business growth
- New technology
- Remote working
- Cloud migration
- Organizational changes
12. Direct Cause
The direct cause is the immediate technical or operational condition that allowed the incident to occur.
Example
An exposed access key was used to authenticate to the AWS environment.
Direct Cause:
13. Contributing Factors
Contributing factors are conditions that increased the likelihood or impact of the incident.
Examples:
- MFA was not enforced for the affected identity.
- Access permissions were broader than required.
- Security alerts were not configured for the activity.
- Credential rotation was not performed.
- Access reviews were overdue.
- Logging existed but was not actively monitored.
- Incident reporting was delayed.
Contributing Factors
| Factor | Evidence | Impact |
|---|---|---|
14. Root Cause
The root cause should explain why the direct cause was possible.
Root Cause Statement
A strong root-cause statement should identify a condition that the organization can actually address.
Example
The compromised AWS identity had broader permissions than required and was not subject to the organization’s expected MFA and privileged-access controls, allowing the stolen credential to be used to access production resources.
15. Control Failure Analysis
Determine whether existing controls:
- Did not exist
- Existed but were not implemented
- Were implemented incorrectly
- Operated inconsistently
- Operated but failed to prevent the event
- Detected the event but too late
- Detected the event effectively
- Were bypassed
- Were not applicable
| Control | Expected Operation | Actual Operation | Failure Type | Evidence |
|---|---|---|---|---|
| MFA | ||||
| IAM | ||||
| Logging | ||||
| Monitoring | ||||
| Vulnerability Management | ||||
| Change Management | ||||
| Incident Response |
16. Preventive Control Analysis
Ask:
What should have prevented this incident?
Examples:
- MFA
- Least privilege
- Network segmentation
- Secure configuration
- Vulnerability management
- Secure coding
- Security monitoring
- Email security
- Endpoint protection
- Backup
- Supplier controls
- Access reviews
Preventive Control
Was the Control Present?
Yes / No / Partial
Was It Effective?
Yes / No / Partial
Improvement Required
17. Detective Control Analysis
Ask:
What should have detected the incident earlier?
Examples:
- SIEM
- CloudTrail
- GuardDuty
- EDR
- WAF
- Application monitoring
- IAM alerts
- DLP
- Vulnerability monitoring
- User reporting
Detection Control
Did It Detect the Activity?
Yes / No / Partially
Detection Delay
Improvement Required
18. Why Was the Incident Not Prevented?
This is one of the most important RCA questions.
Consider:
- Was the control missing?
- Was the control not implemented?
- Was the control misconfigured?
- Was the control bypassed?
- Was the risk incorrectly assessed?
- Was the control not monitored?
- Was an exception approved?
- Was the exception still valid?
- Was the risk underestimated?
- Was the control ineffective?
Finding
19. Why Was the Incident Not Detected Earlier?
Consider:
- Missing logs
- Inadequate monitoring
- Incorrect alert thresholds
- Alert fatigue
- Monitoring coverage gap
- Logging disabled
- No centralized monitoring
- No ownership of alerts
- Alert generated but not investigated
- Detection capability did not cover the attack technique
Finding
20. Why Did the Incident Become Significant?
Determine what allowed the incident to increase in impact.
Consider:
- Excessive privilege
- Lack of segmentation
- Weak containment
- Delayed escalation
- Limited monitoring
- Data concentration
- Lack of backup
- Supplier dependency
- Weak recovery process
- Inadequate incident response
Finding
21. AWS SaaS Example
Consider a SaaS company where a developer’s AWS credential is compromised.
Observed Incident
An attacker uses the developer credential to authenticate to AWS and assumes a role with production access.
Direct Cause
Compromised developer credential was used to access the cloud environment.
Contributing Factors
- Developer identity had broader permissions than required.
- MFA protection was incomplete.
- Long-lived credentials were still active.
- Monitoring did not immediately alert on the unusual activity.
Five Whys
Why did the attacker access production?
→ The compromised identity could assume a production role.
Why could it assume the production role?
→ The role trust and permissions allowed the identity to perform the action.
Why were those permissions available?
→ Access had expanded as the development environment grew.
Why was excessive access not identified?
→ Periodic privileged-access review was incomplete.
Why was the review incomplete?
→ Ownership and review frequency were not clearly defined.
Root Cause
The organization’s cloud access governance did not consistently enforce least privilege and periodic privileged-access review, allowing a compromised development identity to retain broader production access than required.
Corrective Actions
- Remove unnecessary permissions.
- Implement stronger MFA enforcement.
- Replace long-lived credentials with short-lived authentication.
- Review IAM role trust relationships.
- Implement periodic privileged-access reviews.
- Improve detection for unusual role assumption.
- Update cloud access standards.
- Test the revised controls.
22. Recurring Incident Analysis
Determine whether the same or similar issue has happened before.
| Previous Incident | Similarity | Root Cause Similar? | Previous Action Effective? |
|---|---|---|---|
If a similar incident occurred previously, determine why the previous corrective action did not prevent recurrence.
Possible reasons:
- Action was never completed
- Action was completed but ineffective
- Root cause was incorrectly identified
- Corrective action addressed the symptom only
- Risk was accepted
- New technology introduced the same weakness
- Control degraded over time
23. Corrective Actions
Convert the root cause into measurable actions.
| Action ID | Root Cause | Corrective Action | Owner | Priority | Due Date | Evidence | Status |
|---|---|---|---|---|---|---|---|
| RCA-001 | High | ||||||
| RCA-002 | Medium |
A corrective action should address the cause, not merely the visible symptom.
Weak Action
Reset the compromised password.
Stronger Corrective Action
Implement phishing-resistant MFA and privileged-access controls for administrative identities and verify implementation through access testing.
24. Immediate vs Permanent Corrective Action
| Action Type | Purpose | Example |
|---|---|---|
| Immediate Action | Reduce current risk | Disable compromised account |
| Containment | Stop ongoing impact | Isolate affected workload |
| Remediation | Remove weakness | Patch vulnerable system |
| Corrective Action | Address root cause | Redesign access governance |
| Preventive Action | Reduce recurrence | Implement continuous monitoring |
25. Risk Reassessment
Determine whether the root cause changes the organization’s risk assessment.
| Risk | Previous Rating | New Rating | Treatment | Owner |
|---|---|---|---|---|
Consider:
- Likelihood
- Impact
- Threat landscape
- Vulnerability
- Existing controls
- Control effectiveness
- Customer impact
- Regulatory impact
- Business criticality
- Residual risk
Where appropriate, update:
- Risk Register
- Risk Treatment Plan
- Security Controls
- Statement of Applicability considerations
- Policies
- Procedures
- Standards
26. Effectiveness Verification
Corrective action should not automatically be considered successful simply because it was implemented.
Verify:
Action Implemented → Control Tested → Evidence Collected → Effectiveness Assessed → Risk Reassessed
| Action | Implemented? | Tested? | Effective? | Evidence | Verified By |
|---|---|---|---|---|---|
27. Root Cause Validation
Before closing the RCA, ask:
- Does the root cause explain the incident?
- Is it supported by evidence?
- Does it explain why the control failed?
- Does it explain why the incident was possible?
- Does it explain why detection or response was delayed?
- Is it actionable?
- Does the corrective action address the root cause?
- Could the same condition exist elsewhere?
- Has the broader environment been checked?
- Has residual risk been reassessed?
RCA Validation Result
28. Final RCA Findings
Direct Cause
Primary Root Cause
Contributing Factors
Control Failure
Detection Gap
Response Gap
Business Impact
Customer Impact
Data Impact
Residual Risk
29. Management Review
Management should review significant RCA findings.
Management Questions
- Is the identified root cause supported by evidence?
- Does the corrective action address the actual cause?
- Are additional resources required?
- Does the risk assessment need updating?
- Are additional controls required?
- Should other systems be reviewed for the same weakness?
- Should policies or procedures be changed?
- Should additional testing or tabletop exercises be conducted?
- Is residual risk acceptable?
- Does management approve closure of the RCA?
Management Decision
30. RCA Closure Criteria
The RCA may be closed when:
- Incident facts have been established
- Evidence has been reviewed
- Timeline has been validated
- Direct cause has been identified
- Root cause has been documented
- Contributing factors have been identified
- Control effectiveness has been reviewed
- Detection gaps have been assessed
- Corrective actions have been assigned
- Risk has been reassessed
- Corrective actions have appropriate owners and dates
- Required actions have been implemented or formally tracked
- Effectiveness has been verified where required
- Residual risk has been documented
- Management review has occurred where required
31. Recommended RCA Evidence
Maintain supporting evidence such as:
- Incident Report
- Incident Timeline
- Incident Investigation
- Evidence Register
- Security logs
- Cloud logs
- Access records
- Configuration records
- Vulnerability reports
- Change records
- Policies and procedures
- Access review records
- Security alerts
- Corrective Action Tracker
- Risk Register
- Control testing evidence
- Management approval
32. Relationship With Other Security Records
The RCA should connect to the broader incident-management process:
Security Event → Incident → Investigation → Evidence → Timeline → Root Cause Analysis → Corrective Action → Risk Reassessment → Control Improvement → Verification → Closure
Related documents may include:
- Security Event Register
- Event-to-Incident Decision Checklist
- Incident Register
- Incident Severity Matrix
- Incident Investigation Template
- Incident Timeline
- Evidence Preservation Procedure
- Incident Closure Report
- Post-Incident Review
- Corrective Action Tracker
- Risk Register
- Management Review
33. ISO 27001 Alignment
Root Cause Analysis supports the organization’s ISMS by helping it determine why security incidents and control failures occurred and whether corrective actions are necessary.
It can support:
- Incident management
- Information-security risk management
- Control effectiveness evaluation
- Corrective action
- Continual improvement
- Management review
- Risk reassessment
The exact RCA format is an organizational implementation choice. ISO/IEC 27001 does not require every organization to use a specific “Root Cause Analysis Template.”
The depth of the RCA should be proportionate to:
- Incident severity
- Business impact
- Information-security risk
- Customer impact
- Regulatory/contractual requirements
- Recurrence
- Control failure
34. Startup-Friendly RCA
A startup does not need a lengthy RCA for every security alert.
For a significant incident, the minimum practical RCA is:
Problem → Evidence → Timeline → Direct Cause → Contributing Factors → Root Cause → Control Failure → Corrective Action → Risk Reassessment → Verification
For an AWS SaaS company, a useful RCA can often be linked directly to:
- CloudTrail
- IAM
- GuardDuty
- Application logs
- Security alerts
- Incident timeline
- Configuration history
- Access reviews
- Corrective Action Tracker
- Risk Register
This keeps the RCA evidence-driven without creating unnecessary documentation.
35. Final RCA Audit Trail
Problem Identified → Evidence Collected → Timeline Established → Facts Validated → Direct Cause Identified → Contributing Factors Identified → Control Failure Assessed → Root Cause Determined → Corrective Actions Defined → Risk Reassessed → Actions Implemented → Effectiveness Tested → Residual Risk Reviewed → Management Review → RCA Closed
Final Principle
Don’t Stop at What Happened — Determine Why It Happened.
Establish the Facts + Preserve Evidence + Identify the Direct Cause + Understand Contributing Factors + Determine the Root Cause + Assess Control Failure + Correct the Cause + Reassess Risk + Verify Effectiveness + Document the Decision.
