1. Purpose
The Cloud Incident Response Procedure defines how the organization detects, reports, assesses, contains, investigates, eradicates, recovers from, and learns from information-security incidents involving cloud services.
The procedure is designed to ensure that cloud incidents are handled consistently, promptly, and with appropriate evidence while minimizing impact to:
- Information and personal data
- Customer services
- Production systems
- Cloud infrastructure
- Applications and APIs
- Identity and access management
- Confidentiality, integrity, and availability
- Legal, regulatory, and contractual obligations
The procedure supports the organization’s Information Security Management System (ISMS) and should be applied based on the organization’s risk assessment, incident-management requirements, contractual obligations, and applicable Statement of Applicability (SoA).
2. Scope
This procedure applies to security incidents involving:
- IaaS, PaaS, and SaaS services
- Cloud accounts and subscriptions
- Production and non-production environments
- Cloud applications and APIs
- Cloud databases and storage
- Cloud identity and access management
- Privileged accounts
- Service accounts and workload identities
- Cloud networking
- Containers and Kubernetes
- Serverless services
- Cloud-hosted applications
- Cloud backups and recovery systems
- Cloud security services
- Third-party cloud providers
- Cloud subprocessors
- CI/CD and DevOps platforms connected to cloud environments
- Cloud configuration and infrastructure-as-code
- Customer or personal information stored in cloud environments
3. Key Principle
Cloud incident response follows:
Detect → Validate → Classify → Notify → Contain → Investigate → Eradicate → Recover → Verify → Communicate → Learn → Improve
The objective is not simply to close an alert.
The organization should determine:
What happened → What was affected → What information was involved → What access was used → What risk exists → What was done → How recovery was verified → What evidence was retained → What needs to improve
4. What Is a Cloud Security Incident?
A cloud security incident is an event that has caused, or may have caused, unauthorized access, disclosure, alteration, destruction, disruption, compromise, or loss of control involving cloud-based information or systems.
Examples include:
- Compromised cloud administrator account
- Stolen access key
- Unauthorized IAM role assumption
- Public exposure of an S3 bucket
- Unauthorized access to a production database
- Malware or ransomware affecting cloud workloads
- Compromise of an API
- Unauthorized modification of security groups
- Disabled or altered logging
- Unauthorized creation of cloud resources
- Cloud cryptomining
- Compromised CI/CD credentials
- Malicious modification of production infrastructure
- Cloud provider security incident affecting the organization
- Accidental disclosure of customer information
- Compromise of a cloud SaaS account
- Suspicious OAuth application with excessive access
- Compromised container image
- Exploitation of a vulnerable internet-facing cloud application
A security alert is not automatically an incident. The organization should validate the event and determine whether it meets the organization’s incident criteria.
5. Roles and Responsibilities
| Role | Responsibility |
|---|---|
| Incident Manager | Coordinates the overall response |
| CISO / Security Lead | Provides security direction and risk assessment |
| Cloud/DevOps Team | Investigates and contains cloud infrastructure issues |
| IT / IAM Administrator | Handles account, identity, authentication, and access actions |
| Application Owner | Assesses application and business impact |
| Data/Privacy Owner | Assesses personal/customer information impact |
| Legal/Compliance | Determines applicable legal, regulatory, and contractual requirements |
| Business Owner | Assesses business and customer impact |
| Communications Lead | Coordinates approved internal/external communications |
| Supplier/Cloud Provider Owner | Coordinates with cloud providers or affected suppliers |
| Management | Approves major risk decisions and business actions |
For a startup, one individual may perform multiple roles. Responsibilities should nevertheless be clearly assigned.
6. Incident Sources
Cloud incidents may be identified through:
- SIEM alerts
- Cloud-native security monitoring
- IAM alerts
- CloudTrail or equivalent audit logs
- Cloud security services
- Vulnerability scanners
- WAF alerts
- EDR alerts
- Application monitoring
- Database monitoring
- Customer reports
- Employee reports
- Supplier notifications
- Cloud provider notifications
- Threat intelligence
- Penetration testing
- Internal audits
- Access reviews
- Configuration monitoring
- Data-loss prevention alerts
7. Incident Reporting
Anyone identifying a suspected cloud security incident should report it through the organization’s approved incident-reporting channel.
The initial report should capture, where known:
- Date and time detected
- Reporter
- Cloud provider
- Cloud account/subscription/project
- Affected application or service
- Affected environment
- Description of the event
- Source of detection
- Known affected resources
- Known affected users/accounts
- Possible information involved
- Initial business impact
- Evidence available
- Actions already taken
Employees should not independently delete evidence, modify logs, or perform destructive remediation unless authorized by the incident-response team or required for immediate containment.
8. Incident Validation
The Incident Manager or designated security team should determine whether the event represents a genuine security incident.
Validation may include:
- Reviewing the alert.
- Confirming the affected account/resource.
- Checking relevant logs.
- Verifying the activity against authorized changes.
- Identifying whether the activity was expected.
- Determining whether unauthorized access or modification occurred.
- Identifying potentially affected information.
- Establishing the approximate incident timeframe.
The result should be recorded as:
- False Positive
- Security Event
- Confirmed Incident
- Suspected Incident — Investigation Required
9. Incident Classification
Confirmed incidents should be classified according to organizational severity criteria.
An illustrative model is:
| Severity | Example |
|---|---|
| Low | Unauthorized activity with no material impact |
| Medium | Compromised non-production account or limited cloud resource |
| High | Compromised production identity or significant service impact |
| Critical | Confirmed customer-data compromise, major production compromise, ransomware, or prolonged critical-service disruption |
Severity should consider:
- Information sensitivity
- Customer impact
- Personal-data impact
- Production exposure
- Privileged access
- Number of affected systems
- Duration
- Availability impact
- Integrity impact
- Confidentiality impact
- Regulatory obligations
- Contractual obligations
- Business impact
- Ability to contain the incident
The organization’s approved risk/severity methodology should determine the final classification.
10. Immediate Containment
The incident team should take appropriate containment actions based on the incident.
Possible actions include:
Identity compromise
- Disable compromised accounts.
- Revoke active sessions.
- Revoke access keys/tokens.
- Reset credentials.
- Enforce MFA.
- Review recent authentication activity.
- Remove unauthorized roles.
- Review federation or SSO activity.
Compromised cloud workload
- Isolate the workload where practical.
- Restrict network access.
- Preserve relevant evidence.
- Stop malicious processes.
- Block malicious IP addresses where appropriate.
- Prevent further unauthorized access.
Exposed storage
- Remove unintended public access.
- Restrict bucket/container permissions.
- Review access logs.
- Determine whether information was accessed.
- Preserve evidence before making destructive changes where practical.
Compromised API
- Revoke exposed credentials.
- Rotate API keys.
- Restrict affected endpoints.
- Review API logs.
- Investigate unauthorized requests.
CI/CD compromise
- Disable compromised credentials.
- Suspend affected deployment pipelines where necessary.
- Review repository and pipeline activity.
- Rotate secrets.
- Validate recent deployments.
Containment should balance security needs against the risk of destroying evidence or unnecessarily disrupting business services.
11. Evidence Preservation
Relevant evidence should be preserved in a controlled manner.
Potential evidence includes:
- Cloud audit logs
- IAM activity
- Authentication logs
- Network logs
- WAF logs
- Application logs
- Database logs
- Security alerts
- Configuration history
- Cloud resource activity
- API logs
- CI/CD logs
- Git activity
- Container information
- Vulnerability reports
- Screenshots
- Incident tickets
- Provider communications
- Relevant emails
- System timestamps
Evidence should be:
- Protected from unauthorized modification
- Accessible only to authorized personnel
- Traceable to the incident
- Retained according to applicable retention requirements
- Stored securely
Where forensic investigation may be required, evidence should be preserved before unnecessary modification or deletion.
12. Investigation
The investigation should establish, as far as reasonably possible:
What happened?
- What activity occurred?
- When did it start?
- When was it detected?
- When did it stop?
Who or what was involved?
- User account
- Administrator
- Service account
- Workload identity
- API credential
- Application
- Supplier
- Cloud resource
What was affected?
- Applications
- Databases
- Storage
- Virtual machines
- Containers
- Cloud accounts
- Networks
- APIs
What information was involved?
- Internal information
- Confidential information
- Customer information
- Personal information
- Credentials
- Source code
- Security configuration
How did the incident occur?
Examples:
- Stolen credentials
- Phishing
- Excessive privileges
- Vulnerability exploitation
- Misconfiguration
- Compromised supplier
- Exposed secret
- Insecure API
- Malicious insider activity
- Software supply-chain compromise
13. Impact Assessment
The organization should assess:
Confidentiality
Was information accessed or disclosed without authorization?
Integrity
Was information, configuration, code, or infrastructure modified without authorization?
Availability
Was a system or service unavailable or degraded?
Customer impact
Were customer systems, data, or services affected?
Privacy impact
Was personal information potentially accessed, disclosed, altered, or lost?
Regulatory impact
Could the incident trigger regulatory notification or other obligations?
Contractual impact
Do customer or supplier agreements contain notification requirements?
14. Cloud Provider Coordination
Where the incident involves a cloud provider or cloud-hosted supplier, the organization should:
- Identify the affected provider/service.
- Review the provider’s incident-management process.
- Notify the provider through the approved channel where required.
- Request relevant incident information.
- Record provider communications.
- Determine whether the provider has identified affected systems or data.
- Validate the provider’s information against organizational evidence.
- Track provider corrective actions where applicable.
A supplier’s statement should not automatically be treated as the organization’s complete impact assessment. The organization should determine its own exposure and responsibilities.
15. Eradication
After sufficient investigation and containment, the organization should remove the cause or mechanism of compromise.
Actions may include:
- Removing malware
- Deleting unauthorized cloud resources
- Removing malicious IAM policies
- Removing unauthorized users
- Rotating credentials
- Replacing compromised keys
- Patching vulnerabilities
- Correcting cloud configuration
- Removing malicious code
- Rebuilding compromised workloads
- Updating container images
- Removing malicious CI/CD changes
- Closing exposed network paths
Where appropriate, compromised infrastructure should be rebuilt from a trusted configuration rather than simply modified in place.
16. Credential and Secret Rotation
Where credentials may have been exposed, the organization should assess whether rotation is required.
Potentially affected credentials include:
- Cloud access keys
- IAM credentials
- API keys
- Database passwords
- Application secrets
- CI/CD credentials
- SSH keys
- OAuth tokens
- Service-account credentials
- Encryption-related credentials
Rotation should be performed through approved secure mechanisms.
For example, an AWS SaaS environment may use AWS Secrets Manager or another approved secrets-management solution rather than storing credentials directly in application code.
17. Recovery
Recovery should restore systems to a trusted and secure operating state.
Activities may include:
- Restoring from trusted backups
- Rebuilding cloud workloads
- Deploying patched versions
- Restoring databases
- Reconfiguring IAM
- Restoring network controls
- Re-enabling required services
- Validating monitoring
- Validating logging
- Performing security testing
- Confirming configuration baselines
- Monitoring systems closely after recovery
Recovery should not be considered complete merely because the application is available.
18. Recovery Verification
Before closing the incident, the responsible team should verify:
- Unauthorized access has been removed.
- Compromised credentials have been addressed.
- Vulnerabilities have been remediated or appropriately treated.
- Cloud configuration is secure.
- Logging is operational.
- Monitoring is operational.
- Security controls are functioning.
- Affected applications are functioning correctly.
- Data integrity has been verified where applicable.
- Required customer/regulatory actions have been completed.
- Residual risks have been identified and addressed.
19. Communication and Notification
Incident communications should be coordinated through authorized personnel.
Potential stakeholders include:
- Management
- Security team
- IT/Cloud team
- Application owners
- Customers
- Suppliers
- Cloud providers
- Legal counsel
- Regulators
- Law enforcement, where applicable
- Cyber insurance provider
Notifications should be based on applicable:
- Laws and regulations
- Customer contracts
- Supplier contracts
- Data-processing agreements
- Insurance requirements
- Internal escalation requirements
The organization should avoid speculative statements before the facts have been established.
20. Customer and Personal Data Assessment
Where customer or personal information may be involved, the organization should document:
- Categories of information involved
- Approximate number of affected records, where known
- Categories of affected individuals
- Systems involved
- Exposure period
- Evidence of access
- Evidence of exfiltration, if available
- Containment measures
- Risk to affected individuals
- Applicable notification requirements
- Required customer/regulatory communications
The organization should involve the appropriate privacy/legal personnel.
21. Root Cause Analysis
For significant incidents, the organization should determine the underlying cause.
Example:
Incident: Unauthorized access to production AWS environment
Immediate cause: Compromised developer credentials
Contributing factors:
- Excessive privileges
- Inconsistent MFA
- Long-lived credentials
- Insufficient monitoring
Root cause: Inadequate identity and access-management controls for privileged cloud access.
Corrective actions:
- Enforce MFA
- Implement least-privilege roles
- Remove unnecessary access keys
- Implement stronger privileged-access controls
- Improve monitoring
- Conduct periodic access reviews
22. Corrective and Preventive Actions
Incident findings should be transferred into the organization’s corrective-action or risk-management process.
Actions may address:
- IAM
- MFA
- Privileged access
- Network controls
- Secure configuration
- Vulnerability management
- Logging
- Monitoring
- Backup
- Application security
- Supplier management
- Employee awareness
- Cloud architecture
- Security policies
- Procedures
- Technical controls
Each action should have:
- Finding
- Root cause
- Corrective action
- Action owner
- Priority/risk
- Target date
- Status
- Evidence
- Verification
- Closure approval
23. Lessons Learned
For significant incidents, the organization should conduct a post-incident review.
Questions should include:
- How was the incident detected?
- Was detection sufficiently fast?
- Was escalation effective?
- Were responsibilities clear?
- Was sufficient evidence available?
- Were logs adequate?
- Did containment work?
- Were recovery procedures effective?
- Were customers or regulators notified appropriately?
- What controls failed or were missing?
- What should be changed?
- Does the risk assessment need updating?
- Does the SoA or control implementation need updating?
- Should incident-response testing be repeated?
24. Incident Closure
An incident should only be closed after:
- Investigation is sufficiently complete.
- Containment is confirmed.
- Eradication is complete or formally tracked.
- Recovery is verified.
- Required notifications are completed.
- Evidence is retained.
- Corrective actions are assigned.
- Residual risk is assessed.
- Required management approval is obtained.
- Relevant registers are updated.
The incident record should contain the final status and closure date.
25. AWS SaaS Example
Consider a SaaS company hosting its production platform on AWS.
An alert identifies suspicious activity associated with an administrator IAM identity.
Step 1 — Detect
CloudTrail and security monitoring identify unusual API calls.
Step 2 — Validate
The security team confirms that the activity was not part of an approved deployment or administrative activity.
Step 3 — Classify
The account has privileged production access, so the incident is classified as High or Critical according to the organization’s severity methodology.
Step 4 — Contain
The team:
- Disables or restricts the affected identity.
- Revokes active sessions where applicable.
- Rotates potentially compromised credentials.
- Reviews recent IAM changes.
- Restricts suspicious activity.
Step 5 — Investigate
The team reviews:
- CloudTrail
- IAM activity
- VPC/network logs
- Application logs
- WAF logs
- Security alerts
- Configuration changes
Step 6 — Determine Impact
The team establishes whether the attacker accessed:
- EC2/ECS workloads
- RDS databases
- S3 buckets
- Secrets Manager
- Customer information
- Production configuration
Step 7 — Eradicate
Unauthorized IAM policies and resources are removed, vulnerabilities are patched, and affected credentials are rotated.
Step 8 — Recover
Affected services are rebuilt or restored from trusted configurations.
Step 9 — Verify
The team confirms:
- MFA is enforced.
- Least-privilege roles are implemented.
- Logging is operational.
- Monitoring is active.
- No unauthorized access remains.
- Customer data exposure has been assessed.
Step 10 — Improve
The organization updates its:
- Cloud Access Review
- Cloud Security Risk Assessment
- Secure Configuration Standard
- Incident Response controls
- Risk Register
- Corrective Action Register
The incident therefore becomes an input into continual improvement rather than simply a closed ticket.
26. Startup-Friendly Implementation
A startup does not need a large incident-response department to implement this procedure.
A practical model can use:
Minimum capability
- One incident owner
- One technical/cloud responder
- One management escalation point
- One secure incident register
- One communication channel
- Centralized cloud logging
- MFA
- Credential/secret rotation capability
- Backup and recovery process
As the company grows
Add:
- SIEM
- Cloud security monitoring
- Formal incident severity matrix
- On-call coverage
- Forensic capability
- External incident-response provider
- Customer notification workflow
- Tabletop exercises
- Automated alerting
- Security metrics
The level of capability should be proportionate to the organization’s size, cloud exposure, risk, customer commitments, regulatory requirements, and business criticality.
27. Required Records and Evidence
Typical evidence may include:
- Incident reports
- Incident register
- Security alerts
- Cloud audit logs
- IAM logs
- Authentication logs
- Network logs
- Application logs
- Investigation notes
- Screenshots
- Provider communications
- Evidence preservation records
- Containment records
- Credential rotation evidence
- Vulnerability remediation records
- Recovery validation
- Customer/regulatory notifications
- Root-cause analysis
- Corrective-action records
- Lessons-learned report
- Management approval
- Incident-response test records
Actual passwords, access keys, API secrets, private keys, or other credentials should never be stored in the incident record.
28. Relationship With Other ISMS Documents
The Cloud Incident Response Procedure should connect with:
| Document | Relationship |
|---|---|
| Information Security Incident Management Policy | Overall incident-management requirements |
| Incident Register | Central record of incidents |
| Risk Register | Records significant resulting risks |
| Cloud Security Policy | Defines cloud-security requirements |
| Cloud Security Risk Assessment | Identifies cloud-related risks |
| Cloud Secure Configuration Standard | Defines secure configuration expectations |
| Cloud Access Review | Validates identity and access controls |
| Vulnerability Management Procedure | Handles vulnerabilities discovered during investigation |
| Supplier Incident Response Procedure | Handles supplier-related incidents |
| Business Continuity Plan | Supports recovery from major disruption |
| Disaster Recovery Plan | Supports technical recovery |
| Data Breach Procedure | Handles privacy/data-breach requirements |
| Corrective Action Register | Tracks remediation |
| Lessons Learned Register | Tracks improvement actions |
29. Internal Audit Checklist
An auditor can verify:
- Cloud incident-response procedure is approved.
- Roles and responsibilities are defined.
- Incident reporting channels exist.
- Cloud incidents are classified.
- Escalation requirements are defined.
- Cloud logs are available.
- Evidence preservation is addressed.
- Containment procedures exist.
- Credential rotation is addressed.
- Cloud provider escalation is addressed.
- Customer/personal-data impact is assessed.
- Legal/regulatory requirements are considered.
- Recovery is verified.
- Incident records are maintained.
- Root-cause analysis is performed for significant incidents.
- Corrective actions are tracked.
- Lessons learned are documented.
- Incident-response exercises or tests are performed where appropriate.
- Risks and controls are updated following significant incidents.
30. ISO 27001 Connection
This procedure supports the organization’s implementation of applicable information-security incident-management and cloud-security controls.
The exact controls and documented information applicable to the organization should be determined through:
Context → Risk Assessment → Risk Treatment → Applicable Controls → Statement of Applicability → Implementation → Evidence
The procedure should therefore not be treated as proof of compliance by itself.
An auditor will typically want to see evidence that the organization can actually execute the process, such as incident records, alerts, investigation evidence, containment actions, recovery validation, corrective actions, and lessons learned.
31. Final Cloud Incident Response Audit Trail
A complete incident should ideally produce a traceable chain:
Incident Detected
→ Incident Reported
→ Event Validated
→ Incident Classified
→ Stakeholders Notified
→ Evidence Preserved
→ Containment Performed
→ Investigation Conducted
→ Impact Assessed
→ Cloud Provider Engaged
→ Root Cause Identified
→ Eradication Completed
→ Recovery Performed
→ Recovery Verified
→ Notifications Completed
→ Corrective Actions Assigned
→ Residual Risk Assessed
→ Lessons Learned
→ Management Review
→ Incident Closed
→ Controls/Risks Updated
32. Final Principle
Effective Cloud Incident Response = Fast Detection + Controlled Containment + Evidence Preservation + Accurate Impact Assessment + Secure Recovery + Verified Closure + Continual Improvement
The objective is not simply to respond quickly.
The organization should be able to demonstrate:
What happened, what was affected, what information was involved, how the incident was contained, how recovery was verified, what evidence supports the conclusion, what residual risk remains, and what was changed to prevent recurrence.
