SPC Unit 4: Complete Concepts Guide
Unit 4: Monitoring, Auditing and Management -> Generated and Prepared By Thiruselvan (ThiruXD)
1. Introduction to Monitoring, Auditing and Management
1.1 Why Cloud Security Does Not End at Design
Cloud security does not end after designing access controls, encryption mechanisms, and secure architecture. A cloud environment must be continuously monitored, audited, and managed because users, workloads, networks, applications, and threats change every day.
- Monitoring provides live visibility into what is happening in the environment.
- Auditing provides recorded evidence of what happened, who performed an action, when it occurred, and whether the action was successful or failed.
1.2 Why Cloud Monitoring is More Complex
In cloud computing, monitoring is more complex than in a traditional data center because:
- Resources are elastic and distributed
- Resources are often controlled through APIs
- Virtual machines, containers, serverless functions, databases, load balancers, storage buckets, and identity services can be created or modified rapidly
Without proper monitoring and logging, unauthorized access, privilege abuse, misconfiguration, data exfiltration, and malicious traffic may remain hidden until serious damage occurs.
1.3 Did You Know?
In cloud environments, many security incidents are first detected through unusual log patterns, such as:
- Impossible travel sign-ins
- Repeated failed logins
- Suspicious API calls
- Unexpected configuration changes
2. Key Concepts and Glossary
| Term | Meaning in Cloud Security |
|---|---|
| Proactive Monitoring | Continuous observation of cloud assets, users, applications and networks to detect issues before they become serious incidents |
| Incident Response | A structured process for preparing for, detecting, analyzing, containing, eradicating and recovering from security incidents |
| Audit Log | A record of security-relevant activity, such as login attempts, administrative actions, data access, configuration changes and network events |
| Alert | A notification generated when a rule, threshold or anomaly suggests a possible problem or security event |
| Unauthorized Access | Access to a system, account, resource or data without proper permission or outside approved policy |
| Malicious Traffic | Network communication associated with attacks such as scanning, malware command-and-control, DDoS, data exfiltration or exploitation attempts |
| Privilege Abuse | Misuse of elevated permissions by an administrator, compromised account, insider or service identity |
| Tamper-Proof Logs | Logs protected so that unauthorized modification, deletion or concealment is prevented or at least detectable |
| QoS | Quality of Service, referring to performance, availability, latency, reliability and service-level expectations |
| SIEM | Security Information and Event Management, a platform that collects, correlates, analyzes and reports security events |
3. Proactive Activity Monitoring
3.1 What is Proactive Activity Monitoring?
Proactive activity monitoring is the continuous observation of cloud resources, users, services, APIs, and network behavior to identify security and operational problems early. Instead of waiting for a failure or breach report, proactive monitoring uses logs, metrics, traces, alerts, and dashboards to maintain awareness of the cloud environment.
3.2 Activity at Many Layers
In cloud infrastructure, activity happens at many layers:
- A user may sign in through an identity provider
- An administrator may change a firewall rule
- An application may access a database
- A virtual machine may send outbound traffic
- A storage bucket may be made public
Each activity generates telemetry that must be collected and analyzed.
3.3 Questions a Good Monitoring System Answers
- Which users accessed sensitive resources?
- Were there failed login attempts?
- Did a workload communicate with unknown IP addresses?
- Were new administrator roles assigned?
- Did CPU, memory or response time cross a threshold?
- Was a storage policy changed unexpectedly?
3.4 Proactive Monitoring Lifecycle
Collect telemetry → Normalize & enrich → Correlate & detect → Alert & prioritize → Investigate & respond → Improve controls3.5 Monitoring Areas
| Monitoring Area | Examples of Activity to Monitor | Security Value |
|---|---|---|
| Identity activity | Sign-ins, failed logins, MFA failures, password resets, role assignments | Detects account compromise, brute force attacks and privilege changes |
| Control-plane activity | API calls, configuration changes, resource creation or deletion | Reveals unauthorized administration and misconfiguration |
| Network activity | Flow logs, firewall logs, DNS queries, WAF events, VPN logs | Detects scanning, lateral movement, command-and-control and exfiltration |
| Compute activity | VM events, container logs, process behavior, patch status | Detects malware, abnormal processes and insecure hosts |
| Storage activity | Object read/write/delete, policy changes, public access events | Protects sensitive files and supports data-loss investigation |
| Application activity | Authentication events, business transactions, errors, API requests | Supports fraud detection, debugging and user-level investigation |
Best Practice: Start monitoring from the most sensitive assets first: identity services, administrator actions, internet-facing services, critical databases, and storage containing confidential data.
4. Incident Response
4.1 What is Incident Response?
Incident response is the organized approach used to handle cybersecurity incidents. A cloud incident may include:
- Compromised credentials
- Exposed storage
- Malware on a virtual machine
- Data leakage
- DDoS attack
- Suspicious API activity
- Unauthorized database access
- Abuse of system privileges
Cloud incident response must be planned before an incident occurs. Teams should know:
- Who is responsible
- Which logs are required
- How to isolate resources
- How to preserve evidence
- How to communicate with stakeholders
- How to restore service safely
4.2 Cloud-Specific Response Actions
In cloud environments, many response actions are performed through automation and APIs:
- Disabling keys
- Isolating instances
- Changing security groups
- Rotating secrets
- Taking forensic snapshots
4.3 Incident Response Phases
1. Preparation → 2. Detection & Analysis → 3. Containment → 4. Eradication → 5. Recovery → 6. Lessons Learned| Phase | Cloud-Specific Activities | Output |
|---|---|---|
| 1. Preparation | Define playbooks, enable logging, create response roles, prepare forensic storage, train responders | Incident response plan and readiness checklist |
| 2. Detection and Analysis | Review SIEM alerts, IAM logs, network flows, endpoint logs and application evidence | Confirmed incident scope, severity and affected assets |
| 3. Containment | Disable compromised users, revoke tokens, isolate VMs, block IPs, freeze storage access | Attack is stopped from spreading |
| 4. Eradication | Remove malware, patch vulnerability, rotate keys, remove malicious rules or backdoors | Root cause and persistence mechanisms removed |
| 5. Recovery | Restore services, monitor closely, validate backups and business functions | Secure return to normal operations |
| 6. Lessons Learned | Update controls, detection rules, training, architecture and documentation | Improved security posture |
Important Point: During a cloud incident, never delete suspicious resources immediately. Capture evidence such as logs, snapshots, configuration history, and access records before cleanup whenever legal and organizational policies allow.
5. Monitoring for Unauthorized Access
5.1 What is Unauthorized Access?
Unauthorized access occurs when a user, service account, application, or attacker accesses a resource without valid permission or outside approved policy.
5.2 Common Causes
- Stolen credentials
- Weak passwords
- Missing MFA
- Overly permissive IAM policies
- Leaked API keys
- Misconfigured storage
- Exposed management ports
- Compromised service identities
5.3 Detection Approach
Monitoring for unauthorized access requires combining:
- Identity logs
- Authorization decisions
- Resource access logs
- Contextual signals
A single failed login may not be critical, but repeated failures from many countries, a successful login from an unusual location, or a new administrator role assigned after a suspicious login can indicate compromise.
5.4 High-Risk Actions to Monitor
- Root or owner account use
- Creation of access keys
- Disabling logging
- Modifying security policies
- Exporting data
- Changing network rules
- Deleting backups
- Access outside normal working hours or approved locations
5.5 Unauthorized Access Signals
| Unauthorized Access Signal | Possible Meaning | Recommended Response |
|---|---|---|
| Multiple failed login attempts | Password guessing or brute-force attempt | Trigger alert, enforce MFA, rate-limit, investigate source |
| Successful login from unusual location | Credential theft or impossible travel | Require step-up authentication and verify user |
| New admin role assignment | Privilege escalation | Review approver, ticket, identity and timing |
| Access from unknown device | Compromised password or unmanaged endpoint | Check device compliance and conditional access |
| API key used from new IP | Leaked credential or automation drift | Rotate key, restrict source IP and review code repositories |
6. Detection of Malicious Traffic
6.1 What is Malicious Traffic?
Malicious traffic refers to network communication associated with attacks or suspicious behavior. Examples include:
- Port scanning
- Vulnerability exploitation
- Malware command-and-control communication
- DDoS traffic
- Suspicious DNS queries
- TOR or proxy access
- Unexpected outbound connections
- Data exfiltration
- Lateral movement between cloud workloads
6.2 Sources for Traffic Visibility
Cloud platforms provide several sources:
- Virtual network flow logs
- Firewall logs
- Load balancer logs
- DNS logs
- WAF logs
- IDS/IPS alerts
- API gateway logs
- Endpoint telemetry
6.3 Detection Techniques
| Technique | Description |
|---|---|
| Signature-based | Identifies known threats |
| Behavior-based | Identifies unusual patterns (e.g., server sending large volumes of data to unknown country) |
6.4 Traffic Classification
Allowed traffic → Permit
Suspicious traffic → Alert
Malicious traffic → Block6.5 Traffic Types and Detection Sources
| Traffic Type | Common Detection Source | Example Alert |
|---|---|---|
| Port scanning | VPC/VNet flow logs, IDS, firewall logs | Many denied connections to different ports from one source |
| DDoS traffic | Load balancer metrics, CDN/WAF logs, network telemetry | Sudden spike in requests or bandwidth from distributed sources |
| Command-and-control | DNS logs, threat intelligence, outbound proxy logs | Connection to known malicious domain or rare destination |
| Data exfiltration | Flow logs, storage access logs, DLP, CASB | Large outbound transfer from sensitive workload |
| Web attack | WAF logs, application logs, API gateway logs | SQL injection, XSS, path traversal or API abuse pattern |
| Lateral movement | East-west flow logs, endpoint telemetry | Unusual internal connections between unrelated workloads |
Best Practice: Store network-flow logs long enough to support investigations. Many attacks are discovered days or weeks after the first malicious connection.
7. Prevention of Abuse of System Privileges
7.1 What is Privilege Abuse?
System privileges allow users and services to perform powerful actions such as:
- Creating resources
- Changing security settings
- Accessing sensitive data
- Managing encryption keys
- Deleting logs
- Modifying networks
Abuse of privileges may be:
- Intentional insider misuse
- Accidental misuse
- Result of an attacker compromising an administrator account
7.2 Prevention Approach
| Control | How It Prevents Privilege Abuse | Example |
|---|---|---|
| Least privilege | Limits the damage any account can cause | Developer can deploy app but cannot change billing or security logs |
| Separation of duties | Prevents one person from approving and executing sensitive actions alone | Key deletion requires security and operations approval |
| Just-in-time access | Gives elevated rights only for a limited time | Admin role active for two hours after approval |
| Privileged session monitoring | Records commands and actions for accountability | Session log is reviewed after database maintenance |
| MFA for administrators | Reduces risk from stolen passwords | Admin console requires authenticator approval |
| Immutable logging | Prevents attackers from hiding privileged actions | Audit logs written to locked storage |
7.3 Best Practices
- Avoid long-lived static administrator credentials
- Use temporary credentials and managed identities
- Use break-glass accounts for emergencies
- Implement privileged identity management
- Conduct automated policy reviews
8. Events and Alerts Management
8.1 Definitions
| Term | Definition |
|---|---|
| Event | Any recorded activity in a system |
| Alert | A notification generated when one or more events meet a defined condition |
| Incident | A confirmed security event requiring response |
Note: Not every event is an alert, and not every alert is an incident.
8.2 Event and Alert Management Process
- Collect events
- Define alert rules
- Assign severity
- Reduce noise
- Route notifications
- Track actions until closure
8.3 Alert Fatigue
Poor alert management creates alert fatigue. If analysts receive too many low-quality alerts, they may ignore important warnings. Therefore, cloud security teams must:
- Tune rules
- Suppress duplicates
- Enrich alerts with context
- Prioritize alerts based on business impact and threat severity
8.4 Alert Severity Levels
| Severity | Typical Condition | Expected Action |
|---|---|---|
| Critical | Confirmed breach, active exfiltration, root/admin compromise, production outage | Immediate incident response and leadership notification |
| High | Likely compromise, high-risk policy change, malware detection, privilege escalation | Rapid investigation and containment |
| Medium | Suspicious behavior requiring review, abnormal access, repeated failures | Analyze within defined SLA and tune detection if needed |
| Low | Informational anomaly or policy drift | Review during routine monitoring or compliance checks |
| Informational | Normal event recorded for visibility | Store, index and use for reporting or trend analysis |
8.5 Example Alert Rule
Trigger: More than 10 failed sign-in attempts for the same user within 5 minutes
Condition: Source IP is outside approved geography
Severity: High
Action: Notify SOC, lock account temporarily, require password reset and MFA verification9. Auditing in Cloud Systems
9.1 What is Auditing?
Auditing is the systematic review of records, configurations, and activities to verify that cloud systems operate according to policies, standards, contracts, and legal requirements.
- Monitoring is often real-time or near real-time
- Auditing is usually evidence-based and may be periodic, event-driven, or compliance-driven
9.2 What Cloud Audits Examine
- Identity records
- Access policies
- Network rules
- Storage permissions
- Encryption settings
- Backup status
- Vulnerability reports
- Change records
- Service configurations
- Incident history
9.3 Purpose of Auditing
- Prove accountability
- Detect policy violations
- Support forensic investigations
- Demonstrate compliance
9.4 Audit Requirements
A strong audit process requires:
- Complete records
- Synchronized timestamps
- Clear ownership
- Retention policies
- Protected log storage
- Documented review procedures
Auditors should be able to answer:
- Who did what?
- When?
- From where?
- Using which identity?
- Against which resource?
- With what outcome?
9.5 Audit Areas
| Audit Area | Evidence Required | Purpose |
|---|---|---|
| Identity and access | User lists, roles, group membership, MFA status, access reviews | Validate least privilege and user accountability |
| Network security | Firewall rules, security groups, flow logs | Verify segmentation and traffic control |
| Data protection | Encryption settings, key rotation, backup status | Confirm data protection |
| Change management | Change tickets, approvals, configuration history | Verify controlled changes |
| Incident management | Incident reports, timelines, lessons learned | Verify response effectiveness |
| Compliance | Control evidence, audit findings, retention status | Demonstrate regulatory compliance |
10. Record Generation
10.1 What is Record Generation?
Record generation is the creation of structured evidence about events that occur in cloud systems. Records may be generated by:
- Identity providers
- Operating systems
- Applications
- Databases
- Storage services
- Networks
- APIs
- Security tools
- Management platforms
10.2 What Makes a Record Valuable?
A record becomes valuable when it contains enough context for analysis, investigation, and reporting.
10.3 Important Record Fields
| Record Field | Meaning | Example |
|---|---|---|
| Timestamp | When the event occurred | 2026-07-04T10:30:22+05:30 |
| Actor | User, service or process that initiated action | admin@example.com or vm-service-role |
| Action | Operation performed | CreateUser, DeleteBucketPolicy, LoginFailed |
| Resource | Target of the action | database/prod-customer-db |
| Source | Origin of the event | IP address, device ID, region, application |
| Outcome | Result of the action | Success, Failure, Denied |
| Correlation ID | Identifier linking related events | request-id-9c32ab |
| Severity | Risk or importance level | Low, Medium, High, Critical |
10.4 Example Structured Security Log Record
{
"time": "2026-07-04T10:30:22+05:30",
"actor": "cloud-admin@example.com",
"action": "UpdateNetworkSecurityRule",
"resource": "prod-web-subnet",
"source_ip": "203.0.113.25",
"outcome": "success",
"severity": "high",
"correlation_id": "request-id-9c32ab"
}10.5 Best Practices
- Standardize record formats (JSON preferred)
- Avoid unnecessary sensitive data (passwords, tokens, full card numbers, private keys)
- Use structured logs for efficient indexing and querying
- Include correlation IDs for linking related events
11. Reporting and Management
11.1 Purpose of Reporting
Reporting converts monitoring and auditing data into meaningful information for:
- Technical teams
- Management
- Auditors
- Regulators
A report should not simply list raw logs. It should summarize:
- Security posture
- Major risks
- Incidents
- Response performance
- Compliance status
- Trend changes
- Recommended actions
11.2 Types of Reports
| Report Type | Audience | Contents |
|---|---|---|
| Daily SOC report | Security operations team | Open alerts, incidents, blocked attacks, high-risk changes |
| Weekly risk report | Security manager and IT leads | Top risks, vulnerabilities, privilege changes, unresolved actions |
| Monthly compliance report | Auditors, compliance team, management | Control status, audit findings, evidence gaps, retention status |
| Incident report | IR team, leadership, legal if required | Timeline, root cause, impact, containment, recovery and lessons learned |
| QoS report | Operations and service owners | Availability, latency, error rate, capacity, SLA/SLO performance |
11.3 Good Reporting Principles
- Separate technical details from executive summaries
- Security analysts need raw event IDs and packet details
- Management needs trend lines, severity counts, unresolved risks, SLA breaches, and business impact
12. Tamper-Proofing Audit Logs
12.1 What is Tamper-Proofing?
Tamper-proofing audit logs means protecting logs from unauthorized modification, deletion, or concealment. Attackers often try to erase traces after compromising an account or system.
If logs can be changed by the same administrators or workloads being monitored, accountability is weakened.
12.2 Protection Methods
| Protection Method | How It Helps | Example |
|---|---|---|
| Separate log account/project | Prevents compromised workload owners from deleting logs | Production account sends logs to security account |
| Immutable storage | Prevents changes during retention period | Object lock or WORM configuration |
| Encryption | Protects log confidentiality | KMS-managed encryption key |
| Digital signatures or hashes | Detects unauthorized modification | Hash chain for sequential log files |
| Strict access control | Limits who can read, export or delete logs | Only security team can access audit archive |
| Retention and legal hold | Preserves evidence for required period | Keep critical logs for 1 year or as policy requires |
12.3 Tamper-Proof Log Pipeline
Cloud services → Log collector → Normalize & sign → Immutable storage → SIEM / reports12.4 Additional Controls
- Time synchronization
- Access control
- Encryption
- Hash chaining
- Retention policy
- Legal hold
12.5 Important Point
Tamper-proofing does not mean logs can never be deleted. It means deletion or alteration is controlled, detectable, and auditable.
Log confidentiality is also important. Logs may contain usernames, IP addresses, file names, API paths, and business details that should not be exposed unnecessarily.
13. Quality of Service (QoS)
13.1 What is QoS?
Quality of Service (QoS) refers to the expected level of service performance and reliability. In cloud security management, QoS includes:
- Availability
- Latency
- Throughput
- Error rate
- Capacity
- Resilience
- Backup success
- Recovery time
- User experience
13.2 Connection Between Security and QoS
Security and QoS are connected because attacks, misconfigurations, and privilege abuse can directly affect service quality.
Example: Increased latency may indicate:
- Resource exhaustion
- DDoS traffic
- Database failure
- Poor scaling configuration
- Overloaded dependency
13.3 QoS Metrics
| QoS Metric | Meaning | Security Connection |
|---|---|---|
| Availability | Percentage of time service is usable | DDoS, ransomware or misconfiguration may reduce availability |
| Latency | Time taken to respond to a request | Malicious traffic or overloaded security inspection may increase delay |
| Throughput | Volume of requests or data processed | Capacity abuse or exfiltration can distort throughput |
| Error rate | Percentage of failed requests | Attack attempts may cause authentication or application errors |
| Recovery time | Time needed to restore service | Incident response and backups directly affect recovery |
| Backup success | Whether backups complete and can be restored | Backup failure increases impact of ransomware or deletion |
13.4 SLA, SLO, SLI
| Term | Meaning |
|---|---|
| SLA | Service Level Agreement — contract with customer |
| SLO | Service Level Objective — internal target |
| SLI | Service Level Indicator — measured metric |
Best Practice: Security controls should support QoS rather than blindly blocking legitimate business activity.
14. Secure Management Practices
14.1 What is Secure Management?
Secure management practices are the policies, procedures, and technical controls used to administer cloud infrastructure safely. Cloud management includes:
- Provisioning resources
- Changing configurations
- Managing identities
- Applying patches
- Reviewing logs
- Rotating secrets
- Approving changes
- Handling incidents
- Maintaining compliance evidence
14.2 Secure Management Model
A secure management model should use:
- Least privilege
- MFA
- Change control
- Secure administrative workstations
- Separate administrative accounts
- Approved automation
- Configuration baselines
- Vulnerability management
- Backup verification
- Encryption
- Logging
- Periodic access reviews
14.3 Management-Plane Sensitivity
Management-plane access is especially sensitive because cloud APIs can create, delete, or modify resources at scale. A single compromised administrator token can affect the entire environment. Therefore, cloud management should be treated as a critical security boundary.
14.4 Secure Management Practices
| Practice | Description | Example |
|---|---|---|
| Change control | Approve and document important configuration changes | Firewall rule change linked to ticket ID |
| Configuration baseline | Maintain approved secure settings | Default encryption, private storage, logging enabled |
| Patch management | Update OS, applications and agents | Monthly critical patch window |
| Secret rotation | Regularly rotate passwords, keys and tokens | Rotate database password and API key after staff change |
| Backup testing | Verify that backups can actually be restored | Quarterly restore drill |
| Administrative isolation | Protect admin access from normal browsing or email risks | Use privileged access workstation or hardened admin device |
Best Practice: Automate repetitive management tasks through approved infrastructure-as-code and policy-as-code pipelines. Manual console changes should be limited and audited.
15. User Management
15.1 What is User Management?
User management is the process of creating, modifying, disabling, reviewing, and removing user accounts in a cloud environment. It covers:
- Employees
- Contractors
- Administrators
- Developers
- Auditors
- Temporary users
- External partners
15.2 Why User Management Matters
Poor user management causes:
- Orphaned accounts
- Excessive permissions
- Unmanaged access to sensitive resources
15.3 User Lifecycle
| Lifecycle Stage | Security Action | Reason |
|---|---|---|
| Onboarding | Create identity, assign group, enable MFA, accept policy | Gives controlled initial access |
| Role assignment | Map job responsibility to approved roles | Supports least privilege |
| Periodic review | Review access with manager and resource owner | Removes unnecessary permissions |
| Role change | Update groups and revoke old permissions | Prevents privilege accumulation |
| Suspension | Temporarily disable account during leave or investigation | Reduces risk from inactive identity |
| Offboarding | Disable account, revoke sessions, rotate shared secrets | Prevents former users from accessing systems |
15.4 Best Practices
- Integrate with HR or organizational processes
- Group accounts by role and responsibility
- Minimize direct permissions assigned to individuals
- Treat dormant accounts, shared accounts, and accounts without MFA as risks
16. Identity Management
16.1 What is Identity Management?
Identity management is broader than user management. It includes:
- Human users
- Service accounts
- Workloads
- Devices
- APIs
- Federated identities
In a modern cloud system, identity is the new security perimeter because access decisions depend heavily on who or what is requesting access and under what conditions.
16.2 Components of Identity Management
- Authentication
- Authorization
- MFA
- SSO
- Federation
- Conditional access
- Identity governance
- Privileged access management
- Credential rotation
- Identity monitoring
16.3 Identity Types
| Identity Type | Example | Management Requirement |
|---|---|---|
| Human user | Employee, contractor, student administrator | MFA, role assignment, periodic review |
| Privileged user | Cloud administrator, security engineer | Just-in-time access, session monitoring, strong approval |
| Service account | Application identity used by workload | Least privilege, key rotation, no interactive login |
| Device identity | Managed laptop, server, mobile device | Compliance check, certificate, endpoint protection |
| Federated identity | External user authenticated by partner IdP | Trust policy, claims mapping, limited access |
| Workload identity | VM, container, function or pod identity | Managed identity and scoped resource permissions |
16.4 Why Identity Monitoring Matters
Security teams must monitor identity events continuously because many cloud attacks begin with identity compromise. Strong identity management reduces the chance that stolen credentials can be used to access sensitive systems.
17. Security Information and Event Management (SIEM)
17.1 What is SIEM?
A Security Information and Event Management (SIEM) system:
- Collects security events and logs from many sources
- Normalizes them into a searchable format
- Correlates related activity
- Detects threats
- Generates alerts
- Supports investigation
- Produces reports
SIEM is a core tool used by Security Operations Centers (SOCs).
17.2 SIEM Data Sources in Cloud
A SIEM may ingest:
- Identity logs
- Audit logs
- Network flow logs
- DNS logs
- Endpoint logs
- Application logs
- Database logs
- Container logs
- Firewall logs
- WAF logs
- Vulnerability data
It may also enrich events with:
- Threat intelligence
- Asset criticality
- User context
- Geolocation
17.3 SIEM Functions
| SIEM Function | Description | Example |
|---|---|---|
| Log collection | Ingest events from many sources | Cloud audit logs, firewall logs, application logs |
| Normalization | Convert different formats into common fields | Map source_ip, user, action and outcome |
| Correlation | Connect related events across sources | Login from new country followed by admin role change |
| Alerting | Notify when rules or analytics identify risk | Critical alert for disabled logging service |
| Investigation | Search, pivot and build timeline | Trace user activity across multiple services |
| Reporting | Produce dashboards and compliance evidence | Monthly report on privileged access |
| Automation | Trigger playbooks for standard response | Disable suspicious account and open ticket |
17.4 SIEM and SOAR
Modern SIEM systems often integrate with SOAR (Security Orchestration, Automation and Response). SOAR playbooks can automatically:
- Disable a user
- Block an IP address
- Open a ticket
- Notify a team
- Collect evidence
- Enrich an alert
17.5 SIEM Architecture
Identity → Network → Endpoint → Application → Cloud Services
↓
Log Collection
↓
Normalization
↓
Correlation
↓
Alerting
↓
Investigation
↓
Reporting
↓
Automation (SOAR)17.6 SIEM Reminder
A SIEM is only as useful as the quality of the logs and detection rules feeding it. Missing logs, noisy alerts, and poor asset context reduce detection value.
18. Unit Summary
MONITORING, AUDITING AND MANAGEMENT
│
├── Introduction
│ ├── Why monitoring is needed
│ ├── Cloud complexity
│ └── Log patterns for detection
│
├── Key Concepts
│ ├── Proactive monitoring
│ ├── Incident response
│ ├── Audit logs
│ ├── Alerts
│ ├── Unauthorized access
│ ├── Malicious traffic
│ ├── Privilege abuse
│ ├── Tamper-proof logs
│ ├── QoS
│ └── SIEM
│
├── Proactive Activity Monitoring
│ ├── Continuous observation
│ ├── Activity layers
│ ├── Monitoring lifecycle
│ └── Monitoring areas
│
├── Incident Response
│ ├── Preparation
│ ├── Detection and analysis
│ ├── Containment
│ ├── Eradication
│ ├── Recovery
│ └── Lessons learned
│
├── Monitoring for Unauthorized Access
│ ├── Causes
│ ├── Detection approach
│ ├── High-risk actions
│ └── Signals and responses
│
├── Detection of Malicious Traffic
│ ├── Traffic types
│ ├── Detection sources
│ ├── Signature vs behavior
│ └── Traffic classification
│
├── Prevention of Abuse of System Privileges
│ ├── Least privilege
│ ├── Separation of duties
│ ├── Just-in-time access
│ ├── Session monitoring
│ ├── MFA for admins
│ └── Immutable logging
│
├── Events and Alerts Management
│ ├── Events vs alerts vs incidents
│ ├── Alert fatigue
│ ├── Severity levels
│ └── Example alert rule
│
├── Auditing in Cloud Systems
│ ├── Definition
│ ├── Audit areas
│ ├── Evidence required
│ └── Purpose
│
├── Record Generation
│ ├── Record sources
│ ├── Important fields
│ ├── Structured logs
│ └── Example record
│
├── Reporting and Management
│ ├── Report types
│ ├── Audiences
│ └── Good reporting principles
│
├── Tamper-Proofing Audit Logs
│ ├── Protection methods
│ ├── Log pipeline
│ └── Additional controls
│
├── Quality of Service (QoS)
│ ├── QoS metrics
│ ├── Security connection
│ └── SLA/SLO/SLI
│
├── Secure Management Practices
│ ├── Management-plane sensitivity
│ ├── Secure management model
│ └── Practices
│
├── User Management
│ ├── Lifecycle stages
│ └── Best practices
│
├── Identity Management
│ ├── Identity types
│ ├── Components
│ └── Why it matters
│
└── SIEM
├── Definition
├── Data sources
├── Functions
├── SOAR integration
└── Architecture19. Key Terms — Glossary
| Term | Meaning |
|---|---|
| Proactive Monitoring | Continuous observation to detect issues early |
| Incident Response | Structured process for handling security incidents |
| Audit Log | Record of security-relevant activity |
| Alert | Notification when a rule/threshold/anomaly suggests a problem |
| Unauthorized Access | Access without proper permission |
| Malicious Traffic | Network communication associated with attacks |
| Privilege Abuse | Misuse of elevated permissions |
| Tamper-Proof Logs | Logs protected from unauthorized modification |
| QoS | Quality of Service — performance, availability, reliability |
| SIEM | Security Information and Event Management |
| SOAR | Security Orchestration, Automation and Response |
| SLA | Service Level Agreement |
| SLO | Service Level Objective |
| SLI | Service Level Indicator |
| WORM | Write Once Read Many |
| Event | Any recorded activity |
| Incident | Confirmed security event requiring response |
| Correlation ID | Identifier linking related events |
20. Exam-Focused Points
- Proactive monitoring — Continuous observation using logs, metrics, traces, alerts, dashboards.
- Incident response phases — Preparation, detection, containment, eradication, recovery, lessons learned.
- Cloud incident response — Automation and APIs for disabling keys, isolating instances, snapshots.
- Unauthorized access — Causes, signals, responses.
- High-risk actions — Root account use, key creation, disabling logs, policy changes, data export.
- Malicious traffic — Port scanning, DDoS, C2, exfiltration, web attacks, lateral movement.
- Detection techniques — Signature-based and behavior-based.
- Privilege abuse prevention — Least privilege, separation of duties, JIT access, session monitoring, MFA, immutable logging.
- Events vs alerts vs incidents — Event (activity), alert (condition met), incident (confirmed).
- Alert severity — Critical, high, medium, low, informational.
- Auditing — Systematic review of records, configurations, activities.
- Audit areas — Identity, network, data, change, incident, compliance.
- Record fields — Timestamp, actor, action, resource, source, outcome, correlation ID, severity.
- Reporting — Daily SOC, weekly risk, monthly compliance, incident, QoS.
- Tamper-proofing — Separate account, immutable storage, encryption, signatures, access control, retention.
- QoS metrics — Availability, latency, throughput, error rate, recovery time, backup success.
- Secure management — Change control, baseline, patch, secret rotation, backup testing, admin isolation.
- User management — Onboarding, role assignment, review, role change, suspension, offboarding.
- Identity management — Human, privileged, service, device, federated, workload identities.
- SIEM — Collection, normalization, correlation, alerting, investigation, reporting, automation.
- SOAR — Automation of response actions.
- Log integrity — Hash chaining, digital signatures, WORM storage.
- Alert fatigue — Too many low-quality alerts; tune rules, suppress duplicates, enrich context.
- Management-plane — Critical security boundary; protect admin access.
- Identity is the new perimeter — Access decisions depend on who/what is requesting.