BTCE | 5th Sem
SPC SubjectUnit 4

SPC Unit 4: Complete Concepts Guide

Unit 4: Monitoring, Auditing and Management -> Generated and Prepared By Thiruselvan (ThiruXD)

1. Introduction to Monitoring, Auditing and Management

1.1 Why Cloud Security Does Not End at Design

Cloud security does not end after designing access controls, encryption mechanisms, and secure architecture. A cloud environment must be continuously monitored, audited, and managed because users, workloads, networks, applications, and threats change every day.

  • Monitoring provides live visibility into what is happening in the environment.
  • Auditing provides recorded evidence of what happened, who performed an action, when it occurred, and whether the action was successful or failed.

1.2 Why Cloud Monitoring is More Complex

In cloud computing, monitoring is more complex than in a traditional data center because:

  • Resources are elastic and distributed
  • Resources are often controlled through APIs
  • Virtual machines, containers, serverless functions, databases, load balancers, storage buckets, and identity services can be created or modified rapidly

Without proper monitoring and logging, unauthorized access, privilege abuse, misconfiguration, data exfiltration, and malicious traffic may remain hidden until serious damage occurs.

1.3 Did You Know?

In cloud environments, many security incidents are first detected through unusual log patterns, such as:

  • Impossible travel sign-ins
  • Repeated failed logins
  • Suspicious API calls
  • Unexpected configuration changes

2. Key Concepts and Glossary

TermMeaning in Cloud Security
Proactive MonitoringContinuous observation of cloud assets, users, applications and networks to detect issues before they become serious incidents
Incident ResponseA structured process for preparing for, detecting, analyzing, containing, eradicating and recovering from security incidents
Audit LogA record of security-relevant activity, such as login attempts, administrative actions, data access, configuration changes and network events
AlertA notification generated when a rule, threshold or anomaly suggests a possible problem or security event
Unauthorized AccessAccess to a system, account, resource or data without proper permission or outside approved policy
Malicious TrafficNetwork communication associated with attacks such as scanning, malware command-and-control, DDoS, data exfiltration or exploitation attempts
Privilege AbuseMisuse of elevated permissions by an administrator, compromised account, insider or service identity
Tamper-Proof LogsLogs protected so that unauthorized modification, deletion or concealment is prevented or at least detectable
QoSQuality of Service, referring to performance, availability, latency, reliability and service-level expectations
SIEMSecurity Information and Event Management, a platform that collects, correlates, analyzes and reports security events

3. Proactive Activity Monitoring

3.1 What is Proactive Activity Monitoring?

Proactive activity monitoring is the continuous observation of cloud resources, users, services, APIs, and network behavior to identify security and operational problems early. Instead of waiting for a failure or breach report, proactive monitoring uses logs, metrics, traces, alerts, and dashboards to maintain awareness of the cloud environment.

3.2 Activity at Many Layers

In cloud infrastructure, activity happens at many layers:

  • A user may sign in through an identity provider
  • An administrator may change a firewall rule
  • An application may access a database
  • A virtual machine may send outbound traffic
  • A storage bucket may be made public

Each activity generates telemetry that must be collected and analyzed.

3.3 Questions a Good Monitoring System Answers

  1. Which users accessed sensitive resources?
  2. Were there failed login attempts?
  3. Did a workload communicate with unknown IP addresses?
  4. Were new administrator roles assigned?
  5. Did CPU, memory or response time cross a threshold?
  6. Was a storage policy changed unexpectedly?

3.4 Proactive Monitoring Lifecycle

Collect telemetry → Normalize & enrich → Correlate & detect → Alert & prioritize → Investigate & respond → Improve controls

3.5 Monitoring Areas

Monitoring AreaExamples of Activity to MonitorSecurity Value
Identity activitySign-ins, failed logins, MFA failures, password resets, role assignmentsDetects account compromise, brute force attacks and privilege changes
Control-plane activityAPI calls, configuration changes, resource creation or deletionReveals unauthorized administration and misconfiguration
Network activityFlow logs, firewall logs, DNS queries, WAF events, VPN logsDetects scanning, lateral movement, command-and-control and exfiltration
Compute activityVM events, container logs, process behavior, patch statusDetects malware, abnormal processes and insecure hosts
Storage activityObject read/write/delete, policy changes, public access eventsProtects sensitive files and supports data-loss investigation
Application activityAuthentication events, business transactions, errors, API requestsSupports fraud detection, debugging and user-level investigation

Best Practice: Start monitoring from the most sensitive assets first: identity services, administrator actions, internet-facing services, critical databases, and storage containing confidential data.


4. Incident Response

4.1 What is Incident Response?

Incident response is the organized approach used to handle cybersecurity incidents. A cloud incident may include:

  • Compromised credentials
  • Exposed storage
  • Malware on a virtual machine
  • Data leakage
  • DDoS attack
  • Suspicious API activity
  • Unauthorized database access
  • Abuse of system privileges

Cloud incident response must be planned before an incident occurs. Teams should know:

  • Who is responsible
  • Which logs are required
  • How to isolate resources
  • How to preserve evidence
  • How to communicate with stakeholders
  • How to restore service safely

4.2 Cloud-Specific Response Actions

In cloud environments, many response actions are performed through automation and APIs:

  • Disabling keys
  • Isolating instances
  • Changing security groups
  • Rotating secrets
  • Taking forensic snapshots

4.3 Incident Response Phases

1. Preparation → 2. Detection & Analysis → 3. Containment → 4. Eradication → 5. Recovery → 6. Lessons Learned
PhaseCloud-Specific ActivitiesOutput
1. PreparationDefine playbooks, enable logging, create response roles, prepare forensic storage, train respondersIncident response plan and readiness checklist
2. Detection and AnalysisReview SIEM alerts, IAM logs, network flows, endpoint logs and application evidenceConfirmed incident scope, severity and affected assets
3. ContainmentDisable compromised users, revoke tokens, isolate VMs, block IPs, freeze storage accessAttack is stopped from spreading
4. EradicationRemove malware, patch vulnerability, rotate keys, remove malicious rules or backdoorsRoot cause and persistence mechanisms removed
5. RecoveryRestore services, monitor closely, validate backups and business functionsSecure return to normal operations
6. Lessons LearnedUpdate controls, detection rules, training, architecture and documentationImproved security posture

Important Point: During a cloud incident, never delete suspicious resources immediately. Capture evidence such as logs, snapshots, configuration history, and access records before cleanup whenever legal and organizational policies allow.


5. Monitoring for Unauthorized Access

5.1 What is Unauthorized Access?

Unauthorized access occurs when a user, service account, application, or attacker accesses a resource without valid permission or outside approved policy.

5.2 Common Causes

  • Stolen credentials
  • Weak passwords
  • Missing MFA
  • Overly permissive IAM policies
  • Leaked API keys
  • Misconfigured storage
  • Exposed management ports
  • Compromised service identities

5.3 Detection Approach

Monitoring for unauthorized access requires combining:

  • Identity logs
  • Authorization decisions
  • Resource access logs
  • Contextual signals

A single failed login may not be critical, but repeated failures from many countries, a successful login from an unusual location, or a new administrator role assigned after a suspicious login can indicate compromise.

5.4 High-Risk Actions to Monitor

  • Root or owner account use
  • Creation of access keys
  • Disabling logging
  • Modifying security policies
  • Exporting data
  • Changing network rules
  • Deleting backups
  • Access outside normal working hours or approved locations

5.5 Unauthorized Access Signals

Unauthorized Access SignalPossible MeaningRecommended Response
Multiple failed login attemptsPassword guessing or brute-force attemptTrigger alert, enforce MFA, rate-limit, investigate source
Successful login from unusual locationCredential theft or impossible travelRequire step-up authentication and verify user
New admin role assignmentPrivilege escalationReview approver, ticket, identity and timing
Access from unknown deviceCompromised password or unmanaged endpointCheck device compliance and conditional access
API key used from new IPLeaked credential or automation driftRotate key, restrict source IP and review code repositories

6. Detection of Malicious Traffic

6.1 What is Malicious Traffic?

Malicious traffic refers to network communication associated with attacks or suspicious behavior. Examples include:

  • Port scanning
  • Vulnerability exploitation
  • Malware command-and-control communication
  • DDoS traffic
  • Suspicious DNS queries
  • TOR or proxy access
  • Unexpected outbound connections
  • Data exfiltration
  • Lateral movement between cloud workloads

6.2 Sources for Traffic Visibility

Cloud platforms provide several sources:

  • Virtual network flow logs
  • Firewall logs
  • Load balancer logs
  • DNS logs
  • WAF logs
  • IDS/IPS alerts
  • API gateway logs
  • Endpoint telemetry

6.3 Detection Techniques

TechniqueDescription
Signature-basedIdentifies known threats
Behavior-basedIdentifies unusual patterns (e.g., server sending large volumes of data to unknown country)

6.4 Traffic Classification

Allowed traffic → Permit
Suspicious traffic → Alert
Malicious traffic → Block

6.5 Traffic Types and Detection Sources

Traffic TypeCommon Detection SourceExample Alert
Port scanningVPC/VNet flow logs, IDS, firewall logsMany denied connections to different ports from one source
DDoS trafficLoad balancer metrics, CDN/WAF logs, network telemetrySudden spike in requests or bandwidth from distributed sources
Command-and-controlDNS logs, threat intelligence, outbound proxy logsConnection to known malicious domain or rare destination
Data exfiltrationFlow logs, storage access logs, DLP, CASBLarge outbound transfer from sensitive workload
Web attackWAF logs, application logs, API gateway logsSQL injection, XSS, path traversal or API abuse pattern
Lateral movementEast-west flow logs, endpoint telemetryUnusual internal connections between unrelated workloads

Best Practice: Store network-flow logs long enough to support investigations. Many attacks are discovered days or weeks after the first malicious connection.


7. Prevention of Abuse of System Privileges

7.1 What is Privilege Abuse?

System privileges allow users and services to perform powerful actions such as:

  • Creating resources
  • Changing security settings
  • Accessing sensitive data
  • Managing encryption keys
  • Deleting logs
  • Modifying networks

Abuse of privileges may be:

  • Intentional insider misuse
  • Accidental misuse
  • Result of an attacker compromising an administrator account

7.2 Prevention Approach

ControlHow It Prevents Privilege AbuseExample
Least privilegeLimits the damage any account can causeDeveloper can deploy app but cannot change billing or security logs
Separation of dutiesPrevents one person from approving and executing sensitive actions aloneKey deletion requires security and operations approval
Just-in-time accessGives elevated rights only for a limited timeAdmin role active for two hours after approval
Privileged session monitoringRecords commands and actions for accountabilitySession log is reviewed after database maintenance
MFA for administratorsReduces risk from stolen passwordsAdmin console requires authenticator approval
Immutable loggingPrevents attackers from hiding privileged actionsAudit logs written to locked storage

7.3 Best Practices

  • Avoid long-lived static administrator credentials
  • Use temporary credentials and managed identities
  • Use break-glass accounts for emergencies
  • Implement privileged identity management
  • Conduct automated policy reviews

8. Events and Alerts Management

8.1 Definitions

TermDefinition
EventAny recorded activity in a system
AlertA notification generated when one or more events meet a defined condition
IncidentA confirmed security event requiring response

Note: Not every event is an alert, and not every alert is an incident.

8.2 Event and Alert Management Process

  1. Collect events
  2. Define alert rules
  3. Assign severity
  4. Reduce noise
  5. Route notifications
  6. Track actions until closure

8.3 Alert Fatigue

Poor alert management creates alert fatigue. If analysts receive too many low-quality alerts, they may ignore important warnings. Therefore, cloud security teams must:

  • Tune rules
  • Suppress duplicates
  • Enrich alerts with context
  • Prioritize alerts based on business impact and threat severity

8.4 Alert Severity Levels

SeverityTypical ConditionExpected Action
CriticalConfirmed breach, active exfiltration, root/admin compromise, production outageImmediate incident response and leadership notification
HighLikely compromise, high-risk policy change, malware detection, privilege escalationRapid investigation and containment
MediumSuspicious behavior requiring review, abnormal access, repeated failuresAnalyze within defined SLA and tune detection if needed
LowInformational anomaly or policy driftReview during routine monitoring or compliance checks
InformationalNormal event recorded for visibilityStore, index and use for reporting or trend analysis

8.5 Example Alert Rule

Trigger: More than 10 failed sign-in attempts for the same user within 5 minutes
Condition: Source IP is outside approved geography
Severity: High
Action: Notify SOC, lock account temporarily, require password reset and MFA verification

9. Auditing in Cloud Systems

9.1 What is Auditing?

Auditing is the systematic review of records, configurations, and activities to verify that cloud systems operate according to policies, standards, contracts, and legal requirements.

  • Monitoring is often real-time or near real-time
  • Auditing is usually evidence-based and may be periodic, event-driven, or compliance-driven

9.2 What Cloud Audits Examine

  • Identity records
  • Access policies
  • Network rules
  • Storage permissions
  • Encryption settings
  • Backup status
  • Vulnerability reports
  • Change records
  • Service configurations
  • Incident history

9.3 Purpose of Auditing

  1. Prove accountability
  2. Detect policy violations
  3. Support forensic investigations
  4. Demonstrate compliance

9.4 Audit Requirements

A strong audit process requires:

  • Complete records
  • Synchronized timestamps
  • Clear ownership
  • Retention policies
  • Protected log storage
  • Documented review procedures

Auditors should be able to answer:

  • Who did what?
  • When?
  • From where?
  • Using which identity?
  • Against which resource?
  • With what outcome?

9.5 Audit Areas

Audit AreaEvidence RequiredPurpose
Identity and accessUser lists, roles, group membership, MFA status, access reviewsValidate least privilege and user accountability
Network securityFirewall rules, security groups, flow logsVerify segmentation and traffic control
Data protectionEncryption settings, key rotation, backup statusConfirm data protection
Change managementChange tickets, approvals, configuration historyVerify controlled changes
Incident managementIncident reports, timelines, lessons learnedVerify response effectiveness
ComplianceControl evidence, audit findings, retention statusDemonstrate regulatory compliance

10. Record Generation

10.1 What is Record Generation?

Record generation is the creation of structured evidence about events that occur in cloud systems. Records may be generated by:

  • Identity providers
  • Operating systems
  • Applications
  • Databases
  • Storage services
  • Networks
  • APIs
  • Security tools
  • Management platforms

10.2 What Makes a Record Valuable?

A record becomes valuable when it contains enough context for analysis, investigation, and reporting.

10.3 Important Record Fields

Record FieldMeaningExample
TimestampWhen the event occurred2026-07-04T10:30:22+05:30
ActorUser, service or process that initiated actionadmin@example.com or vm-service-role
ActionOperation performedCreateUser, DeleteBucketPolicy, LoginFailed
ResourceTarget of the actiondatabase/prod-customer-db
SourceOrigin of the eventIP address, device ID, region, application
OutcomeResult of the actionSuccess, Failure, Denied
Correlation IDIdentifier linking related eventsrequest-id-9c32ab
SeverityRisk or importance levelLow, Medium, High, Critical

10.4 Example Structured Security Log Record

{
  "time": "2026-07-04T10:30:22+05:30",
  "actor": "cloud-admin@example.com",
  "action": "UpdateNetworkSecurityRule",
  "resource": "prod-web-subnet",
  "source_ip": "203.0.113.25",
  "outcome": "success",
  "severity": "high",
  "correlation_id": "request-id-9c32ab"
}

10.5 Best Practices

  • Standardize record formats (JSON preferred)
  • Avoid unnecessary sensitive data (passwords, tokens, full card numbers, private keys)
  • Use structured logs for efficient indexing and querying
  • Include correlation IDs for linking related events

11. Reporting and Management

11.1 Purpose of Reporting

Reporting converts monitoring and auditing data into meaningful information for:

  • Technical teams
  • Management
  • Auditors
  • Regulators

A report should not simply list raw logs. It should summarize:

  • Security posture
  • Major risks
  • Incidents
  • Response performance
  • Compliance status
  • Trend changes
  • Recommended actions

11.2 Types of Reports

Report TypeAudienceContents
Daily SOC reportSecurity operations teamOpen alerts, incidents, blocked attacks, high-risk changes
Weekly risk reportSecurity manager and IT leadsTop risks, vulnerabilities, privilege changes, unresolved actions
Monthly compliance reportAuditors, compliance team, managementControl status, audit findings, evidence gaps, retention status
Incident reportIR team, leadership, legal if requiredTimeline, root cause, impact, containment, recovery and lessons learned
QoS reportOperations and service ownersAvailability, latency, error rate, capacity, SLA/SLO performance

11.3 Good Reporting Principles

  • Separate technical details from executive summaries
  • Security analysts need raw event IDs and packet details
  • Management needs trend lines, severity counts, unresolved risks, SLA breaches, and business impact

12. Tamper-Proofing Audit Logs

12.1 What is Tamper-Proofing?

Tamper-proofing audit logs means protecting logs from unauthorized modification, deletion, or concealment. Attackers often try to erase traces after compromising an account or system.

If logs can be changed by the same administrators or workloads being monitored, accountability is weakened.

12.2 Protection Methods

Protection MethodHow It HelpsExample
Separate log account/projectPrevents compromised workload owners from deleting logsProduction account sends logs to security account
Immutable storagePrevents changes during retention periodObject lock or WORM configuration
EncryptionProtects log confidentialityKMS-managed encryption key
Digital signatures or hashesDetects unauthorized modificationHash chain for sequential log files
Strict access controlLimits who can read, export or delete logsOnly security team can access audit archive
Retention and legal holdPreserves evidence for required periodKeep critical logs for 1 year or as policy requires

12.3 Tamper-Proof Log Pipeline

Cloud services → Log collector → Normalize & sign → Immutable storage → SIEM / reports

12.4 Additional Controls

  • Time synchronization
  • Access control
  • Encryption
  • Hash chaining
  • Retention policy
  • Legal hold

12.5 Important Point

Tamper-proofing does not mean logs can never be deleted. It means deletion or alteration is controlled, detectable, and auditable.

Log confidentiality is also important. Logs may contain usernames, IP addresses, file names, API paths, and business details that should not be exposed unnecessarily.


13. Quality of Service (QoS)

13.1 What is QoS?

Quality of Service (QoS) refers to the expected level of service performance and reliability. In cloud security management, QoS includes:

  • Availability
  • Latency
  • Throughput
  • Error rate
  • Capacity
  • Resilience
  • Backup success
  • Recovery time
  • User experience

13.2 Connection Between Security and QoS

Security and QoS are connected because attacks, misconfigurations, and privilege abuse can directly affect service quality.

Example: Increased latency may indicate:

  • Resource exhaustion
  • DDoS traffic
  • Database failure
  • Poor scaling configuration
  • Overloaded dependency

13.3 QoS Metrics

QoS MetricMeaningSecurity Connection
AvailabilityPercentage of time service is usableDDoS, ransomware or misconfiguration may reduce availability
LatencyTime taken to respond to a requestMalicious traffic or overloaded security inspection may increase delay
ThroughputVolume of requests or data processedCapacity abuse or exfiltration can distort throughput
Error ratePercentage of failed requestsAttack attempts may cause authentication or application errors
Recovery timeTime needed to restore serviceIncident response and backups directly affect recovery
Backup successWhether backups complete and can be restoredBackup failure increases impact of ransomware or deletion

13.4 SLA, SLO, SLI

TermMeaning
SLAService Level Agreement — contract with customer
SLOService Level Objective — internal target
SLIService Level Indicator — measured metric

Best Practice: Security controls should support QoS rather than blindly blocking legitimate business activity.


14. Secure Management Practices

14.1 What is Secure Management?

Secure management practices are the policies, procedures, and technical controls used to administer cloud infrastructure safely. Cloud management includes:

  • Provisioning resources
  • Changing configurations
  • Managing identities
  • Applying patches
  • Reviewing logs
  • Rotating secrets
  • Approving changes
  • Handling incidents
  • Maintaining compliance evidence

14.2 Secure Management Model

A secure management model should use:

  • Least privilege
  • MFA
  • Change control
  • Secure administrative workstations
  • Separate administrative accounts
  • Approved automation
  • Configuration baselines
  • Vulnerability management
  • Backup verification
  • Encryption
  • Logging
  • Periodic access reviews

14.3 Management-Plane Sensitivity

Management-plane access is especially sensitive because cloud APIs can create, delete, or modify resources at scale. A single compromised administrator token can affect the entire environment. Therefore, cloud management should be treated as a critical security boundary.

14.4 Secure Management Practices

PracticeDescriptionExample
Change controlApprove and document important configuration changesFirewall rule change linked to ticket ID
Configuration baselineMaintain approved secure settingsDefault encryption, private storage, logging enabled
Patch managementUpdate OS, applications and agentsMonthly critical patch window
Secret rotationRegularly rotate passwords, keys and tokensRotate database password and API key after staff change
Backup testingVerify that backups can actually be restoredQuarterly restore drill
Administrative isolationProtect admin access from normal browsing or email risksUse privileged access workstation or hardened admin device

Best Practice: Automate repetitive management tasks through approved infrastructure-as-code and policy-as-code pipelines. Manual console changes should be limited and audited.


15. User Management

15.1 What is User Management?

User management is the process of creating, modifying, disabling, reviewing, and removing user accounts in a cloud environment. It covers:

  • Employees
  • Contractors
  • Administrators
  • Developers
  • Auditors
  • Temporary users
  • External partners

15.2 Why User Management Matters

Poor user management causes:

  • Orphaned accounts
  • Excessive permissions
  • Unmanaged access to sensitive resources

15.3 User Lifecycle

Lifecycle StageSecurity ActionReason
OnboardingCreate identity, assign group, enable MFA, accept policyGives controlled initial access
Role assignmentMap job responsibility to approved rolesSupports least privilege
Periodic reviewReview access with manager and resource ownerRemoves unnecessary permissions
Role changeUpdate groups and revoke old permissionsPrevents privilege accumulation
SuspensionTemporarily disable account during leave or investigationReduces risk from inactive identity
OffboardingDisable account, revoke sessions, rotate shared secretsPrevents former users from accessing systems

15.4 Best Practices

  • Integrate with HR or organizational processes
  • Group accounts by role and responsibility
  • Minimize direct permissions assigned to individuals
  • Treat dormant accounts, shared accounts, and accounts without MFA as risks

16. Identity Management

16.1 What is Identity Management?

Identity management is broader than user management. It includes:

  • Human users
  • Service accounts
  • Workloads
  • Devices
  • APIs
  • Federated identities

In a modern cloud system, identity is the new security perimeter because access decisions depend heavily on who or what is requesting access and under what conditions.

16.2 Components of Identity Management

  • Authentication
  • Authorization
  • MFA
  • SSO
  • Federation
  • Conditional access
  • Identity governance
  • Privileged access management
  • Credential rotation
  • Identity monitoring

16.3 Identity Types

Identity TypeExampleManagement Requirement
Human userEmployee, contractor, student administratorMFA, role assignment, periodic review
Privileged userCloud administrator, security engineerJust-in-time access, session monitoring, strong approval
Service accountApplication identity used by workloadLeast privilege, key rotation, no interactive login
Device identityManaged laptop, server, mobile deviceCompliance check, certificate, endpoint protection
Federated identityExternal user authenticated by partner IdPTrust policy, claims mapping, limited access
Workload identityVM, container, function or pod identityManaged identity and scoped resource permissions

16.4 Why Identity Monitoring Matters

Security teams must monitor identity events continuously because many cloud attacks begin with identity compromise. Strong identity management reduces the chance that stolen credentials can be used to access sensitive systems.


17. Security Information and Event Management (SIEM)

17.1 What is SIEM?

A Security Information and Event Management (SIEM) system:

  • Collects security events and logs from many sources
  • Normalizes them into a searchable format
  • Correlates related activity
  • Detects threats
  • Generates alerts
  • Supports investigation
  • Produces reports

SIEM is a core tool used by Security Operations Centers (SOCs).

17.2 SIEM Data Sources in Cloud

A SIEM may ingest:

  • Identity logs
  • Audit logs
  • Network flow logs
  • DNS logs
  • Endpoint logs
  • Application logs
  • Database logs
  • Container logs
  • Firewall logs
  • WAF logs
  • Vulnerability data

It may also enrich events with:

  • Threat intelligence
  • Asset criticality
  • User context
  • Geolocation

17.3 SIEM Functions

SIEM FunctionDescriptionExample
Log collectionIngest events from many sourcesCloud audit logs, firewall logs, application logs
NormalizationConvert different formats into common fieldsMap source_ip, user, action and outcome
CorrelationConnect related events across sourcesLogin from new country followed by admin role change
AlertingNotify when rules or analytics identify riskCritical alert for disabled logging service
InvestigationSearch, pivot and build timelineTrace user activity across multiple services
ReportingProduce dashboards and compliance evidenceMonthly report on privileged access
AutomationTrigger playbooks for standard responseDisable suspicious account and open ticket

17.4 SIEM and SOAR

Modern SIEM systems often integrate with SOAR (Security Orchestration, Automation and Response). SOAR playbooks can automatically:

  • Disable a user
  • Block an IP address
  • Open a ticket
  • Notify a team
  • Collect evidence
  • Enrich an alert

17.5 SIEM Architecture

Identity → Network → Endpoint → Application → Cloud Services
                    ↓
              Log Collection
                    ↓
              Normalization
                    ↓
              Correlation
                    ↓
              Alerting
                    ↓
              Investigation
                    ↓
              Reporting
                    ↓
              Automation (SOAR)

17.6 SIEM Reminder

A SIEM is only as useful as the quality of the logs and detection rules feeding it. Missing logs, noisy alerts, and poor asset context reduce detection value.


18. Unit Summary

MONITORING, AUDITING AND MANAGEMENT
│
├── Introduction
│   ├── Why monitoring is needed
│   ├── Cloud complexity
│   └── Log patterns for detection
│
├── Key Concepts
│   ├── Proactive monitoring
│   ├── Incident response
│   ├── Audit logs
│   ├── Alerts
│   ├── Unauthorized access
│   ├── Malicious traffic
│   ├── Privilege abuse
│   ├── Tamper-proof logs
│   ├── QoS
│   └── SIEM
│
├── Proactive Activity Monitoring
│   ├── Continuous observation
│   ├── Activity layers
│   ├── Monitoring lifecycle
│   └── Monitoring areas
│
├── Incident Response
│   ├── Preparation
│   ├── Detection and analysis
│   ├── Containment
│   ├── Eradication
│   ├── Recovery
│   └── Lessons learned
│
├── Monitoring for Unauthorized Access
│   ├── Causes
│   ├── Detection approach
│   ├── High-risk actions
│   └── Signals and responses
│
├── Detection of Malicious Traffic
│   ├── Traffic types
│   ├── Detection sources
│   ├── Signature vs behavior
│   └── Traffic classification
│
├── Prevention of Abuse of System Privileges
│   ├── Least privilege
│   ├── Separation of duties
│   ├── Just-in-time access
│   ├── Session monitoring
│   ├── MFA for admins
│   └── Immutable logging
│
├── Events and Alerts Management
│   ├── Events vs alerts vs incidents
│   ├── Alert fatigue
│   ├── Severity levels
│   └── Example alert rule
│
├── Auditing in Cloud Systems
│   ├── Definition
│   ├── Audit areas
│   ├── Evidence required
│   └── Purpose
│
├── Record Generation
│   ├── Record sources
│   ├── Important fields
│   ├── Structured logs
│   └── Example record
│
├── Reporting and Management
│   ├── Report types
│   ├── Audiences
│   └── Good reporting principles
│
├── Tamper-Proofing Audit Logs
│   ├── Protection methods
│   ├── Log pipeline
│   └── Additional controls
│
├── Quality of Service (QoS)
│   ├── QoS metrics
│   ├── Security connection
│   └── SLA/SLO/SLI
│
├── Secure Management Practices
│   ├── Management-plane sensitivity
│   ├── Secure management model
│   └── Practices
│
├── User Management
│   ├── Lifecycle stages
│   └── Best practices
│
├── Identity Management
│   ├── Identity types
│   ├── Components
│   └── Why it matters
│
└── SIEM
    ├── Definition
    ├── Data sources
    ├── Functions
    ├── SOAR integration
    └── Architecture

19. Key Terms — Glossary

TermMeaning
Proactive MonitoringContinuous observation to detect issues early
Incident ResponseStructured process for handling security incidents
Audit LogRecord of security-relevant activity
AlertNotification when a rule/threshold/anomaly suggests a problem
Unauthorized AccessAccess without proper permission
Malicious TrafficNetwork communication associated with attacks
Privilege AbuseMisuse of elevated permissions
Tamper-Proof LogsLogs protected from unauthorized modification
QoSQuality of Service — performance, availability, reliability
SIEMSecurity Information and Event Management
SOARSecurity Orchestration, Automation and Response
SLAService Level Agreement
SLOService Level Objective
SLIService Level Indicator
WORMWrite Once Read Many
EventAny recorded activity
IncidentConfirmed security event requiring response
Correlation IDIdentifier linking related events

20. Exam-Focused Points

  1. Proactive monitoring — Continuous observation using logs, metrics, traces, alerts, dashboards.
  2. Incident response phases — Preparation, detection, containment, eradication, recovery, lessons learned.
  3. Cloud incident response — Automation and APIs for disabling keys, isolating instances, snapshots.
  4. Unauthorized access — Causes, signals, responses.
  5. High-risk actions — Root account use, key creation, disabling logs, policy changes, data export.
  6. Malicious traffic — Port scanning, DDoS, C2, exfiltration, web attacks, lateral movement.
  7. Detection techniques — Signature-based and behavior-based.
  8. Privilege abuse prevention — Least privilege, separation of duties, JIT access, session monitoring, MFA, immutable logging.
  9. Events vs alerts vs incidents — Event (activity), alert (condition met), incident (confirmed).
  10. Alert severity — Critical, high, medium, low, informational.
  11. Auditing — Systematic review of records, configurations, activities.
  12. Audit areas — Identity, network, data, change, incident, compliance.
  13. Record fields — Timestamp, actor, action, resource, source, outcome, correlation ID, severity.
  14. Reporting — Daily SOC, weekly risk, monthly compliance, incident, QoS.
  15. Tamper-proofing — Separate account, immutable storage, encryption, signatures, access control, retention.
  16. QoS metrics — Availability, latency, throughput, error rate, recovery time, backup success.
  17. Secure management — Change control, baseline, patch, secret rotation, backup testing, admin isolation.
  18. User management — Onboarding, role assignment, review, role change, suspension, offboarding.
  19. Identity management — Human, privileged, service, device, federated, workload identities.
  20. SIEM — Collection, normalization, correlation, alerting, investigation, reporting, automation.
  21. SOAR — Automation of response actions.
  22. Log integrity — Hash chaining, digital signatures, WORM storage.
  23. Alert fatigue — Too many low-quality alerts; tune rules, suppress duplicates, enrich context.
  24. Management-plane — Critical security boundary; protect admin access.
  25. Identity is the new perimeter — Access decisions depend on who/what is requesting.

On this page

1. Introduction to Monitoring, Auditing and Management1.1 Why Cloud Security Does Not End at Design1.2 Why Cloud Monitoring is More Complex1.3 Did You Know?2. Key Concepts and Glossary3. Proactive Activity Monitoring3.1 What is Proactive Activity Monitoring?3.2 Activity at Many Layers3.3 Questions a Good Monitoring System Answers3.4 Proactive Monitoring Lifecycle3.5 Monitoring Areas4. Incident Response4.1 What is Incident Response?4.2 Cloud-Specific Response Actions4.3 Incident Response Phases5. Monitoring for Unauthorized Access5.1 What is Unauthorized Access?5.2 Common Causes5.3 Detection Approach5.4 High-Risk Actions to Monitor5.5 Unauthorized Access Signals6. Detection of Malicious Traffic6.1 What is Malicious Traffic?6.2 Sources for Traffic Visibility6.3 Detection Techniques6.4 Traffic Classification6.5 Traffic Types and Detection Sources7. Prevention of Abuse of System Privileges7.1 What is Privilege Abuse?7.2 Prevention Approach7.3 Best Practices8. Events and Alerts Management8.1 Definitions8.2 Event and Alert Management Process8.3 Alert Fatigue8.4 Alert Severity Levels8.5 Example Alert Rule9. Auditing in Cloud Systems9.1 What is Auditing?9.2 What Cloud Audits Examine9.3 Purpose of Auditing9.4 Audit Requirements9.5 Audit Areas10. Record Generation10.1 What is Record Generation?10.2 What Makes a Record Valuable?10.3 Important Record Fields10.4 Example Structured Security Log Record10.5 Best Practices11. Reporting and Management11.1 Purpose of Reporting11.2 Types of Reports11.3 Good Reporting Principles12. Tamper-Proofing Audit Logs12.1 What is Tamper-Proofing?12.2 Protection Methods12.3 Tamper-Proof Log Pipeline12.4 Additional Controls12.5 Important Point13. Quality of Service (QoS)13.1 What is QoS?13.2 Connection Between Security and QoS13.3 QoS Metrics13.4 SLA, SLO, SLI14. Secure Management Practices14.1 What is Secure Management?14.2 Secure Management Model14.3 Management-Plane Sensitivity14.4 Secure Management Practices15. User Management15.1 What is User Management?15.2 Why User Management Matters15.3 User Lifecycle15.4 Best Practices16. Identity Management16.1 What is Identity Management?16.2 Components of Identity Management16.3 Identity Types16.4 Why Identity Monitoring Matters17. Security Information and Event Management (SIEM)17.1 What is SIEM?17.2 SIEM Data Sources in Cloud17.3 SIEM Functions17.4 SIEM and SOAR17.5 SIEM Architecture17.6 SIEM Reminder18. Unit Summary19. Key Terms — Glossary20. Exam-Focused Points