Domain

Monitoring, Logging, Analysis, Remediation, and Performance Optimization

How to see what your workloads are doing and act on it: CloudWatch metrics, logs, alarms, and dashboards, event-driven remediation with EventBridge and Systems Manager, and performance tuning for compute, storage, and databases. One of the three heaviest domains on the exam, and the daily core of the CloudOps job.

Something is wrong with a workload and nobody can say what. Closing the gap between "the app feels slow" and "the EBS volume ran out of burst credits at 09:14" is what this domain is about. It covers what AWS measures for you, what you have to instrument yourself, how to turn that telemetry into alerts and automatic fixes, and how to read it when you are tuning for speed.

This domain carries 22% of SOA-C03 and comes first in the course because everything later reads from it. A scaling policy needs a metric. A remediation runbook needs an event. A performance fix needs a number that proves it worked.

What This Domain Covers

  • CloudWatch metrics and logs, Logs Insights queries, the CloudWatch agent, and CloudTrail for API-level history
  • Container Insights, managed Prometheus and Grafana, and monitoring for serverless and AI workloads
  • CloudWatch alarms, composite alarms, alarm actions, SNS alerting, and dashboards
  • Event-driven remediation with EventBridge rules and Pipes, Systems Manager Automation runbooks, and the common auto-remediation patterns
  • Compute and storage tuning: right-sizing with Compute Optimizer, placement groups, EBS volume types and IOPS, S3 request performance and transfer, EFS and FSx
  • RDS Performance Insights, Enhanced Monitoring, RDS Proxy, and connection-level tuning

Why It Matters

Exam questions in this domain are diagnostic. You get symptoms (an alarm stuck in INSUFFICIENT_DATA, a Lambda function whose logs never appear, an EC2 instance with no memory metric in CloudWatch) and are asked what to check. Pattern-matching service names will not get you there. You need to know which metrics AWS publishes by default, which ones require the agent, where each log stream lands, and what an alarm actually evaluates over its period.

That same knowledge carries the rest of the course. Auto Scaling in the next domain is a CloudWatch alarm with a scaling policy attached. Automated compliance response in the Security domain is an EventBridge rule with a different event source. Learn the telemetry layer properly here and you spend the remaining four domains applying it instead of relearning it.

Topics in this domain

Send us a message

Have a question about a course, a partnership, or the product? Drop us a line, we reply by email.

We reply within 2 business days.

© 2026 Syllaro Academy. All rights reserved.