Topic

CloudWatch Metrics and Logs

How CloudWatch collects metrics and logs across EC2, containers, serverless, and AI workloads, and where CloudTrail fits in.

Every other topic in this course eventually reads a number or a log line out of CloudWatch. A scaling policy needs a metric. A remediation runbook needs an event. A performance fix needs proof that it worked. This topic builds that foundation: what AWS measures for you, what you have to instrument yourself, and how to get an answer out of the result quickly.

It also draws the line that a lot of exam questions turn on. CloudWatch records what your resources are doing. CloudTrail records who asked them to do it. Both are here, because incidents usually need both.

What This Topic Covers

  • Metric identity and storage: namespaces, dimensions, resolution, statistics and percentiles, and the retention rollup that decides what you can still query months later
  • CloudWatch Logs structure and retention, plus the three ways log data becomes useful: metric filters, subscription filters, and Logs Insights queries
  • Deploying the unified CloudWatch agent across a fleet with Systems Manager and Parameter Store, and the IAM boundary between the server policy and the admin policy
  • CloudTrail for operations: event history versus trails, the four event types, organization trails, CloudTrail Lake, and alarming on account activity
  • Container telemetry with Container Insights on ECS and EKS, and when Amazon Managed Service for Prometheus with Amazon Managed Grafana is the better fit
  • Serverless and AI monitoring: Lambda metrics and the throttle-versus-error trap, Lambda Insights, X-Ray, API Gateway latency breakdowns, and Amazon Bedrock usage tracking

Why It Matters

Questions in this part of the exam are diagnostic rather than definitional. You are given a symptom (an alarm stuck in INSUFFICIENT_DATA, an EC2 instance with no memory metric, a Lambda function whose logs never appear, an error rate that looks healthy while users report failures) and asked what to check. Recognizing service names will not get you there.

What does get you there is knowing which metrics AWS publishes by default, which ones need an agent, which features are switched off until someone switches them on, and how long each kind of data survives. That knowledge is also the daily core of the CloudOps job, and it carries straight into alarms, auto scaling, and automated remediation in the topics that follow.

Lessons in this topic

  1. 1CloudWatch Metrics FundamentalsFree
  2. 2CloudWatch Logs and Logs Insights
  3. 3Deploying the CloudWatch Agent
  4. 4CloudTrail for Operations
  5. 5Container Insights, Prometheus, and Grafana
  6. 6Monitoring Serverless and AI Workloads
Send us a message

Have a question about a course, a partnership, or the product? Drop us a line, we reply by email.

We reply within 2 business days.

© 2026 Syllaro Academy. All rights reserved.