Topic

Operational Automation with Systems Manager

Day-2 operations at scale: reaching a fleet through Systems Manager instead of SSH, keeping nodes patched and their configuration from drifting, centralizing configuration in Parameter Store, and wiring routine work to events and schedules.

The previous topics built and shipped infrastructure. This one is about the servers that are already running and the work that never ends: reaching them, patching them, holding their configuration in place, and automating the parts of the week that repeat. It is also the topic where a single foundation carries everything, because no Systems Manager tool works until an instance is a managed node, and most of the failures in this area are that one condition not being met.

What This Topic Covers

  • The 3 conditions that make an EC2 instance a managed node, and the diagnostic order that finds the broken one fastest
  • Instance profiles versus Default Host Management Configuration, including the permission that makes one silently override the other
  • Session Manager as a replacement for SSH keys and bastion hosts, and the sessions its logging cannot reach
  • Run Command targeting by ID, tag, and resource group, with the concurrency and error threshold defaults that decide how a fleet-wide command behaves
  • Fleet Manager and Inventory, and the resource data sync that makes inventory queryable across accounts
  • Patch Manager: Scan versus Install, baseline approval rules and the 7-day default, how patch groups resolve to a baseline, and the compliance states with the reboot option behind them
  • Maintenance windows with duration and cutoff, patch policies for an organization, and State Manager associations for configuration that puts itself back
  • Parameter Store: the 3 types, hierarchies and the IAM path trap, tiers, policies, versions and labels, the 40 TPS ceiling, and CloudFormation dynamic references
  • S3 Event Notifications, EventBridge Scheduler, and the choice of where an operational schedule should live
  • Change Calendar as the guardrail that tells running automation to stand down

Why It Matters

Domain 3 carries 22% of SOA-C03, and this topic supplies the questions that read like a support case rather than a design exercise: an instance that will not appear in the console, a patch that was released yesterday and did not install, a node reporting compliant while a critical patch is missing, an application throttled on configuration reads the moment it scales out. Each has one specific setting behind it.

The job pressure is the same shape. The difference between an operations team that scales and one that does not is whether routine work is described once and applied by a service, or performed by a person who has to remember it.

Lessons in this topic

  1. 1Systems Manager Fleet ManagementFree
  2. 2Patch and State Management
  3. 3Parameter Store for Configuration
  4. 4Event-Driven Operations Automation
Send us a message

Have a question about a course, a partnership, or the product? Drop us a line, we reply by email.

We reply within 2 business days.

© 2026 Syllaro Academy. All rights reserved.