[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"cheat-sheet---en":3,"domain-info---en":3,"topic-info----en":3,"prev-cloud-computing-fundamentals-security-and-reliability-reliability-and-operations-high-availability-and-fault-tolerance-en":4,"next-cloud-computing-fundamentals-security-and-reliability-reliability-and-operations-high-availability-and-fault-tolerance-en":326,"lesson-cloud-computing-fundamentals-security-and-reliability-reliability-and-operations-high-availability-and-fault-tolerance-en":600},null,{"locked":5,"reason":3,"meta":6,"item":16},false,{"title":7,"description":8,"isFree":5,"estimatedMinutes":9,"difficulty":10,"learningObjectives":11},"Governance and Compliance","How cloud governance sets an organization's own rules for resource use, how compliance proves those rules satisfy an external framework or regulation, and where a provider's certifications end and a customer's own evidence has to begin.",18,"intermediate",[12,13,14,15],"Distinguish governance from compliance and explain what each one is answering","Compare ISO 27001, SOC 2, PCI DSS, HIPAA, and GDPR by what each covers and who typically needs it","Explain what a provider's compliance report proves and what it leaves for the customer to demonstrate separately","Distinguish data residency from data sovereignty and identify which one a scenario is actually testing",{"id":17,"title":7,"body":18,"description":8,"difficulty":10,"estimatedMinutes":9,"extension":261,"infographics":262,"isFree":5,"learningObjectives":263,"meta":264,"navigation":265,"path":266,"quiz":267,"seo":323,"stem":324,"__hash__":325},"courses/courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/01-cloud-security/04-governance-and-compliance.md",{"type":19,"value":20,"toc":247},"minimark",[21,26,30,33,37,40,43,47,143,146,150,153,157,160,163,167,170,174,177,181,240,244],[22,23,25],"h2",{"id":24},"a-certificate-that-didnt-cover-what-they-thought","A certificate that didn't cover what they thought",[27,28,29],"p",{},"A healthcare startup signs up with a major cloud provider, sees the SOC 2 and ISO 27001 badges on the provider's compliance page, and tells its board the platform is HIPAA-compliant. Months later, an actual HIPAA audit asks for the startup's own access logs, its encryption configuration, and records of staff security training, none of which the provider's certifications ever claimed to cover.",[27,31,32],{},"This is the same gap the first lesson in this topic named as a misconception, now playing out in full: a provider's certification proves its own infrastructure meets a standard. Whether the application built on top of it meets that same standard is a separate question, and it's exactly what governance and compliance, as a discipline, exist to answer.",[22,34,36],{"id":35},"governance-asks-inward-compliance-points-outward","Governance asks inward, compliance points outward",[27,38,39],{},"Governance is an organization's own rulebook for how its cloud resources get configured and used: who reviews access requests, how resources get tagged, what configuration baseline every new service has to meet before it goes live. It's internally driven, and a company can have strong governance with no external framework in play at all.",[27,41,42],{},"Compliance is different: it's proving, to an auditor, a regulator, or a customer's own procurement team, that an organization's practices meet a specific external bar. That bar might be a certifiable standard a company chooses to pursue, or a law it has no choice but to satisfy. Good governance makes compliance easier to demonstrate, but the two answer different questions: governance asks \"are we doing this consistently,\" compliance asks \"can we prove it to someone outside the organization.\"",[22,44,46],{"id":45},"the-compliance-landscape-what-each-framework-actually-covers","The compliance landscape: what each framework actually covers",[48,49,50,69],"table",{},[51,52,53],"thead",{},[54,55,56,60,63,66],"tr",{},[57,58,59],"th",{},"Framework",[57,61,62],{},"What it covers",[57,64,65],{},"Who typically needs it",[57,67,68],{},"Certification or law",[70,71,72,87,101,115,129],"tbody",{},[54,73,74,78,81,84],{},[75,76,77],"td",{},"ISO 27001",[75,79,80],{},"A full information security management system: risk assessment, access control, incident management",[75,82,83],{},"Any organization, internationally recognized",[75,85,86],{},"Certification, valid 3 years",[54,88,89,92,95,98],{},[75,90,91],{},"SOC 2",[75,93,94],{},"Five trust service criteria: security, availability, processing integrity, confidentiality, privacy",[75,96,97],{},"US-focused, especially SaaS and tech vendors",[75,99,100],{},"Attestation report, an auditor's opinion",[54,102,103,106,109,112],{},[75,104,105],{},"PCI DSS",[75,107,108],{},"Protecting payment card data specifically, 12 requirements from network security to access restriction",[75,110,111],{},"Any organization that stores, processes, or transmits cardholder data",[75,113,114],{},"Certifiable standard",[54,116,117,120,123,126],{},[75,118,119],{},"HIPAA",[75,121,122],{},"Protecting patient health information (PHI)",[75,124,125],{},"US healthcare organizations and their vendors",[75,127,128],{},"US federal law",[54,130,131,134,137,140],{},[75,132,133],{},"GDPR",[75,135,136],{},"Personal data of individuals in the EU, including international transfer rules",[75,138,139],{},"Any organization handling EU residents' personal data, regardless of where the organization is based",[75,141,142],{},"EU law",[27,144,145],{},"ISO 27001 and SOC 2 overlap substantially in the controls they check, enough that a well-run compliance program can often satisfy both without duplicating the work, but they serve different audiences: SOC 2 is what a US enterprise customer usually asks for, ISO 27001 is what an international one usually expects.",[22,147,149],{"id":148},"where-to-find-the-providers-half-of-the-evidence","Where to find the provider's half of the evidence",[27,151,152],{},"A cloud provider doesn't email its compliance reports on request; it publishes them through a self-service portal. AWS's version is called AWS Artifact, giving any account on-demand access to AWS's own SOC 1/2/3, ISO, and PCI reports, each generated with a unique watermark for the requester. That portal is the customer's starting point when an auditor asks \"show me the provider's controls,\" and it's free to use. What it can never hand over is the other half of the evidence, the customer's own configuration, access logs, and internal processes, because AWS Artifact only documents AWS's side of the shared responsibility line.",[22,154,156],{"id":155},"data-residency-versus-data-sovereignty","Data residency versus data sovereignty",[27,158,159],{},"These two terms get used interchangeably, and the exam boundary between them is worth naming explicitly. Data residency is about physical location: where, geographically, is the data stored. Data sovereignty is about legal jurisdiction: whose laws govern that data, regardless of where it happens to sit.",[27,161,162],{},"GDPR is a useful case study because it clarifies which one it actually cares about. GDPR does not require EU personal data to physically stay inside the EU. What it requires, under its rules on international transfers, is that any transfer of personal data outside the EU or EEA receive protection that's essentially equivalent to what GDPR itself provides, through mechanisms like an adequacy decision covering the destination country, or standard contractual clauses between sender and receiver. A company can choose an EU cloud region for good operational reasons, lower latency to EU users, a simpler story for customers, but doing so addresses residency, a location question. It does not, by itself, satisfy the legal transfer question GDPR is actually asking if that same data later gets copied, backed up, or processed somewhere outside the EU.",[22,164,166],{"id":165},"a-worked-example-choosing-a-region-for-eu-personal-data","A worked example: choosing a region for EU personal data",[27,168,169],{},"Say a team is building a service that stores personal data belonging to EU residents. Picking an EU-based region keeps that data physically close to its users and gives a simple answer to \"where does our data live.\" But the team still has to check every place that data travels afterward: a support tool hosted outside the EU, an analytics pipeline in a US region, a backup replicated to another continent. Each of those is a transfer under GDPR's Chapter 5 rules, and each one needs its own legal basis, an adequacy decision for that destination, or contractual clauses, regardless of how carefully the primary region was chosen.",[22,171,173],{"id":172},"misconception-revisited-inheriting-a-badge","Misconception, revisited: inheriting a badge",[27,175,176],{},"The first lesson in this topic named this misconception in the context of security generally; here it is the specific, recurring compliance mistake: assuming a provider's certification transfers automatically to whatever a customer builds on top of it. It never does. AWS Artifact, or its equivalent on any provider, hands over evidence of the provider's own controls. The customer still has to assemble evidence of its own: access reviews, encryption configuration from the previous lesson, employee training, incident response processes, everything that HIPAA, PCI DSS, or a customer's own security questionnaire actually asks to see.",[22,178,180],{"id":179},"exam-cues-matching-a-scenario-to-a-framework","Exam cues: matching a scenario to a framework",[48,182,183,193],{},[51,184,185],{},[54,186,187,190],{},[57,188,189],{},"The scenario says...",[57,191,192],{},"Points to",[70,194,195,202,209,216,224,232],{},[54,196,197,200],{},[75,198,199],{},"\"Patient health information,\" \"healthcare provider\"",[75,201,119],{},[54,203,204,207],{},[75,205,206],{},"\"Credit card data,\" \"cardholder data environment\"",[75,208,105],{},[54,210,211,214],{},[75,212,213],{},"\"Personal data of EU individuals,\" \"international transfer\"",[75,215,133],{},[54,217,218,221],{},[75,219,220],{},"\"Where our data is physically stored\"",[75,222,223],{},"Data residency",[54,225,226,229],{},[75,227,228],{},"\"Whose laws apply to our data\"",[75,230,231],{},"Data sovereignty",[54,233,234,237],{},[75,235,236],{},"\"Need the provider's own audit reports\"",[75,238,239],{},"AWS Artifact, or the provider's equivalent portal",[22,241,243],{"id":242},"where-this-leaves-you","Where this leaves you",[27,245,246],{},"Compliance is never something a customer inherits wholesale from a provider's badge page; it's something assembled from the provider's certified infrastructure plus the customer's own controls layered on top, the exact same boundary the shared responsibility model drew at the start of this topic, now applied specifically to proving it to an outsider. That closes out Cloud Security: you can now name who owns which security task, how identity and encryption enforce that ownership, and how governance and compliance prove it holds up under audit. The next topic in this domain leaves security behind and asks a different question entirely: once a system is secure, how does it stay running.",{"title":248,"searchDepth":249,"depth":249,"links":250},"",3,[251,253,254,255,256,257,258,259,260],{"id":24,"depth":252,"text":25},2,{"id":35,"depth":252,"text":36},{"id":45,"depth":252,"text":46},{"id":148,"depth":252,"text":149},{"id":155,"depth":252,"text":156},{"id":165,"depth":252,"text":166},{"id":172,"depth":252,"text":173},{"id":179,"depth":252,"text":180},{"id":242,"depth":252,"text":243},"md",[],[12,13,14,15],{},true,"/courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/01-cloud-security/04-governance-and-compliance",{"passingScore":268,"questions":269},70,[270,279,287,291,297,305,313],{"question":271,"type":272,"options":273,"correctAnswer":276,"explanation":278},"A healthcare startup sees the SOC 2 and ISO 27001 badges on its cloud provider's compliance page and tells its board the platform is fully HIPAA-compliant. An auditor later asks for the startup's own access logs, encryption configuration, and staff training records. What did the startup get wrong?","single",[274,275,276,277],"Nothing; the provider's certifications cover HIPAA automatically","SOC 2 and ISO 27001 aren't real compliance frameworks","It assumed the provider's infrastructure certifications covered its own application-level controls, which HIPAA also requires evidence for","HIPAA does not apply to healthcare data stored in the cloud","This is the same boundary the shared responsibility model draws for security, applied to compliance: a provider's certification proves its own infrastructure meets a standard, but a customer's access controls, encryption configuration, and internal processes need their own separate compliance evidence.",{"question":280,"type":272,"options":281,"correctAnswer":283,"explanation":286},"What is the clearest distinction between governance and compliance in a cloud context?",[282,283,284,285],"They are two names for the same activity","Governance is an organization's own rules for how cloud resources are configured and used; compliance is proving those rules satisfy an external framework or regulation","Governance only applies to cost management, and compliance only applies to security","Compliance is optional, while governance is legally required","Governance is internally driven, an organization deciding its own policies for access review, tagging, or configuration standards. Compliance points outward, demonstrating to an auditor, regulator, or customer that those internal practices meet a specific external bar, whether that bar is a certifiable standard or a legal requirement.",{"question":288,"type":272,"options":289,"correctAnswer":105,"explanation":290},"A company processes credit card payments directly and needs to prove it protects cardholder data. Which framework specifically governs this?",[119,133,105,77],"PCI DSS is built specifically around protecting payment card data, with 12 requirements covering everything from network security to restricting access by business need to know. HIPAA covers health data, GDPR covers personal data of EU individuals, and ISO 27001 is a general information security management standard, not payment-specific.",{"question":292,"type":272,"options":293,"correctAnswer":295,"explanation":296},"ISO 27001 and SOC 2 are functionally identical, so a company that has one never needs to pursue the other.",[294,295],"True","False","ISO 27001 is an internationally recognized certification for an information security management system, valid for three years and widely expected outside the US. SOC 2 is a US-focused attestation report evaluated against five trust service criteria. Their controls overlap substantially, but a company selling into both US and international markets often ends up needing both.",{"question":298,"type":272,"options":299,"correctAnswer":300,"explanation":304},"A team wants to hand its auditor AWS's own SOC 2 and PCI reports as part of a compliance review, without contacting an AWS sales representative. Which AWS service is built exactly for this?",[300,301,302,303],"AWS Artifact","AWS IAM","AWS KMS","Amazon CloudWatch","AWS Artifact is a self-service portal giving any AWS account on-demand access to AWS's own compliance reports, including SOC 1/2/3, ISO, and PCI documentation, each generated with a unique watermark. It's the customer's way to retrieve the provider's half of the compliance evidence without a manual request process.",{"question":306,"type":272,"options":307,"correctAnswer":310,"explanation":312},"A company decides to store all EU customer data in an EU-based cloud region and concludes this alone satisfies GDPR's requirements for that data. What does this conclusion overlook?",[308,309,310,311],"GDPR does not apply to data stored in the EU","GDPR requires all data to be stored outside the EU","Data residency, where data physically sits, is different from GDPR's actual requirement, which governs international transfers and requires adequate protection wherever data eventually moves","Storing data in a specific region has no relationship to any compliance framework","GDPR itself does not mandate that personal data stay physically inside the EU. What it mandates is that any transfer outside the EU or EEA receive essentially equivalent protection, through mechanisms like an adequacy decision or standard contractual clauses. Choosing an EU region is a reasonable practice, but it addresses data residency, a physical-location question, not the legal-transfer question GDPR is actually asking.",{"question":314,"type":315,"options":316,"correctAnswers":321,"explanation":322},"Which of the following are true about NIST SP 800-144, the guidelines on security and privacy in public cloud computing? (Select all that apply.)","multiple",[317,318,319,320],"It is published by the US National Institute of Standards and Technology","It is a legally binding regulation enforced worldwide","It discusses security and privacy challenges and safeguards for organizations moving to public cloud","It is aimed at people making decisions about cloud computing initiatives, including security professionals and IT managers",[317,319,320],"NIST publications like SP 800-144 are guidance documents, widely referenced and often required for US federal systems, but not a global legal mandate in the way GDPR is. They're written for the people actually deciding how an organization approaches cloud security, from executives to system administrators.",{"title":7,"description":8},"courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/01-cloud-security/04-governance-and-compliance","_PCoeERQGlRTd0TQ-nUyl6XNcQqfriWuYbxryLMEJ28",{"locked":5,"reason":3,"meta":327,"item":335},{"title":328,"description":329,"isFree":5,"estimatedMinutes":9,"difficulty":10,"learningObjectives":330},"Scalability and Elasticity","Why a system that can grow and a system that grows and shrinks itself automatically are 2 different skills, and how an Auto Scaling group's capacity settings and scaling policies deliver the second one.",[331,332,333,334],"Distinguish scalability from elasticity and explain why they answer different questions","Explain how an Auto Scaling group's minimum, desired, and maximum capacity settings work together","Compare target tracking, step, and scheduled scaling policies and identify which fits a given demand pattern","Apply elasticity to a worked example that shows the cost impact of matching capacity to real-time demand",{"id":336,"title":328,"body":337,"description":329,"difficulty":10,"estimatedMinutes":9,"extension":261,"infographics":528,"isFree":5,"learningObjectives":539,"meta":540,"navigation":265,"path":541,"quiz":542,"seo":597,"stem":598,"__hash__":599},"courses/courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/02-reliability-and-operations/02-scalability-and-elasticity.md",{"type":19,"value":338,"toc":516},[339,343,346,350,353,357,360,364,367,370,374,377,380,385,389,439,442,446,449,453,456,460,511,513],[22,340,342],{"id":341},"the-2-ways-a-plan-for-more-traffic-can-go-wrong","The 2 ways a plan for more traffic can go wrong",[27,344,345],{},"A retailer sizes its checkout servers for an average shopping day and gets crushed the first time a flash sale doubles traffic overnight; requests queue, pages time out, and the sale becomes the outage story instead of the revenue story. A different team overcorrects: it provisions enough servers to survive its single biggest sale day of the year, then pays for that same fleet the other 364 days, most of them running at a fraction of their capacity. Neither team has solved the actual problem, which has 2 separate parts: can the system grow at all, and does it grow (and shrink) itself to match what is actually happening right now.",[22,347,349],{"id":348},"recall-growth-you-already-know-how-to-do","Recall: growth you already know how to do",[27,351,352],{},"The compute lesson earlier in this course covered horizontal scaling, adding more instances of the same size behind a load balancer, as one answer to running out of capacity. What it did not cover is who decides when to add those instances, or when to take them away again once the load passes. That gap, the decision-making layer sitting on top of horizontal scaling, is exactly what this lesson fills in.",[22,354,356],{"id":355},"scalability-the-capacity-to-grow","Scalability: the capacity to grow",[27,358,359],{},"Scalability is a system's ability to handle more load by adding resources, whether that means resizing 1 machine (vertical scaling) or adding more machines of the same size (horizontal scaling). Nothing in that definition says the growth has to be automatic. A team that notices rising traffic and manually launches 3 more instances has scaled the system. So has an on-premises IT department that orders and racks 2 new physical servers over the course of a month. Both improved the system's capacity to handle load; neither one did it by itself, in real time, or gave any capacity back once it stopped being needed.",[22,361,363],{"id":362},"elasticity-matching-capacity-to-demand-automatically","Elasticity: matching capacity to demand, automatically",[27,365,366],{},"Elasticity is what happens when that decision-making gets automated in both directions: capacity increases as demand rises and decreases as demand falls, without a person watching a dashboard and clicking a button each time. It is a genuinely cloud-native idea. A physical server you already own and paid for doesn't shrink your bill when it sits idle overnight; a cloud instance that gets terminated during a quiet period stops costing you anything the moment it's gone. Elasticity is the reason \"pay only for what you use\" is more than a slogan: it requires the system to actually notice when \"what you use\" has changed and to act on that, continuously, without waiting for a human.",[27,368,369],{},"Picture a highway that engineers can widen with a construction crew over several months versus a highway with a reversible lane that opens automatically when sensors detect rush-hour traffic and closes again once it clears. Widening the highway is scalability: real capacity, real effort, and no reason to ever undo it. The reversible lane is elasticity: the same road, resized instantly and automatically to the traffic actually on it right now, and closed back up the moment it isn't needed. Push the comparison further and it breaks, a reversible lane is a fixed piece of infrastructure that gets toggled, not created and destroyed, while a cloud instance is provisioned and torn down entirely each time, which is exactly what makes it cost nothing while it's gone.",[22,371,373],{"id":372},"worked-example-1-day-inside-an-auto-scaling-group","Worked example: 1 day inside an Auto Scaling group",[27,375,376],{},"An Auto Scaling group is the AWS mechanism that automates this. Say a group is configured with a minimum of 4 instances, a desired capacity of 6, and a maximum of 12. Minimum and maximum are hard boundaries the group will never cross in either direction; desired capacity is the starting point a scaling policy adjusts from within that range.",[27,378,379],{},"At 9am, traffic climbs and average CPU utilization across the group rises past its target. A target-tracking policy launches new instances, 1 or 2 at a time, until CPU utilization settles back near target, capacity landing around 10 instances by mid-morning. Traffic holds through the day and eases off by 6pm; the same policy now terminates instances as CPU utilization drops below target, easing capacity back down toward 6, and further toward the 4-instance minimum overnight. The group never touches its 12-instance ceiling that day, and it never runs below 4. The team pays for roughly 10 instances during the 9 busy hours that needed them and far fewer overnight, instead of running 12 (or even a flat 10) for all 24.",[381,382],"infographic",{"alt":383,"slug":384},"A chart showing customer traffic rising and falling over a 24-hour day as a smooth curve, with an Auto Scaling group's stepped capacity line tracking it automatically between a minimum and a maximum boundary.","scalability-and-elasticity-capacity-vs-demand",[22,386,388],{"id":387},"scaling-policies-3-ways-to-tell-the-group-when-to-act","Scaling policies: 3 ways to tell the group when to act",[48,390,391,404],{},[51,392,393],{},[54,394,395,398,401],{},[57,396,397],{},"Policy",[57,399,400],{},"How it decides",[57,402,403],{},"Best fit",[70,405,406,417,428],{},[54,407,408,411,414],{},[75,409,410],{},"Target tracking",[75,412,413],{},"Continuously adjusts capacity to hold a chosen metric, like average CPU utilization, near a target value",[75,415,416],{},"Live, somewhat unpredictable demand",[54,418,419,422,425],{},[75,420,421],{},"Step scaling",[75,423,424],{},"Adds or removes capacity in defined steps sized to how far a metric has moved past its threshold",[75,426,427],{},"Demand that can spike sharply and needs a proportionally larger response",[54,429,430,433,436],{},[75,431,432],{},"Scheduled scaling",[75,434,435],{},"Sets capacity ahead of a known time, independent of any live metric",[75,437,438],{},"Predictable patterns, like a retailer's daily 9am traffic or a monthly batch job",[27,440,441],{},"These are not mutually exclusive. A retail team might run scheduled scaling to pre-warm capacity just before its known 9am spike, then let target tracking handle the rest of the day's smaller fluctuations on top of that head start.",[22,443,445],{"id":444},"elasticity-also-means-self-healing","Elasticity also means self-healing",[27,447,448],{},"Auto Scaling groups do more than resize for demand. They continuously health-check every instance and replace any that fail, automatically, to keep the group at its desired capacity even when nothing about traffic has changed. That connects directly back to the previous lesson: an Auto Scaling group spread across multiple Availability Zones is part of how a highly available architecture gets built in the first place, not a separate concern from it.",[22,450,452],{"id":451},"misconception-auto-scaling-is-not-just-about-scaling-up","Misconception: \"Auto Scaling\" is not just about scaling up",[27,454,455],{},"It's tempting to hear \"Auto Scaling\" and picture only the exciting half: automatically adding capacity to survive a traffic spike. The half that actually saves money is scaling in, removing capacity once demand drops. A group with no minimum worth respecting or no scale-in policy attached just becomes an expensive, self-launching version of plain scalability: it grows on its own, but it never gives anything back, and the cost advantage elasticity is supposed to deliver never shows up on the bill.",[22,457,459],{"id":458},"exam-cues-reading-a-scalability-versus-elasticity-question","Exam cues: reading a scalability-versus-elasticity question",[48,461,462,470],{},[51,463,464],{},[54,465,466,468],{},[57,467,189],{},[57,469,192],{},[70,471,472,480,488,496,503],{},[54,473,474,477],{},[75,475,476],{},"\"Handle a growing amount of load by adding resources\"",[75,478,479],{},"Scalability",[54,481,482,485],{},[75,483,484],{},"\"Automatically adjusts to real-time demand,\" \"pay only for what you use,\" \"scales in when idle\"",[75,486,487],{},"Elasticity",[54,489,490,493],{},[75,491,492],{},"\"Minimum,\" \"desired,\" \"maximum capacity\"",[75,494,495],{},"Auto Scaling group settings",[54,497,498,501],{},[75,499,500],{},"\"Known, predictable spike at a specific time\"",[75,502,432],{},[54,504,505,508],{},[75,506,507],{},"\"Maintain a target CPU utilization\"",[75,509,510],{},"Target tracking scaling",[22,512,243],{"id":242},[27,514,515],{},"Scalability answers whether a system can grow at all; elasticity answers whether it grows, and shrinks, itself, automatically, in step with demand that changes minute to minute. Every cloud workload needs the first. Only workloads with genuinely variable demand get real value from the second, which is most of them, since traffic that never varies is rare outside a handful of steady internal systems. The next lesson assumes all of this, redundancy, health checks, right-sized capacity, and asks what happens anyway: when a failure is bigger than any of it can absorb, and a team has to fall back on backups and a disaster recovery plan.",{"title":248,"searchDepth":249,"depth":249,"links":517},[518,519,520,521,522,523,524,525,526,527],{"id":341,"depth":252,"text":342},{"id":348,"depth":252,"text":349},{"id":355,"depth":252,"text":356},{"id":362,"depth":252,"text":363},{"id":372,"depth":252,"text":373},{"id":387,"depth":252,"text":388},{"id":444,"depth":252,"text":445},{"id":451,"depth":252,"text":452},{"id":458,"depth":252,"text":459},{"id":242,"depth":252,"text":243},[529],{"slug":384,"concept":530,"style":531,"aspectRatio":532,"labels":533},"A single quantitative chart over a 24-hour horizontal axis. A smooth curved line shows customer traffic rising through the morning, peaking at midday, and falling overnight. A stepped line overlays it, showing Auto Scaling group capacity rising and falling in blocks to track the traffic curve, bounded above and below by dashed horizontal lines marking maximum and minimum capacity. A footer strip carries the cost takeaway.","diagram","16:9",[534,535,536,537,538],"Traffic: rises in the morning, peaks midday, falls overnight","Auto Scaling group capacity: steps up and down to track traffic","Maximum capacity: 12 instances","Minimum capacity: 4 instances","Elasticity means paying for the stepped line, not a flat line held at the maximum all day.",[331,332,333,334],{},"/courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/02-reliability-and-operations/02-scalability-and-elasticity",{"passingScore":268,"questions":543},[544,552,560,568,577,585,589],{"question":545,"type":272,"options":546,"correctAnswer":548,"explanation":551},"A team manually doubles the size of its database server after a product launch permanently increases traffic. What is this an example of?",[547,548,549,550],"Elasticity, since capacity changed to meet demand","Scalability, since the system's capacity to handle load grew","Fault tolerance, since the change prevents future outages","A single point of failure being removed","Scalability is the ability to grow, whether by resizing 1 machine or adding more of them, and whether that growth is manual, scripted, or automatic. This example is scalability because the capacity grew; it is not elasticity because nothing here happened automatically in response to a live demand signal, and the change was not reversed once it was no longer needed.",{"question":553,"type":272,"options":554,"correctAnswer":556,"explanation":559},"Which of the following best distinguishes elasticity from scalability?",[555,556,557,558],"Elasticity only applies to compute, while scalability applies to storage and databases as well","Scalability is the system's capacity to grow at all; elasticity is that growth (and shrinkage) happening automatically in response to real-time demand","Scalability requires downtime, while elasticity never does","They are the same concept, described with 2 different marketing terms","A system can be scalable without being elastic, for example an on-premises data center that can add more physical servers, slowly and manually, but can never give any of them back once demand drops. Elasticity specifically adds automation in both directions, scaling out and scaling back in, which is why it is a cloud-era concept more than an on-premises one.",{"question":561,"type":272,"options":562,"correctAnswer":565,"explanation":567},"An Auto Scaling group is configured with a minimum of 4, a desired capacity of 6, and a maximum of 12. What does the maximum of 12 guarantee?",[563,564,565,566],"The group will always run exactly 12 instances","The group will never run fewer than 12 instances","The group will never scale beyond 12 instances, no matter how high demand climbs","The group will terminate all instances once traffic exceeds 12 requests per second","Maximum capacity is a ceiling, not a target: the group can run anywhere from its minimum up to its maximum, and scaling policies decide the actual count moment to moment within that range. Desired capacity, 6 in this example, is the starting point a scaling policy adjusts from, not a fixed value the group is locked to.",{"question":569,"type":315,"options":570,"correctAnswers":575,"explanation":576},"Which of the following are true about Amazon EC2 Auto Scaling? (Select all that apply.)",[571,572,573,574],"It automatically replaces instances that fail their health checks to maintain the desired capacity","It balances instances across the Availability Zones configured for the group","It can only scale a group up; scaling back in requires manual intervention","Target tracking scaling adjusts capacity to hold a chosen metric, such as average CPU utilization, near a target value",[571,572,574],"Auto Scaling groups scale in both directions automatically once a policy is attached; a scale-out-only group would keep adding instances forever and never realize the cost savings elasticity is meant to deliver. Health-check-based replacement and AZ balancing are both automatic parts of what makes an Auto Scaling group reliable, not just elastic.",{"question":578,"type":272,"options":579,"correctAnswer":581,"explanation":584},"A scheduled scaling policy is best suited for which situation?",[580,581,582,583],"Traffic spikes that are unpredictable and happen at random times","A demand pattern that is known in advance, such as a retailer's traffic spiking every day at 9am","Keeping a single metric, like CPU utilization, at a constant target at all times","Replacing an unhealthy instance immediately after it fails a health check","Scheduled scaling sets capacity ahead of a known pattern, so instances are already warmed up before the predictable spike hits, instead of reacting after the fact. Target tracking is the better fit for unpredictable, live demand, since it continuously adjusts to hold a metric near its target regardless of when the change happens.",{"question":586,"type":272,"options":587,"correctAnswer":295,"explanation":588},"A well-configured Auto Scaling group only ever adds capacity in response to rising demand; it never removes it once traffic drops.",[294,295],"Scaling in, removing instances as demand falls, is where elasticity actually pays for itself. A group with no minimum-respecting scale-in behavior is not elastic, it is just an expensive, manually-triggered version of scalability that happens to launch itself but never shrinks back down.",{"question":590,"type":272,"options":591,"correctAnswer":593,"explanation":596},"A team sizes its Auto Scaling group's maximum capacity for its historical worst-case traffic day, but sets a low minimum and a target-tracking policy for everyday demand. What is the main cost benefit of this setup compared to running the maximum capacity permanently?",[592,593,594,595],"None; AWS charges the same regardless of how many instances are actually running","The team pays only for the capacity it actually needs at each moment, instead of paying for peak capacity around the clock","It removes the need for a load balancer entirely","It guarantees the system becomes fault-tolerant","This is the entire economic case for elasticity: instances cost money by the hour they run, so a group that tracks real demand and scales back in during quiet periods avoids paying for idle peak-day capacity 24 hours a day. A fixed fleet sized for the worst case would handle the same traffic but at a much higher steady-state bill.",{"title":328,"description":329},"courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/02-reliability-and-operations/02-scalability-and-elasticity","0PTUatXygnuSsmAEr9ZZClZQBUy9H7qc9Z4RJq5xNXs",{"locked":5,"reason":3,"meta":601,"item":610},{"title":602,"description":603,"isFree":265,"estimatedMinutes":9,"difficulty":604,"learningObjectives":605},"High Availability and Fault Tolerance","Why a system that recovers from a failure in under 2 minutes and a system that never blinks at all are solving the same problem with 2 different budgets, and the redundancy patterns cloud teams use to build each.","beginner",[606,607,608,609],"Define high availability and fault tolerance and explain what separates them","Explain how spreading compute and database capacity across Availability Zones removes a single point of failure","Compare active-active and active-passive redundancy using a worked failover example","Apply the high-availability-versus-fault-tolerance distinction to a scenario and identify which pattern its requirement actually calls for",{"id":611,"title":602,"body":612,"description":603,"difficulty":604,"estimatedMinutes":9,"extension":261,"infographics":862,"isFree":265,"learningObjectives":872,"meta":873,"navigation":265,"path":874,"quiz":875,"seo":930,"stem":931,"__hash__":932},"courses/courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/02-reliability-and-operations/01-high-availability-and-fault-tolerance.md",{"type":19,"value":613,"toc":851},[614,618,621,625,628,632,635,638,642,645,648,652,656,659,720,723,727,730,779,782,786,789,793,843,846,848],[22,615,617],{"id":616},"the-2am-hardware-failure","The 2am hardware failure",[27,619,620],{},"A retailer runs its checkout API on a single virtual machine. At 2am, the physical host underneath that instance fails, silently and completely, the way hardware eventually does no matter whose logo is on the rack. The instance is gone, and so is checkout, until someone notices and launches a replacement. That lone instance was a single point of failure: 1 resource whose failure alone took the whole system down with it. Removing single points of failure like it is what the rest of this lesson is about, and it turns out there are 2 different levels of \"removed,\" not 1.",[22,622,624],{"id":623},"recall-the-building-blocks-you-already-have","Recall: the building blocks you already have",[27,626,627],{},"You already have the 2 tools that make removal possible. Availability Zones give you physically separate locations to spread instances across, so 1 data center's bad night doesn't take out everything at once. A load balancer sits in front of those instances, health-checks each one, and routes traffic only to the ones still responding. Put the 2 together, instances spread across multiple AZs behind a load balancer, and a repeat of the 2am failure stops being catastrophic. What it becomes next, a brief hiccup or nothing at all, depends on how much spare capacity was already running the moment the failure hit.",[22,629,631],{"id":630},"high-availability-staying-up-through-a-brief-interruption","High availability: staying up through a brief interruption",[27,633,634],{},"Amazon RDS Multi-AZ deployments show this pattern clearly. A primary database instance handles every read and write, while a standby in a different Availability Zone stays synchronously updated in the background, ready but idle. If the primary fails, RDS detects it, promotes the standby, and repoints the database's DNS record to it, typically in 60 to 120 seconds, with 0 data loss. For those 60 to 120 seconds, though, connections to the database do drop or queue.",[27,636,637],{},"That is high availability: the system recovers automatically and quickly, but the recovery itself is a visible, if brief, gap. Most applications only need this. A minute of connection errors during a rare zone failure is a very different problem than the outright outage the single-instance retailer suffered.",[22,639,641],{"id":640},"fault-tolerance-staying-up-with-no-interruption-at-all","Fault tolerance: staying up with no interruption at all",[27,643,644],{},"Now go back to the checkout API, but redesigned. Say it needs 6 running instances to handle its normal load. Spread across 2 Availability Zones, 3 instances each, that is already highly available: lose 1 zone, and the surviving 3 instances keep checkout running, just at half capacity until new instances launch behind them. Spread across 3 Availability Zones instead, with 3 instances in each, 9 instances total, and losing any 1 zone still leaves 6 instances live, exactly the capacity checkout needs. Nothing about the user experience changes. That is fault tolerance: the system was never actually short of capacity, because the \"spare\" instances were already running and already serving traffic before the failure happened.",[27,646,647],{},"The comparison is the entire distinction in 1 image: a highly available design has enough redundancy to survive a failure, and a fault-tolerant design has enough redundancy that surviving it produces no dip to survive. Fault tolerance is closer to a self-sealing tire that keeps a car driving at full speed through a puncture than to a spare tire that gets the car moving again after a stop to change it, which is what high availability's recovery window looks like. Push the analogy further than that single point, though, and it breaks: a spare tire is cheap to carry unused, while fault-tolerant capacity bills you every hour it sits idle, waiting for a failure that may never come.",[381,649],{"alt":650,"slug":651},"A side-by-side comparison showing a highly available setup of 6 instances across 2 zones losing half its capacity when 1 zone fails, next to a fault-tolerant setup of 9 instances across 3 zones keeping its full 6-instance capacity when 1 zone fails.","high-availability-and-fault-tolerance-redundancy-capacity",[22,653,655],{"id":654},"where-redundancy-has-to-reach-every-layer-not-just-compute","Where redundancy has to reach: every layer, not just compute",[27,657,658],{},"A system is only as available as its least redundant layer. Spreading the compute tier across zones does nothing if the database, the load balancer, or DNS is still a single point of failure underneath it.",[48,660,661,674],{},[51,662,663],{},[54,664,665,668,671],{},[57,666,667],{},"Layer",[57,669,670],{},"Single point of failure",[57,672,673],{},"Redundancy pattern",[70,675,676,687,698,709],{},[54,677,678,681,684],{},[75,679,680],{},"Compute",[75,682,683],{},"1 instance",[75,685,686],{},"Multiple instances across AZs behind a load balancer",[54,688,689,692,695],{},[75,690,691],{},"Load balancer",[75,693,694],{},"1 load balancer node",[75,696,697],{},"Managed load balancers run redundantly across AZs by default",[54,699,700,703,706],{},[75,701,702],{},"Database",[75,704,705],{},"1 database instance",[75,707,708],{},"Multi-AZ deployment with automatic failover",[54,710,711,714,717],{},[75,712,713],{},"DNS",[75,715,716],{},"1 static record with no health awareness",[75,718,719],{},"Health-checked DNS routing that stops pointing at a failed target",[27,721,722],{},"Cloud providers build managed load balancers and managed DNS health checks to already be redundant, so the layers a team usually has to design for explicitly are compute and database.",[22,724,726],{"id":725},"active-active-versus-active-passive","Active-active versus active-passive",[27,728,729],{},"The 2 redundancy patterns above have names, and the difference between them shows up again later in this topic. Active-active means every replica is serving traffic at the same time, the way a load balancer spreads requests across all healthy compute instances right now, not just when something fails. Active-passive means 1 side does the work while the other stays synchronized and idle, ready to take over, the way an RDS Multi-AZ standby behaves.",[48,731,732,744],{},[51,733,734],{},[54,735,736,738,741],{},[57,737],{},[57,739,740],{},"Active-active",[57,742,743],{},"Active-passive",[70,745,746,757,768],{},[54,747,748,751,754],{},[75,749,750],{},"Who serves traffic normally",[75,752,753],{},"All replicas, simultaneously",[75,755,756],{},"1 primary only",[54,758,759,762,765],{},[75,760,761],{},"What a failure looks like",[75,763,764],{},"Remaining replicas absorb the load already",[75,766,767],{},"A failover promotes the standby",[54,769,770,773,776],{},[75,771,772],{},"Typical use",[75,774,775],{},"Load-balanced compute tiers",[75,777,778],{},"Multi-AZ managed databases",[27,780,781],{},"Neither pattern is inherently better. Compute tiers are usually active-active because splitting stateless requests across many identical instances is straightforward. Databases are more often active-passive because keeping every replica simultaneously writable raises hard consistency questions that a single active writer avoids.",[22,783,785],{"id":784},"misconception-multi-az-does-not-automatically-mean-fault-tolerant","Misconception: \"multi-AZ\" does not automatically mean \"fault-tolerant\"",[27,787,788],{},"It is tempting to hear \"deployed across multiple Availability Zones\" and assume that settles the reliability question. It settles only half of it. Multi-AZ tells you a failure in 1 zone will not take down every instance. It says nothing about whether the surviving capacity is enough to keep serving your full load without a dip. A team that spreads 6 instances across 2 zones has genuinely improved its availability over a single-AZ deployment, but it has not made itself fault-tolerant, and assuming otherwise is exactly the gap that shows up as an unplanned capacity shortage during the next zone-level event.",[22,790,792],{"id":791},"exam-cues-reading-a-high-availability-versus-fault-tolerance-question","Exam cues: reading a high-availability-versus-fault-tolerance question",[48,794,795,803],{},[51,796,797],{},[54,798,799,801],{},[57,800,189],{},[57,802,192],{},[70,804,805,813,821,829,836],{},[54,806,807,810],{},[75,808,809],{},"\"Brief interruption is acceptable,\" \"automatic recovery,\" \"minimal downtime\"",[75,811,812],{},"High availability",[54,814,815,818],{},[75,816,817],{},"\"No interruption at all,\" \"seamless,\" \"users must never notice\"",[75,819,820],{},"Fault tolerance",[54,822,823,826],{},[75,824,825],{},"\"1 instance,\" \"1 data center,\" \"no redundancy\"",[75,827,828],{},"A single point of failure to remove first",[54,830,831,834],{},[75,832,833],{},"\"Standby,\" \"failover,\" \"promoted\"",[75,835,743],{},[54,837,838,841],{},[75,839,840],{},"\"All instances serve traffic simultaneously\"",[75,842,740],{},[27,844,845],{},"The trap to watch for: a scenario that describes a solid multi-AZ setup and asks whether it is fault-tolerant. Check the capacity math, not just the zone count, before answering.",[22,847,243],{"id":242},[27,849,850],{},"Both patterns start from the same move, removing a single point of failure, and then diverge on how much spare capacity you are willing to run and pay for around the clock. High availability tolerates a short, automated recovery window; fault tolerance eliminates that window by keeping the redundant capacity already live. Most applications only need the first. The next lesson turns from staying up during a failure to a related but separate question: once you have decided how much capacity a system needs, who decides when to add more of it, and when to take it away again.",{"title":248,"searchDepth":249,"depth":249,"links":852},[853,854,855,856,857,858,859,860,861],{"id":616,"depth":252,"text":617},{"id":623,"depth":252,"text":624},{"id":630,"depth":252,"text":631},{"id":640,"depth":252,"text":641},{"id":654,"depth":252,"text":655},{"id":725,"depth":252,"text":726},{"id":784,"depth":252,"text":785},{"id":791,"depth":252,"text":792},{"id":242,"depth":252,"text":243},[863],{"slug":651,"concept":864,"style":865,"aspectRatio":532,"labels":866},"A 2-panel side-by-side comparison. Left panel titled Highly Available shows 2 Availability Zone boxes, each holding 3 instance icons, 6 total, with 1 zone grayed out as failed and a capacity bar dropping from 6 to 3. Right panel titled Fault Tolerant shows 3 Availability Zone boxes, each holding 3 instance icons, 9 total, with 1 zone grayed out as failed and a capacity bar staying flat at 6. A footer strip carries the cost takeaway.","comparison",[867,868,869,870,871],"Highly Available: 6 instances across 2 zones","1 zone fails: capacity drops to 3 until new instances launch","Fault Tolerant: 9 instances across 3 zones","1 zone fails: 6 instances keep serving traffic, no drop","Fault tolerance means paying for capacity that sits idle, just to buy zero interruption.",[606,607,608,609],{},"/courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/02-reliability-and-operations/01-high-availability-and-fault-tolerance",{"passingScore":268,"questions":876},[877,885,893,901,910,914,922],{"question":878,"type":272,"options":879,"correctAnswer":881,"explanation":884},"A retailer runs its checkout API on a single virtual machine. At 2am, the physical host underneath it fails and checkout goes down until someone notices and launches a replacement. What term describes the instance in this scenario?",[880,881,882,883],"A load balancer target","A single point of failure","A fault-tolerant resource","An Availability Zone","A single point of failure is any one resource whose failure alone takes the whole system down. The instance qualifies because nothing else was running to absorb checkout traffic when it disappeared. Removing single points of failure, not making any one resource perfectly reliable, is the actual goal of this lesson.",{"question":886,"type":272,"options":887,"correctAnswer":889,"explanation":892},"An Amazon RDS Multi-AZ database fails over from its primary to its standby in 90 seconds, with zero data loss but a brief drop in database connections during the switch. This is an example of which pattern?",[888,889,890,891],"Fault tolerance, since no data was lost","High availability, since the interruption was brief rather than absent","Elasticity, since capacity changed automatically","A single point of failure, since only 1 instance was serving traffic beforehand","High availability tolerates a brief interruption while a standby takes over. The RDS standby was not already serving traffic, so the switch itself, however fast, is a visible gap. Fault tolerance requires that spare capacity already be live and absorbing load, so a failure produces no gap at all, not just a short one.",{"question":894,"type":272,"options":895,"correctAnswer":897,"explanation":900},"A team needs 6 running instances to handle its normal checkout load. Which of the following setups is fault-tolerant against the loss of any single Availability Zone, not just highly available?",[896,897,898,899],"6 instances split across 2 Availability Zones, 3 in each","9 instances split across 3 Availability Zones, 3 in each","6 instances all placed in a single Availability Zone","6 instances split across 2 Availability Zones is already fault-tolerant, since it uses more than 1 zone","Losing 1 of 3 zones in the 9-instance setup still leaves 6 instances running, exactly the capacity the checkout load needs, so nothing about the user experience changes. The 6-instance, 2-zone setup is highly available (it survives the failure) but not fault-tolerant: losing 1 zone drops capacity to 3, a visible degradation until new instances launch.",{"question":902,"type":315,"options":903,"correctAnswers":908,"explanation":909},"Which of the following are true about fault tolerance compared to high availability? (Select all that apply.)",[904,905,906,907],"Fault tolerance requires redundant capacity that is already running and absorbing load before a failure happens","Fault tolerance always costs less than high availability, since it prevents outages entirely","High availability allows a brief, automatic recovery window that users may notice","Both patterns require removing single points of failure as a starting point",[904,906,907],"Fault tolerance costs more, not less, because it means running extra capacity around the clock that only earns its keep during a failure. High availability's defining trait is that a failure is survivable but not invisible. Neither pattern is possible without first removing the single point of failure that a lone instance or lone AZ represents.",{"question":911,"type":272,"options":912,"correctAnswer":295,"explanation":913},"A highly available architecture and a fault-tolerant architecture both always use exactly the same number of running instances.",[294,295],"A fault-tolerant architecture needs enough spare, already-running capacity that losing 1 zone still leaves full capacity in place, which typically means more total instances than a highly available setup built for the same normal load. The 6-versus-9 instance comparison in this lesson is exactly that gap.",{"question":915,"type":272,"options":916,"correctAnswer":918,"explanation":921},"In an active-passive database setup like Amazon RDS Multi-AZ, what is the standby instance doing while the primary is healthy?",[917,918,919,920],"Serving read traffic to reduce load on the primary","Nothing visible to users; it stays synchronized but does not serve read or write traffic","Running a separate, unrelated workload to avoid wasting capacity","Actively load-balancing writes with the primary","Active-passive means exactly 1 side is doing the work at any moment. The standby stays synchronously updated so it can take over quickly, but it sits idle from a traffic standpoint until a failover promotes it. Active-active is the pattern where multiple replicas serve traffic simultaneously, which is what a load balancer already does across your compute instances.",{"question":923,"type":272,"options":924,"correctAnswer":926,"explanation":929},"An exam scenario describes a payment-processing system that cannot show users any interruption, even a few seconds, during a single Availability Zone outage. Which pattern does the scenario call for?",[925,926,927,928],"High availability, since a short automated recovery is normally acceptable","Fault tolerance, since the requirement rules out any visible gap at all","Elasticity, since the system needs to grow with demand","A backup and restore strategy, since data loss is the main concern","The phrase \"cannot show users any interruption, even a few seconds\" is the exact keyword pattern that rules out high availability, which by definition tolerates a brief automated recovery. Only a fault-tolerant design, with spare capacity already live before the failure, meets a zero-visible-gap requirement.",{"title":602,"description":603},"courses/cloud-computing-fundamentals/en/domains/03-security-and-reliability/02-reliability-and-operations/01-high-availability-and-fault-tolerance","5Xh9o9PQt0Yepp0K4AUTkTESOZ1VsXntBg5JLaS6IlI"]