5) Elastic Load Balancing (ELB) & Auto Scaling Groups (ASG)
Fundamentals of Scalability and High Availability
Scalability Definition: Scalability refers to the capability of an application or system to handle increased workloads by adapting its resource allocation.
Two Primary Kinds of Scalability:
Vertical Scalability: Scaling by increasing the hardware capacity of an individual system.
Horizontal Scalability: Scaling by increasing the total number of systems or instances, also referred to as elasticity.
Scalability vs. High Availability: While closely linked, scalability and high availability address different architectural requirements:
Scalability focuses on capacity and load management.
High Availability focuses on system resilience and continuous operational uptime.
Call Center Analogy:
Scaling vertically is equivalent to upgrading a call center staff member from a junior operator to a senior operator who can handle more complex or higher volume tasks individually.
Scaling horizontally is equivalent to hiring multiple operators who work concurrently to handle increased call volume.
Vertical Scalability
Definition: Vertical scalability means increasing the structural size and hardware specs of a single running compute instance.
Amazon EC2 Example:
An application running on a small instance size such as a
t2.microis vertically scaled up by moving it to run on a larger instance size such as at2.large.
EC2 Instance Size Spectrum:
Smallest Scale Example:
t2.nanowith 0.5\n\text{ GB} of RAM and .Largest Scale Example:
u-12tb1.metalwith of RAM and .
Common Use Cases: Vertical scalability is standard for non-distributed software systems and applications, most notably traditional relational databases.
Hardware Limits: Vertical scaling is inherently capped by a hardware ceiling because there is a physical limit to the maximum memory and processing power that can be configured on a single machine.

Horizontal Scalability
Definition: Horizontal scalability means increasing the total number of instances or systems running an application in parallel.
Distributed Systems Requirement: Horizontal scaling inherently requires distributed software systems capable of processing work across multiple nodes.
Common Use Cases: Highly standard for web applications and modern cloud-native microservice architectures.
Cloud Infrastructure Integration: Cloud infrastructure providers make horizontal scaling straightforward using managed services such as Amazon EC2.
Scaling Terms:
Scale Out: Adding additional compute instances to process load.
Scale In: Removing compute instances when load decreases.

High Availability
Definition: High availability (HA) means operating an application or system across at least Availability Zones (AZs) to prevent single points of failure.
Relationship to Scalability: High availability typically works in tandem with horizontal scaling architectures.
Primary Objective: To survive the failure or total loss of a physical data center during a disaster event.
Geographic / Facility Concept:
Rather than deploying resources in a single physical location (e.g., a single building in New York), high availability distributes infrastructure across multiple independent physical locations (e.g., a building in New York and a second building in San Francisco or distinct Availability Zones within a region).
AWS EC2 Implementation:
Deploying an Auto Scaling Group across multiple Availability Zones (Auto Scaling Group multi AZ).
Deploying an Elastic Load Balancer across multiple Availability Zones (Load Balancer multi AZ).
Scalability vs. Elasticity vs. Agility
Scalability:
The structural capacity of a system to handle higher workloads by upgrading hardware power (scale up) or adding nodes/instances (scale out).
Elasticity:
The property of a system that is already scalable to automatically adjust its active resource capacity based on real-time changes in demand.
Represents a core cloud principle: pay-per-use, matching demand dynamically, and optimizing operational costs.
Agility:
A conceptual distractor unrelated to system capacity scaling.
Refers to the operational ability to provision new IT infrastructure resources rapidly with a single click.
Reduces the time required to make technical resources available to software developers from weeks down to mere minutes.
Elastic Load Balancing (ELB)
Definition: Load balancers are dedicated servers that receive incoming internet traffic and forward it upstream/downstream across multiple application targets, such as Amazon EC2 instances.
Core Architectural Benefits:
Spreads application traffic evenly across multiple downstream instances.
Exposes a single consolidated access point via Domain Name System (DNS) for the entire application.
Seamlessly isolates and handles downstream instance failures without service disruption.
Executes regular, automated health checks on all registered downstream instances.
Provides centralized SSL/TLS termination for encrypted HTTPS traffic.
Enforces high availability across multiple Availability Zones.
AWS Managed Service Advantages:
AWS provides high availability and operational guarantees for Elastic Load Balancers.
AWS handles infrastructure maintenance, system upgrades, and seamless scaling.
Provides streamlined configuration options (knobs).
While building a custom self-managed load balancer on raw compute may cost less in direct server hosting fees, managed ELB significantly lowers engineering effort and operational complexity.

Types of AWS Load Balancers
Application Load Balancer (ALB):
OSI Operating Layer: Layer (Application Layer).
Protocols Supported: HTTP, HTTPS, and gRPC.
Features: Features advanced HTTP request routing capabilities (e.g., path-based or host-based routing) and provides a static DNS URL.
Network Load Balancer (NLB):
OSI Operating Layer: Layer (Transport Layer).
Protocols Supported: TCP and UDP.
Performance: Ultra-high performance capable of processing millions of requests per second with ultra-low latency.
Features: Supports assigning static IP addresses using Elastic IP addresses.
Gateway Load Balancer (GWLB):
OSI Operating Layer: Layer (Network Layer).
Protocols Supported: GENEVE protocol encapsulation on IP packets.
Functionality: Inspects and routes network traffic through third-party security virtual appliances (e.g., firewalls, deep packet inspection, intrusion detection systems) hosted on EC2 instances before delivering traffic to destination applications.
Classic Load Balancer (CLB):
OSI Operating Layer: Layer and Layer .
Status: Legacy load balancer retired by AWS in year .

Auto Scaling Groups (ASG)
Purpose: Application workloads fluctuate over time. Auto Scaling Groups capitalize on cloud elasticity to dynamically adjust server capacity.
Core Responsibilities:
Scale Out: Automatically provision additional EC2 instances during load increases.
Scale In: Automatically terminate EC2 instances during load decreases.
Capacity Control: Enforce explicit limits for minimum, desired, and maximum instance counts.
Load Balancer Integration: Automatically register newly launched EC2 instances to an attached Elastic Load Balancer.
Self-Healing: Automatically detect and replace unhealthy compute instances.
Cost Optimization: Ensure infrastructure runs at optimal capacity to reduce unnecessary cloud operational expenditure.
Capacity Boundaries:
Minimum Size: Lower boundary limit for active EC2 instances.
Desired Capacity / Actual Size: Target count of currently running EC2 instances.
Maximum Size: Upper boundary limit for active EC2 instances.
Auto Scaling Group Scaling Strategies
Manual Scaling: Updating the desired capacity of an ASG manually via the AWS Console or CLI.
Dynamic Scaling: Automatically scaling in response to live operational metrics and demand:
Simple / Step Scaling: Triggered by Amazon CloudWatch alarm thresholds.
Example 1: If a CloudWatch alarm detects CPU utilization , add instances.
Example 2: If a CloudWatch alarm detects CPU utilization , remove instance.
Target Tracking Scaling: Automatically adds or removes instances to keep a specified metric at a fixed target value.
Example: Maintain average ASG CPU utilization at .
Scheduled Scaling: Adjusts instance capacity based on predictable usage timelines. * Example: Increase minimum instance capacity to every Friday at .
Predictive Scaling:
Leverages Machine Learning (ML) models to forecast workload shifts ahead of time based on historical usage data.
Proactively provisions EC2 instances prior to expected traffic surges.
Ideal for systems with well-defined daily or weekly recurring usage patterns.

Key Takeaways Summary
Scalability vs. High Availability vs. Elasticity vs. Agility:
Vertical Scalability: Upgrading individual server size (limited by hardware bounds).
Horizontal Scalability: Adding server nodes (enables distributed architectures).
High Availability: Operating across multiple Availability Zones to ensure survival against data center disasters.
Elasticity: Automated scaling out/in matching real-time system demand.
Agility: Rapid provisioning of cloud resources in minutes rather than weeks.
Elastic Load Balancers (ELB):
Managed service distributing network traffic across Multi-AZ EC2 backends.
Monitors instance health and performs automatic target registration/deregistration.
Active types include ALB (Layer - HTTP/HTTPS/gRPC), NLB (Layer - TCP/UDP), and GWLB (Layer - GENEVE for virtual security appliances).
Auto Scaling Groups (ASG):
Implements cloud elasticity by automatically adding or removing EC2 instances across Multi-AZ configurations.
Supports Manual, Dynamic (Simple/Step, Target Tracking), Scheduled, and Machine Learning-powered Predictive scaling strategies.
Fully integrated with Elastic Load Balancers for seamless capacity expansion and health management.