4) EC2 Instance Storage

Amazon EBS (Elastic Block Store) Volumes

  • Definition and Core Concept:

    • An Amazon Elastic Block Store (EBS) volume is a network-attached storage drive that can be attached to Amazon EC2 instances while they are running.

    • It allows EC2 instances to persist data across reboots and instance terminations.

    • At the Cloud Practitioner (CCP) certification level, an EBS volume can only be mounted to a single EC2 instance at a time.

    • EBS volumes are logically and physically bound to a specific Availability Zone (AZ).

    • Analogy: EBS volumes operate conceptually like a high-performance network USB stick.

  • Operational Characteristics and Limitations:

    • Network Communication: EBS is a network drive rather than a directly attached physical disk. Data moves between the EC2 instance and the EBS volume over the network, introducing a slight degree of network latency.

    • Detachment and Reattachment: An EBS volume can be detached from an existing EC2 instance and rapidly attached to another EC2 instance within the same Availability Zone.

    • Availability Zone Binding: An EBS volume residing in us-east-1a cannot be attached directly to an instance located in us-east-1b.

    • Cross-AZ / Cross-Region Migration Strategy: Moving data on an EBS volume across Availability Zones or AWS Regions requires taking an EBS Snapshot of the volume and restoring that snapshot into a new volume in the targeted AZ or Region.

  • Provisioning and Billing:

    • Provisioned Capacity: EBS volumes require configuring a specific drive capacity (measured in gigabytes, GBs) and performance profile (measured in Input/Output Operations Per Second, IOPS).

    • Billing Basis: Billing is based strictly on total provisioned capacity (GBs and IOPS), regardless of how much space is actively occupied by files.

    • Dynamic Scaling: The storage capacity of an existing EBS volume can be increased dynamically over time without unmounting or re-creating the drive.

  • EBS Volume Binding and Attachment Architecture:

EBS Volume Availability Zone Binding and Instance Attachment Diagram
  • Availability Zone us-east-1a Examples:

    • One EC2 instance attached to a single 10 GB10\text{ GB} EBS volume.

    • A second EC2 instance attached simultaneously to a 100 GB100\text{ GB} EBS volume and a 50 GB50\text{ GB} EBS volume.

  • Availability Zone us-east-1b Examples:

    • One EC2 instance attached to a 50 GB50\text{ GB} EBS volume.

    • An unattached 10 GB10\text{ GB} EBS volume present within the AZ.

Delete on Termination Attribute

  • Purpose and Functionality:

    • The Delete on Termination attribute controls whether an attached EBS volume is automatically deleted or preserved when its parent EC2 instance is terminated.

  • Default System Behaviors:

    • Root EBS Volume: The Delete on Termination attribute is enabled by default (the root volume containing the operating system is automatically deleted upon instance termination).

    • Additional Attached EBS Volumes: The Delete on Termination attribute is disabled by default (any non-root secondary EBS volumes persist after the instance terminates).

  • Configuration and Management:

    • Attribute settings can be altered prior to or after instance launch through the AWS Management Console or AWS Command Line Interface (CLI).

    • Primary Use Case: Preserving critical OS configurations or data logs by preventing the automatic deletion of the root volume when an instance terminates.

EBS Delete on Termination Console Configuration

EBS Snapshots and Snapshot Management Features

  • Point-in-Time Backups:

    • An EBS Snapshot creates a point-in-time backup copy of an EBS volume.

    • Although detaching the EBS volume before triggering a snapshot is not mandatory, it is strongly recommended to ensure complete data consistency and prevent uncommitted data loss.

    • Snapshots are portable objects that can be copied across Availability Zones or AWS Regions to recreate new EBS volumes anywhere within an AWS infrastructure.

EBS Snapshot Creation and Cross-AZ Restoration Diagram
  • EBS Snapshot Archive:

    • A specialized storage tier designed for long-term storage of infrequently accessed snapshots.

    • Cost Savings: Archiving a snapshot reduces storage costs by up to 75%75\% compared to standard snapshot storage.

    • Restoration Lead Time: Restoring an archived snapshot back to an active EBS snapshot tier takes between 2424 and 7272 hours.

  • Recycle Bin for EBS Snapshots:

    • Provides a safety mechanism to prevent data loss resulting from accidental snapshot deletion.

    • Retention Rules: Rules can be configured to retain deleted snapshots in a Recycle Bin prior to permanent purging.

    • Retention Window: Custom retention durations can be set anywhere from 1 day1\text{ day} to 1 year1\text{ year}.

Amazon Machine Images (AMIs)

  • Overview and Concept:

    • An Amazon Machine Image (AMI) is a standardized, pre-packaged customization template used to launch EC2 instances.

    • AMIs encapsulate user-defined software configurations, operating systems, applications, dependencies, security settings, and monitoring parameters.

    • Efficiency: Using AMIs significantly reduces instance boot and deployment configuration times because software is pre-installed.

    • Regional Scope: AMIs are created and bound within a specific AWS Region but can be copied across regions for global deployment.

  • Sources for AMIs:

    • Public AMIs: Default base images provided directly by AWS (such as Amazon Linux 2 or Windows Server).

    • Custom / Your Own AMIs: Custom images created, managed, and maintained independently by users for organizational standards.

    • AWS Marketplace AMIs: Pre-built vendor images offered or commercialized by third-party software creators.

  • AMI Creation Process from an EC2 Instance:

    • Step 1: Launch an EC2 instance and apply all necessary software customizations and updates.

    • Step 2: Stop the EC2 instance to ensure storage volume snapshot consistency and data integrity.

    • Step 3: Initiate the AMI creation operation (this automatically triggers creation of underlying EBS snapshots for all attached volumes).

    • Step 4: Launch new identical EC2 instances in any Availability Zone using the newly minted Custom AMI.

Custom AMI Creation and Cross-AZ Instance Deployment Diagram

EC2 Image Builder

  • Purpose and Capabilities:

    • A fully managed AWS service designed to automate the creation, maintenance, validation, testing, and deployment of Virtual Machine (AMI) images and container images.

    • Automates operating system updates, application patching, and compliance verification pipelines.

  • Scheduling and Pricing:

    • Image building pipelines can be configured to execute automatically on a scheduled interval (e.g., weekly) or triggered upon software package updates.

    • Pricing Model: The EC2 Image Builder service itself is completely free; users only pay for the underlying AWS infrastructure resources utilized during image compilation and automated testing.

  • End-to-End Image Pipeline Process:

    • Step 1: EC2 Image Builder provisions a temporary Builder EC2 Instance.

    • Step 2: Defined Build Components are executed on the instance to install, configure, and customize software.

    • Step 3: A New AMI image is compiled from the Builder instance.

    • Step 4: A temporary Test EC2 Instance is launched from the new AMI.

    • Step 5: Automated test suites run to check security standards, system stability, and functional requirements.

    • Step 6: Upon passing tests, the AMI is distributed to targeted target deployment Regions.

EC2 Image Builder Automated Workflow Diagram

EC2 Instance Store

  • Hardware Profile and Performance:

    • EC2 Instance Store provides physically attached local host-server block storage drives for EC2 instances.

    • Standard EBS volumes are network-bound with finite performance limits; Instance Store drives offer direct host hardware attachment for maximum disk I/O performance.

    • Delivers extremely high IOPS and exceptionally high read/write throughput.

  • Ephemeral Data Lifecycle and Risks:

    • Ephemeral Nature: Instance Store volumes are temporary storage devices. If an EC2 instance is stopped or terminated, all data stored on its local Instance Store drive is permanently lost.

    • Data persists across system software reboots, but power-off states cause complete physical disk purge.

    • Hardware Risk: Hardware drive failure results in immediate data loss on that host server.

    • User Responsibility: Data replication, operational backups, and redundancy procedures are entirely the responsibility of the user.

  • Primary Workloads:

    • Ideal for temporary file buffering, in-memory caches, scratch data, temporary processing space, or distributed database clusters that handle data replication at the application layer.

  • Local EC2 Instance Store Performance Metrics:

Instance Size

100% Random Read IOPS

Write IOPS

i3.large *

100,125

35,000

i3.xlarge *

206,250

70,000

i3.2xlarge

412,500

180,000

i3.4xlarge

825,000

360,000

i3.8xlarge

1.65 million

720,000

i3.16xlarge

3.3 million

1.4 million

i3.metal

3.3 million

1.4 million

i3en.large *

42,500

32,500

i3en.xlarge *

85,000

65,000

i3en.2xlarge *

170,000

130,000

i3en.3xlarge

250,000

200,000

i3en.6xlarge

500,000

400,000

i3en.12xlarge

1 million

800,000

i3en.24xlarge

2 million

1.6 million

i3en.metal

2 million

1.6 million

Amazon Elastic File System (EFS)

  • Architecture and Multi-AZ Capabilities:

    • Amazon EFS is a fully managed Network File System (NFS) that can be attached concurrently to hundreds of EC2 compute instances.

    • Supports multi-AZ deployment, allowing Linux EC2 instances across multiple Availability Zones in an AWS Region to read and write to the file system simultaneously.

    • Operating System Compatibility: Compatible exclusively with Linux-based EC2 instances.

    • Scalability and Cost Structure: Highly available, automatically scalable, and billed based strictly on consumed storage (no manual capacity provisioning required). Costs approximately 3×3\times as much as standard gp2 EBS volumes.

Amazon EFS Architecture and Multi-AZ Concurrent Mounting Diagram
  • Architectural Comparison: EBS vs. EFS:

    • EBS Volumes: Bound to a single Availability Zone. Must snapshot and restore across AZs to migrate data. Attached to 1 instance at a time at the CCP level.

    • EFS File Systems: Region-wide shared storage accessible concurrently by instances in Availability Zone 1, Availability Zone 2, etc., via designated EFS Mount Targets.

EBS versus EFS Architectural Comparison
  • EFS Infrequent Access (EFS-IA) and Lifecycle Rules:

    • Storage Class: A cost-optimized storage class engineered for files that are not accessed on a daily basis.

    • Cost Savings: Provides up to 92%92\% cost savings relative to the EFS Standard tier.

    • Lifecycle Policies: Automated rules manage moving files from EFS Standard to EFS-IA after a specified duration without access (e.g., transition files after 60 days60\text{ days} of inactivity).

    • Application Transparency: Data tiering transitions take place transparently without interrupting connected user applications.

EFS Lifecycle Policy and Storage Class Transition Diagram

Amazon FSx Managed File Systems

  • Service Overview:

    • Amazon FSx provides fully managed third-party high-performance file systems natively on AWS.

    • Includes specific engine optimizations: FSx for Windows File Server, FSx for Lustre, FSx for NetApp ONTAP, and FSx for OpenZFS.

  • Amazon FSx for Windows File Server:

    • Fully managed, highly reliable, scalable Windows-native shared file storage.

    • Built directly on Windows File Server technology.

    • Full native support for Server Message Block (SMB) protocol and Windows NTFS file system standards.

    • Native integration with Microsoft Active Directory authentication.

    • Accessible concurrently from AWS cloud instances or on-premises infrastructure across corporate data centers.

Amazon FSx for Windows File Server Architecture Diagram
  • Amazon FSx for Lustre:

    • Fully managed high-performance file system designed specifically for High Performance Computing (HPC) Linux workloads.

    • Etymology: The name Lustre is derived as a portmanteau from "Linux" and "cluster".

    • Key Workloads: Machine Learning, Big Data Analytics, Video Processing, Financial Modeling, and High-Performance Compute simulations.

    • Performance Scale: Scales up to hundreds of GB/s throughput, millions of IOPS, and sub-millisecond latencies.

    • Amazon S3 Integration: Seamlessly integrates with Amazon S3 buckets, allowing persistent data to be loaded into FSx for Lustre, processed, and written back to S3.

Amazon FSx for Lustre HPC and S3 Integration Architecture Diagram

Shared Responsibility Model for EC2 Storage

  • AWS Responsibility (Security and Reliability OF the Cloud):

    • Maintaining physical storage hardware infrastructure.

    • Managing physical hardware data replication for EBS volumes and EFS file systems.

    • Replacing faulty physical host server storage hardware.

    • Ensuring strict security access controls preventing unauthorized employee access to customer physical storage media.

  • Customer Responsibility (Security and Reliability IN the Cloud):

    • Setting up system backup routines, snapshots, and lifecycle policies.

    • Configuring data encryption at rest and data in transit.

    • Managing and maintaining all application data stored on drives.

    • Assessing operational risk when electing to utilize ephemeral EC2 Instance Store volumes.

Summary Comparison of EC2 Storage Options

  • EBS Volumes: Network drives bound to a single AZ; attached to 1 EC2 instance at a time (CCP level); backed up and migrated via EBS Snapshots.

  • AMI (Amazon Machine Image): Standardized deployment templates containing customized OS and software for launching EC2 instances fast.

  • EC2 Image Builder: Fully automated service to build, test, validate, and distribute custom AMIs across regions on a schedule.

  • EC2 Instance Store: Physically attached host-server storage disks offering ultra-high throughput and IOPS; data is ephemeral and lost if the instance stops or terminates.

  • Amazon EFS: Shared network file system (NFS) for Linux instances; accessible concurrently across multiple AZs within an AWS Region.

  • EFS-IA: Low-cost storage class for EFS controlled by Lifecycle Policies to move inactive files automatically.

  • FSx for Windows: Fully managed native Windows file server utilizing SMB protocol, NTFS, and Active Directory integration.

  • FSx for Lustre: Massively scalable, parallel file system built for high-performance compute Linux clusters and direct Amazon S3 integration.