1.3 Virtual Machines
Virtual Machines
VMs are a fundamental infrastructure component in GCP, provided by Compute Engine. In Compute Engine, a VM is a networked service that simulates the features of a computer.
Similar to hardware computers but not identical.
Composed of:
Virtual CPU
Memory
Disk Storage
IP Address
Compute Engine offers flexibility, including options not available in physical hardware.
Micro VMs: Share a CPU with other VMs.
Burst capability: Virtual CPU runs above rated capacity for a short period.
Main VM options: CPUs, memory, disks, and networking.
Agenda
Compute Engine Lab: Creating Virtual Machines
Compute Options (vCPU and Memory)
Images
Disk Options
Common Compute Engine Actions Lab: Working with Virtual Machines
Google Cloud Compute Options
Google Cloud offers a spectrum of compute and processing options.
Compute Engine: Offers maximum flexibility.
Supports any language.
Infrastructure as a Service (IaaS) model.
Provides a VM and an operating system.
Allows management and autoscaling configuration.
Autoscaling: Configuring rules for adding more VMs in specific situations (covered later).
Primary use case: Generic workloads, especially enterprise applications designed for server infrastructure.
Other services like Google Kubernetes Engine (GKE) may not be as easily transferable as on-premises solutions.
Compute Engine: Infrastructure as a Service (IaaS)
Compute Engine is based on physical servers within the Google Cloud environment.
Offers predefined and custom machine types.
Allows choosing memory and CPU.
Storage options:
Zonal or regional persistent disk (HDD or SSD)
Local SSD
Cloud Storage
Networking configuration.
Supports Linux and Windows machines.
Compute Engine Features
Preemptible and Spot VMs:
Up to 91% discount
No SLA
Availability policies:
Live migrate
Auto restart
Per-second billing
Sustained use discounts
Committed use discounts
Global load balancing:
Multiple regions for availability
Machine rightsizing
Recommendation engine for optimum machine size
Cloud Monitoring statistics
New recommendation 24 hrs after VM create or resize
OS patch management:
Create patch approvals
Set up flexible scheduling
Apply advanced patch configuration settings
Instance metadata
Startup and shutdown scripts
Hardware Limitations and TPUs
Hardware manufacturers face limitations in scaling CPUs and GPUs to meet the demands of machine learning (ML).
CPU: Central Processing Unit
GPU: Graphics Processing Unit
TPU (Tensor Processing Unit): Google's custom-developed application-specific integrated circuits (ASICs) for accelerating machine learning workloads (introduced in 2016).
Domain-specific hardware.
Tailored architecture for computation needs (e.g., matrix multiplication in machine learning).
Faster and more energy-efficient than GPUs and CPUs for AI and ML applications.
Integrated across Google products.
Recommended for models that train for long durations and large models with large effective batch sizes.
Compute Options (vCPU and Memory)
vCPU: Virtual CPU
Compute Engine provides several machine types.
Custom machine configuration is possible.
Network throughput scales with CPU.
2 Gbps per vCPU (small exceptions)
Theoretical max of 200 Gbps with 176 vCPUs (C3 machine series).
vCPU is implemented as a single hardware hyper-thread.
For an up-to-date list of all the available CPU platforms, refer to the CPU platforms documentation https://cloud.google.com/compute/docs/cpu-platforms.
Disks
Disk options include Standard, SSD, or Local SSD.
Standard (HDD): Higher capacity for the dollar.
SSD: Higher Input/Output Operations Per Second IOPS per dollar.
Local SSDs: Higher throughput and lower latency, attached to the physical hardware.
Data persists only until instance stop or delete.
Used as a swap disk or for additional capacity.
Standard and non-local SSD disks can be sized up to 257 TB for each instance.
Performance scales with space allocated.
Networking
Robust networking features demonstrated in previous modules.
Auto, custom networks.
Inbound/outbound firewall rules:
IP based
Instance/group tags
Regional HTTPS load balancing.
Network load balancing:
Does not require pre-warming
Global and multi-regional subnetworks
Load balancer is essentially a set of traffic engineering rules.
VM Access
Windows: RDP (Remote Desktop Protocol)
RDP clients
Powershell terminal
Requires setting the Windows password
Requires firewall rule to allow tcp:3389
Linux: SSH
SSH from Google Cloud console or Cloud Shell via Google Cloud SDK
SSH from computer or third-party client and generate key pair
Requires firewall rule to allow tcp:22
Instance creator has full root privileges.
Linux: Creator can grant SSH capability to other users.
Windows: Creator can generate a username and password for RDP access.
VM Lifecycle
VM lifecycle includes provisioning, staging, running, repairing, stopping, terminated, and suspending states.
Provisioning: Resources (CPU, memory, disks) are reserved.
Staging: Resources acquired, instance prepared for launch.
IP addresses added, system image and system booted up.
Running: Startup scripts executed, SSH or RDP access enabled.
Live migration: VM moved to another host without reboot.
Stopping: Shutdown scripts executed, ends in terminated state.
Terminated: Instance can be restarted or deleted.
Reset: Wipes memory contents and resets the VM to its initial state.
Repairing: VM encounters an internal error or the underlying machine is unavailable due to maintenance (VM is free during repair).
Suspending: VM enters into the suspending state, before being suspended. The VM can then be resumed or deleted.
Changing VM State from Running
Methods to change VM state include Google Cloud console, gcloud command, and OS commands.
Rebooting, stopping, or deleting an instance takes about 90 seconds.
Preemptible VMs receive an ACPI G3 Mechanical Off signal after 30 seconds.
Availability Policy: Automatic changes
Automatic restart: Automatic VM restart due to crash or maintenance event.
On host maintenance:
Determines whether host is live-migrated or terminated during maintenance.
Live migration is the default.
OS Patch Management
Manage OSes easily through Google Cloud.
Keep infrastructures up-to-date.
Reduce the risk of security vulnerabilities.
Patch compliance reporting
Patch deployment
Patch management is an essential part of managing an infrastructure
Tasks that can be performed with patch management:
Create patch approvals
Set up flexible scheduling
Apply advanced patch configuration settings
Manage these patch jobs or updates from a centralized location.
Charges for Terminated VMs
No charge for memory and CPU resources.
Charges apply for attached disks and reserved IP addresses.
Actions Supported on a Terminated VM
Change the machine type
Migrate the VM instance to another network
Add or remove attached disks; change auto-delete settings
Modify instance tags
Modify custom VM or project-wide metadata
Remove or set a new static IP
Modify VM availability policy
Cannot change the image of a stopped VM
VM availability policies can be changed while the VM is running
Creating Virtual Machines Review for Lab
Explore VM instance options by creating several standard VMs and a custom VM.
Connect to those VMs using both SSH for Linux machines and RDP for Windows machines.
Start with smaller VMs when prototyping to keep the costs down.
Trade up to larger VMs based on capacity for production.
Consider using custom VMs when application's requirements fit between the features of the standard types.
Compute Options (vCPU and Memory) Details
Creating a VM:
console.google.com
command line including Cloudshell
REST API
Many VM options:
Project
Region
Zone
Subnetwork
Machine type
Disk options
Image
IP options
Three options for creating a VM: Cloud Console, Cloud Shell command line, or RESTful API.
Using the Cloud Console first helps avoid typos and provides dropdown lists of options when using the command line or RESTful API.
Machine Type Structure
Machine Family
Machine Series
Machine Type
Machine type selection involves choosing a machine family, series, and predefined type.
Machine family: Curated sets of processor and hardware configurations.
Custom machine types: Specify the number of vCPUs and the amount of memory.
Compute Engine Machine Families
Four machine families:
General-purpose
Compute-optimized
Memory-optimized
Accelerator-optimized
General-Purpose Machine Family (page 26 look at table)
Series: E2 (Cost-optimized), N2, N2D, N1 (Balanced), Tau T2D, Tau T2A (Scale-out optimized)
Workload: Day-to-day computing at a lower cost
Applications: Web serving, App serving, Back office apps, Small-medium databases, Microservices, Virtual desktops, Development environments
Best price-performance with flexible vCPU to memory ratios.
E2: Lowest price on Compute Engine with committed-use discounts; 2-32 vCPUs; 0.5 GB to 8 GB memory per vCPU.
Shared-core E2: 0.25 to 1 vCPUs with 0.5 GB to 8 GB of memory.
N2 and N2D: Flexible VM types with balanced price and performance; enterprise applications, medium-to-large databases.
N2: Intel scalable processor with up to 128 vCPUs, 0.5 to 8 GB memory per vCPU.
N2D: AMD-based, up to 224 vCPUs per node.
Tau T2D and Tau T2A VMs are optimized for cost-effective performance of demanding scale-out workloads.
General purpose machines: https://cloud.google.com/compute/docs/general-purpose-machines
Compute-Optimized Machine Family
Series: C2, C2D Ultra high performance for compute-intensive workloads H3
Applications: Compute-bound workloads, High-performance web serving, Gaming (AAA game servers), Ad serving, High-performance computing (HPC), Media transcoding, AI/ML
Highest performance per core on Compute Engine.
C2 VMs: High-frequency Intel-scalable processors, up to 3.8 GHz, 4 to 60 vCPUs, up to 240 GB memory.
C2D: Largest VM sizes, best for HPC; 2 to 112 vCPUs, 4 GB memory per vCPU.
H3 series offer 88 cores and 352 GB of DDR5 memory and are available on the Intel Sapphire Rapids CPU platform and Google's custom Intel Infrastructure Processing Unit (IPU).
Compute-optimized machines: https://cloud.google.com/compute/docs/compute-optimized-machines
Memory-Optimized Machine Family
Series: M1, M2, M3
Workload: Ultra high-memory workloads
Applications: In-memory databases, in-memory analytics, business warehousing, genomics analysis, SQL analysis services.
Provides the most compute and memory resources.
M1: Up to 4 TB of memory.
M2: Up to 12 TB of memory.
M3 VMs offer up to 128 vCPUs, with up to 30.5 GB of memory per vCPU, and are available on the Intel Ice Lake CPU platform
Memory-optimized machines: https://cloud.google.com/compute/docs/memory-optimized-machines
Accelerator-Optimized Machine Family
Series: A2, G2
Workload: Optimized for high-performance computing workloads
Applications: CUDA-enabled ML training and inference, HPC, Massive parallelized computation, video transcoding, remote visualization workstation.
Ideal for massively parallelized CUDA compute workloads.
A2: 12 to 96 vCPUs, up to 1360 GB memory, up to 16 NVIDIA Ampere A100 GPUs with 40 GB GPU memory.
G2 VMs offer 4 to 96 vCPUs, up to 432 GB of memory, and are available on the Intel Cascade Lake CPU platform.
Accelerator-optimized machine family: https://cloud.google.com/compute/docs/accelerator-optimized-machines
Custom Machine Types
When to select custom:
Requirements fit between the predefined types
Need more memory or more CPU
Customize the amount of memory and vCPU for your machine:
Either 1 vCPU or even number of vCPU
Up to 8 GB per vCPU
Total memory must be multiple of 256 MB
Ideal when predefined machine types don't fit workload needs.
Slightly more expensive than equivalent predefined types.
Limitations:
Only 1 vCPU or an even number of vCPUs.
Memory between 1 GB and 8 GB per vCPU.
Total memory must be a multiple of 256 MB.
Extended memory: Get more memory per vCPU beyond the 8 GB limit (at an additional cost).
Choosing Region and Zone
Consider geographical location to run resources.
Each zone supports a combination of Ivy Bridge, Sandy Bridge, Haswell, Broadwell, and Skylake platforms. When you create an instance in the zone, your instance will use the default processor supported in that zone. For example, if you create an instance in the us-central1-a zone, your instance will use a Sandy Bridge processor.
Available regions and zones: https://cloud.google.com/compute/docs/regions-zones/#available
Pricing
Per-second billing (minimum 1 minute) for vCPUs, GPUs, and memory.
Resource-based pricing: Each vCPU and GB of memory billed separately.
Discounts:
Sustained use
Committed use
Preemptible and Spot VM instances
Recommendation Engine: Notifies of underutilized instances (24 hours after instance creation).
Free usage limits.
Sustained Use Discounts
Automatic discounts for running Compute Engine resources for a significant portion of the billing month.
Increase with usage (up to 30% net discount for instances that run the entire month).
Sustained use discounts for up to 30%
General-purpose N2 and N2D predefined and custom machine types, and compute-optimized machine types Sustained use discounts for up to 20%
Compute Engine calculates sustained use discounts based on vCPU and memory usage across each region and separately for each of the following categories:
Predefined machine types
Custom machine type
Sustained use discounts: https://cloud.google.com/compute/docs/sustained-use-discounts
Preemptible VMs
As I mentioned earlier, a preemptible VM is an instance that you can create and run at a much lower cost than normal instances.
Lower price for interruptible service (up to 91%).
VM might be terminated at any time.
No charge if terminated in the first minute.
24 hours max.
30-second terminate warning (not guaranteed).
Time for a shutdown script.
No live migrate; no auto restart.
Batch processing is a major use case.
https://cloud.google.com/compute/docs/instances/preemptible#whatisa_preemptible instance
Spot VMs
Spot VMs are the latest version of preemptible VMs
Spot VMs and preemptible VMs share the same pricing model
No minimum or maximum runtime
Spot VMs are finite Compute Engine resources, so they might not always be available
No live migrate; no auto restart
Resources for Spot VMs come out of excess and backup Google Cloud capacity. Capacity for Spot VMs is often easier to get for smaller machine types, meaning machine types with less resources like vCPUs and memory.
Best practice use cases help you get the most of using Spot VMs
For example, preemptible VMs can only run for up to 24 hours at a time, but Spot VMs do not have a maximum runtime.
For more information on best practices, see https://cloud.google.com/compute/docs/instances/create-use-spot#best-practices
For more information on Spot VMs, see https://cloud.google.com/compute/docs/instances/spot
Sole-Tenant Nodes
Physically isolate workloads for compliance requirements.
Sole-tenant node: Physical Compute Engine server dedicated to hosting VM instances for a specific project.
Existing OS licenses can be brought to Compute Engine using sole-tenant nodes.
https://cloud.google.com/compute/docs/nodes/create-nodes
Shielded VMs
Offer verifiable integrity of VM instances.
Secure Boot
Virtual trusted platform module (vTPM)
Integrity monitoring
Requires shielded image.
Confidential VMs
Encrypts data while it’s being processed.
Easy to use with no changes to code or performance compromise.
N2D Compute Engine VM running on second generation AMD Epyc processors.
Provides high memory capacity, high throughput, and supports parallel and compute heavy workloads.
You can select Confidential VM service when creating a new VM.
Images
Images include the boot loader, the operating system, the file system structure, any pre-configured software, and any other customizations.
Image Types
Public base images offered by Google, third-party vendors, and the community.
Premium images (p)
CentOS, CoreOS, Debian, RHEL(p), SUSE(p), Ubuntu, openSUSE, and FreeBSD
Windows Server 2019(p), 2016(p), 2012-r2(p)
SQL Server pre-installed on Windows(p)
Custom images: Create new image from VM.
Import from on-prem, workstation, or another cloud.
Management features: image sharing, image family, deprecation
Premium images have per-second charges (1-minute minimum), SQL Server images charged per minute (10-minute minimum).
Share custom images within a project or among other projects.
Machine Images (page 52 for table)
A machine image is a Compute Engine resource that stores all the configuration, metadata, permissions, and data from one or more disks required to create a virtual machine (VM) instance. You can use a machine image in many system maintenance scenarios, such as creation, backup and recovery, and instance cloning.
Machine images are the most ideal resources for disk backups as well as instance cloning and replication.
Disk Options
Discussion of various disk options and features in Compute Engine.
Boot Disk
Every single VM comes with a single root persistent disk, because you're choosing a base image to have that loaded on.
Image is loaded onto root disk during first boot:
Bootable: you can attach to a VM and boot from it
Durable: can survive VM terminate.
Some OS images are customized for Compute Engine.
Survives VM deletion if "Delete boot disk when instance is deleted" is disabled.
Persistent Disks
Network storage appearing as a block device
Attached to a VM through the network interface
Durable storage: can survive VM terminate
Bootable: you can attach to a VM and boot from it
Snapshots: incremental backups
Performance: Scales with size
HDD (magnetic) or SSD (faster, solid-state) options
Disk resizing: even running and attached!
Can be attached in read-only mode to multiple VMs
Zonal or Regional
pd-standard
pd-ssd
pd-balanced
pd-extreme (zonal only)
Encryption keys:
Google-managed
Customer-managed
Customer-supplied
Standard persistent disks (pd-standard). These types of disks are backed by standard hard disk drives (HDD) and are suitable for large data processing workloads that primarily use sequential I/Os.
Performance SSD persistent disks (pd-ssd). These types of disks are backed by solid-state drives (SSD) and are suitable for enterprise applications and high-performance databases that require lower latency and more IOPS than standard persistent disks provide.
Balanced persistent disks (pd-balanced). These types of disks are also backed by solid-state drives (SSD). They are an alternative to SSD persistent disks that balance performance and cost. These disks have the same maximum IOPS as SSD persistent disks and lower IOPS per GB. For most VM shapes, except very large ones, this disk type offers performance levels suitable for most general-purpose applications at a price point between that of standard and performance persistent disks.
Extreme persistent disks (pd-extreme) are zonal persistent disks and are also backed by solid-state drives (SSD). Extreme persistent disks are designed for high-end database workloads, providing consistently high performance for both random access workloads and bulk throughput. Unlike other disk types, you can provision your desired IOPS
Local SSDs
Physically attached to a VM
More IOPS, lower latency, and higher throughput than persistent disk
375-GB disk up to 24, total of 9 TB
Data survives a reset, but not a VM stop or terminate
VM-specific: cannot be reattached to a different VM
RAM Disk
tmpfs
Faster than local disk, slower than memory
Use when your application expects a file system structure and cannot directly store its data in memory
Fast scratch disk, or fast cache
Very volatile; erase on stop or restart
May need a larger machine type if RAM was sized for the application
Consider using a persistent disk to back up RAM disk data
Summary of Disk Options (p59 for table)
Comparison of Persistent disk HDD, Persistent Disk SSD, Local SSD disk, and RAM disk.
I recommend choosing a persistent HDD disk when you don't need performance but just need capacity. If you have high performance needs, start looking at the SSD options. The persistent disks offer data redundancy because the data on each persistent disk is distributed across several physical disks.
Maximum Persistent Disks
Disk number limit depends on machine type.
Shared-core: 16
Standard: 128
High-memory: 128
High-CPU: 128
Memory-optimized: 128
Compute-optimized: 128
Throughput is limited by the number of cores and shares the same bandwidth with Disk IO.
Persistent Disk Management Differences (p61) for table
Differences between Cloud Persistent Disk and Computer Hardware Disk
Cloud Persistent Disk has features such as Single file system is best, Resize (grow) disks, Resize file system. Built-in snapshot service, Automatic encryption
Common Compute Engine Actions
Overview of common actions one can perform with Compute Engine instances and related resources.
Metadata and Scripts
Time startup-script-url=URL shutdown-script-url=URL
Boot Metadata
Run Metadata
Maintenance Metadata
Shutdown Metadata
Every VM instance stores its metadata on a metadata server. The metadata server is useful in combination with startup and shutdown scripts. For example, you can write a startup script that gets the metadata key/value pair for an instance's external IP address and use that IP address in your script to set up a database
Moving an Instance to a New Zone
Discussed automated (within region) and manual (between regions) processes for moving VM instances.
Zone 1 Zone 2
gcloud compute instances move
Moving an Instance to a New Zone Automated Process
Automated process (moving within region):
gcloud compute instances move
Update references to VM; not automatic
Moving an Instance to a New Zone Manual Process
Manual process (moving between regions):
Snapshot all persistent disks on the source VM.
Create new persistent disks in destination zone restored from snapshots.
Create new VM in the destination zone and attach new persistent disks.
Assign static IP to new VM.
Update references to VM.
Delete the snapshots, original disks, and original VM.
Disk Snapshots (Use Cases)
Snapshots: Backup critical data
Cloud Storage
Snapshot Service
Compute Engine root data
Migrate data between zones
Discuss the use case where snapshots are used to easily migrate data between zones to reduce latency.
Zone 1 Zone 2
Compute Engine Compute Engine
Snapshot Service
Transfer to SSD to improve performance
Cloud Storage
Compute Engine root PD HDD to PD SSD
Snapshot service
Persistent Disk Snapshots
Snapshot is not available for local SSD.
Creates an incremental backup to Cloud Storage.
Not visible in your buckets; managed by the snapshot service.
Create scheduled snapshots.
Regularly and automatically back up your zonal and regional persistent disks.
Snapshots can be restored to a new persistent disk.
New disk can be in another region or zone in the same project.
Resize Persistent Disk
Grow disks, but never shrink them.
Working with Virtual Machines Review of Lab
Demonstration of customized VM creation, software installation (Minecraft), high-speed SSD attachment, static external IP reservation, backup system setup, and automation.
Lab Review Main Points
The VM was customized by preparing and attaching a high-speed SSD, and reserved a static external IP address so that the address would remain consistent
You then Automated backups using cron script
Set up maintenance scripts using metadata for graceful startup and shutdown of the server
Review: Virtual Machines
Recap of the topics covered in Compute Engine, including compute, image, and disk options, along with common actions.
Final Recommendation
Recommend enrolling in the “Essential Cloud Infrastructure: Core Services” course of the “Architecting with Google Compute Engine” series. Topics covered will be Cloud IAM, different data storage services in GCP, resource management, and lastly, resource monitoring.