1.3 Virtual Machines

Virtual Machines

VMs are a fundamental infrastructure component in GCP, provided by Compute Engine. In Compute Engine, a VM is a networked service that simulates the features of a computer.

  • Similar to hardware computers but not identical.

  • Composed of:

    • Virtual CPU

    • Memory

    • Disk Storage

    • IP Address

  • Compute Engine offers flexibility, including options not available in physical hardware.

    • Micro VMs: Share a CPU with other VMs.

    • Burst capability: Virtual CPU runs above rated capacity for a short period.

  • Main VM options: CPUs, memory, disks, and networking.

Agenda

  1. Compute Engine Lab: Creating Virtual Machines

  2. Compute Options (vCPU and Memory)

  3. Images

  4. Disk Options

  5. Common Compute Engine Actions Lab: Working with Virtual Machines

Google Cloud Compute Options

Google Cloud offers a spectrum of compute and processing options.

  • Compute Engine: Offers maximum flexibility.

    • Supports any language.

    • Infrastructure as a Service (IaaS) model.

    • Provides a VM and an operating system.

    • Allows management and autoscaling configuration.

  • Autoscaling: Configuring rules for adding more VMs in specific situations (covered later).

  • Primary use case: Generic workloads, especially enterprise applications designed for server infrastructure.

  • Other services like Google Kubernetes Engine (GKE) may not be as easily transferable as on-premises solutions.

Compute Engine: Infrastructure as a Service (IaaS)

Compute Engine is based on physical servers within the Google Cloud environment.

  • Offers predefined and custom machine types.

    • Allows choosing memory and CPU.

  • Storage options:

    • Zonal or regional persistent disk (HDD or SSD)

    • Local SSD

    • Cloud Storage

  • Networking configuration.

  • Supports Linux and Windows machines.

Compute Engine Features

  • Preemptible and Spot VMs:

    • Up to 91% discount

    • No SLA

  • Availability policies:

    • Live migrate

    • Auto restart

  • Per-second billing

  • Sustained use discounts

  • Committed use discounts

  • Global load balancing:

    • Multiple regions for availability

  • Machine rightsizing

    • Recommendation engine for optimum machine size

    • Cloud Monitoring statistics

    • New recommendation 24 hrs after VM create or resize

  • OS patch management:

    • Create patch approvals

    • Set up flexible scheduling

    • Apply advanced patch configuration settings

  • Instance metadata

  • Startup and shutdown scripts

Hardware Limitations and TPUs

Hardware manufacturers face limitations in scaling CPUs and GPUs to meet the demands of machine learning (ML).

  • CPU: Central Processing Unit

  • GPU: Graphics Processing Unit

  • TPU (Tensor Processing Unit): Google's custom-developed application-specific integrated circuits (ASICs) for accelerating machine learning workloads (introduced in 2016).

    • Domain-specific hardware.

    • Tailored architecture for computation needs (e.g., matrix multiplication in machine learning).

    • Faster and more energy-efficient than GPUs and CPUs for AI and ML applications.

    • Integrated across Google products.

    • Recommended for models that train for long durations and large models with large effective batch sizes.

Compute Options (vCPU and Memory)

  • vCPU: Virtual CPU

  • Compute Engine provides several machine types.

  • Custom machine configuration is possible.

  • Network throughput scales with CPU.

    • 2 Gbps per vCPU (small exceptions)

    • Theoretical max of 200 Gbps with 176 vCPUs (C3 machine series).

  • vCPU is implemented as a single hardware hyper-thread.

  • For an up-to-date list of all the available CPU platforms, refer to the CPU platforms documentation https://cloud.google.com/compute/docs/cpu-platforms.

Disks

Disk options include Standard, SSD, or Local SSD.

  • Standard (HDD): Higher capacity for the dollar.

  • SSD: Higher Input/Output Operations Per Second IOPS per dollar.

  • Local SSDs: Higher throughput and lower latency, attached to the physical hardware.

    • Data persists only until instance stop or delete.

    • Used as a swap disk or for additional capacity.

  • Standard and non-local SSD disks can be sized up to 257 TB for each instance.

  • Performance scales with space allocated.

Networking

  • Robust networking features demonstrated in previous modules.

  • Auto, custom networks.

  • Inbound/outbound firewall rules:

    • IP based

    • Instance/group tags

  • Regional HTTPS load balancing.

  • Network load balancing:

    • Does not require pre-warming

  • Global and multi-regional subnetworks

  • Load balancer is essentially a set of traffic engineering rules.

VM Access

  • Windows: RDP (Remote Desktop Protocol)

    • RDP clients

    • Powershell terminal

    • Requires setting the Windows password

    • Requires firewall rule to allow tcp:3389

  • Linux: SSH

    • SSH from Google Cloud console or Cloud Shell via Google Cloud SDK

    • SSH from computer or third-party client and generate key pair

    • Requires firewall rule to allow tcp:22

  • Instance creator has full root privileges.

  • Linux: Creator can grant SSH capability to other users.

  • Windows: Creator can generate a username and password for RDP access.

VM Lifecycle

VM lifecycle includes provisioning, staging, running, repairing, stopping, terminated, and suspending states.

  • Provisioning: Resources (CPU, memory, disks) are reserved.

  • Staging: Resources acquired, instance prepared for launch.

    • IP addresses added, system image and system booted up.

  • Running: Startup scripts executed, SSH or RDP access enabled.

    • Live migration: VM moved to another host without reboot.

  • Stopping: Shutdown scripts executed, ends in terminated state.

  • Terminated: Instance can be restarted or deleted.

  • Reset: Wipes memory contents and resets the VM to its initial state.

  • Repairing: VM encounters an internal error or the underlying machine is unavailable due to maintenance (VM is free during repair).

  • Suspending: VM enters into the suspending state, before being suspended. The VM can then be resumed or deleted.

Changing VM State from Running

Methods to change VM state include Google Cloud console, gcloud command, and OS commands.

  • Rebooting, stopping, or deleting an instance takes about 90 seconds.

  • Preemptible VMs receive an ACPI G3 Mechanical Off signal after 30 seconds.

Availability Policy: Automatic changes

  • Automatic restart: Automatic VM restart due to crash or maintenance event.

  • On host maintenance:

    • Determines whether host is live-migrated or terminated during maintenance.

    • Live migration is the default.

OS Patch Management

  • Manage OSes easily through Google Cloud.

  • Keep infrastructures up-to-date.

  • Reduce the risk of security vulnerabilities.

  • Patch compliance reporting

  • Patch deployment

  • Patch management is an essential part of managing an infrastructure

  • Tasks that can be performed with patch management:

    • Create patch approvals

    • Set up flexible scheduling

    • Apply advanced patch configuration settings

    • Manage these patch jobs or updates from a centralized location.

Charges for Terminated VMs

  • No charge for memory and CPU resources.

  • Charges apply for attached disks and reserved IP addresses.

Actions Supported on a Terminated VM

  • Change the machine type

  • Migrate the VM instance to another network

  • Add or remove attached disks; change auto-delete settings

  • Modify instance tags

  • Modify custom VM or project-wide metadata

  • Remove or set a new static IP

  • Modify VM availability policy

  • Cannot change the image of a stopped VM

  • VM availability policies can be changed while the VM is running

Creating Virtual Machines Review for Lab

  • Explore VM instance options by creating several standard VMs and a custom VM.

  • Connect to those VMs using both SSH for Linux machines and RDP for Windows machines.

  • Start with smaller VMs when prototyping to keep the costs down.

  • Trade up to larger VMs based on capacity for production.

  • Consider using custom VMs when application's requirements fit between the features of the standard types.

Compute Options (vCPU and Memory) Details

  • Creating a VM:

    • console.google.com

    • command line including Cloudshell

      • gcloudcomputeinstancescreate[instance−name]gcloud compute instances create [instance-name]

    • REST API

  • Many VM options:

    • Project

    • Region

    • Zone

    • Subnetwork

    • Machine type

    • Disk options

    • Image

    • IP options

  • Three options for creating a VM: Cloud Console, Cloud Shell command line, or RESTful API.

  • Using the Cloud Console first helps avoid typos and provides dropdown lists of options when using the command line or RESTful API.

Machine Type Structure

  • Machine Family

  • Machine Series

  • Machine Type

  • Machine type selection involves choosing a machine family, series, and predefined type.

  • Machine family: Curated sets of processor and hardware configurations.

  • Custom machine types: Specify the number of vCPUs and the amount of memory.

Compute Engine Machine Families

Four machine families:

  1. General-purpose

  2. Compute-optimized

  3. Memory-optimized

  4. Accelerator-optimized

General-Purpose Machine Family (page 26 look at table)

  • Series: E2 (Cost-optimized), N2, N2D, N1 (Balanced), Tau T2D, Tau T2A (Scale-out optimized)

  • Workload: Day-to-day computing at a lower cost

  • Applications: Web serving, App serving, Back office apps, Small-medium databases, Microservices, Virtual desktops, Development environments

  • Best price-performance with flexible vCPU to memory ratios.

  • E2: Lowest price on Compute Engine with committed-use discounts; 2-32 vCPUs; 0.5 GB to 8 GB memory per vCPU.

  • Shared-core E2: 0.25 to 1 vCPUs with 0.5 GB to 8 GB of memory.

  • N2 and N2D: Flexible VM types with balanced price and performance; enterprise applications, medium-to-large databases.

  • N2: Intel scalable processor with up to 128 vCPUs, 0.5 to 8 GB memory per vCPU.

  • N2D: AMD-based, up to 224 vCPUs per node.

  • Tau T2D and Tau T2A VMs are optimized for cost-effective performance of demanding scale-out workloads.

  • General purpose machines: https://cloud.google.com/compute/docs/general-purpose-machines

Compute-Optimized Machine Family

  • Series: C2, C2D Ultra high performance for compute-intensive workloads H3

  • Applications: Compute-bound workloads, High-performance web serving, Gaming (AAA game servers), Ad serving, High-performance computing (HPC), Media transcoding, AI/ML

  • Highest performance per core on Compute Engine.

  • C2 VMs: High-frequency Intel-scalable processors, up to 3.8 GHz, 4 to 60 vCPUs, up to 240 GB memory.

  • C2D: Largest VM sizes, best for HPC; 2 to 112 vCPUs, 4 GB memory per vCPU.

  • H3 series offer 88 cores and 352 GB of DDR5 memory and are available on the Intel Sapphire Rapids CPU platform and Google's custom Intel Infrastructure Processing Unit (IPU).

  • Compute-optimized machines: https://cloud.google.com/compute/docs/compute-optimized-machines

Memory-Optimized Machine Family

  • Series: M1, M2, M3

  • Workload: Ultra high-memory workloads

  • Applications: In-memory databases, in-memory analytics, business warehousing, genomics analysis, SQL analysis services.

  • Provides the most compute and memory resources.

  • M1: Up to 4 TB of memory.

  • M2: Up to 12 TB of memory.

  • M3 VMs offer up to 128 vCPUs, with up to 30.5 GB of memory per vCPU, and are available on the Intel Ice Lake CPU platform

  • Memory-optimized machines: https://cloud.google.com/compute/docs/memory-optimized-machines

Accelerator-Optimized Machine Family

  • Series: A2, G2

  • Workload: Optimized for high-performance computing workloads

  • Applications: CUDA-enabled ML training and inference, HPC, Massive parallelized computation, video transcoding, remote visualization workstation.

  • Ideal for massively parallelized CUDA compute workloads.

  • A2: 12 to 96 vCPUs, up to 1360 GB memory, up to 16 NVIDIA Ampere A100 GPUs with 40 GB GPU memory.

  • G2 VMs offer 4 to 96 vCPUs, up to 432 GB of memory, and are available on the Intel Cascade Lake CPU platform.

  • Accelerator-optimized machine family: https://cloud.google.com/compute/docs/accelerator-optimized-machines

Custom Machine Types

  • When to select custom:

    • Requirements fit between the predefined types

    • Need more memory or more CPU

  • Customize the amount of memory and vCPU for your machine:

    • Either 1 vCPU or even number of vCPU

    • Up to 8 GB per vCPU

    • Total memory must be multiple of 256 MB

  • Ideal when predefined machine types don't fit workload needs.

  • Slightly more expensive than equivalent predefined types.

  • Limitations:

    • Only 1 vCPU or an even number of vCPUs.

    • Memory between 1 GB and 8 GB per vCPU.

    • Total memory must be a multiple of 256 MB.

  • Extended memory: Get more memory per vCPU beyond the 8 GB limit (at an additional cost).

Choosing Region and Zone

  • Consider geographical location to run resources.

  • Each zone supports a combination of Ivy Bridge, Sandy Bridge, Haswell, Broadwell, and Skylake platforms. When you create an instance in the zone, your instance will use the default processor supported in that zone. For example, if you create an instance in the us-central1-a zone, your instance will use a Sandy Bridge processor.

  • Available regions and zones: https://cloud.google.com/compute/docs/regions-zones/#available

Pricing

  • Per-second billing (minimum 1 minute) for vCPUs, GPUs, and memory.

  • Resource-based pricing: Each vCPU and GB of memory billed separately.

  • Discounts:

    • Sustained use

    • Committed use

    • Preemptible and Spot VM instances

  • Recommendation Engine: Notifies of underutilized instances (24 hours after instance creation).

  • Free usage limits.

Sustained Use Discounts

  • Automatic discounts for running Compute Engine resources for a significant portion of the billing month.

  • Increase with usage (up to 30% net discount for instances that run the entire month).

  • Sustained use discounts for up to 30%

  • General-purpose N2 and N2D predefined and custom machine types, and compute-optimized machine types Sustained use discounts for up to 20%

  • Compute Engine calculates sustained use discounts based on vCPU and memory usage across each region and separately for each of the following categories:

    • Predefined machine types

    • Custom machine type

  • Sustained use discounts: https://cloud.google.com/compute/docs/sustained-use-discounts

Preemptible VMs

    As I mentioned earlier, a preemptible VM is an instance that you can create and run at a much lower cost than normal instances.

  • Lower price for interruptible service (up to 91%).

  • VM might be terminated at any time.

    • No charge if terminated in the first minute.

    • 24 hours max.

    • 30-second terminate warning (not guaranteed).

    • Time for a shutdown script.

  • No live migrate; no auto restart.

  • Batch processing is a major use case.

  • https://cloud.google.com/compute/docs/instances/preemptible#whatisa_preemptible instance

Spot VMs

  • Spot VMs are the latest version of preemptible VMs

  • Spot VMs and preemptible VMs share the same pricing model

  • No minimum or maximum runtime

  • Spot VMs are finite Compute Engine resources, so they might not always be available

  • No live migrate; no auto restart

  • Resources for Spot VMs come out of excess and backup Google Cloud capacity. Capacity for Spot VMs is often easier to get for smaller machine types, meaning machine types with less resources like vCPUs and memory.

  • Best practice use cases help you get the most of using Spot VMs

  • For example, preemptible VMs can only run for up to 24 hours at a time, but Spot VMs do not have a maximum runtime.

  • For more information on best practices, see https://cloud.google.com/compute/docs/instances/create-use-spot#best-practices

  • For more information on Spot VMs, see https://cloud.google.com/compute/docs/instances/spot

Sole-Tenant Nodes

  • Physically isolate workloads for compliance requirements.

  • Sole-tenant node: Physical Compute Engine server dedicated to hosting VM instances for a specific project.

  • Existing OS licenses can be brought to Compute Engine using sole-tenant nodes.

  • https://cloud.google.com/compute/docs/nodes/create-nodes

Shielded VMs

  • Offer verifiable integrity of VM instances.

  • Secure Boot

  • Virtual trusted platform module (vTPM)

  • Integrity monitoring

  • Requires shielded image.

Confidential VMs

  • Encrypts data while it’s being processed.

  • Easy to use with no changes to code or performance compromise.

  • N2D Compute Engine VM running on second generation AMD Epyc processors.

  • Provides high memory capacity, high throughput, and supports parallel and compute heavy workloads.

  • You can select Confidential VM service when creating a new VM.

Images

  • Images include the boot loader, the operating system, the file system structure, any pre-configured software, and any other customizations.

Image Types

  • Public base images offered by Google, third-party vendors, and the community.

    • Premium images (p)

      • CentOS, CoreOS, Debian, RHEL(p), SUSE(p), Ubuntu, openSUSE, and FreeBSD

      • Windows Server 2019(p), 2016(p), 2012-r2(p)

      • SQL Server pre-installed on Windows(p)

  • Custom images: Create new image from VM.

    • Import from on-prem, workstation, or another cloud.

    • Management features: image sharing, image family, deprecation

  • Premium images have per-second charges (1-minute minimum), SQL Server images charged per minute (10-minute minimum).

  • Share custom images within a project or among other projects.

Machine Images (page 52 for table)

  • A machine image is a Compute Engine resource that stores all the configuration, metadata, permissions, and data from one or more disks required to create a virtual machine (VM) instance. You can use a machine image in many system maintenance scenarios, such as creation, backup and recovery, and instance cloning.

  • Machine images are the most ideal resources for disk backups as well as instance cloning and replication.

Disk Options

Discussion of various disk options and features in Compute Engine.

Boot Disk

  • Every single VM comes with a single root persistent disk, because you're choosing a base image to have that loaded on.

  • Image is loaded onto root disk during first boot:

    • Bootable: you can attach to a VM and boot from it

    • Durable: can survive VM terminate.

  • Some OS images are customized for Compute Engine.

  • Survives VM deletion if "Delete boot disk when instance is deleted" is disabled.

Persistent Disks

  • Network storage appearing as a block device

  • Attached to a VM through the network interface

  • Durable storage: can survive VM terminate

  • Bootable: you can attach to a VM and boot from it

  • Snapshots: incremental backups

  • Performance: Scales with size

  • HDD (magnetic) or SSD (faster, solid-state) options

  • Disk resizing: even running and attached!

  • Can be attached in read-only mode to multiple VMs

  • Zonal or Regional

    • pd-standard

    • pd-ssd

    • pd-balanced

    • pd-extreme (zonal only)

  • Encryption keys:

    • Google-managed

    • Customer-managed

    • Customer-supplied

  • Standard persistent disks (pd-standard). These types of disks are backed by standard hard disk drives (HDD) and are suitable for large data processing workloads that primarily use sequential I/Os.

  • Performance SSD persistent disks (pd-ssd). These types of disks are backed by solid-state drives (SSD) and are suitable for enterprise applications and high-performance databases that require lower latency and more IOPS than standard persistent disks provide.

  • Balanced persistent disks (pd-balanced). These types of disks are also backed by solid-state drives (SSD). They are an alternative to SSD persistent disks that balance performance and cost. These disks have the same maximum IOPS as SSD persistent disks and lower IOPS per GB. For most VM shapes, except very large ones, this disk type offers performance levels suitable for most general-purpose applications at a price point between that of standard and performance persistent disks.

  • Extreme persistent disks (pd-extreme) are zonal persistent disks and are also backed by solid-state drives (SSD). Extreme persistent disks are designed for high-end database workloads, providing consistently high performance for both random access workloads and bulk throughput. Unlike other disk types, you can provision your desired IOPS

Local SSDs

  • Physically attached to a VM

  • More IOPS, lower latency, and higher throughput than persistent disk

  • 375-GB disk up to 24, total of 9 TB

  • Data survives a reset, but not a VM stop or terminate

  • VM-specific: cannot be reattached to a different VM

RAM Disk

  • tmpfs

  • Faster than local disk, slower than memory

    • Use when your application expects a file system structure and cannot directly store its data in memory

    • Fast scratch disk, or fast cache

  • Very volatile; erase on stop or restart

  • May need a larger machine type if RAM was sized for the application

  • Consider using a persistent disk to back up RAM disk data

Summary of Disk Options (p59 for table)

Comparison of Persistent disk HDD, Persistent Disk SSD, Local SSD disk, and RAM disk.

  • I recommend choosing a persistent HDD disk when you don't need performance but just need capacity. If you have high performance needs, start looking at the SSD options. The persistent disks offer data redundancy because the data on each persistent disk is distributed across several physical disks.

Maximum Persistent Disks

Disk number limit depends on machine type.

  • Shared-core: 16

  • Standard: 128

  • High-memory: 128

  • High-CPU: 128

  • Memory-optimized: 128

  • Compute-optimized: 128

  • Throughput is limited by the number of cores and shares the same bandwidth with Disk IO.

Persistent Disk Management Differences (p61) for table

Differences between Cloud Persistent Disk and Computer Hardware Disk

Cloud Persistent Disk has features such as Single file system is best, Resize (grow) disks, Resize file system. Built-in snapshot service, Automatic encryption

Common Compute Engine Actions

Overview of common actions one can perform with Compute Engine instances and related resources.

Metadata and Scripts

  • Time startup-script-url=URL shutdown-script-url=URL

  • Boot Metadata

  • Run Metadata

  • Maintenance Metadata

  • Shutdown Metadata

  • Every VM instance stores its metadata on a metadata server. The metadata server is useful in combination with startup and shutdown scripts. For example, you can write a startup script that gets the metadata key/value pair for an instance's external IP address and use that IP address in your script to set up a database

Moving an Instance to a New Zone

Discussed automated (within region) and manual (between regions) processes for moving VM instances.

Zone 1 Zone 2
gcloud compute instances move

Moving an Instance to a New Zone Automated Process

  • Automated process (moving within region):

    • gcloud compute instances move

    • Update references to VM; not automatic

Moving an Instance to a New Zone Manual Process

  • Manual process (moving between regions):

    • Snapshot all persistent disks on the source VM.

    • Create new persistent disks in destination zone restored from snapshots.

    • Create new VM in the destination zone and attach new persistent disks.

    • Assign static IP to new VM.

    • Update references to VM.

    • Delete the snapshots, original disks, and original VM.

Disk Snapshots (Use Cases)

Snapshots: Backup critical data

Cloud Storage

Snapshot Service

Compute Engine root data

Migrate data between zones

Discuss the use case where snapshots are used to easily migrate data between zones to reduce latency.

Zone 1 Zone 2

Compute Engine Compute Engine

Snapshot Service

Transfer to SSD to improve performance

Cloud Storage

Compute Engine root PD HDD to PD SSD

Snapshot service

Persistent Disk Snapshots

  • Snapshot is not available for local SSD.

  • Creates an incremental backup to Cloud Storage.

    • Not visible in your buckets; managed by the snapshot service.

  • Create scheduled snapshots.

    • Regularly and automatically back up your zonal and regional persistent disks.

  • Snapshots can be restored to a new persistent disk.

    • New disk can be in another region or zone in the same project.

Resize Persistent Disk

Grow disks, but never shrink them.

Working with Virtual Machines Review of Lab

Demonstration of customized VM creation, software installation (Minecraft), high-speed SSD attachment, static external IP reservation, backup system setup, and automation.

Lab Review Main Points

  • The VM was customized by preparing and attaching a high-speed SSD, and reserved a static external IP address so that the address would remain consistent

  • You then Automated backups using cron script

  • Set up maintenance scripts using metadata for graceful startup and shutdown of the server

Review: Virtual Machines

Recap of the topics covered in Compute Engine, including compute, image, and disk options, along with common actions.

Final Recommendation

  • Recommend enrolling in the “Essential Cloud Infrastructure: Core Services” course of the “Architecting with Google Compute Engine” series. Topics covered will be Cloud IAM, different data storage services in GCP, resource management, and lastly, resource monitoring.