Lecture 5a GPU description
Introduction to Graphics Processing Units (GPUs)
Prepared by Francesco Terrosi for the Italian Workshop on Embedded Systems (IWES) on 8-9 February 2021.
Key Concepts
Parallel Computing:
Definition: The execution of multiple instructions simultaneously across multiple computing cores, improving speed and computational efficiency.
Implementation: Can occur on a single machine with multicore processors or scaled over a network through cloud clustering, allowing for distributed processing tasks.
Computing Evolution:
Moore’s Law: A foundational principle which states that the number of transistors on a microchip doubles approximately every two years, resulting in increased performance and decreased costs. However, the exponential growth of transistors means that reliance on improvements to single-core processors becomes unfeasible, necessitating a shift to parallel computing methods.
Impact: This evolution in chip technology has driven the need for advanced architectures that can handle greater workloads without solely increasing clock speeds.
Importance of Parallel Architectures:
Challenge Addressed: Parallel architectures provide viable solutions to the performance limitations faced by sequential computing, particularly in applications requiring extensive data processing and analysis.
GPU vs CPU: Architectural Differences
CPUs (Central Processing Units):
Core Structure: CPUs typically contain a small number of complex cores (ranging from 2 to 16) designed to optimize performance on single-threaded tasks. They are ideal for managing diverse tasks, executing intricate algorithms, and handling operations that benefit from strong sequential performance.
Example: For instance, the Intel Core i7 dynamically adjusts clock speeds based on the workload across its cores to deliver optimized performance.
GPUs (Graphics Processing Units):
Core Structure: In contrast, GPUs are built with hundreds to thousands of simpler cores designed for high throughput and efficient parallel processing. They excel in processing large blocks of data in parallel, which is crucial for rendering graphics and performing computations on large datasets.
Data Parallelism: GPUs focus primarily on data parallelism, making them highly effective for applications such as scientific simulations, machine learning, and any other tasks that involve repetitive calculations over large volumes of data.
CAPACITY: Performance Metrics
Core Performance: Each GPU core operates with specific configurations of cache memory to enhance memory bandwidth and reduce latency during processing, essential for maintaining high performance in compute-heavy tasks.
Benchmarking: Performance statistics from various Intel Xeon and Core processors showcase significant improvements in benchmark tests highlighting GPU advantages in parallel tasks.
GPU Architecture: Components and Levels
Single Instruction Multiple Data (SIMD): A key feature of GPU architecture that allows a single instruction to operate on multiple data points concurrently, taking full advantage of available processing power.
Components Include:
Global Memory: Comprising off-chip DRAM and various caches optimized for fast data access.
Streaming Multiprocessors (SMs): The core processing units within GPUs that consist of multiple identical components, enhancing computational throughput.
Architectural Generations
FERMI Architecture (2010): Each Streaming Multiprocessor contains a single processor with specified memory hierarchies aimed at improving performance and efficiency.
AMPERE Architecture (2020): Each SM is capable of housing four processors, significantly increasing computational throughput. Also introduces TENSOR cores, designed for optimized performance on deep learning tasks, thereby broadening the spectrum of applications.
Memory Hierarchy in GPUs
Memory Structures: GPUs utilize an intricate memory hierarchy consisting of:
Global Memory: The off-chip DRAM storage allows for significant data storage, essential for large applications.
Caches: Including L1, L2, and texture caches, which are optimized for bandwidth filtering to minimize data access times.
Register Files: Temporary storage within processors that boosts efficiency and speed during computation cycles, ensuring that working data is readily available.
GPU Execution Model
Threading Model:
Organizational Structure: Threads are strategically organized into blocks, which are then further organized into grids, optimizing execution across the processors based on workload distribution.
Warp Concept: Each block can contain 32 threads known as a Warp, which allows for efficient management and scheduling of threads, significantly improving processing speed.
Kernel Definition:
Kernel: A kernel is a program that runs on the GPU, designed to allow concurrent execution of multiple threads, facilitating the processing of large data sets effectively and efficiently.
Programming with CUDA
CUDA (Compute Unified Device Architecture): Developed by NVIDIA, this parallel computing platform and programming model allows developers to leverage the computational power of GPUs for general-purpose computing.
Language Features: These features enable a smoother transition from traditional CPU-centric programming to GPU-oriented approaches, enhancing accessibility for developers.
Example Code for Vector Addition: Illustrates the memory allocation process and kernel execution setup. In this example, vectors are allocated on the host and device, data is copied from the host to the device, and a kernel is launched to perform vector addition.
size_t bytes = n * sizeof(int); vector<int> host_a(n); vector<int> host_b(n); vector<int> host_c(n); cudaMalloc(&device_a, bytes); cudaMemcpy(device_a, host_a.data(), bytes, cudaMemcpyHostToDevice); vectAdd<<<blocks, THREADS>>>(device_a, device_b, n, device_c);
Performance Comparison: GPU vs CPU
Advantages of Using GPUs: Emphasizes their suitability for tasks that can effectively harness parallel processing, particularly for complex mathematical computations and large-scale data analysis. GPUs are particularly beneficial in fields such as artificial intelligence, machine learning, and real-time graphics rendering.
Challenges: Despite their advantages, some challenges remain, particularly concerning data transfer between CPU and GPU, which can incur significant overheads that need to be managed for optimal performance. Understanding these intricacies is essential for software developers and engineers in optimizing application performance.