1/34
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Why are GPUs well suited for AI?
They perform large numbers of computations in parallel, which is ideal for matrix/tensor operations used in neural networks.
CPU vs GPU?
CPU = general-purpose/sequential processing. GPU = massively parallel processing.
What is a Tensor Core?
Specialized NVIDIA GPU hardware optimized for matrix/tensor operations used heavily in AI.
What is HBM?
High Bandwidth Memory — high-speed memory used by data-center GPUs for model parameters, activations and other working data.
Why is GPU memory capacity important for LLMs?
Model weights and runtime data must fit in available memory; larger models/workloads require more memory or techniques such as quantization/distribution.
What is PCIe?
A general-purpose high-speed interconnect connecting CPUs, GPUs, NICs and other devices.
What is NVLink?
NVIDIA's high-bandwidth interconnect for fast communication between accelerators; think scale-up.
What is NVSwitch?
A switched NVLink fabric allowing multiple GPUs to communicate efficiently.
Scale-up vs scale-out?
Scale-up: connect GPUs within/larger compute domains. Scale-out: connect multiple GPU systems/nodes over a network.
What networking technologies are associated with scale-out AI clusters?
High-performance Ethernet and InfiniBand.
What is InfiniBand?
A high-performance, low-latency network fabric widely used for HPC and AI clusters.
What is RDMA?
Remote Direct Memory Access — transfers data between systems with reduced CPU/OS involvement.
What is RoCE?
RDMA over Converged Ethernet — RDMA communication over Ethernet.
What is GPUDirect RDMA?
Technology enabling efficient data movement between GPU memory and network devices, reducing unnecessary CPU/system-memory involvement.
Why is RDMA valuable in distributed AI?
It reduces CPU overhead and data-copying costs, improving communication efficiency between nodes.
What is CUDA?
NVIDIA's accelerated-computing platform/programming model for using NVIDIA GPUs.
What is NCCL?
NVIDIA Collective Communications Library, optimized for GPU-to-GPU communication across single and multiple nodes.
CUDA vs NCCL?
CUDA → GPU computation. NCCL → GPU communication.
Name common NCCL collective operations.
AllReduce, Broadcast, Reduce, AllGather and ReduceScatter.
Where does PyTorch fit?
AI/ML framework above CUDA and NVIDIA GPU libraries.
What is cuDNN?
NVIDIA GPU-accelerated library containing primitives optimized for deep neural networks.
What does Kubernetes do in GPU infrastructure?
Schedules and orchestrates containerized AI workloads across GPU-capable nodes.
How can a Kubernetes pod request an NVIDIA GPU?
Commonly through a resource request/limit such as nvidia.com/gpu: 1, with the NVIDIA/Kubernetes GPU stack configured.
What is Slurm?
A workload manager/scheduler widely used in HPC and AI compute clusters.
Why can networking become an AI bottleneck?
Distributed training requires frequent communication among GPUs; slow communication can leave GPUs waiting instead of computing.
Eight GPUs in one system need fast communication. What technology should come to mind?
NVLink / NVSwitch.
Hundreds of GPU servers need high-performance communication. What should come to mind?
Scale-out networking: InfiniBand or high-performance Ethernet/RDMA.
Need RDMA over Ethernet. What technology?
RoCE.
Need efficient GPU-memory-to-network transfers between servers?
GPUDirect RDMA.
Need software for collective communication across GPUs?
NCCL.
Need to execute accelerated computations on NVIDIA GPUs?
CUDA.
A model doesn't fit into one GPU's memory. What resource is the immediate constraint?
GPU memory capacity, such as HBM.
GPU utilization is low while GPUs wait for synchronization. What subsystem should you investigate?
GPU communication/networking, along with the distributed workload configuration.
Need to schedule an AI container onto an available GPU node?
Kubernetes scheduler + NVIDIA GPU resource/device integration.
What are the major physical AI-infrastructure resources?
Compute, GPU memory, networking, storage, power and cooling.