NCA-AIIO Lesson 1

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/34

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 8:52 PM on 9/22/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

35 Terms

1
New cards

Why are GPUs well suited for AI?

They perform large numbers of computations in parallel, which is ideal for matrix/tensor operations used in neural networks.

2
New cards

CPU vs GPU?

CPU = general-purpose/sequential processing. GPU = massively parallel processing.

3
New cards

What is a Tensor Core?

Specialized NVIDIA GPU hardware optimized for matrix/tensor operations used heavily in AI.

4
New cards

What is HBM?

High Bandwidth Memory — high-speed memory used by data-center GPUs for model parameters, activations and other working data.

5
New cards

Why is GPU memory capacity important for LLMs?

Model weights and runtime data must fit in available memory; larger models/workloads require more memory or techniques such as quantization/distribution.

6
New cards

What is PCIe?

A general-purpose high-speed interconnect connecting CPUs, GPUs, NICs and other devices.

7
New cards

What is NVLink?

NVIDIA's high-bandwidth interconnect for fast communication between accelerators; think scale-up.

8
New cards

What is NVSwitch?

A switched NVLink fabric allowing multiple GPUs to communicate efficiently.

9
New cards

Scale-up vs scale-out?

Scale-up: connect GPUs within/larger compute domains. Scale-out: connect multiple GPU systems/nodes over a network.

10
New cards

What networking technologies are associated with scale-out AI clusters?

High-performance Ethernet and InfiniBand.

11
New cards

What is InfiniBand?

A high-performance, low-latency network fabric widely used for HPC and AI clusters.

12
New cards

What is RDMA?

Remote Direct Memory Access — transfers data between systems with reduced CPU/OS involvement.

13
New cards

What is RoCE?

RDMA over Converged Ethernet — RDMA communication over Ethernet.

14
New cards

What is GPUDirect RDMA?

Technology enabling efficient data movement between GPU memory and network devices, reducing unnecessary CPU/system-memory involvement.

15
New cards

Why is RDMA valuable in distributed AI?

It reduces CPU overhead and data-copying costs, improving communication efficiency between nodes.

16
New cards

What is CUDA?

NVIDIA's accelerated-computing platform/programming model for using NVIDIA GPUs.

17
New cards

What is NCCL?

NVIDIA Collective Communications Library, optimized for GPU-to-GPU communication across single and multiple nodes.

18
New cards

CUDA vs NCCL?

CUDA → GPU computation. NCCL → GPU communication.

19
New cards

Name common NCCL collective operations.

AllReduce, Broadcast, Reduce, AllGather and ReduceScatter.

20
New cards

Where does PyTorch fit?

AI/ML framework above CUDA and NVIDIA GPU libraries.

21
New cards

What is cuDNN?

NVIDIA GPU-accelerated library containing primitives optimized for deep neural networks.

22
New cards

What does Kubernetes do in GPU infrastructure?

Schedules and orchestrates containerized AI workloads across GPU-capable nodes.

23
New cards

How can a Kubernetes pod request an NVIDIA GPU?

Commonly through a resource request/limit such as nvidia.com/gpu: 1, with the NVIDIA/Kubernetes GPU stack configured.

24
New cards

What is Slurm?

A workload manager/scheduler widely used in HPC and AI compute clusters.

25
New cards

Why can networking become an AI bottleneck?

Distributed training requires frequent communication among GPUs; slow communication can leave GPUs waiting instead of computing.

26
New cards

Eight GPUs in one system need fast communication. What technology should come to mind?

NVLink / NVSwitch.

27
New cards

Hundreds of GPU servers need high-performance communication. What should come to mind?

Scale-out networking: InfiniBand or high-performance Ethernet/RDMA.

28
New cards

Need RDMA over Ethernet. What technology?

RoCE.

29
New cards

Need efficient GPU-memory-to-network transfers between servers?

GPUDirect RDMA.

30
New cards

Need software for collective communication across GPUs?

NCCL.

31
New cards

Need to execute accelerated computations on NVIDIA GPUs?

CUDA.

32
New cards

A model doesn't fit into one GPU's memory. What resource is the immediate constraint?

GPU memory capacity, such as HBM.

33
New cards

GPU utilization is low while GPUs wait for synchronization. What subsystem should you investigate?

GPU communication/networking, along with the distributed workload configuration.

34
New cards

Need to schedule an AI container onto an available GPU node?

Kubernetes scheduler + NVIDIA GPU resource/device integration.

35
New cards

What are the major physical AI-infrastructure resources?

Compute, GPU memory, networking, storage, power and cooling.