1/61
62 cards: vocabulary, MPI argument examples, and high-yield exam practice on choosing collectives, tracing ranks, and tree time/work.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Process rank
A process's unique ID within a communicator, numbered 0 through p-1. Use it to choose that process's work.
Process count p
The number of MPI processes in the communicator. MPI_Comm_size stores it in p.
MPI_COMM_WORLD
The communicator containing all processes launched for the program.
Root process
The designated process for a collective operation; often rank 0. It may own the input or receive the answer.
MPI_Init
Starts the MPI environment. Call it before other MPI operations.
MPI_Finalize
Ends the MPI environment after the parallel work is done.
MPI_Comm_rank
Writes this process's rank into the supplied integer variable.
MPI_Comm_size
Writes the number of processes into the supplied integer variable.
MPI_Send
Sends a specified buffer to one destination rank with a tag.
MPI_Recv
Receives a message from a specified source and tag; it waits for a matching message.
Message tag
An integer label used to distinguish messages. A receive must match the intended send's tag.
Blocking receive
An MPI_Recv call that cannot return until its receive buffer contains a matching message.
Deadlock
Processes wait in a cycle or for messages that will never arrive, so the program cannot progress.
Collective operation
An MPI operation involving every process in a communicator. All participating ranks must call it compatibly.
MPI_Bcast
Broadcasts a value from one root process to every process in the communicator.
MPI_Scatter
Distributes equal-sized consecutive chunks of a root array among all processes.
MPI_Gather
Collects equal-sized chunks from all processes into an array on the root.
MPI_Reduce
Combines one value or array from every process using an operation such as sum, max, or min; the result is at the root.
MPI_Allreduce
Like MPI_Reduce, but every process receives the combined result.
MPI_SUM / MPI_MAX / MPI_MIN
Predefined reduction operations for sum, maximum, and minimum. Choose the one that matches the required answer.
Blocking data allocation
Give each process a consecutive chunk of the input, such as indices 0-3 to rank 0 and 4-7 to rank 1.
Striping
Give rank r indices r, r+p, r+2p, and so on. It can balance uneven work but may hurt data locality.
Load balancing
Distributing work so processes finish at roughly the same time instead of leaving some idle.
Data locality
Keeping a process's data near or contiguous to its other data, reducing costly access and transfer.
Distributed memory
Each process has its own memory; processes exchange data explicitly using messages such as MPI sends and receives.
Shared memory
Processes or threads can access a common memory space; coordination is needed to avoid conflicting access.
Parallel overhead
Extra time spent dividing work, communicating, synchronizing, and combining results.
Speedup
Sequential running time divided by parallel running time. Overhead and serial work limit it.
Work W(n)
The total number of operations performed across all processors and rounds in the work-depth model.
Depth T(n)
The number of sequential parallel rounds on the longest dependency path in the work-depth model.
Parallel reduction
Combine values up a balanced tree to get one answer; for n values, depth O(log n) and work O(n).
Prefix computation / scan
Compute one cumulative answer for every prefix; tree up-sweep and down-sweep give depth O(log n) and work O(n).
Associative operation
An operation where grouping does not change the result, such as sum, min, or max. This lets a tree combine subproblems safely.