Central Processing Unit Architecture, Operation, and Parallelism

Central Processing Unit (CPU) Architectural Fundamentals

  • Definition and Primary Function:

    • The Central Processing Unit (CPU) is the primary electronic circuitry within a computer system that carries out the instructions of a computer program.
    • It performs fundamental arithmetic, logical, control, and input/output (I/O) operations specified by program instructions.
    • The term "CPU" has been standard in the computing industry since at least the early 1960s.
    • Traditionally, the term "CPU" refers specifically to the processing unit and the control unit (CU), distinguishing these core elements from external components such as main memory and I/O circuitry.
  • Core Internal Components:

    • Arithmetic Logic Unit (ALU): Executes integer arithmetic (such as addition and subtraction) and bitwise logic operations.
    • Processor Registers: High-speed internal memory storage elements that supply operands directly to the ALU and store the output results of ALU operations.
    • Control Unit (CU): Fetches instructions from memory, decodes them, and directs the coordinated operations of the ALU, registers, and other components.
  • System Interconnections and Functional Architecture:

    • Data paths transfer data and instructions between input, main memory, processor registers, combinational logic, and output units.
    • Control pathways originate from the Control Unit to direct the operations of input, main memory, registers, combinational logic, and output units.

Block diagram of a basic uniprocessor-CPU computer showing data flow in black and control flow in red

  • Physical Integration and Implementations:
    • Microprocessors: CPUs contained entirely on a single integrated circuit (IC) chip die.
    • Microcontrollers / System on a Chip (SoC): Integrated circuit devices that contain a CPU alongside memory, peripheral interfaces, and additional hardware subsystems.
    • Multi-Core Processors: Single integrated circuit chips (sometimes called sockets) containing two or more individual CPU cores.
    • Vector / Array Processors: Parallel processing systems with multiple processors operating in parallel, where no single unit is considered central.

Historical Evolution of Processing Units

  • Fixed-Program Systems vs. Stored-Program Architecture:

    • Early computing devices such as ENIAC were "fixed-program computers" that required physical rewiring to execute different tasks or algorithms.
    • Physical rewiring created severe operational limitations, requiring significant time and effort to reconfigure the computer for a new task.
    • Modern CPU definition requires software program execution capability, making stored-program computers the origin of modern CPUs.
  • Foundational Stored-Program Milestones:

    • ENIAC: Designed by J. Presper Eckert and John William Mauchly; stored-program capability was initially omitted to hasten completion.
    • EDVAC (Electronic Discrete Variable Automatic Computer):
    • Mathematician John von Neumann authored and distributed First Draft of a Report on the EDVAC on June 30, 1945.
    • Outlined a design where program instructions were stored in high-speed computer memory rather than physical wiring, allowing programs to be changed simply by changing memory contents.
    • EDVAC was completed in August 1949.
    • Manchester Small-Scale Experimental Machine: Small prototype computer that became the first operational stored-program computer, executing its first program on June 21, 1948.
    • Manchester Mark 1: Ran its first program during the night of June 16–17, 1949.
    • Konrad Zuse: Conceptualized and implemented stored-program computer designs prior to or alongside von Neumann's publications.
  • Architectural Paradigms:

    • Von Neumann Architecture: Employs a single unified memory space to store both program instructions and operational data. Most modern general-purpose CPUs are primarily von Neumann in design.
    • Harvard Architecture: Uses physically separate storage and treatment for CPU instructions and data.
    • Originated with the Harvard Mark I, which used punched paper tape for instructions and electronic memory for data.
    • Common in modern embedded microcontrollers, such as Atmel AVR processors.
  • Hardware Switching Technologies:

    • Relays vs. Vacuum Tubes (Thermionic Tubes):
    • Early digital computers required thousands or tens of thousands of individual switching devices.
    • Relay-based computers (e.g., Harvard Mark I) were slower but highly reliable, failing very rarely.
    • Tube-based computers (e.g., EDVAC) provided major speed advantages but suffered from low reliability, averaging approximately eight hours between failures.
    • Tube-based CPUs became dominant because speed advantages outweighed reliability issues.
    • Early synchronous CPUs ran at low clock signal frequencies ranging from 100 kHz\text{100\,kHz} to 4 MHz\text{4\,MHz}, limited by switching device speeds.
    • Transistorization (1950s–1960s):
    • Replaced bulky, fragile vacuum tubes and relays with discrete transistors on printed circuit boards.
    • Significantly increased switching speed, reduced power consumption, improved physical reliability, and enabled clock rates in the tens of megahertz (MHz\text{MHz}).
    • Experimental designs like Single Instruction Multiple Data (SIMD) vector processors emerged during this period, giving rise to specialized supercomputers (e.g., Cray Inc.).
  • Major Historical Systems and Standardization:

    • IBM System/360 (1964):
    • Introduced a standardized computer architecture capable of running identical software across different models with varying performance levels.
    • Popularized the concept of a microprogram (microcode), which translates high-level operations into hardware signals and remains in widespread use.
    • Dominated mainframe computing for decades, evolving into modern IBM zSeries systems.
    • DEC PDP-8 (1965): Influential transistor-based computer targeted at scientific and research markets.
  • Integration Eras (SSI, MSI, LSI, and Microprocessors):

    • Small-Scale Integration (SSI):
    • Miniaturized basic digital circuits (e.g., NOR gates) into integrated circuits containing up to a few score transistors.
    • Building a CPU required thousands of individual SSI ICs.
    • Examples: Apollo guidance computer; IBM System/370 (replacing Solid Logic Technology discrete-transistor modules); DEC PDP-8/I and KI10 PDP-10; original DEC PDP-11 models.
    • Medium-Scale Integration (MSI) & Large-Scale Integration (LSI):
    • Lee Boysel's 1967 manifesto detailed building the equivalent of a 32-bit mainframe CPU using a small number of LSI circuits.
    • LSI chips contained 100 or more logic gates, manufactured using Metal-Oxide-Semiconductor (MOS) processes (PMOS logic, NMOS logic, CMOS logic).
    • Companies building high-speed computers continued using bipolar junction transistors (e.g., TTL / 7400 series logic gates) through the 1970s and early 1980s (e.g., Datapoint processors) because early MOS chips were slow.
    • By 1968, CPU chip counts were reduced to 24 ICs of eight different types (each containing roughly 1,000 MOSFETs).
    • The first LSI implementation of the DEC PDP-11 reduced the CPU to four integrated circuits.
    • Microprocessors:
    • Pioneered in the 1970s by Federico Faggin via Silicon Gate MOS ICs with self-aligned gates and a random logic design methodology.
    • Intel 4004 (1970): First commercially available microprocessor.
    • Intel 8080 (1974): First widely used general-purpose microprocessor.
    • Example Die Size: An Intel 80486DX2 microprocessor die measures 12×6.75 mm12 \times 6.75\,\text{mm}.
    • Physical Limits of Silicon Scaling:
    • Miniaturization described by Moore's law encounters physical limitations such as electromigration and subthreshold leakage.
    • Prompted research into quantum computing and expanded architectural parallelism to extend performance beyond classical von Neumann limitations.

CPU Instruction Execution Cycle

  • Overview of the Instruction Cycle:

    • The operational sequence of a CPU consists of executing stored instructions (a program) through three main phases: Fetch, Decode, and Execute.
    • Following execution, the sequence repeats, fetching the next instruction determined by the incremented value in the Program Counter (PC).
  • Phase 1: Fetch:

    • Retrieves an instruction (represented as a number or sequence of numbers) from program memory using the address stored in the Program Counter (PC).
    • After fetching, the PC is incremented by the length of the instruction to point to the next instruction in sequence.
    • Memory latency can cause the CPU to stall while waiting for instructions; modern processors address this using caches and pipelined architectures.
  • Phase 2: Decode:

    • Performed by the instruction decoder circuitry to convert the instruction into control signals for other CPU components.
    • Interpreted according to the CPU's Instruction Set Architecture (ISA).
    • Opcode (Operation Code): A specific bit field within the instruction indicating the exact operation to be performed.
    • Operands & Addressing Modes: Remaining instruction fields specify operands, which may be provided as:
    • Immediate values (constants within the instruction).
    • Processor register locations.
    • Memory addresses.
    • Implementation approaches:
    • Hardwired Decoder: Unchangeable direct circuit.
    • Microprogrammed Decoder: A microprogram applies configuration signals sequentially over multiple clock pulses; rewritable microprogram memory allows changing instruction decoding.
  • Phase 3: Execute:

    • Electrically connects relevant CPU components to perform the specified operation, typically in response to a clock pulse.
    • Results are written back to an internal CPU register for immediate access or to slower main memory.
    • Arithmetic Overflow: If an addition or arithmetic result exceeds the output word size of the ALU, an arithmetic overflow flag is set in a dedicated flags register.
  • Control Flow and Program Redirection:

    • Jump / Branch Instructions: Modify the Program Counter directly rather than producing data, enabling loops, conditional execution, and function calls.
    • Flags Register (Status Register):
    • Bits inside the flags register change state based on operation outcomes (e.g., zero, equal, greater than).
    • Example: A compare instruction evaluates two values and sets status flags; subsequent conditional jump instructions read these flags to determine program flow.

Detailed CPU Internal Components and Mechanisms

  • Control Unit (CU):

    • Uses electrical signals to direct the entire computer system in carrying out stored program instructions.
    • Directs functional units rather than executing program instructions directly.
    • Communicates directly with both the ALU and system memory.
  • Arithmetic Logic Unit (ALU) & Floating Point Unit (FPU):

    • ALU: Digital circuit performing integer arithmetic and bitwise logic operations.
    • Operands originate from internal registers, external memory, or internal constants.
    • Output consists of result data words and status flag updates.
    • FPU: Dedicated internal circuit or coprocessor responsible for floating-point calculations.
  • Data Representation, Integer Range, and Word Size:

    • Modern CPUs represent numbers in binary form using two-valued physical quantities (high/low voltages).
    • Word Size (Bit Width / Data Path Width / Integer Size): The number of binary bits a CPU processes in a single operation.
    • Direct Integer Range Calculation:
    • An nn-bit CPU directly processes 2n2^n discrete integer values.
    • An 8-bit CPU directly manipulates integers represented by 8 bits, yielding a range of 28=2562^8 = 256 values.
    • Direct Memory Addressing Calculation:
    • An nn-bit address bus directly addresses 2n2^n memory locations.
    • A 32-bit address bus directly addresses 2322^{32} distinct memory locations.
    • Mechanisms like bank switching allow addressing additional memory beyond native bit limits.
    • Mixed Bit Width Architectures:
    • CPUs often mix bit widths across different functional units to balance cost, performance, and accuracy.
    • Example: IBM System/370 operated with a primarily 32-bit CPU, but implemented 128-bit precision within its floating-point units.
  • Clock Rate, Timing, and Power Considerations:

    • Synchronous Circuits: Sequential operations are paced by a global clock signal generated by an external oscillator as a periodic square wave.
    • Timing Margins: Clock periods are set strictly longer than the worst-case propagation delay across the CPU to ensure all signals settle before state changes.
    • Clock Distribution Drawbacks:
    • Keeping clock signals in phase (synchronized) across complex die areas becomes difficult at high frequencies, requiring multiple identical clock signals.
    • Unused switching logic dissipates energy as heat on every clock pulse.
    • Mitigation Strategies:
    • Clock Gating: Turns off clock signals to idle components to eliminate switching energy (e.g., used extensively in the IBM PowerPC Xenon CPU of the Xbox 360).
    • Asynchronous (Clockless) CPUs: Completely eliminate the global clock signal, relying on internal handshaking protocols.
      • Benefits: Marked advantages in lower power consumption and reduced heat dissipation.
      • Examples: ARM-compliant AMULET, MIPS R3000-compatible MiniMIPS.
      • Hybrid Implementations: Asynchronous ALUs combined with superscalar pipelining.

Instruction-Level and Data-Level Parallelism

  • Subscalar Execution Performance:

    • Subscalar CPUs execute less than one instruction per clock cycle, yielding Cycles Per Instruction of CPI>1\text{CPI} > 1.
    • When an instruction requires multiple cycles to complete, the execution units stall, leaving transistors idle.
  • Instruction-Level Parallelism (ILP):

    • Instruction Pipelining:
    • Breaks instruction execution pathways into discrete stages operating concurrently like an assembly line.
    • Ideal peak completion rate is scalar execution: 1 instruction per cycle (CPI=1\text{CPI} = 1).
    • Data Dependency Conflicts: Occur when a stage requires the output of a preceding instruction still in the pipeline, causing pipeline stalls.

Basic five-stage instruction pipeline showing IF, ID, EX, MEM, and WB stages

  • Classic Five-Stage RISC Pipeline Stages:

    1. IF: Instruction Fetch
    2. ID: Instruction Decode
    3. EX: Execute
    4. MEM: Memory Access
    5. WB: Write Back
  • Superscalar Execution:

    • Features a long instruction pipeline paired with multiple identical execution units.
    • A dispatcher reads and dispatches multiple instructions simultaneously, achieving execution rates above 1 instruction per cycle (IPC>1\text{IPC} > 1 or CPI<1\text{CPI} < 1).
    • Requires significant CPU cache and performance techniques: branch prediction (predicting conditional paths), speculative execution (executing potential code paths in advance), and out-of-order execution (reordering execution to avoid data dependencies).
    • Asymmetric Superscalar Historical Example:
      • Intel P5 Pentium: Included two superscalar ALUs accepting 1 integer instruction per clock each, but its FPU was non-superscalar.
      • Intel P6: Added superscalar capabilities to floating-point operations.
  • Very Long Instruction Word (VLIW):

    • Moves instruction parallelization logic out of hardware and into compiler software, reducing hardware implementation complexity.

    • Thread-Level Parallelism (TLP) and Parallel Computing:

  • Classified under Flynn's Taxonomy as MIMD (Multiple Instruction Stream, Multiple Data Stream).

  • Symmetric Multiprocessing (SMP): Multiple CPUs share a uniform coherent view of main memory; limited to a small number of CPUs.

  • Non-Uniform Memory Access (NUMA) & Directory-Based Coherence: Enables scaling to thousands of cooperating processors.

  • Multi-Core / Chip-Level Multiprocessing (CMP): Integrates multiple complete CPU cores on a single silicon die.

  • Multi-Threading (MT): Replicates specific internal hardware components to allow a single CPU core to run multiple software threads concurrently while sharing execution units and caches.

    • Block Multithreading: Executes a thread until a stall occurs (e.g., waiting for external memory), then rapidly switches to another thread in 1 clock cycle (e.g., UltraSPARC).
    • Simultaneous Multithreading (SMT): Executes instructions from multiple threads in parallel within a single clock cycle.
  • Historical Shift toward Throughput Computing:

    • High-frequency ILP maximization stalled in the early 2000s due to memory latency gaps and severe power dissipation (e.g., Intel Pentium 4).

    • Designers adopted multi-thread throughput computing (dual/multi-core CMP, revived P6-like efficient pipelines).

    • Notable CMP Implementations: x86-64 Opteron, Athlon 64 X2, SPARC UltraSPARC T1, IBM POWER4, IBM POWER5, Xbox 360 (triple-core PowerPC), PlayStation 3 (7-core Cell microprocessor).

    • Data Parallelism and Vector Processors:

  • Classified under Flynn's Taxonomy as SIMD (Single Instruction, Multiple Data) versus SISD (Single Instruction, Single Data).

  • Vector Processors: Operate on large vectors/arrays of data using a single instruction (e.g., Cray-1 supercomputer).

  • General-Purpose SIMD Extensions:

    • Accelerate repetitive operations in digital multimedia (audio, video, images) and scientific computations.
    • Early Integer-Only SIMD: HP MAX (Multimedia Acceleration eXtensions), Intel MMX.
    • Modern Floating-Point SIMD: Intel SSE (Streaming SIMD Extensions), PowerPC AltiVec (VMX).

Performance Metrics and Hardware Utilization

  • Core Performance Equations and Metrics:

    • Determinant factors of processing speed: Clock Rate (in Hz\text{Hz}, MHz\text{MHz}, GHz\text{GHz}) and Instructions Per Clock (IPC).
    • Instructions Per Second calculation:     IPS=Clock Rate (Hz)×IPC=Clock Rate (Hz)CPI\text{IPS} = \text{Clock Rate (Hz)} \times \text{IPC} = \frac{\text{Clock Rate (Hz)}}{\text{CPI}}
    • Limitations of Peak MIPS / IPS:
    • Represent artificial peak rates on optimal instruction sequences with minimal branching.
    • Do not reflect real-world workloads or memory hierarchy delays.
  • Standardized Benchmarking:

    • Standardized benchmark suites (e.g., SPECint) evaluate effective operational performance on real applications.
  • Multi-Core Scaling Realities:

    • Multi-core processors handle asynchronous events and interrupts more efficiently.
    • Dual-core processors deliver approximately a 50% performance increase over single-core processors (rather than a theoretical 100%) due to software parallelization limits and overhead.
  • Hardware Utilization Monitoring:

    • Shared resources in modern features (e.g., hyper-threading and uncore) make tracking software utilization complex.
    • Hardware counters integrated into modern CPUs monitor module utilization directly (e.g., Intel Performance Counter Monitor technology).