Control Unit Fundamentals: Architecture, Microprogramming, and Data Path Operations

Fundamental Role and Concept of the Control Unit

  • The Orchestral Conductor Metaphor:

    • The hardware components of a Central Processing Unit (CPU)—such as registers, the Arithmetic Logic Unit (ALU), memory, and internal data buses—act like individual musicians in an orchestra (e.g., violinists, trombonists).

    • If each hardware component operates independently without coordination, the system generates chaotic signals ("noise") rather than structured execution ("a symphony").

    • The Control Unit (CU) functions as the conductor. It does not perform actual computational work or play an instrument directly; instead, it coordinates and directs the other components.

  • Definition of Control Signals:

    • Control signals are precise electrical instructions generated by the control unit.

    • These signals tell hardware components exactly what action to perform and at what specific time instance.

  • Coordination of the Data Path:

    • Directs registers when to output stored data onto the data bus.

    • Commands the ALU on which operation (e.g., addition, subtraction) to execute.

    • Signals destination registers when to load and save results from the data bus.

Data Path Execution and Step-by-Step Register Operations

  • Standard Running Example:

    • Consider an addition operation involving registers r1r_1, r2r_2, and r0r_0 represented as:         r0=r1+r2r_0 = r_1 + r_2

    • Initial states:

      • Register r1r_1 holds the value 55

      • Register r2r_2 holds the value 33

    • Target final state:

      • Register r0r_0 holds the calculated sum 88

  • Hardware Execution Sequence:

    • From a high-level perspective, the operation appears as a single action. From a hardware perspective, the CPU breaks the instruction into fine-grained sequential actions:

      1. Retrieve/output the stored value from register r1r_1.

      2. Retrieve/output the stored value from register r2r_2.

      3. Pass both values into the ALU inputs and trigger the addition operation.

      4. Capture and store the resulting output from the ALU into register r0r_0

  • Instruction vs. Control Unit Function:

    • The machine instruction defines what output is desired.

    • The control unit determines how the hardware accomplishes the request through discrete, timed control actions.

Timing Systems, System Clock, and Control Signal Generation

  • Sequential Execution Requirement:

    • Hardware operations cannot occur simultaneously in a single instantaneous step.

    • Data movement must be strictly sequenced so data becomes available at a specific moment and is captured by destination components only when stable.

  • The System Clock:

    • Emits a continuous sequence of electrical pulses establishing discrete timing steps or checkpoints represented as t0,t1,t2,etc.t_0, t_1, t_2, \text{etc.} (or tsub0,tsub1,tsub2t_{\text{sub}0}, t_{\text{sub}1}, t_{\text{sub}2}).

    • At timing checkpoint t0t_0, a specific initial subset of control actions occurs.

    • At timing checkpoint t1t_1, the next discrete subset of control actions occurs.

    • At timing checkpoint t2t_2, subsequent actions occur.

  • Control Logic Inputs and Outputs:

    • The control unit combines three primary information sources inside its control logic:

      1. Opcode (Operation Code): Identifies the specific macro-instruction being executed.

      2. Timing Signals: Indicates the exact current step in the execution cycle (t0,t1,t2t_0, t_1, t_2).

      3. Status/Coordination Signals: Provides flag conditions and system state info.

    • Output: Generates the specific set of active control signals directed to the data path for that clock step.

Hardwired vs. Microprogrammed Control Architectures

  • Hardwired Control Unit:

    • Mechanism: Decision-making logic is built directly into fixed digital logic circuits (combining gates such as AND, OR, NOT).

    • Speed: Extremely fast signal generation due to direct propagation delays through hardware gates.

    • Cost/Complexity: Economical and efficient for small instruction sets; becomes extremely complex and difficult to manage as the CPU instruction set expands.

    • Flexibility: Rigid; any modifications or additions to the instruction set require a physical redesign of the hardware circuit layout.

  • Microprogrammed Control Unit:

    • Mechanism: Stores control signal patterns as binary microinstruction words inside a dedicated internal memory called Control Memory (CMCM

    • Function: Operates as a miniature instruction interpreter embedded within the CPU.

    • Speed: Slower than hardwired control due to the memory read overhead required to fetch each microinstruction.

    • Flexibility: Highly flexible and easier to adapt; complex control routines and new instructions can be implemented or modified by updating microprogram memory without altering hardware logic gates.

Hardwired Control Logic and Boolean Expression Derivation

  • Deriving Control Signals using Boolean Logic:

    • If a specific control signal AA must activate whenever instruction instance XX or instruction instance ZZ is executing at clock timing step t1t_1, it is formally expressed in Boolean algebra as:         A=(X+Z)×t1A = (X + Z) \times t_1

    • In this expression, ++ represents the logical OR operation, and ×\times represents the logical AND operation.

    • Signal AA turns active if and only if the timing condition t1t_1 AND at least one of the instruction conditions (XX OR ZZ) are true.

  • Control Signals for Three-Bus Addition Example (r0=r1+r2r_0 = r_1 + r_2):

    • r1 out=1r_1\text{ out} = 1: Enables register r1r_1 to output its operand to the first bus.

    • r2 out=1r_2\text{ out} = 1: Enables register r2r_2 to output its operand to the second bus.

    • ALU add=1ALU\text{ add} = 1: Commands the ALU to perform an addition operation.

    • r0 in=1r_0\text{ in} = 1: Enables destination register r0r_0 to capture the input value from the result bus.

  • Consequences of Incorrect Control Signals:

    • If ALU add=0ALU\text{ add} = 0, the ALU will not execute the addition operation.

    • If r0 in=0r_0\text{ in} = 0, the result produced by the ALU will fail to store in register r0r_0.

  • Table-to-Logic Design Technique:

    • Engineers map out instructions across timing cycles against control lines in a truth table.

    • Example control derivations from instruction timing tables:

      • A=(X+Z)×t1A = (X + Z) \times t_1

      • B=X×t0+Y×t2B = X \times t_0 + Y \times t_2

      • C=(X+Z)×t1+(X+Y)×t2C = (X + Z) \times t_1 + (X + Y) \times t_2

    • The truth table acts as the system specification, and the derived Boolean expressions form the literal hardware logic gate implementation.

Finite State Machine (FSM) View of Instruction Execution

  • FSM Concept:

    • Instead of evaluating purely Boolean equations, instruction execution can be modeled as state transitions within a Finite State Machine.

    • The CPU progresses through cyclic functional states: Fetch →\rightarrow Decode →\rightarrow Execute →\rightarrow Write Back →\rightarrow Fetch.

  • State-Driven Signal Assertion:

    • An FSM state represents a specific control step in the instruction sequence rather than the data itself.

    • At each state, the control unit asserts the precise control signals defined for that state.

  • Example FSM Execution Flow for Instruction Instance YY:

    1. Fetch State: CPU retrieves the machine instruction from primary memory.

    2. Decode State: CU reads the opcode and identifies the operation as Instance YY

    3. Timing State t0t_0: CU enters state t0t_0 and asserts control signals FF, HH, and GG

    4. Timing State t1t_1: CU transitions to state t1t_1 and asserts control signal GG

    5. Timing State t2t_2: CU transitions to state t2t_2 and asserts control signals BB and CC

    6. Return State: Execution completes and state control loops back to the Fetch state.

  • FSM Model for Addition (add r1,r2,r0add\, r_1, r_2, r_0):

    • Fetch: Retrieve addition instruction.

    • Decode: Recognize the addadd opcode.

    • Execute: Assert register output enables for r1r_1 and r2r_2; assert ALU addALU\text{ add}.

    • Write Back: Assert input enable r0 inr_0\text{ in} to store result.

    • Fetch: Prepare for next instruction.

Microprogrammed Control Architecture and Hierarchy

  • Hierarchical Levels of Abstraction:

    1. Macro Instruction: High-level assembly/machine instruction requested by the programmer (e.g., add r1,r2,r0add\, r_1, r_2, r_0).

    2. Microprogram: A sequence of microinstructions stored in Control Memory that implements a single macro instruction.

    3. Microinstruction: A single binary control word residing in Control Memory that specifies one or more compatible microoperations to be performed in a single clock step.

    4. Microoperation: The fundamental hardware-level elementary action performed on the data path (e.g., placing r1r_1 onto ALU input A).

    5. Data Path Activity: Physical signal movement and pulse propagation across circuits.

  • Step-by-Step Microprogram Execution:

    1. CPU fetches the macro instruction and decodes the opcode.

    2. The opcode acts as a lookup pointer to locate the starting address of its corresponding microprogram inside the Control Memory (CMCM).

    3. The control unit fetches the first microinstruction from CMCM

    4. The microinstruction asserts its stored control signals to execute the specified microoperations.

    5. The control unit determines the address of the next microinstruction (sequential increment or branching).

    6. Steps repeat until the microprogram sequence completes, returning control to the instruction fetch loop.

  • Microprogram Branching Conditions:

    • If condition code bits indicate a branch, an explicit address field within the microinstruction specifies the non-sequential location of the next microinstruction.

    • Otherwise, the microprogrammed control unit automatically increments to fetch the next sequential microinstruction.

Horizontal vs. Vertical Microinstruction Formats

  • Horizontal Microinstructions:

    • Structure: Very wide microinstruction words where every bit directly corresponds to a specific physical control line (analogous to a wide control panel with individual dedicated toggle switches).

    • Parallelism: Maximum parallelism; allows many control lines and microoperations to activate simultaneously within a single cycle.

    • Decoding: Requires no external decoding logic.

    • Disadvantage: Requires extremely long control words and massive Control Memory capacity.

  • Vertical Microinstructions:

    • Structure: Short, compact microinstruction words where control signals are grouped and encoded into binary bit fields (e.g., 4 bits for ALU operation, 5 bits for register selection).

    • Parallelism: Limited parallelism; control lines encoded within the exact same bit field cannot be activated at the same time.

    • Decoding: Requires external decoders to interpret the encoded bit fields into individual physical control line activations.

    • Advantage: Significantly reduces word length and conserves Control Memory space.

Advanced Microinstruction Control Schemes

  • Residual Control:

    • Static or repetitive control information is established by an early microinstruction and stored in a setup configuration register.

    • Subsequent microinstructions reuse this residual configuration without repeatedly specifying those control bits, saving control word bandwidth.

  • Nanoprogramming:

    • Employs two distinct levels of control memory: Micro Store and Nano Store.

    • Microinstructions in the Micro Store hold pointers to control words in the Nano Store (which houses the actual unique control signal combinations).

    • Dramatically reduces total control memory footprint when duplicate control signal combinations exist across microinstructions.

Microinstruction Encoding, Decoding, and Binary Field Mapping

  • 19-Bit Vertical Microinstruction Worked Example (add r1,r2,r0add\, r_1, r_2, r_0):

    • Instruction: add r1,r2,r0add\, r_1, r_2, r_0

    • Binary Code: 001000010000100100000001000010000100100000

    • Field Breakdown (19 bits total):

      • Opcode Field (7 bits): 0010000100100001 (decodes to addadd operation)

      • Source Register 1 Field (5 bits): 0000100001 (decodes to register r1r_1)

      • Source Register 2 Field (4 bits): 00100010 (decodes to register r2r_2)

      • Destination Register Field (4 bits): 00000000 (decodes to register r0r_0

  • Decoding Process:

    • A binary string has no inherent meaning to the processor unless decoding logic knows how to divide and interpret the exact bit positions.

    • The decoder reads the bit streams, parses the fields according to defined boundaries, and generates the hardware signals needed to execute the specified parameters.

Instruction Fetch Sequence and the Instruction Register

  • Register Transfer Steps for Instruction Fetch:

    1. PC→MARPC \rightarrow MAR: Copy the address in the Program Counter (PCPC) into the Memory Address Register (MARMAR

    2. Memory→MBR/MDRMemory \rightarrow MBR / MDR: Read data from memory address into the Memory Buffer Register / Memory Data Register (MDRMDR

    3. MDR→IRMDR \rightarrow IR: Transfer the fetched instruction from MDRMDR into the Instruction Register (IRIR

  • Role and Functions of the Instruction Register (IRIR):

    • Dedicated hardware register located inside the Control Unit.

    • Holds the actual binary machine instruction currently being decoded and executed (rather than a memory address).

    • Supplies the opcode directly to the control unit decoders to initiate microprogram sequencing or hardwired execution logic.

Design Trade-Offs, Historical Context, and Instruction Set Emulation

  • Hardwired vs. Microprogrammed Trade-Offs:

    • Hardwired: Optimized for raw speed and minimal execution latency; rigid, difficult to modify, and complex to design.

    • Microprogrammed: Optimized for flexibility, design simplicity, and ease of modification; incurs speed penalties due to memory read operations.

  • Historical Processors using Microprogramming:

    • Intel 8,080

    • Zilong z 80

    • Deckvox (DEC VAX)

  • RISC Preference:

    • Reduced Instruction Set Computer (RISC) architectures generally prefer hardwired control units to maximize execution speed and simplify hardware layout.

  • Instruction Set Emulation:

    • Because microprogrammed control relies on software-like routines in Control Memory, one physical hardware machine can emulate a completely different computer architecture.

    • By swapping the microprograms stored in Control Memory, a physical processor can execute a totally different instruction set interface without modifying its underlying physical data path.

Manual Data Path Circuit Simulation and Demonstration

  • Circuit Walkthrough Elements:

    • Components: System clock, toggle switches, 8-bit registers (r1,r2,r0r_1, r_2, r_0 ), ALU, and interconnecting data buses.

    • Control Input Connections: Manual logic switches attached to enable input pins (r1 inr_1\text{ in}, r2 inr_2\text{ in}, r0 inr_0\text{ in}) to simulate control unit assertions.

  • Step-by-Step Manual Operation:

    1. Set data bus input switch to binary 1012101_2 (decimal value 55).

    2. Assert control pin r1 in=1r_1\text{ in} = 1: Clock pulse loads value 55 into register r1r_1

    3. Set data bus input switch to binary 0112011_2 (decimal value 33).

    4. Assert control pin r2 in=1r_2\text{ in} = 1: Clock pulse loads value 33 into register r2r_2

    5. ALU processes inputs 55 and 33, outputting binary sum 100021000_2 (decimal value 88).

    6. Assert control pin r0 in=1r_0\text{ in} = 1: Clock pulse captures and stores the ALU sum 88 into destination register r0r_0

    7. Demonstrates how controlled assertions direct precise data movement step-by-step through the physical data path.