1/95
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
PROGRAM EXECUTION IS CARRIED OUT AS FOLLOWS:
The CPU transfers instruction and when necessary their input data (operands) from the main memory to registers in the CPU.
The CPU executes the instruction in their stored sequence except when the execution sequence is explicitly altered by a branch instruction.
When necessary, the CPU transfers output data (results) from the CPU registers to main memory.
CPU-IO DEVICES COMMUNICATION APPROACHES:
PROGRAMS EXECUTED BY A GENERAL-PURPOSE COMPUTER:
ALL INSTRUCTIONS REQUIRE TWO MAJOR STEPS:
A fetch step during which a new instruction is read from the external memory M.
An execute step during which the operations specified by the instructions are executed.
TWO ESSENTIAL MEMORY-ADDRESSING INSTRUCTIONS ARE:
MOST RECENT CPUS CONTAIN THE FOLLOWING EXTENSIONS WHICH SIGNIFICANTLY IMPROVE THEIR PERFORMANCE AND EASE OF PROGRAMMING.
Multipurpose register set for storing data and addresses - replaces the accumulator AC and the auxiliary register DR and AR of our basic CPU. The set of general registers is now usually referred to as register files.
Additional data, instruction, and addresses types - most CPUs have instructions to handle data and addresses with several different word sizes and formats. Call and return instructions also simplify program design.
Register to indicate computation status - A status register indicates infrequent or exceptional conditions resulting from the instruction execution. It also indicates the user and supervisor states.
Program control stack - various special registers and instructions facilitate the transfer control among programs due to procedure calling or external interrupts. A CPU address register called a sack pointer automatically keeps track of the stack's entry point.
TWO MAIN NUMBER FORMATS:
Fixed-Point - It takes the form babBbc…bK where each bi is 0 or 1 and a binary point is present in some fixed but implicit position. This allows limited range of values and have relatively simple hardware requirements.
Floating-Point - it is consist of a pair of fixed-point numbers M,E which denote the number M x BE, where B is a predetermined base. This allow a much larger range of values but require either a costly processing hardware or lengthy software implementation.[cite: 33]
TWO BASIC BYTE STORAGE METHODS:
Big endian - the most significant byte of a word is assigned to the lowest address and the least significant byte is assigned to the highest address.
Little-endian - the lowest address is assigned to the least significant byte and the highest address is assigned to the most significant byte.
ADVANTAGES OF TAGS:
Determine the type of operand
Tag inspection permits the hardware to check for software errors, such as an attempt to add operands whose types are incompatible.
DISADVANTAGES OF TAGS:
Increase the memory size
Add to the system hardware costs without increasing computing performance.
IN SELECTING A NUMBER REPRESENTATION TO BE USED IN A COMPUTER, THE FOLLOWING FACTORS SHOULD BE TAKEN INTO ACCOUNT:
The number types to be represented; for example, integers or real numbers
The range of values (number magnitudes) likely to be encountered
The precision of the numbers, which refers to the maximum accuracy of the representation
The cost of the hardware required to store and process the numbers.
The primary advantage of the complement codes is that
subtraction can be performed by logical complementation and addition only.
ADVANTAGE OF THE EXCESS-THREE CODE
It may be processed using the same logic used for binary codes.
ADVANTAGES OF DECIMAL CODES:
Ease of conversion between the internal computer representation that allows only symbols 0, 1
External representation using the 10 decimal symbols 0,1,2…9.
DISADVANTAGES OF DECIMAL CODES:
THE REPRESENTATION OF ZERO POSES SOME SPECIAL PROBLEM:
The mantissa must, of course, be zero, but the exponent can have any value, since 0 x BE = 0 for all values of E.
The desirability of representing zero by a sequence of 0-bits only.
RISC formats
serve to reduce both program length and what has been called the semantic gap between the user and the computer languages.
COMPLEX INSTRUCTIONS LEAD TO SEVERAL DIFFICULTIES, WHICH RISCS WITH THEIR SMALLER AND STREAMLINED INSTRUCTION SETS ATTEMPT TO MINIMIZE
The purpose of an address field is to
point to the current value V(X) of some operand X used by an instruction.
THE ADDRESSING MODE OF X AFFECTS THE FOLLOWING ISSUES:
MAIN DRAWBACK OF RELATIVE ADDRESSING:
Are the extra logic circuits and processing time needed to compute addresses.
THE REQUIREMENTS TO BE SATISFIED BY AN INSTRUCTION SET CAN BE STATED IN THE FOLLOWING GENERAL, BUT RATHER IMPRECISE, TERMS:
It should be complete in the sense that we should be able to construct a machine-language program to evaluate any function that is computable using a reasonable amount of memory space.
It should be efficient in that frequently required functions can be performed rapidly using relatively few instructions.
It should be regular in that the instruction set should contain expected opcodes and addressing modes
To reduce both hardware and software design costs, the instructions may be required to be compatible with those of existing machine.
Instructions are conveniently divided into following five types:
Data transfer instructions - which copy information from one location to another either in the processor's internal register set or in the external main memory.
Arithmetic instructions - which perform operations on numerical data.
Logical instruction - which include Boolean and other nonnumerical operations.
Program-control instructions - such as branch instructions, which change the sequence in which programs are executed.
Input-output (IO) instructions, which cause information to be transferred between the processor or its main memory and external IO devices.
THE MAJOR ATTRIBUTES OF RISCS
Relatively few instructions types and addressing modes
Fixed and easily decoded instruction formats
Fast, single-cycle instruction execution
Hardwired rather than microprogrammed control
Memory access limited mainly to load and store instructions
Use of compilers to optimize object-code performance.
INSTRUCTION CAN BE GROUPED INTO SEVERAL MAJOR TYPES:
TWO USEFUL TOOLS FOR SIMPLIFYING PROGRAM DESIGN BY ALLOWING GROUP OF INSTRUCTIONS TO BE TREATED AS SINGLE ENTITIES:
Macros- it is defined by placing a portion of assembly-language code between appropriate directives.
Subroutines - is also a sequence of instructions that can invoked by name, much like a single (macro) instructions. Unlike a macro, a subroutine definition is assembled into object code (CALL, JMP and RET)
EACH OPERAND SPECIFICATION IS DIVIDED INTO TWO PARTS:
An address field that points to the location of the first word of the operand
A length field L that indicates the number of words in the operand.
HIGH LEVEL VIEW OF A SERIAL ADDER THAT HAS D F/F AS THE CARRY STORE. ONE SUM BIT AND CARRY IS GENERATED PER CLOCK CYCLE.
Parallel Adders - in one clock cycle add all bits of two n-bit numbers as well as an external carry-in signal.
Ripple-carry adder - type of parallel adder. connecting full adders one full adders to generate 1 on its carry signal.
Subtracters
Subtraction is relatively simple with two's complement code because negation is very easy to implement. Adding -X to Y is equivalent to subtracting X from Y, so the ability to add negative numbers implies the ability to do subtraction.
Overflow
When the result of an arithmetic operation exceeds the standard word size n, overflow occurs.[cite: 33]
Carry-lookahead adders
A high-speed adder Compute the input carry needed by stage /directly from carry-like signals obtained from all the preceding stages.[cite: 33]
TWO AUXILIARY SIGNALS FOR CARRY-LOOKAHEAD ADDER
Multiplication
usually implemented by some form of addition.
TWO MULTIPLICATION ALGORITHMS FOR TWOS COMPLEMENT NUMBERS
Combinational array multiplier
can multiply large scale of numbers[cite: 33]
SEVERAL DIVISION DIFFICULTIES
Division by repeated multiplication
division is performed efficiently and low cost.
TO SIMPLIFY THE DISCUSSION, WE MAKE THE FOLLOWING REALISTIC ASSUMPTIONS:
XM is an nM-bit binary (twos-complement or sign magnitude) fraction.
XE is an ne-bit integer in excess-2nE-1 code, implying an exponent bias of 2nE-1.
B = 2
FLOATING-POINT ADDITION AND SUBTRACTION HAVE THREE MAIN STEPS:
Compute YE-Xe a fixed-point subtraction.
Shift XM by YE- XE places to the right to form XM2XE-YE
Compute XM2XE-YE + YM a fixed-point addition or subtraction.
SEVERAL MINOR PROBLEMS ARE ASSOCIATED WITH EXPONENT BIASING
If biased exponent are added or subtracted using fixed-point arithmetic in the course of a floating-point calculation, the resulting exponent is doublly biased and must be corrected by subtracting the bias.
Another problem arises from the all-0 representation usually required of zero. If X x Y is computed as (XM x YM) x 2XE+ YE and either XM or YM is zero, the resulting product has an all 0-mantissa but may not have an all-0 exponent.
Overflow and Underflow - A floating point operation causes overflow if the result is too large or too small to be represented. However, the exponent overflows or underflows, an error signal indicating floating-point overflow or underflow is generated.
Guard Bits - to preserve accuracy during floating point calculations, one or more extra bits called guard bits are temporarily attached to the right end of the mantissa.
A COPROCESSOR INSTRUCTION TYPICALLY CONTAINS THE FOLLOWING THREE FIELDS:
An opcode Fo that distinguishes coprocessor instructions from other CPU instructions
The address Fi of the particular coprocessor to be used if several coprocessors are allowed
The type F2 of the particular operation to be executed by the coprocessor.
Drawback of Coprocessor
Unlike the CPU, it does not know the contents of the registers defining the current memory addressing mode.
Pipelining
is a general technique for increasing processor throughput without requiring large amounts of extra hardware. It is applied to the design of the complex datapath units such as multipliers and floating-point adders It is also used to improve the overall throughput of an instruction set processor.
Stages or segments
a pipeline processor consist of a sequence of m data-processing circuits, which collectively perform a single operation on a stream of data operands passing through them.
Advantage of pipeline
An m-stage pipeline can simultaneously process up to m independent sets of data operands[cite: 33]
T
pipeline's clock period
MT
Delay or latency of the pipeline[cite: 33]
1/T
Pipeline's throughput[cite: 33]
Latency For a non-pipelined processor:
NmT[cite: 33]
Latency For a pipelined processor:
[m + (N - 1)]T Where: N = number of Tasks, m = number of stages, T = pipeline's clock period[cite: 33]
ADDITION OF TWO NORMALIZED FLOATING-POINT NUMBERS X AND Y CAN BE IMPLEMENTED USING FOUR-STEP SEQUENCE:
Feedback:
THE MAJOR CHARACTERISTIC OF A SYSTOLIC ARRAY CAN BE DEDUCED FROM THE PRECEDING EXAMPLE
It provides a high degree of parallelism by processing many sets of operands concurrently.
Partially processed data sets flow synchronously through the array in pipeline fashion, but possibly in several directions at once, with complete results eventually appearing at the array boundary.
The use of uniform cells and interconnection simplifies implementation.
The control of the array is simple, since all cells perform the same operations: however care must be taken to supply the data in the correct sequences for the operation being implemented.
If the X and Y matrices are generated in real time, it is unnecessary to store them before computing X x Y, as with most sequential or parallel processing techniques.
The amount of hardware needed to implement a systolic array.
SEPARATE A DIGITAL SYSTEM INTO TWO PARTS:
Datapath - is a network of functional and storage units capable of performing certain (micro) operations on data words
Control Unit - selects the functions to be performed at specific times and route the data through appropriate parts of the Datapath unit. Logically reconfigures the datapath to implement some specified instructions or program.
Multicycle operations
Single-cycle execution is a central goal of RISC design.
Single precision floating point
4 bytes[cite: 33]
Double-precision floating point
8 bytes[cite: 33]
Microprogram
an associated set of microinstructions in digital computers[cite: 33]
What are Microinstructions?
control the CPU at a very fundamental level of hardware circuitry
IMPLEMENTATION METHOD Two general approaches control unit design have evolved:
Hardwired - views the controller as a sequential logic circuit or fsm that generates specific sequences of control signals in response to externally supplied instructions. - designed with the usual goals of minimizing the number of components used and maximizing the speed of operation
Microprogrammed control unit - built around a storage unit called control memory, where all the control signals are stored in a program-like format resembling
Control Memory
stores set of microprograms designed to implement or emulate the behavior of the given instruction set.[cite: 33]
Microprogramming
makes control unit design more systematic by organizing signals into formatted words (microinstructions).
ON THE NEGATIVE SIDE, MICROPROGRAMMED CONTROL UNIT:
More costly to manufacture than hardwired due to the presence of control memory and its access circuitry.
Microprogrammed also tend to be slower because of the extra time required to fetch microinstructions from control memory. Hardwired control units use RISC, the reason it has fast instruction set.
DESIGN METHODS
STATE TABLES
Classical Method
Construct a P-Row state table that defines the desired input-output behavior
Select the minimum number p of D-type flip-flop and assign a p-bit binary code to each state
Design a combinational circuit C that generates the primary output signals {zi} and secondary outputs {Di} that must be applied to the flip-flops[cite: 33] Defined states: S0 = 0 0 - Begin, S1 = 0 1 - Swap, S2 = 1 0 - Sub, S3 = 1 1 - End
One-hot method
Binary state assignment always contains a single 1 - the "hot" bit - while all the remaining bits are 0.
1. Construct a P-row state table that defines the desired input-output behavior
2. Associate a separate D-type flip-flop Di with each state Si and assign the P-bit onehot binary code to each state.
3. Design a combinational circuit that generates the primary and secondary output signals {Di} and {zk} respectively. Di+ is defined by the logic equation.
Defined states: S0 = 0 0 0 1 - Begin, S1 = 0 0 1 0 - Swap, S2 = 0 1 0 0 - Sub, S3 = 1 0 0 0 - End
CPU Control Unit
DPU - datapath unit designed to execute the set of 10 basic single-address[cite: 33] PCU - program control unit is responsible for managing the control signals linking PCU to the DPU, as well as the control signals between the CPU and external memory M.[cite: 33]
RISC processors
are usually designed so that all instruction execution times are equalized to one CPU clock period Tc in length, making the cycles associated with the registertransfer operations into subcycles of Tc.[cite: 33]
Microprogramming is
a method of control-unit design in which the control signal selection and sequencing information is stored in a ROM or RAM called control memory, CM.
Microinstructions
activates the control signals at any time which is fetched from CM in much the same way an instruction is fetched from main memory.
Microprogram
is a set of related microinstructions.[cite: 33]
Emulator
is the set of microprograms that interpret a particular instruction set or machine language.[cite: 33]
Microassembler
is necessary to translate microprograms into executable programs that can be stored in the control memory.[cite: 33]
MICROINSTRUCTION'S TWO PARTS:
Control fields that specify the control signals to be activated
Address fields that contains the address in the CM of the next microinstruction to be executed.
CMAR
control memory address register[cite: 33]
MICROINSTRUCTION LENGTH IS DETERMINED BY THREE FACTORS:
The maximum number of simultaneous microoperations that must be specified, that is, the degree of parallelism required at the microoperation level
The way in which the control information is represented or encoded.
The way in which the next microinstruction address is specified.
WCM
Writable control memory allows us to change a processor's instruction set by changing the microprograms that interpret the instruction set.
Dynamically microprogrammable
if the control memory contents can be altered under program control, e.g. WCM[cite: 33]
Parallelism in microinstructions
Microinstruction formats take advantage of the fact that, at the microprogramming level, many operations can be performed in parallel.[cite: 33]
HORIZONTAL MICROINSTRUCTIONS:
Long formats
Ability to express a high degree of parallelism
Little encoding of the control information; Allows no encoding of control information. Specifies many microoperation.
VERTICAL MICROINSTRUCTIONS:
Microinstruction addressing
uPC, microprogram counter is the primary source of microinstruction address and can also be used as CMAR[cite: 33]
Microoperation timing
A single clock signal synchronizes the control signals, and its period can be the same as the microinstruction cycle period; this mode of control has been termed monophase.
Design of a typical microprogrammed control unit using the microinstruction format:
A condition-select field specifies the external condition to be tested in the case of conditional branch microinstructions
An address field contains the next-address field to be used when a branch condition is satisfied. A microprogram counter uPC provides the next microinstruction address when no branching is needed.
The rest of the microinstruction specifies in encoded or unencoded format the control signals that are activated to perform the desired microoperations.
THIS OPERATION CAN BE PERFORMED IN SEVERAL PHASES; THE FOLLOWING FOUR-PHASE INTERPRETATION IS REPRESENTATIVE:
WE USE THE MICROINSTRUCTION WHICH HAS THREE PARTS ARRANGED AS FOLLOWS:
A condition-selct field specifies the external condition to be tested in the case of conditional branch microinstruction.
An address field contains the next-address field to be used when a branch condition is satisfied.
The rest of the microinstruction specifies in encoded or unencoded format the control signals that are activated to perform the desired microoperations.
Instruction Pipeline
is a multifunction, reconfigurable pipeline designed to speed up a computer's performance by efficiently overlapping the processing of instructions.[cite: 33]
SIMPLEST INSTRUCTION PIPELINE:
FOUR-STAGE PIPELINE:
IF: instruction fetching and decoding using the I-cache.
OL: operand loading from the D-cache to RF
EX: data processing using the ALU and RF
OS: operand storing to the D-cache from RF
THE FACTORS THAT THE PCU OF A SUPERSCALAR COMPUTER MACHINE THAT NEEDS TO BE TAKEN INTO ACCOUNT FOR THE SAID PART OF THE MACHINE TO DO ITS TASK:
Instruction type. For example, a floating-point add instruction has to be issued to a floating-point E-unit and not to an integer E-unit[cite: 33]
E-unit availability. An instruction can be issued to a pipelined E-unit only if no collisions will result, as determined by the pipeline's reservation table[cite: 33]
Data dependencies. To avoid conflicting use of registers, data-dependency constraints among the operands of the active instructions must be satisfied[cite: 33]
Control dependencies. To maintain high performance levels, techniques are needed to be reduce the impact of branch instructions on pipeline efficiency
Program order. Instructions must eventually produce results in the order specified by the program being executed. The results may be computed out-of-order[cite: 33]
ENUMERATE (IN BULLET FORM) AND DISCUSS BRIEFLY THE FACTORS THAT DETERMINE THE LENGTH OF MICROINSTRUCTIONS
The maximum number of simultaneous microoperations that must be specified.
The way in which the control information is represented or encoded.
The way in which the next microinstruction address is specified.
IT IS QUITE FEASIBLE TO MANUFACTURE AN ENTIRE SEQUENTIAL ALU FOR FIXED-POINT M-BIT NUMBER ON A SINGLE IC CHIP. ENUMERATE AND DISCUSS BRIEFLY THE WAYS ON HOW THE ALU COULD EASILY BE DESIGNED FOR EXPANSION TO HANDLE OPERANDS OF SIZE N = KM, OR ANY WORD SIZE N > M.
Spatial Expansion - connect k copies of the m-bit ALU in the manner of a ripple-carry to form a single ALU capable of processing km-bit words directly. The resulting array-like circuit is said to be bit sliced because each component ALU concurrently processes a separate "slice" of mbits from each km-bit operand.
Temporal Expansion - use one copy of the m-bit ALU chip in the manner of a serial adder to perform an operation on km-bit words in k consecutive steps (clock cycles). In each step the ALU processes a separate m-bit slice of each operand. This processing is called multicycle or multiple-precision Processing.
Processor-memory communication: (a) without a cache and (b) with a cache
Figure[cite: 33]
Overview of CPU behavior.
Figure[cite: 33]
Four-stage floating-point adder pipeline.
Figure[cite: 33]