Computer Design Introduction - HPC 2024/25

Computer Evolution

  • The first general-purpose computer was created in the late 1940s.
  • A personal computer that costs around $500 today has roughly the same performance and memory as a machine that cost $1 million in 1985.
  • This evolution is due to:
    • Advances in semiconductor technology
    • Innovations in computer design
    • Advances in software

Chip-Design History

  • Evolution of chip design from 1970 to today:
    • 1970: 2,000 transistors, simple design entry.
    • 1980: 30,000 transistors, design complexity increases, Design Rule Checking introduced.
    • 1990: 1,000,000 transistors, P&R (Placement and Routing), Timing closure adopted, Logic Synthesis introduced.
    • 2000: 40,000,000 transistors, IP-Based Design introduced, Design platform and collaboration become important.
    • 2015-Today: >10,000,000,000 transistors, AI-Based Design.

Annual Size of the Global Datasphere

  • Data is growing exponentially.
  • zettaByte (ZB) -> 102110^{21}

Microprocessor Performance Growth

  • The annual increase in computer performance was about 25-30% during the 1970s, when mainframes and minicomputers dominated.
  • The annual increase raised to more than 50% for the RISC architectures in the 1980s.
  • Starting from 2002, the yearly processor performance increase dropped to about 20% due to:
    • Power issues
    • Lower instruction-level parallelism
    • Unchanged memory latency
  • Since 2004, major industries have shifted from the race for high-performance single-processor projects, and embrace multicore CPUs.
  • Currently at approximately 12% improvement or doubling every 8 years, mainly due to limits on instruction parallelism.
  • Recent performance improvement is around 3.5%, doubling every 20 years, which raises the question: Is this the end of Moore’s Law?
  • This growth results from advancements in semiconductor technology, innovations in computer design, and software improvements.
  • CPU performance is positively impacted by architecture innovation by a factor of approximately 15.

Trends in GPU/CPU Price-Performance

  • Referenced Marius Hobbhahn and Tamay Besiroglu (2022) for trends in GPU price performance.

Microprocessor Architecture Examples

  • 8086
    • Execution Unit (EU)
    • Bus Interface Unit (BIU)
    • Instruction Queue (IQ)
  • Core 2

Trend in Microprocessor Market

  • Major players (Intel, AMD, IBM, ARM) are investing in multiprocessor single-chip systems (multicore devices) rather than faster processors.

The Computer Market

  • Split into 5 different areas:
    • Personal Mobile Device (PMD)
    • Desktop computing
    • Servers
    • Supercomputers / Warehouse Scale Computers (WSC)
    • Embedded computers

Personal Mobile Device (PMD)

  • Includes smartphones and tablets.
  • Emphasis on energy efficiency and real-time applications.
  • System price: $100 - $1,000
  • Microprocessor price: $10 - $100

Desktop Computing

  • Covers PCs to workstations.
  • Main target is to optimize the price-performance ratio.
  • System price: $300 - $2,500
  • Microprocessor price: $50 - $500

Servers

  • Provide larger-scale and more reliable computing services.
  • Main parameters are availability, scalability, and throughput.
  • System price: $5,000 - $10,000,000
  • Microprocessor price: $200 - $2,000

Supercomputers / Warehouse-Scale Computers (WSC)

  • High-Performance Computers (HPCs) are designed to compute a vast number of user applications as fast as possible.
  • Emphasis on availability, price-performance, and power consumption: floating-point performance and fast internal networks.
  • System price: $100,000 - $200,000,000
  • Microprocessor price: $50 - $250
  • Difference between HPC and conventional computer is its organization, interconnectivity, scale of electronic and software components.
  • Organization:
    • Global Interconnection Network
    • Accelerator Core Array
    • Scratch pad Memory
    • Memory Banks
    • Multicore Sockets
    • NIC
  • Examples of HPC Performance:
    • FRONTIER: 1.2 EFLOP/s
    • El Capitan: 1.7 EFLOP/s

Embedded Computers

  • Fastest-growing portion of the computer market.
  • Covers all special-purpose computer-based applications (microwaves, coffee machines, automotive, videogames).
  • Microprocessors vary from cheap low-end 8-bit processors to efficient high-end processors, but usually do not run third-party software.
  • System price: $10 - $100,000
  • Microprocessor price: $0.01 - $100
  • Special requirements:
    • Real-time performance requirements
    • Memory minimization
    • Power consumption minimization
    • Reliability constraints

Classes of Parallelism

  • Data-level Parallelism (DLP): Operating on many data items simultaneously.
  • Task-level Parallelism (TaskLP): Different independent tasks operating concurrently.

Parallel Architectures

  • Instruction-level Parallelism (ILP): Modestly exploits Data-level Parallelism.
  • Vector Architectures and Graphic processor unit (GPUs): Exploits Data-level Parallelism.
  • Thread-level Parallelism (TLP): Exploits Data-level Parallelism and Task-level Parallelism.
  • Request-level Parallelism (RLP): Exploits parallelism among decoupled tasks.

Designing a Computer

  • Involves choosing important attributes and designing a machine to:
    • Maximize performance
    • Match cost and power constraints
  • Computer architect must consider:
    • Functional requirements
    • Price
    • Power
    • Performance
    • Dependability
  • Average performance P<em>avgP<em>{avg}: P</em>avg=e×S×aR×µ(E)P</em>{avg} = e \times S \times \frac{a}{R} \times µ(E)
    • e: efficiency
    • S: scaling
    • a: availability (depends on Reliability)
    • µ: rate of completed instructions as a function of Power E

Computer Architecture

  • Includes three aspects of computer design:
    • Instruction set architecture
    • Organization
    • Hardware

Moore's Law

  • The number of devices (transistors) that can be integrated into a single chip doubles every 18/24 months.

IC Manufacturing Cost

  • Impacted by yield, i.e., the percentage of products that pass the test phase.
  • The production process undergoes an evolution that normally leads to an improvement in yield (learning curve).
  • When yield increases, the cost decreases.
  • More than 50% of manufacturing cost is due to validation and testing procedures.

Yield Behavior

  • Yield changes as the time changes

Power Consumption

  • Continuous increase in system complexity and device integration leads to problems with power consumption.
  • Critical under two aspects:
    • Power (static and dynamic)
    • Energy (mainly for portable devices)

Power

  • Until now it has been dominated by dynamic power, i.e., that consumed by each transistor when switching between different states.
  • Dynamic power for each transistor:
    Powerdynamic=12×capacitive load×voltage2×frequencyPower_{dynamic} = \frac{1}{2} \times capacitive \ load \times voltage^2 \times frequency
  • Static power:
    Powerstatic=V×IPower_{static}= V \times I (25% of total power consumption)
  • Voltage continuously dropped in recent years to manage power.

Energy

  • Given by:
    Energydynamic=capacitive load×voltage2Energy_{dynamic} = capacitive \ load \times voltage^2
  • Mainly of interest for mobile devices.

Dependability

  • The quality of the system to deliver a correct service.
  • Traditionally very high, but can be lowered by:
    • Bugs in the design of the hardware
    • Bugs in the software
    • Defects in the hardware (introduced by the manufacturing process)
    • Faults happening during the product operation

Safety-Critical Applications

  • In the past:
    • Space
    • Avionics
    • Nuclear plants control
  • More recently:
    • Rail-road traffic control
    • Automotive
    • Biomedical
    • Telecommunications

Importance of Dependability

  • In several areas, it is crucial to guarantee that the system matches the dependability constraints, e.g., in terms of probability of behaving as expected for long periods.

Dependability Evaluation

  • Often measured using:
    • Mean Time To Failure (MTTF) or Failures In Time (FIT), which is its reciprocal.
    • 1 FIT = 1 failure in one billion hours
    • Mean Time Between Failures (MTBF)
    • Mean Time To Repair (MTTR)
  • The three measures are related by:
    MTBF=MTTF+MTTRMTBF = MTTF + MTTR
  • Availability is the probability that a system works correctly at a generic time instant.

Computer Performance

  • Measured by counting the number of operations executed or the time needed to execute them.
  • User point of view: performance = response time (time between start and completion of an operation).
  • System manager point of view: performance = throughput (total amount of work done in the time unit).

Time

  • Needs to be considered for performance computation:
    • Elapsed time
    • CPU time
      • user CPU time
      • system CPU time
  • UNIX provides all of them through the time command.

Performance Evaluation

  • Performed by letting the computer execute applications and observing its behavior.
  • The choice of applications impacts the performance.
  • Ideally, use the mix of applications the user will run as a workload.
  • Benchmarks are selected to mimic real cases.

Program Benchmarks

  • Contain several different types of programs.
  • Real programs (e.g., C compilers, text processors, special-purpose tools), possibly modified.
  • Kernels (e.g., Livermore Loops, Linpack).
  • Toy benchmarks (e.g., Quicksort, Sieve of Eratostenes).
  • Synthetic benchmarks (e.g., Whetstone, Dhrystone).

SPEC Evolution

  • Standard Performance Evaluation Corporation.

MiBench Benchmarks

  • Benchmarks for embedded systems that are not HPC.

HPC Benchmarks

  • Help decide the size and type of supercomputer.
  • Estimate the performance of a user application in an HPC.
  • Compare new technologies against mature ones.
  • Examples:
    • HPL: Dense linear algebra, estimates system’s effective flops.
    • STREAM: Synthetic, estimates suitable memory Bandwidth (GB/s).
    • Random Access: Synthetic, estimates system’s rate of integer random updates of memory.
    • HPCG: Sparse linear algebra, estimates System’s effective flops different than HPL.
    • SPEC CPU 2006: Varies, estimates system’s effective processor, memory and compiler performance.

Reproducibility

  • Information about execution times on benchmarks should allow reproducibility.
  • Report detailed information about:
    • hardware (system configuration)
    • software (OS, compiler, program)
    • program input

Comparing and Summarizing Performance

  • Problem 1: I know the performance of one machine on a set of programs: which is its global performance?
  • Problem 2: I know the performance of two machines on the same set of programs: which is their relative performance?
  • A number of metrics have been proposed.

Total Execution Time

  • Adopt a reference machine (e.g., VAX-11/780) and execution times are normalized with respect to it.
    Time<em>i=∑</em>i=1nTimeiTime<em>i = \sum</em>{i=1}^{n} Time_i

Arithmetic Mean

  • Arithmetic mean:
    1n∑<em>i=1nTime</em>i\frac{1}{n} \sum<em>{i=1}^{n} Time</em>i

Weighted Mean

  • Weighted arithmetic mean:
    ∑<em>i=1nWeight</em>i×Timei\sum<em>{i=1}^{n} Weight</em>i \times Time_i

Suggested Solution

  • Measure a real workload and weight the programs according to their frequency of execution.
  • Program inputs should be carefully specified.

Guidelines and Principles for Computer Design

  • Amdahl’s law
  • CPU performance equation

Amdahl’s Law: Preliminaries

  • The speedup resulting from an enhancement depends on two factors:
    • fractionenhanced: the fraction of the computation time that takes advantage of the enhancement
    • speedupenhanced: the size of the enhancement on the parts it affects.
  • Speedup:
    speedup=performance with enhancementperformance without enhancementspeedup = \frac{performance \ with \ enhancement}{performance \ without \ enhancement}

Amdahl’s Law: Equations

  • Execution time:
    execution time<em>new=execution time</em>old×((1−fraction<em>enhanced)+fraction</em>enhancedspeedupenhanced)execution \ time<em>{new} = execution \ time</em>{old} \times ((1 - fraction<em>{enhanced}) + \frac{fraction</em>{enhanced}}{speedup_{enhanced}})
  • Overall speedup:
    speedup<em>overall=execution time</em>oldexecution time<em>new=1(1−fraction</em>enhanced)+fraction<em>enhancedspeedup</em>enhancedspeedup<em>{overall} = \frac{execution \ time</em>{old}}{execution \ time<em>{new}} = \frac{1}{(1 - fraction</em>{enhanced}) + \frac{fraction<em>{enhanced}}{speedup</em>{enhanced}}}

Amdahl’s Law: Example

  • An enhancement makes one machine 10 times faster for 40% of the programs the machine runs.
  • The overall speedup is 1.56

Amdahl’s Law: Choosing Between Two Solutions

  • Two solutions are available for increasing the floating-point performance of one machine.
    • Solution 1: Increasing by 10 the performance of square root operations (responsible for 20% of the execution time) by adding specialized hardware.
    • Solution 2: Increasing by 2 the performance of all the floating-point operations (responsible for 50% of the execution time).

Amdahl’s Law: applications to solutions

  • Solution 1 overall speedup: 1.22
  • Solution 2 overall speedup: 1.33

Measuring Time

  • Measure the time required to execute a program
  • Possible approaches:
    • by observing the real system
    • by simulation
    • by applying the CPU performance equation

The CPU Performance Equation

  • CPU time equation:
     CPU time=∑<em>i=1n(CPI</em>i×ICi)×Clock cycle time\ CPU \ time = \sum<em>{i=1}^{n} (CPI</em>i \times IC_i) \times Clock \ cycle \ time
  • Where:
    • CPIiCPI_i is the number of clock cycles required by instruction i
    • ICiIC_i is the number of times instruction i is executed in the program
    • Clock cycle time is the inverse of clock frequency
  • CPI depends on hardware organization and instruction set architecture.
  • Instruction Count depends on instruction set architecture and compiler technology.
  • Clock cycle time depends on the hardware technology and organization.

The CPU Performance Equation: limitations

  • In pipelined processors, CPIi may vary for a given instruction, depending on different parameters.
  • Instructions executed before and after
  • Memory system behavior (e.g., cache miss or hit)
  • Evaluating the execution time analytically becomes much harder.