Computer Design Introduction - HPC 2024/25
Computer Evolution
- The first general-purpose computer was created in the late 1940s.
- A personal computer that costs around $500 today has roughly the same performance and memory as a machine that cost $1 million in 1985.
- This evolution is due to:
- Advances in semiconductor technology
- Innovations in computer design
- Advances in software
Chip-Design History
- Evolution of chip design from 1970 to today:
- 1970: 2,000 transistors, simple design entry.
- 1980: 30,000 transistors, design complexity increases, Design Rule Checking introduced.
- 1990: 1,000,000 transistors, P&R (Placement and Routing), Timing closure adopted, Logic Synthesis introduced.
- 2000: 40,000,000 transistors, IP-Based Design introduced, Design platform and collaboration become important.
- 2015-Today: >10,000,000,000 transistors, AI-Based Design.
Annual Size of the Global Datasphere
- Data is growing exponentially.
- zettaByte (ZB) -> 1021
- The annual increase in computer performance was about 25-30% during the 1970s, when mainframes and minicomputers dominated.
- The annual increase raised to more than 50% for the RISC architectures in the 1980s.
- Starting from 2002, the yearly processor performance increase dropped to about 20% due to:
- Power issues
- Lower instruction-level parallelism
- Unchanged memory latency
- Since 2004, major industries have shifted from the race for high-performance single-processor projects, and embrace multicore CPUs.
- Currently at approximately 12% improvement or doubling every 8 years, mainly due to limits on instruction parallelism.
- Recent performance improvement is around 3.5%, doubling every 20 years, which raises the question: Is this the end of Moore’s Law?
- This growth results from advancements in semiconductor technology, innovations in computer design, and software improvements.
- CPU performance is positively impacted by architecture innovation by a factor of approximately 15.
- Referenced Marius Hobbhahn and Tamay Besiroglu (2022) for trends in GPU price performance.
Microprocessor Architecture Examples
- 8086
- Execution Unit (EU)
- Bus Interface Unit (BIU)
- Instruction Queue (IQ)
- Core 2
Trend in Microprocessor Market
- Major players (Intel, AMD, IBM, ARM) are investing in multiprocessor single-chip systems (multicore devices) rather than faster processors.
The Computer Market
- Split into 5 different areas:
- Personal Mobile Device (PMD)
- Desktop computing
- Servers
- Supercomputers / Warehouse Scale Computers (WSC)
- Embedded computers
Personal Mobile Device (PMD)
- Includes smartphones and tablets.
- Emphasis on energy efficiency and real-time applications.
- System price: $100 - $1,000
- Microprocessor price: $10 - $100
Desktop Computing
- Covers PCs to workstations.
- Main target is to optimize the price-performance ratio.
- System price: $300 - $2,500
- Microprocessor price: $50 - $500
Servers
- Provide larger-scale and more reliable computing services.
- Main parameters are availability, scalability, and throughput.
- System price: $5,000 - $10,000,000
- Microprocessor price: $200 - $2,000
Supercomputers / Warehouse-Scale Computers (WSC)
- High-Performance Computers (HPCs) are designed to compute a vast number of user applications as fast as possible.
- Emphasis on availability, price-performance, and power consumption: floating-point performance and fast internal networks.
- System price: $100,000 - $200,000,000
- Microprocessor price: $50 - $250
- Difference between HPC and conventional computer is its organization, interconnectivity, scale of electronic and software components.
- Organization:
- Global Interconnection Network
- Accelerator Core Array
- Scratch pad Memory
- Memory Banks
- Multicore Sockets
- NIC
- Examples of HPC Performance:
- FRONTIER: 1.2 EFLOP/s
- El Capitan: 1.7 EFLOP/s
Embedded Computers
- Fastest-growing portion of the computer market.
- Covers all special-purpose computer-based applications (microwaves, coffee machines, automotive, videogames).
- Microprocessors vary from cheap low-end 8-bit processors to efficient high-end processors, but usually do not run third-party software.
- System price: $10 - $100,000
- Microprocessor price: $0.01 - $100
- Special requirements:
- Real-time performance requirements
- Memory minimization
- Power consumption minimization
- Reliability constraints
Classes of Parallelism
- Data-level Parallelism (DLP): Operating on many data items simultaneously.
- Task-level Parallelism (TaskLP): Different independent tasks operating concurrently.
Parallel Architectures
- Instruction-level Parallelism (ILP): Modestly exploits Data-level Parallelism.
- Vector Architectures and Graphic processor unit (GPUs): Exploits Data-level Parallelism.
- Thread-level Parallelism (TLP): Exploits Data-level Parallelism and Task-level Parallelism.
- Request-level Parallelism (RLP): Exploits parallelism among decoupled tasks.
Designing a Computer
- Involves choosing important attributes and designing a machine to:
- Maximize performance
- Match cost and power constraints
- Computer architect must consider:
- Functional requirements
- Price
- Power
- Performance
- Dependability
- Average performance P<em>avg:
P</em>avg=e×S×Ra×µ(E)
- e: efficiency
- S: scaling
- a: availability (depends on Reliability)
- µ: rate of completed instructions as a function of Power E
Computer Architecture
- Includes three aspects of computer design:
- Instruction set architecture
- Organization
- Hardware
Moore's Law
- The number of devices (transistors) that can be integrated into a single chip doubles every 18/24 months.
IC Manufacturing Cost
- Impacted by yield, i.e., the percentage of products that pass the test phase.
- The production process undergoes an evolution that normally leads to an improvement in yield (learning curve).
- When yield increases, the cost decreases.
- More than 50% of manufacturing cost is due to validation and testing procedures.
Yield Behavior
- Yield changes as the time changes
Power Consumption
- Continuous increase in system complexity and device integration leads to problems with power consumption.
- Critical under two aspects:
- Power (static and dynamic)
- Energy (mainly for portable devices)
Power
- Until now it has been dominated by dynamic power, i.e., that consumed by each transistor when switching between different states.
- Dynamic power for each transistor:
Powerdynamic=21×capacitive load×voltage2×frequency - Static power:
Powerstatic=V×I (25% of total power consumption) - Voltage continuously dropped in recent years to manage power.
Energy
- Given by:
Energydynamic=capacitive load×voltage2 - Mainly of interest for mobile devices.
Dependability
- The quality of the system to deliver a correct service.
- Traditionally very high, but can be lowered by:
- Bugs in the design of the hardware
- Bugs in the software
- Defects in the hardware (introduced by the manufacturing process)
- Faults happening during the product operation
Safety-Critical Applications
- In the past:
- Space
- Avionics
- Nuclear plants control
- More recently:
- Rail-road traffic control
- Automotive
- Biomedical
- Telecommunications
Importance of Dependability
- In several areas, it is crucial to guarantee that the system matches the dependability constraints, e.g., in terms of probability of behaving as expected for long periods.
Dependability Evaluation
- Often measured using:
- Mean Time To Failure (MTTF) or Failures In Time (FIT), which is its reciprocal.
- 1 FIT = 1 failure in one billion hours
- Mean Time Between Failures (MTBF)
- Mean Time To Repair (MTTR)
- The three measures are related by:
MTBF=MTTF+MTTR - Availability is the probability that a system works correctly at a generic time instant.
- Measured by counting the number of operations executed or the time needed to execute them.
- User point of view: performance = response time (time between start and completion of an operation).
- System manager point of view: performance = throughput (total amount of work done in the time unit).
Time
- Needs to be considered for performance computation:
- Elapsed time
- CPU time
- user CPU time
- system CPU time
- UNIX provides all of them through the time command.
- Performed by letting the computer execute applications and observing its behavior.
- The choice of applications impacts the performance.
- Ideally, use the mix of applications the user will run as a workload.
- Benchmarks are selected to mimic real cases.
Program Benchmarks
- Contain several different types of programs.
- Real programs (e.g., C compilers, text processors, special-purpose tools), possibly modified.
- Kernels (e.g., Livermore Loops, Linpack).
- Toy benchmarks (e.g., Quicksort, Sieve of Eratostenes).
- Synthetic benchmarks (e.g., Whetstone, Dhrystone).
SPEC Evolution
- Standard Performance Evaluation Corporation.
MiBench Benchmarks
- Benchmarks for embedded systems that are not HPC.
HPC Benchmarks
- Help decide the size and type of supercomputer.
- Estimate the performance of a user application in an HPC.
- Compare new technologies against mature ones.
- Examples:
- HPL: Dense linear algebra, estimates system’s effective flops.
- STREAM: Synthetic, estimates suitable memory Bandwidth (GB/s).
- Random Access: Synthetic, estimates system’s rate of integer random updates of memory.
- HPCG: Sparse linear algebra, estimates System’s effective flops different than HPL.
- SPEC CPU 2006: Varies, estimates system’s effective processor, memory and compiler performance.
Reproducibility
- Information about execution times on benchmarks should allow reproducibility.
- Report detailed information about:
- hardware (system configuration)
- software (OS, compiler, program)
- program input
- Problem 1: I know the performance of one machine on a set of programs: which is its global performance?
- Problem 2: I know the performance of two machines on the same set of programs: which is their relative performance?
- A number of metrics have been proposed.
Total Execution Time
- Adopt a reference machine (e.g., VAX-11/780) and execution times are normalized with respect to it.
Time<em>i=∑</em>i=1nTimei
Arithmetic Mean
- Arithmetic mean:
n1∑<em>i=1nTime</em>i
Weighted Mean
- Weighted arithmetic mean:
∑<em>i=1nWeight</em>i×Timei
Suggested Solution
- Measure a real workload and weight the programs according to their frequency of execution.
- Program inputs should be carefully specified.
Guidelines and Principles for Computer Design
- Amdahl’s law
- CPU performance equation
Amdahl’s Law: Preliminaries
- The speedup resulting from an enhancement depends on two factors:
- fractionenhanced: the fraction of the computation time that takes advantage of the enhancement
- speedupenhanced: the size of the enhancement on the parts it affects.
- Speedup:
speedup=performance without enhancementperformance with enhancement
Amdahl’s Law: Equations
- Execution time:
execution time<em>new=execution time</em>old×((1−fraction<em>enhanced)+speedupenhancedfraction</em>enhanced) - Overall speedup:
speedup<em>overall=execution time<em>newexecution time</em>old=(1−fraction</em>enhanced)+speedup</em>enhancedfraction<em>enhanced1
Amdahl’s Law: Example
- An enhancement makes one machine 10 times faster for 40% of the programs the machine runs.
- The overall speedup is 1.56
Amdahl’s Law: Choosing Between Two Solutions
- Two solutions are available for increasing the floating-point performance of one machine.
- Solution 1: Increasing by 10 the performance of square root operations (responsible for 20% of the execution time) by adding specialized hardware.
- Solution 2: Increasing by 2 the performance of all the floating-point operations (responsible for 50% of the execution time).
Amdahl’s Law: applications to solutions
- Solution 1 overall speedup: 1.22
- Solution 2 overall speedup: 1.33
Measuring Time
- Measure the time required to execute a program
- Possible approaches:
- by observing the real system
- by simulation
- by applying the CPU performance equation
- CPU time equation:
CPU time=∑<em>i=1n(CPI</em>i×ICi)×Clock cycle time - Where:
- CPIi is the number of clock cycles required by instruction i
- ICi is the number of times instruction i is executed in the program
- Clock cycle time is the inverse of clock frequency
- CPI depends on hardware organization and instruction set architecture.
- Instruction Count depends on instruction set architecture and compiler technology.
- Clock cycle time depends on the hardware technology and organization.
- In pipelined processors, CPIi may vary for a given instruction, depending on different parameters.
- Instructions executed before and after
- Memory system behavior (e.g., cache miss or hit)
- Evaluating the execution time analytically becomes much harder.