Computer Simulation Module 9: Output Data Analysis

Computer Simulation Module 9: Output Data Analysis

Dave Goldsman, Ph.D. Professor Stewart School of Industrial and Systems Engineering

Module Overview

  • Last Module: Discussed proper modeling techniques for input random variables (RVs) that drive a simulation.

  • This Module: Focuses on analyzing the output from the simulation.

    • Key Point: Output is rarely independent and identically distributed (i.i.d.).

Overview of the Module

  1. Introduction

  2. A Mathematical Interlude

  3. Finite-Horizon Analysis

  4. Finite-Horizon Extensions

  5. Simulation Initialization Issues

  6. Steady-State Analysis

  7. Properties of Batch Means

  8. Other Steady-State Methods

Introduction

Steps in a Simulation Study
  • Preliminary analysis of system

  • Model building

  • Verification & validation

  • Experimental design & simulation runs

  • Statistical analysis of output

  • Implementation

Why Worry about Output?

  • The input processes driving a simulation are random variables (e.g., interarrival times, service times, breakdown times).

  • Must regard the output from the simulation as random.

  • Simulation runs yield estimates of measures of system performance (e.g., mean customer waiting time), but these estimators are themselves random variables, thus subject to sampling error.

  • Conclusion: Sampling error must be accounted for when making valid inferences concerning system performance.

Measures of Interest

  • Means: (e.g., what is the mean customer waiting time?)

    • Caution: Means alone aren't enough (e.g., "If I have one foot in boiling water and one foot in freezing water, on average, I’m fine").

  • Variances: (e.g., how much is the waiting time expected to vary?)

  • Quantiles: (e.g., what’s the 99% quantile of line length in a queue?)

  • Success probabilities: (e.g., will my job be completed on time?)

  • Important: We would like point estimators and confidence intervals for the above measures.

Problem of Non-i.i.d. Output

  • Simulations seldom produce raw output that's independent and identically distributed (i.i.d.).

  • Example: Customer waiting times in a queueing system:

    • (1) Not independent; typically, they are serially correlated. If one customer at the post office waits long, the next is likely to do the same.

    • (2) Not identically distributed; customers arriving early in the morning have shorter waits than those just before closing time.

    • (3) Not normally distributed; generally, they are skewed right (never negative).

Implications of Non-i.i.d. Output

  • Traditional statistical techniques fail when analyzing simulation output.

  • Importance of Proper Statistical Analysis:

    • Improper analysis can invalidate all results.

    • Correct analysis can lead to significant applications.

    • Interesting research problems stem from these complexities.

Types of Simulations

To facilitate presentation, we identify two types of simulations concerning output analysis:

  1. Finite-Horizon (Terminating) Simulations:

    • Focus on short-term performance.

  2. Steady-State Simulations:

    • Focus on long-term performance.

Finite-Horizon Simulations

  • Definition: A finite-horizon simulation terminates at a specific time or upon occurrence of a specific event.

  • Examples include:

    • Mass transit system during rush hour

    • Distribution system over a month

    • Production system until a set of machines break down

    • Start-up phase of any system, either stationary or nonstationary

Steady-State Simulations

  • Purpose: Study long-run behavior of a system.

    • A performance measure is a steady-state parameter if it characterizes the equilibrium distribution of an output process.

  • Examples:

    • Continuously operating communication systems aiming to compute the average packet delay over the long term.

    • Distribution systems assessed over extended periods.

    • Usage of Markov chains.

    • Note: Steady-state simulations may be less viewed due to the absence of a transient state (i.e., “you’re always dead”).

Summary of the Lesson

  • The lesson provided an overview of simulation output data analysis, outlining issues and expectations.

  • Next Time: A brief mathematical excursion explaining the emphasis on robust analysis techniques and an exploration of finite-horizon simulation.

Mathematical Interlude

Lesson Overview
  • Last Lesson: Introduction to Output Analysis

  • This Lesson: Explore why classical analysis techniques are unsuitable.

  • Follow-up: Finite-horizon simulation analysis will be discussed subsequently.

Conceptual Visualization of Non-i.i.d. Observations

  • Discussing assumptions regarding identically distributed observations while acknowledging the presence of dependent observations will be fundamental in steady-state simulations.

Properties of the Sample Mean

  • First Property:
    E[X]=rac1nimesextsumofXi=extconstanttermE[X] = rac{1}{n} imes ext{sum of } X_i = ext{constant term}

    • The sample mean remains unbiased for true mean ( ext{mean} ).

  • Formula Transition:

    • Summation of terms leads us to a matrix of covariance leading to equation transitions indicating normalization under certain conditions.

Variance Implications in Non-i.i.d. Context

  • Referring to estimation of sample mean variance under dependent scenarios, linking the variance parameters to covariance.

  • Essential observation: The variance parameter heta2heta^2 manifests frequently across various analyses, particularly in queueing applications where the covariance is positive.

  • Key Warning: High dependence may cause the classical confidence interval (CI) to misbehave.

Example: Autoregressive Processes

  • Define first-order autoregressive process as an illustration for dependent RVs.

    • Situation where current value relates to past values introduces complexity in estimation.

Properties of the Sample Variance

  • Basic equation:
    racn2extParameterrelatedconditionsrac{n^2}{ ext{Parameter related conditions}}

  • Unbiased when conditions define identities in relations.

    • Challenges arise in dependency scenarios yielding suboptimal estimates for variance estimators.

Steps to find Expected Values

  • General strategies for estimating function outputs under sequential relations derived from theory-defined conditions.

Confidence Intervals and Their Usefulness

  • Classical CI methodologies neglecting dependencies can lead to significant errors, especially with non-i.i.d samples.

Summary of Mathematical Interlude

  • Competently built statistical foundations underpinning potential pitfalls in simulation output data analysis are essential.

  • Next time: Introduction to finite-horizon simulations to streamline error reduction practices.

Finite-Horizon Analysis

Lesson Overview
  • Last Discussion: Math issues pertinent to simulation output analysis

  • Current Focus: Navigational strategies for resolving issues in terminating simulations.

  • The Core Method: Independent replications.

Goal and Practical Application

  • Simulate a system of interest over a fixed horizon followed by output analysis.

  • Example criteria: Collection of waiting times for a defined number of customers.

Easiest Goal: Estimating Sample Mean

Understanding how expected values differ through time will help enhance forecast accuracy.

Methods of Replication and Statistical Independence

  • Notation clarifying means of deriving averages from independent simulations initialized correctly shows robust models for obtaining reliable output characterization.

Grand Sample Mean Definition

Establishing grand sample mean leads to formulating effective variance measures for statistical outcomes.

Central Limit Theorem Application

  • Using standard statistical principles, large sample sizes enforce independence, yielding CI for estimators.

Example in Finite-Horizon Context

  • Estimating wait times in a defined queue through independent simulation supports methodical data analysis for practical decision-making.

Summary of Finite-Horizon Analysis

  • Application of independent replications for generating confidence intervals characterized the essence of good practice.

Advancements in Finite-Horizon Extensions

Lesson Overview
  • Previous Lesson: Independent replications methodology.

  • Focus of Current Lesson: Enhancing CI specifications and obtaining new analytic measures, particularly for quantiles.

Strategies for Smaller Confidence Intervals

  • Increased replication for narrowing CI lengths:

  • Mathematical evaluation:
    a/2r1imesZa / 2^{r-1} imes Z

  • Conclusion: Reduction of CI size directly relates to increasing independent runs exponentially.

Introduction to Quantiles Definition

  • The concept of p-quantiles of random variables connects distributions via cumulative density functions (c.d.f).

Practical Example for Quantile Estimation

  • Assessing maximum wait time through methods yields deep insights into service efficiency at dynamic venues.

Recommended Approach to Choosing Parameters

  • Various strategies outlined for optimizing the analysis as depth increases in data acquired.

Summary of Finite-Horizon Extensions

  • Looked into acquiring narrower CIs and methodologies for quantiles.

Simulation Initialization Issues

Lesson Overview
  • Current Lesson: Addressing initial conditions in simulations contributing to bias.

  • Core Philosophy: Combining art and science to achieve modeling fidelity.

Initialization Challenges

  • Establish systematic processes to derive initial system states which heavily influence model outputs.

Detection of Initialization Bias

  • Procedures developed to identify initialization bias utilize both visual inspection and data transformation techniques.

Strategies for Managing Bias

  • Output truncation: Allow simulation to warm up before data retention to enhance accuracy.

  • Long Run Approach: Making extended runs to overcome initialization biases, though may be inefficient.

Summary of Initialization Issues

  • Identifying and addressing bias critical for understanding dynamic simulations.

Transition to Steady State Analysis

Overview of Steady-State Analysis
  • Building on the previous initialized practices leading into an analytical scope concerning long-term performance.

Interplay Between Batch Means and Varied Estimators

  • Establishing methodologies for steady-state estimator derivations and implications on output variance.

Properties of Batch Means

  • Validating estimators through direct relation to the established variance properties discussed previously, emphasizing unbiasing techniques.

Techniques Accessed for Returning Estimated Parameters

  • Optimal batch sizes and dimensional controls for result accuracy will hold continued emphasis.

Conclusive Findings from Steady-State Analysis

  • Necessary explorations into related methodologies enhance frameworks while evolving practice around simulation outputs.

Final Thoughts to Transition to Other Techniques

  • Observations of overlapping batch means will prominently showcase the next evolution in statistical analysis.

Next Module: Transition into Comparing Systems