Variability/Graphing Data

Class Overview

  • Start of Class

    • Instructor greets students and wishes them a lovely weekend.

    • Acknowledgment of previous class discussions about variability.

    • Mention of uploaded PDFs for Chapter 3 and Chapter 4.

Types of Statistics

  • Descriptive Statistics

    • Define: Statistics that describe and summarize data.

    • Focus: Currently concentrating on descriptive statistics.

  • Inferential Statistics

    • Define: Statistics that allow us to infer or draw conclusions about a population based on a sample.

    • Current focus is on descriptive statistics rather than inferential.

Measures of Central Tendency

  • Definition: Indicates the average score of a dataset.

  • Includes the following elements:

    • Mean: Average score.

    • Median: Middle score when data is arranged in order.

    • Mode: Most frequently occurring score.

Measures of Variability

  • Overview: Measures how much scores differ from each other.

  • Evaluation of data differences through standard deviation.

    • Explains how different student scores are within a class.

  • Range: Difference between the highest and lowest score.

    • Calculate by subtracting the lowest score from the highest score.

  • Standard Deviation: Indicates on average how much scores differ.

    • If standard deviation is low, scores are close to the mean; higher standard deviation indicates greater variation in scores.

    • Calculation context:

    • For a test out of 100 points:

      • Standard deviation of 1% = scores are typically within 1 point of the mean.

      • Standard deviation of 10 = scores typically vary by 10 points from the mean.

      • Standard deviation of 20 reflects a much wider range of scores.

Calculation Practice

  • Standard Deviation Calculation

    • Example provided: Scores vary greatly between lower and upper limits.

    • In-class exercise to calculate standard deviation for two sets of scores.

  • Understanding Variability

    • Class discussion about which dataset has a higher or lower standard deviation by comparing visual distributions of data points.

Example Contextualization

  • Football Game Scores

    • Standard deviation context of scores during a game:

    • Example: Standard deviation of 14 in scores compared to standard deviation of 13.37 indicates variance.

    • Discussion on practical implications of these numbers in competitive contexts (touchdowns).

Deviation and Squaring

  • Deviations Explained

    • Calculation: Subtract scores from the mean to find deviations.

    • Reason for squaring deviations:

    • Ensures that average deviations do not equal zero; negatives won't cancel positives.

    • Allows for handling the magnitude of deviations without losing directionality (negative vs. positive).

Sample Size Penalty Concept

  • n - 1 Adjustment

    • Explanation of why we take one less in the denominator during standard deviation calculations: ensures it compensates smaller samples for representation issues.

  • Implications

    • Having a smaller sample may misrepresent the true population behavior; thus, adjusting the denominator helps mitigate this bias.

Variance and Its Relation to Standard Deviation

  • Fundamental Understanding

    • Variance: Square of standard deviation.

    • Calculation provided: Given standard deviation of 14, variance can be computed as 14.1² = 198.81.

    • Importance of variance in later calculations despite it being less intuitive than standard deviation.

Outlier Detection

  • Definition: Outlier as a data point that significantly deviates from other observations.

  • Identification Techniques

    • A common rule of thumb for identifying outliers involves checking for values that lie more than two standard deviations from the mean.

  • Application to Data Analysis

    • Discussion on using this method for analyzing football game scores to determine if games were outliers with respect to team performance.

Identifying Games Above/Below Standard Deviation

  • Practical Application

    • Mean and standard deviation are both calculated; identifying whether game scores fall above or below established cutoffs helps to detect anomalies.

  • Final Analysis

    • Discussion prompts regarding the implications of outlier analysis in understanding team performance during seasons.

Deeper Understanding of Data Variability

  • Comparing Classes A and B

    • Exploring how standard deviation impacts understanding of performance relative to the mean.

  • Differences between scores in classes indicate variability that standard deviation can reveal.

Histograms and Data Distribution

  • Definition and Importance

    • Histograms allow for visualization of score distributions along a continuous variable.

    • Explanation of axes and what they represent in terms of frequency distribution.

Measures of Central Tendency, Variability, and Skewness

  • Comparison of Different Distributions

  • Skewness Types

    • Positive Skew: Longer tail on the right.

    • Negative Skew: Longer tail on the left.

    • Example provided of understanding skew through illustrative analogies (Mr. Stokovich's foot analogy).

  • Ceiling and Floor Effects

    • Description of where data clusters function with respect to outliers in skewed distributions. Such as employee wages (floor effect) and test scores (ceiling effect).

Conclusion and Next Steps

  • Reminder of concepts discussed and anticipation for next class sessions focused on histograms further.