Variability/Graphing Data
Class Overview
Start of Class
Instructor greets students and wishes them a lovely weekend.
Acknowledgment of previous class discussions about variability.
Mention of uploaded PDFs for Chapter 3 and Chapter 4.
Types of Statistics
Descriptive Statistics
Define: Statistics that describe and summarize data.
Focus: Currently concentrating on descriptive statistics.
Inferential Statistics
Define: Statistics that allow us to infer or draw conclusions about a population based on a sample.
Current focus is on descriptive statistics rather than inferential.
Measures of Central Tendency
Definition: Indicates the average score of a dataset.
Includes the following elements:
Mean: Average score.
Median: Middle score when data is arranged in order.
Mode: Most frequently occurring score.
Measures of Variability
Overview: Measures how much scores differ from each other.
Evaluation of data differences through standard deviation.
Explains how different student scores are within a class.
Range: Difference between the highest and lowest score.
Calculate by subtracting the lowest score from the highest score.
Standard Deviation: Indicates on average how much scores differ.
If standard deviation is low, scores are close to the mean; higher standard deviation indicates greater variation in scores.
Calculation context:
For a test out of 100 points:
Standard deviation of 1% = scores are typically within 1 point of the mean.
Standard deviation of 10 = scores typically vary by 10 points from the mean.
Standard deviation of 20 reflects a much wider range of scores.
Calculation Practice
Standard Deviation Calculation
Example provided: Scores vary greatly between lower and upper limits.
In-class exercise to calculate standard deviation for two sets of scores.
Understanding Variability
Class discussion about which dataset has a higher or lower standard deviation by comparing visual distributions of data points.
Example Contextualization
Football Game Scores
Standard deviation context of scores during a game:
Example: Standard deviation of 14 in scores compared to standard deviation of 13.37 indicates variance.
Discussion on practical implications of these numbers in competitive contexts (touchdowns).
Deviation and Squaring
Deviations Explained
Calculation: Subtract scores from the mean to find deviations.
Reason for squaring deviations:
Ensures that average deviations do not equal zero; negatives won't cancel positives.
Allows for handling the magnitude of deviations without losing directionality (negative vs. positive).
Sample Size Penalty Concept
n - 1 Adjustment
Explanation of why we take one less in the denominator during standard deviation calculations: ensures it compensates smaller samples for representation issues.
Implications
Having a smaller sample may misrepresent the true population behavior; thus, adjusting the denominator helps mitigate this bias.
Variance and Its Relation to Standard Deviation
Fundamental Understanding
Variance: Square of standard deviation.
Calculation provided: Given standard deviation of 14, variance can be computed as 14.1² = 198.81.
Importance of variance in later calculations despite it being less intuitive than standard deviation.
Outlier Detection
Definition: Outlier as a data point that significantly deviates from other observations.
Identification Techniques
A common rule of thumb for identifying outliers involves checking for values that lie more than two standard deviations from the mean.
Application to Data Analysis
Discussion on using this method for analyzing football game scores to determine if games were outliers with respect to team performance.
Identifying Games Above/Below Standard Deviation
Practical Application
Mean and standard deviation are both calculated; identifying whether game scores fall above or below established cutoffs helps to detect anomalies.
Final Analysis
Discussion prompts regarding the implications of outlier analysis in understanding team performance during seasons.
Deeper Understanding of Data Variability
Comparing Classes A and B
Exploring how standard deviation impacts understanding of performance relative to the mean.
Differences between scores in classes indicate variability that standard deviation can reveal.
Histograms and Data Distribution
Definition and Importance
Histograms allow for visualization of score distributions along a continuous variable.
Explanation of axes and what they represent in terms of frequency distribution.
Measures of Central Tendency, Variability, and Skewness
Comparison of Different Distributions
Skewness Types
Positive Skew: Longer tail on the right.
Negative Skew: Longer tail on the left.
Example provided of understanding skew through illustrative analogies (Mr. Stokovich's foot analogy).
Ceiling and Floor Effects
Description of where data clusters function with respect to outliers in skewed distributions. Such as employee wages (floor effect) and test scores (ceiling effect).
Conclusion and Next Steps
Reminder of concepts discussed and anticipation for next class sessions focused on histograms further.