Ch. 2
OPAN 2101 - Business Statistics
Course and Instructor Information
Course Name: OPAN 2101 - Subject: Business Statistics
Instructor: Professor Amrita Kundu
Institution: Georgetown University
Types of Data Variables
Data Variable: A characteristic or attribute that can be measured or observed.
Quantitative: Numerical values that represent counts or measurements. These can be subjected to mathematical operations.
Discrete: Can only take on specific, distinct values, often integers resulting from counting. There are gaps between possible values.
Example: Number of words in a Tweet ( words maximum), Number of students in a class ( students), Number of cars passing a point in an hour ( cars).
Continuous: Can take on any value within a given range, typically results from measuring. There are no gaps between possible values.
Example: Temperature (), Height of a person ( meters), Runtime of a movie ( minutes).
Qualitative: Non-numerical categories or labels, representing qualities or characteristics.
Categorical: Values that fall into distinct groups or categories. These can often be represented by text or codes.
Example: Gender (Male, Female), Type of car (Sedan, SUV, Truck), Opinions (Agree, Disagree, Neutral).
Identifier: Unique labels or codes used to distinguish individual observations. They are not meant for mathematical operations and mainly serve for identification.
Example: Social Security Number, Country of Residence (e.g., USA, CAN), Student ID, an individual Tweet's unique ID.
Numeric Analysis: The process of examining data with numerical methods to extract meaningful insights. Example: Counting the number of words in a Tweet to understand its length distribution.
Lab Assignments Information
Submission Method: Submit answers through Canvas -> Assignments -> Lab
Extra Credit: Lab assignments may count for extra credit
Classifying Datasets
Types of Datasets:
Cross Sectional: Data collected from multiple subjects (e.g., individuals, firms, states) at a single point in time.
Example: Number of residents in each U.S. state for a specific year (e.g., 2021). All states are observed at the same point in time.
Example: Number of emails received yesterday for each individual student in this class. Each student's email count is captured for a single day.
Panel: Data collected from multiple subjects over multiple time periods. This combines elements of both cross-sectional and time-series data.
Example: Number of residents in each U.S. state for each year from 2000 to 2021. Here, 'state' represents cross-sectional units, and 'year' represents the time dimension.
Example: Total parking citations issued within each zip code in D.C. for each month from October to December 2021. Each zip code is observed across several months.
Example: Number of tornadoes that touched down in each U.S. county in 2021. This is actually a cross-sectional dataset because it observes multiple counties (subjects) at a single time point (year 2021).
Kaggle Dataset Example
Kaggle Variables:
Year = Quantitative, Discrete (Represents a specific count of a year)
Runtime = Quantitative, Continuous (Duration of something, like a movie, measured in units like minutes)
Text = Qualitative (The content itself is descriptive, like a movie plot or a user review)
Descriptive Statistics
Purpose: First step in analyzing data by summarizing its main features. It helps to understand patterns, distributions, and relationships.
Selection of Descriptive Statistics: Depends on the type of data.
Qualitative or Categorical Data
Tools: Frequency tables, Bar charts (for visualizing data distribution)
Quantitative Data
Metrics:
Measures of Centre (e.g., Mean, Median, Mode to describe the typical value)
Measures of Spread (e.g., Standard Deviation, Range, Interquartile Range to describe variability)
Measures of Shape (e.g., Skewness, Kurtosis to describe the distribution's form)
Tools: Histograms, Boxplots (to visualize data distribution)
Descriptive Statistics – Frequency Tables
Definitions:
Frequency: The count of data observations that fall into each specific category or numerical range. (e.g., If a class has freshmen, sophomores, juniors, and seniors, the frequency for freshmen is ).
Relative Frequency: The proportion or percentage of data observations in each category relative to the total number of observations. It's calculated as . (e.g., With total students and freshmen, the relative frequency for freshmen is or . This means of the students are freshmen.)
Usage:
Frequency tables and bar charts summarize categorical data effectively, providing a clear overview of the distribution of items across different categories.
A frequency table records the count, proportion, or percentage of data in each category, e.g., different age groups, educational levels, or types of products.
Descriptive Statistics – Bar Charts
Definition: A graphical representation using rectangular bars to show the frequency or relative frequency of observations in different categories of a qualitative variable. The height or length of the bars corresponds to the values they represent.
Significance of Data Distribution:
It is the list of all possible values for a variable and their relative frequency, showing how often each value occurs.
It illustrates the likelihood of different outcomes for a variable, providing insights into the typical values and spread of the data.
Contingency Tables
Definition: Also known as a cross-tabulation table, a contingency table is a type of matrix table that shows the frequency distribution of two or more categorical variables simultaneously. It helps to observe the relationship between these variables.
Example Analysis:
Helps observe how workers in various age groups differ depending on their payment type (At or Below Minimum Wage versus Hourly Workers).
Displays total observations in age categories based on a GROUP variable, allowing for a comparison of proportions across different groups.
Understanding Conditional Distribution
Definition: The distribution of one categorical variable, conditioned on at least one category of another categorical variable, expressed as a percentage. It shows how the distribution of one variable changes when another variable is held constant at a specific category.
Independence Check:
To determine if age depends on GROUP, you would compare the marginal distribution of AGE with the conditional distributions of AGE for each category of GROUP.
If they are the same (or very similar), then the variables (e.g., AGE and GROUP) are considered independent, meaning there is no statistically significant relationship between them. If they differ significantly, the variables are dependent.
Application Example with Financial Data
Hypothetical Study: Aggregate wage data of hourly workers at a large U.S. firm, collected to analyze employment patterns and wage structures.
Variable Definitions:
ID: Worker ID - A unique identifier assigned to each worker, typically a quantitative discrete variable for identification purposes.
AGE: Age category of workers - A categorical variable grouping workers into distinct age ranges (e.g., 16-24, 25-34, 35-44, 45-54, 55-64, 65+).
GROUP: Type of wage workers receive - A categorical variable indicating the wage structure.
Hourly Workers: Workers receiving an hourly wage greater than the minimum wage, representing a specific employment group.
At or Below Minimum Wage: Workers earning at or below the minimum wage, highlighting another employment group.
SURVEY: Indicates data collection from survey one or survey two - A categorical variable identifying the source or wave of data collection.
Independence and Distribution Analysis
Marginal Distribution Example: Represents the distribution of a single variable without regard to other variables. For instance, workers categorized by AGE and GROUP with detailed counts and totals shown in a table would first present the overall counts for AGE categories before considering GROUP.
Conditional Distribution Analysis for various age groups (16-24, 25-34, etc.): This involves calculating the percentage distribution of AGE within each GROUP category. For example, what percentage of "At or Below Minimum Wage" workers fall into the "16-24" age group?
Example Calculations for conditional distributions of AGE for workers earning at or below minimum wage and hourly workers presented in a table form for clarity. This would involve calculating: for each age group.
Conclusion on Independence:
If the conditional distribution of AGE is identical across all GROUP categories (e.g., the percentage of 16-24 year olds is the same for both "At or Below Minimum Wage" and "Hourly Workers"), then AGE and GROUP are independent.
Example Percentages Presented: Suppose for the 16-24 age group:
At or Below Minimum Wage: Approximately of all workers in the "At or Below Minimum Wage" group are 16-24 years old.
Hourly Workers: Approximately of all workers in the "Hourly Workers" group are 16-24 years old.
Calculation example for 16-24 AGE group provided for clarity: If there are total workers earning at or below minimum wage, and of them are 16-24 years old, this signifies their proportion within that specific group. This would be compared to the proportion of 16-24 year olds in the "Hourly Workers" group.
Lab Assignment Instructions
Task:
Find the marginal distribution of GROUP variable (i.e., the overall percentage of workers in "At or Below Minimum Wage" vs. "Hourly Workers" regardless of age).
Find the conditional distribution of GROUP for 45-54 year old workers (i.e., among only 45-54 year olds, what percentage are "At or Below Minimum Wage" vs. "Hourly Workers").
Determine if AGE and GROUP are independent, providing quantitative (using calculated percentages) and qualitative