Ch. 1 Data and Data Sources
Introduction to Business Statistics
Importance of understanding, applying, analyzing, and evaluating data and data sources.
Definition of data: Facts and figures collected, analyzed, summarized for their presentation and interpretation, essentially forming information that we aim to learn from.
Key Concepts
Elements
Definition: Entities upon which data is collected (e.g., degree candidates).
Population
Definition: The entire set of elements of interest (e.g., all degree candidates worldwide).
Size of populations can be vast (e.g., approximately 7.97 billion people in the world).
Direct measurement of a population is often impractical, leading to the use of samples.
Samples
Definition: A subset of the population used for analysis (e.g., degree candidates enrolled in Business Statistics).
Types of Data
Qualitative Data
Also referred to as categorical data.
Nominal Data:
Definition: Labels without a specific order (e.g., bank, credit union, savings and loan).
No judgment of better or worse is implied.
Ordinal Data:
Definition: Data that can be ordered (e.g., service ratings: excellent, good, poor).
Infers a ranking (excellent > good > poor).
Quantitative Data
Two subtypes:
Interval Data:
Definition: Differences that are meaningful (e.g., temperature differences).
Ratio Data:
Definition: Ratios that are meaningful (e.g., a starting salary of $60,000 being twice that of $30,000).
Data Sources
Cross-Sectional Data
Definition: Collected at a single point in time (e.g., average rainfall in 50 states in 2022).
Time Series Data
Definition: Collected over a period of time (e.g., average rainfall in Arizona from 1988 to 2022).
Panel or Pool Data
Definition: A combination of cross-sectional and time series data (e.g., average rainfall in 50 states from 1988 to 2022).
Conclusion
Overview of the first learning objective in Business Statistics.
Emphasis on understanding different types of data and their sources to apply them effectively in managerial decisions.
Descriptive Statistics
Descriptive statistics provide a format to present data that is understandable to the general public.
The goal is to share data effectively, which can be done through:
Tabular summaries
Graphical representations
Numerical summaries
Key Concepts
Class
A class is a set of items or categories that can describe certain characteristics
(e.g., hot, cool, service ratings).
Classes group items based on their characteristics.
Frequency
Frequency refers to the count of items or observations in a particular class.
Counting is a fundamental skill in statistics, essential for data representation.
Relative Frequency
Definition: Relative frequency is a proportion of the total observations in a class.
Example: If 30% of the days are cool, this could be expressed as 0.3.
Relationship: Relative frequency can be expressed as a percentage or a proportion.
Cumulative Relative Frequency
Definition: Cumulative relative frequency is the successive addition of relative frequencies.
Importance: This provides insight into the total proportion of observations that fall below a certain class.
Sample Dataset Example
Example: Analyze average high temperatures over 20 randomly sampled days.
Importance of Random Sampling: Ensures unbiased data collection.
Temperature Data
A series of high temperatures recorded randomly. Example values include:
72, 85, 89, with a range specified for classes.
Constructing a Frequency Table
Table Structure:
Columns: Class (temperature ranges), Count (frequency), Relative Frequency, Cumulative Relative Frequency
Rows: Number of groups (e.g., temperature ranges)
Example Classes and Counts
Average high temp for 20 randomly sampled days:
72 91 91 89 90 98 85 82 85 89 87 89 66 77 51 89 75 47 54 89
40-49: Frequency = 1
50-59: Frequency = 2
60-69: Frequency = 1
70-79: Frequency = 3
80-89: Frequency = 9
90-99: Frequency = 4
Calculating Relative Frequencies
Relative Frequency Calculation:
Class 40-49: 1/20 = 0.05
Class 50-59: 2/20 = 0.10
Class 60-69: 1/20 = 0.05
Class 70-79: 3/20 = 0.15
Class 80-89: 9/20 = 0.45
Class 90-99: 4/20 = 0.20
Sum of Relative Frequencies: Should total 1.0.
Calculating Cumulative Relative Frequencies
Cumulative Calculation:
40-49: 0.05
50-59: 0.05 + 0.10 = 0.15
60-69: 0.15 + 0.05 = 0.20
70-79: 0.20 + 0.15 = 0.35
80-89: 0.35 + 0.45 = 0.80
90-99: 0.80 + 0.20 = 1.00
Importance: Cumulative frequencies provide insight into how data accumulates over classes.
Introduction to Statistical Inference
Statistical inference involves estimating, predicting, or generalizing about a population based on information from a random sample.
Decisions should not be made based on feelings or hunches but rather on statistical data.
The Process of Making Inferences
Data Collection and Learning
Collect data to make informed decisions about an unknown population from a known random sample.
Example: To assess climatologists' views on global warming, a random sample of 100 expert climatologists can be taken.
Analyzing Expert Opinions
If 99% of climatologists support the reality of global warming, this is seen as a strong consensus.
The contrary opinion of the remaining 1% should be critically evaluated and often disregarded in decision-making.
Components of Statistical Inference
Fundamental Elements of Inferential Statistics:
Population or Sample of Interest:
Define what or who you want to learn about;
Ex: all degree candidates worldwide.
Variables or Characteristics of the Population:
Identify concerns such as age, height, or weight of the population.
Random Sample of Population Units:
Select a manageable sample size to analyze, e.g., 30 degree candidates enrolled in Business Statistics.
Inference about the Population Based on the Sample:
Draw conclusions or gain knowledge based on the data collected from the sample.
Measure of Reliability for the Inference:
Establish the level of significance, which reflects the confidence in the statistical inference process.