Ch. 1 Data and Data Sources

Introduction to Business Statistics

  • Importance of understanding, applying, analyzing, and evaluating data and data sources.

  • Definition of data: Facts and figures collected, analyzed, summarized for their presentation and interpretation, essentially forming information that we aim to learn from.

Key Concepts

Elements
  • Definition: Entities upon which data is collected (e.g., degree candidates).

Population
  • Definition: The entire set of elements of interest (e.g., all degree candidates worldwide).

  • Size of populations can be vast (e.g., approximately 7.97 billion people in the world).

  • Direct measurement of a population is often impractical, leading to the use of samples.

Samples
  • Definition: A subset of the population used for analysis (e.g., degree candidates enrolled in Business Statistics).

Types of Data

Qualitative Data
  • Also referred to as categorical data.

  • Nominal Data:

    • Definition: Labels without a specific order (e.g., bank, credit union, savings and loan).

    • No judgment of better or worse is implied.

  • Ordinal Data:

    • Definition: Data that can be ordered (e.g., service ratings: excellent, good, poor).

    • Infers a ranking (excellent > good > poor).

Quantitative Data
  • Two subtypes:

    • Interval Data:

      • Definition: Differences that are meaningful (e.g., temperature differences).

    • Ratio Data:

      • Definition: Ratios that are meaningful (e.g., a starting salary of $60,000 being twice that of $30,000).

Data Sources

Cross-Sectional Data
  • Definition: Collected at a single point in time (e.g., average rainfall in 50 states in 2022).

Time Series Data
  • Definition: Collected over a period of time (e.g., average rainfall in Arizona from 1988 to 2022).

Panel or Pool Data
  • Definition: A combination of cross-sectional and time series data (e.g., average rainfall in 50 states from 1988 to 2022).

Conclusion

  • Overview of the first learning objective in Business Statistics.

  • Emphasis on understanding different types of data and their sources to apply them effectively in managerial decisions.


Descriptive Statistics

  • Descriptive statistics provide a format to present data that is understandable to the general public.

  • The goal is to share data effectively, which can be done through:

    • Tabular summaries

    • Graphical representations

    • Numerical summaries

Key Concepts

Class
  • A class is a set of items or categories that can describe certain characteristics

    • (e.g., hot, cool, service ratings).

  • Classes group items based on their characteristics.

Frequency
  • Frequency refers to the count of items or observations in a particular class.

    • Counting is a fundamental skill in statistics, essential for data representation.

Relative Frequency
  • Definition: Relative frequency is a proportion of the total observations in a class.

    • Example: If 30% of the days are cool, this could be expressed as 0.3.

  • Relationship: Relative frequency can be expressed as a percentage or a proportion.

Cumulative Relative Frequency
  • Definition: Cumulative relative frequency is the successive addition of relative frequencies.

  • Importance: This provides insight into the total proportion of observations that fall below a certain class.

Sample Dataset Example

  • Example: Analyze average high temperatures over 20 randomly sampled days.

  • Importance of Random Sampling: Ensures unbiased data collection.

Temperature Data
  • A series of high temperatures recorded randomly. Example values include:

    • 72, 85, 89, with a range specified for classes.

Constructing a Frequency Table
  • Table Structure:

    • Columns: Class (temperature ranges), Count (frequency), Relative Frequency, Cumulative Relative Frequency

    • Rows: Number of groups (e.g., temperature ranges)

Example Classes and Counts

Average high temp for 20 randomly sampled days:
72 91 91 89 90 98 85 82 85 89 87 89 66 77 51 89 75 47 54 89

  1. 40-49: Frequency = 1

  2. 50-59: Frequency = 2

  3. 60-69: Frequency = 1

  4. 70-79: Frequency = 3

  5. 80-89: Frequency = 9

  6. 90-99: Frequency = 4

Calculating Relative Frequencies
  • Relative Frequency Calculation:

    • Class 40-49: 1/20 = 0.05

    • Class 50-59: 2/20 = 0.10

    • Class 60-69: 1/20 = 0.05

    • Class 70-79: 3/20 = 0.15

    • Class 80-89: 9/20 = 0.45

    • Class 90-99: 4/20 = 0.20

  • Sum of Relative Frequencies: Should total 1.0.

Calculating Cumulative Relative Frequencies
  • Cumulative Calculation:

    • 40-49: 0.05

    • 50-59: 0.05 + 0.10 = 0.15

    • 60-69: 0.15 + 0.05 = 0.20

    • 70-79: 0.20 + 0.15 = 0.35

    • 80-89: 0.35 + 0.45 = 0.80

    • 90-99: 0.80 + 0.20 = 1.00

  • Importance: Cumulative frequencies provide insight into how data accumulates over classes.

Class

Frequency

Relative Frequency

Cumulative Relative Frequency

40-49

1

0.05

0.05

50-59

2

0.10

0.15 (0.05 from row 1 + 0.10 from row 2 = 0.15)

60-69

1

0.05

0.20 (0.15 from row 1 and 2 + 0.05 from row 3 = 0.20)

70-79

3

0.15

0.35

80-89

9

0.45

0.80

90-99

4

0.20

1.00

Total

20

1.00




Introduction to Statistical Inference

  • Statistical inference involves estimating, predicting, or generalizing about a population based on information from a random sample.

  • Decisions should not be made based on feelings or hunches but rather on statistical data.

The Process of Making Inferences

  • Data Collection and Learning

    • Collect data to make informed decisions about an unknown population from a known random sample.

    • Example: To assess climatologists' views on global warming, a random sample of 100 expert climatologists can be taken.

  • Analyzing Expert Opinions

    • If 99% of climatologists support the reality of global warming, this is seen as a strong consensus.

    • The contrary opinion of the remaining 1% should be critically evaluated and often disregarded in decision-making.

Components of Statistical Inference

  • Fundamental Elements of Inferential Statistics:

    1. Population or Sample of Interest:

      • Define what or who you want to learn about;

        • Ex: all degree candidates worldwide.

    2. Variables or Characteristics of the Population:

      • Identify concerns such as age, height, or weight of the population.

    3. Random Sample of Population Units:

      • Select a manageable sample size to analyze, e.g., 30 degree candidates enrolled in Business Statistics.

    4. Inference about the Population Based on the Sample:

      • Draw conclusions or gain knowledge based on the data collected from the sample.

    5. Measure of Reliability for the Inference:

      • Establish the level of significance, which reflects the confidence in the statistical inference process.