1.2: Data, Sampling, and Variation in Data and Sampling

Categorization of Data

  • Statistics involves the collection, analysis, interpretation, and presentation of data. The types of data collected are fundamental to how statistics is performed.

  • Qualitative Data (Categorical Data):

    • Definition: Information that is not numbers-based. It is based on descriptions or types (qualities or categories).
    • Examples: Eye color, favorite animal, or brand of shoe.
  • Quantitative Data:

    • Definition: Information that is numbers-based. Unlike qualitative data, mathematical calculations can be performed on quantitative data.
    • Examples: Shoe size, height, or number of pets.
    • Quantitative Discrete Data:
      • The possible data values are limited to specific numerical values.
      • Example: Shoe size is discrete because only half- or full-sizes are recorded (e.g., 77 or 7.57.5). There is no possibility of a shoe size existing between these specific recorded values, such as a size 7.237.23.
      • Other examples of discrete counts include a student’s class load (number of credits), a state’s capital city population, or the number of fish caught at a fishing tournament.
    • Quantitative Continuous Data:
      • The data values are not limited to certain values and any possibility can occur anywhere on the number line between the smallest and largest values.
      • Example: Height is continuous because it can be measured to any level of precision along the number line.
      • Other examples include the length of fish at a fishing tournament.

Displaying Qualitative Data

  • The type of data collected dictates the appropriate methods for analysis, interpretation, and presentation.
  • Effective tools for displaying qualitative (categorical) data include:
    • Pie Charts: A circular chart divided into sectors illustrating proportions.
    • Bar Graphs: Diagrams in which the numerical values of variables are represented by the height or length of lines or rectangles of equal width.
    • Pareto Charts: A specific type of bar graph where the bars are organized in order from largest to smallest frequency.
  • Presentation clarity is critical. Choosing the correct tool is only part of the process; the delivery must be clear enough to effectively convey the desired information.

Principles of Sampling

  • Population and Sample:

    • A population consists of every member of a group.
    • Because collecting information from every member of a population is often impossible or too labor-intensive, a sample (a smaller portion of the population) is used.
    • Sampling is the process of collecting data from this sample. The characteristics of the sample are then assumed to represent the characteristics of the entire population.
  • Goal of Sampling:

    • The sample must be representative of the entire population.
    • The primary goal is to remove as much potential for bias as possible. Bias occurs when certain members of the population are more likely to be selected for the sample than others, leading to skewed results.

Random Sampling Methods

In random sampling, every member of the population has an equal chance of being selected. There are several specific methods:

  • Simple Random Sampling: The sample is selected from the population without any additional organization or complex requirements.
  • Stratified Sampling:
    • The population is first organized into groups called strata.
    • A simple random sample is then taken from each individual group to build the overall sample.
    • The selection from each stratum is proportional to the overall population.
  • Cluster Sampling:
    • The population is organized into groups called clusters.
    • Researchers randomly select some of these clusters.
    • Every member within the randomly chosen clusters becomes part of the final sample.
  • Systematic Sampling:
    • The population is organized into a "line" or a list.
    • A mathematical pattern is followed to select members (e.g., every nn-th member).

Non-Random Sampling and Scenarios

  • Convenience Sampling: A non-random method where the sample consists of members of the population who are easiest to contact or reach. This method typically results in bias.

  • Sampling Identification Examples (Jane’s Pencil Brand Preference Study):

    • Systematic: Jane lists all students alphabetically and asks every 10th10^{\text{th}} student.
    • Convenience: Jane waits at the library doors on a Friday and asks any students who visit.
    • Stratified: Jane randomly selects some students from every grade (strata) and asks them.
    • Simple Random: Jane uses a computer program to randomly select individuals from the whole population.
    • Cluster: Jane selects entire classes to visit during 1st1^{\text{st}} period and asks everyone in those specific classes.

Sampling and Nonsampling Errors

  • Sampling Errors: Errors caused by the actual process of selecting the sample from the population. Bias is a type of sampling error.

  • Nonsampling Errors: Errors that occur when the means by which data is collected from the sample is incorrect or flawed.

  • Common Problems to Identify:

    • Problems with samples: The sample is biased and not representative of the population.
    • Self-selected samples: Members of the population decide for themselves whether or not to belong to the sample (e.g., call-in polls).
    • Sample size issues: If the sample size is too small, it might fail to represent the whole population.
    • Undue influence: The method of data collection (such as the wording of a question) influences the sample's responses.
    • Non-response/Refusal to participate: Selected individuals may choose not to provide data.
    • Causality: Determining a relationship where none exists; two unrelated concepts may appear related.
    • Self-funded/Self-interest studies: Data collected by an entity with a biased intent or a stake in the outcome.
    • Misleading use of data: Presenting data in an untruthful or deceptive manner.
    • Confounding: Outside factors that lead to incorrect conclusions by interfering with the variables being studied.

Variation in Data

Variation describes how data values differ from one another, which can prevent a sample from reaching the exact same conclusion every time.

  • Variation between individual members: Natural differences exist within a population. For instance, a factory making "99-inch nails" may produce nails ranging from 8.88.8 to 9.29.2 inches, with the goal that the average length is exactly 99 inches.
  • Variation between samples: Since a sample is only a part of the population, choosing a different sample will often result in different data and slightly different conclusions. This is natural and does not automatically imply bias.
  • Sample Size Impact: Increasing the sample size helps to minimize the effects of sample variation.

Statistical Notation and Approximations

  • It is critical to distinguish between population values and sample values using different variable names:
    • Population Average (μ\mu): Represented by the Greek letter mu. This value is often difficult or impossible to compute directly for an entire population.
    • Sample Average (Xˉ\bar{X}): Calculated from a sample and used to approximate the true population average (μ\mu).
  • The goal is for the sample average (Xˉ\bar{X}) to be "close enough" to the true population average (μ\mu) to be useful for analysis.