Sampling Bias and Word Length Analysis

Sampling Bias and Word Length Analysis

  • Primary Objective: Demonstrate the presence and effect of human sampling bias when visually estimating parameters of a population, specifically using word length in a written text.
  • Population Text: Abraham Lincoln's Gettysburg Address (beginning with the phrase "Four score and seven years ago").
  • Observational Unit: Individual words within the address.
  • Variable of Interest: Word length, defined as the total number of letters contained within a single word.
    • Example: The word "what" contains 44 letters, giving it a word length of 44

Judgment Sampling Procedure and Data Collection

  • Subjective Selection Procedure:
    • Select a sample size of n=10n = 10 words from the text that appear visually representative of the overall average word length of the address.
    • Calculate the sample mean word length:
    • Add up the individual word lengths of all 1010 selected words.
    • Divide the total sum of letters by 1010 (arithmetically accomplished by moving the decimal point one position to the left).
  • Class Data Visualization:
    • Construct a dot plot on the board.
    • Plot each calculated sample average by placing a physical dot marker directly above the corresponding numerical value on the axis.

Mechanics of Visual Sampling Bias

  • Sample Student Results:
    • Calculated sample mean word length: 5.85.8 letters per word.
    • Selection approach: Sampled every other line down the middle of the text, selecting long words at the end to complete the sample size of 1010.
    • Specific included words: Included visually prominent words such as "government" (length of 1010 letters) and "unfinished" (length of 1010 letters).
  • Root Cause of Bias in Human Selection:
    • Human eyes naturally gravitate toward larger, visually prominent words (e.g., 1010-letter words like "government" or "unfinished") while routinely overlooking smaller, high-frequency words (e.g., 22-letter or 33-letter words).
    • Subjective visual selection ("judgment sampling") consistently leads to an overestimation of the true population mean, directly demonstrating sampling bias.

Transition to Random Sampling Schemes

  • Random Selection Method:
    • To eliminate human visual bias, words must be selected strictly through a random process.
    • Assign a distinct numerical index to every word in the text population.
    • Use a random number generator to select specific numerical indices.
    • Extract the words corresponding to those selected indices, count their letter lengths, and calculate the unbiased sample average by dividing the total letter count by the sample size.
  • Curricular Scope:
    • Focus is placed strictly on determining the presence or absence of randomness in a sampling scheme.
    • Evaluating whether a scheme incorporates genuine chance takes precedence over categorizing complex variations of random sampling designs.