Lesson 4: The Gettysburg Address and Sampling Methods

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/16

flashcard set

Earn XP

Description and Tags

Flashcards covering literary computing, population vs. sample concepts, statistical analysis of Gettysburg Address word lengths, random vs. nonrandom sampling bias, and media article evaluation checklist.

Last updated 2:21 AM on 10/1/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

17 Terms

1
New cards

What is literary computing?

Literary computing is a field that numerically analyzes authors' works by examining variables such as sentence length and rates of occurrence of specific words.

2
New cards

What historical debate is given as an example of a topic studied in literary computing?

The debate over whether works attributed to William Shakespeare were actually written by Francis Bacon or Christopher Marlowe.

3
New cards

When and where was Abraham Lincoln's Gettysburg Address delivered?

It was delivered on November 19, 1863, on the battlefield near Gettysburg, PA.

4
New cards

In the Gettysburg Address activity, what defines the population and a sample?

The population is the entire Gettysburg Address passage consisting of 268268 words, while a sample is a subset of 1010 words chosen from that population.

5
New cards

What is the primary goal of sampling in statistical studies?

The goal of sampling is to learn about a very large population by studying a smaller subset that is carefully selected to be representative of the larger population.

6
New cards
<p>What distribution and parameters are displayed in this histogram of the Gettysburg Address?</p>

What distribution and parameters are displayed in this histogram of the Gettysburg Address?

It displays the word lengths of all 268268 words in the Gettysburg Address, which range from 22 to 1111 letters, with a mean word length of 4.2954.295 letters indicated by the vertical red dashed line.

7
New cards

What were the Summer 2019 sample averages and 95%95\% confidence interval for nonrandom samples?

For 2525 student nonrandom samples, sample averages ranged from 4.44.4 to 9.09.0 letters, with a 95%95\% confidence interval of (6.38,7.26)(6.38, 7.26).

8
New cards

What were the Summer 2019 sample averages and 95%95\% confidence interval for random samples?

For 99 random sample groups, sample averages ranged from 3.43.4 to 6.76.7 letters, with a 95%95\% confidence interval of (3.94,5.46)(3.94, 5.46).

9
New cards
<p>What summary statistics does this table show for nonrandom versus random samples in the 2021-2022 study?</p>

What summary statistics does this table show for nonrandom versus random samples in the 2021-2022 study?

For nonrandom samples (n=148n = 148), the mean was 6.6086.608, median was 6.5506.550, and std. dev. was 1.2671.267. For random samples (n=135n = 135), the mean was 4.5414.541, median was 4.5004.500, and std. dev. was 0.9240.924, yielding t=15.773t = 15.773, p<.001p < .001, and Cohen's d=1.83d = 1.83.

10
New cards

What does the pp-value (p<.001p < .001) indicate in the Gettysburg Address sampling experiment?

It indicates that the differences in average word length between nonrandom and random samples are due to more than random chance.

11
New cards

Which word was most commonly selected in Fall 2021 nonrandom samples, and how does it compare to its actual text frequency?

"dedicated" (length 99) was selected 2121 times in nonrandom samples, whereas in the actual text it appears 44 times.

12
New cards

Which words are the most frequent and the longest in the actual text of the Gettysburg Address?

The most frequent word is "that" (count 1212, length 44), and the longest words (length 1111) are "battlefield", "consecrated", and "proposition".

13
New cards

What are the key takeaways regarding human sampling performance from the Gettysburg Address activity?

People are terrible at selecting representative samples even when explicitly trying to avoid sampling bias, and overall, random samples are more accurate than nonrandom samples.

14
New cards

How is the "Evaluating Media Articles" checklist scored and interpreted?

The checklist consists of 1010 questions where a "yes" answer indicates cause for concern; 1010 "no" answers signify a perfect study perfectly reported, while 1010 "yes" answers mean the article should be immediately discarded.

15
New cards

Why was the headline "Every hour of TV watching shortens life by 22 minutes" rated as exaggerated in the media evaluation activity?

Because it implied a causal connection between television watching and shortened lifespan, which went beyond the observational evidence of the study.

16
New cards

Why was the advice in the TV watching media article considered unjustified?

The article gave advice specifically to avoid watching television, but the study outcomes evaluated overall sedentary behavior rather than supporting a TV-specific cause.

17
New cards

What was the overall score and conclusion for the article evaluating TV watching and lifespan?

The article received a score of 55 "yes" answers out of 1010, indicating many unsatisfactory features originating from both the original study and how it was reported.