1/16
Flashcards covering literary computing, population vs. sample concepts, statistical analysis of Gettysburg Address word lengths, random vs. nonrandom sampling bias, and media article evaluation checklist.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is literary computing?
Literary computing is a field that numerically analyzes authors' works by examining variables such as sentence length and rates of occurrence of specific words.
What historical debate is given as an example of a topic studied in literary computing?
The debate over whether works attributed to William Shakespeare were actually written by Francis Bacon or Christopher Marlowe.
When and where was Abraham Lincoln's Gettysburg Address delivered?
It was delivered on November 19, 1863, on the battlefield near Gettysburg, PA.
In the Gettysburg Address activity, what defines the population and a sample?
The population is the entire Gettysburg Address passage consisting of 268 words, while a sample is a subset of 10 words chosen from that population.
What is the primary goal of sampling in statistical studies?
The goal of sampling is to learn about a very large population by studying a smaller subset that is carefully selected to be representative of the larger population.

What distribution and parameters are displayed in this histogram of the Gettysburg Address?
It displays the word lengths of all 268 words in the Gettysburg Address, which range from 2 to 11 letters, with a mean word length of 4.295 letters indicated by the vertical red dashed line.
What were the Summer 2019 sample averages and 95% confidence interval for nonrandom samples?
For 25 student nonrandom samples, sample averages ranged from 4.4 to 9.0 letters, with a 95% confidence interval of (6.38,7.26).
What were the Summer 2019 sample averages and 95% confidence interval for random samples?
For 9 random sample groups, sample averages ranged from 3.4 to 6.7 letters, with a 95% confidence interval of (3.94,5.46).

What summary statistics does this table show for nonrandom versus random samples in the 2021-2022 study?
For nonrandom samples (n=148), the mean was 6.608, median was 6.550, and std. dev. was 1.267. For random samples (n=135), the mean was 4.541, median was 4.500, and std. dev. was 0.924, yielding t=15.773, p<.001, and Cohen's d=1.83.
What does the p-value (p<.001) indicate in the Gettysburg Address sampling experiment?
It indicates that the differences in average word length between nonrandom and random samples are due to more than random chance.
Which word was most commonly selected in Fall 2021 nonrandom samples, and how does it compare to its actual text frequency?
"dedicated" (length 9) was selected 21 times in nonrandom samples, whereas in the actual text it appears 4 times.
Which words are the most frequent and the longest in the actual text of the Gettysburg Address?
The most frequent word is "that" (count 12, length 4), and the longest words (length 11) are "battlefield", "consecrated", and "proposition".
What are the key takeaways regarding human sampling performance from the Gettysburg Address activity?
People are terrible at selecting representative samples even when explicitly trying to avoid sampling bias, and overall, random samples are more accurate than nonrandom samples.
How is the "Evaluating Media Articles" checklist scored and interpreted?
The checklist consists of 10 questions where a "yes" answer indicates cause for concern; 10 "no" answers signify a perfect study perfectly reported, while 10 "yes" answers mean the article should be immediately discarded.
Why was the headline "Every hour of TV watching shortens life by 22 minutes" rated as exaggerated in the media evaluation activity?
Because it implied a causal connection between television watching and shortened lifespan, which went beyond the observational evidence of the study.
Why was the advice in the TV watching media article considered unjustified?
The article gave advice specifically to avoid watching television, but the study outcomes evaluated overall sedentary behavior rather than supporting a TV-specific cause.
What was the overall score and conclusion for the article evaluating TV watching and lifespan?
The article received a score of 5 "yes" answers out of 10, indicating many unsatisfactory features originating from both the original study and how it was reported.