1/153
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What is Statistics?
The science of planning studies and experiments, obtaining data, and organizing, summarizing, analyzing, and interpreting those data, and the drawing conclusions based on them.
What is a Population?
The complete collection of all measurements or data that are being considered. We also call it the population of interest.
What is a Sample?
A subset of members selected FROM a POPULATION.
What is a Parameter?
A numerical measurement describing some characteristic of a population. NOTE: Population and Parameter start with P
What is a Statistic?
A numerical measurement describing some characteristic of a sample. NOTE: Both Statistics and Sample start with S
Is this a parameter or a statistic: Each member of the U.S. House of Representatives is selected, and the average age is found.
This is a parameter; the indication here is that it says each member is chosen rather than a specific group.
Is this a parameter or a statistic: At a movie theater, every seventh person who enters is selected. The heights of the selected individuals are recorded. The mean height is calculated.
Statistic
Tell me the population, sample, parameter, and statistic: A recent survey of college career centers reported that the average starting salary for 1,283 newly graduated petroleum engineers is $83,121.
Population: All newly graduated petroleum engineers
Sample: The 1,283 graduates
Parameter: The true starting salary
Statistic: Mean salary $83,121
What does Quantitative mean?
Consists of numbers representing counts or measurements, such as ages or weights.
What does Categorical mean?
Consists of names or labels NOT number or counts. This could be a college major or hometown.
Quantitative data also has two sub-categories, what are they?
Discrete and Continuous
What does Discrete mean?
Results when the data values are quantitative, and the number of values is finite/countable.
What does Continuous mean?
Results from infinitely many possible quantitative values, where the collection of values is not countable. This does not have to range to infinity but can have infinitely many possible values.
Is this Categorical or Quantitative: A city directory lists telephone numbers for every resident. Every 20th resident’s phone number is recorded.
Categorical
Is this Categorical or Quantitative: The local news station records the daily high temperature for each day in October.
Quantitative
Is this Discrete or Continuous: Each person on a crowded bus is asked how many coins they have in their pockets.
Discrete, this is in a specific area, so we can count.
Is this Discrete or Continuous: The ten tallest trees in Oregon are measured, and their heights are listed.
Continuous, huge range all across Oregon, so this is more difficult to count; numerous values could be listed.
Give me the Population of interest, Sample, Parameter, and Statistic: According to Pew Research Center, only 32% of American adults are against marijuana legalization. What proportion of Wake County residents support marijuana legalization? A researcher is interested in finding out. She plans to ask 200 people.
Population of interest: Wake County residents
Sample: 200 people
Parameter: The true percentage in Wake County
Statistic: What percentage in Wake County
What is the issue with this: Suppose the researcher stops 200 people on the NC State campus and asks if they support marijuana legalization.
Some are not American citizens, and some are not adults; it is also not a good representation of Wake County since the only people being asked are NCSU residents.
What is Bias?
Biased samples are those that are more likely to produce some outcomes than others. The resulting statistic may be too high or low.
What is a convenience sample?
These are easy to collect. These often have some bias or do not represent the population in general.
What is a volunteer response?
A volunteer response sample is a self-selected sample of people who respond to a general appeal.
How do we avoid bias in our samples?
We use random sampling techniques that consider preferences or researcher bias.
What is a Simple Random Sample (SRS)?
A sample of n subjects is selected in such a way that every possible sample of the same size n has the same chance or probability of being chosen.
How do you perform an SRS?
Compile a numbered list of the units in the population (this is called a sampling frame). Then use something to generate random numbers. Those whose numbers are generated are selected for the sample.
What is a Stratified Sample?
where you subdivide the population into at least two different subgroups (or strata) so that the subjects within the same subgroup share the same characteristics. Then draw a sample from each subgroup. The number sampled from each subgroup may be done proportionally with respect to the size of the population.
What is a Cluster Sample?
Where you divide the population area into naturally occurring sections or clusters, then randomly select some of those clusters and choose ALL the members from those selected clusters.
What is a Systematic Sample?
Where you select some starting point and then select every kth element in the population. This works well when units are in some order, like an assembly line, houses on a block, ect.
What is a Multistage sample?
You collect data by using some combination of the basic sampling methods.
Is this Simple Random, Stratified, Cluster, Systematic, Volunteer, or Convenience: Take a sample that includes 200 republicans, 200 democrats, and 200 unaffiliated.
Stratified.
Why? Divided into subgroups that share the same characteristics, then take a sample from each.
Is this Simple Random, Stratified, Cluster, Systematic, Volunteer, or Convenience: Take a sample of 10 randomly selected voting precincts and talk with all the voters in those precincts.
Cluster.
Why? Divide the population into clusters, then randomly select some of those clusters, then question everyone in those clusters.
Is this Simple Random, Stratified, Cluster, Systematic, Volunteer, or Convenience: Pick every 10th person from a list of all registered voters.
Systematic.
Why? Every kth person is chosen.
Is this Simple Random, Stratified, Cluster, Systematic, Volunteer, or Convenience: Find a busy intersection in downtown Charlotte and stop the first 600 people that walk by.
Convenience.
Why? Easy and not random: the first people who walk by.
What is a Bad Sampling Frame?
When attempting to list all members of a population, some subjects are missing. It can be difficult to obtain a full, complete list.
What is Undercoverage?
The sampling frame is missing groups from the population, or the groups have smaller representation in the sample than in the population.
What is a Non-response bias?
Some part of the population chooses not to respond, or subjects were selected but are not able to be contacted.
What is Response bias?
Responses given to questions or surveys are not truthful. This may occur when people are unwilling to reveal personal matters, admit to illegal activity, or otherwise tailor their responses to please the investigator.
What is Wording & Order?
The way questions are worded may be leading or inflammatory to elicit a particular response. The order in which questions are asked may influence the answer.
Is this Bad sampling frame, Undercoverage, Non-response bias, Response bias, Wording & Order: Of the 300 individuals selected, some individuals threw away the forms when they received them, believing that they were junk mail.
Non-response bias.
They threw away the forms since they thought it was just junk mail, so now these subjects cannot be recorded to have an accurate random sample.
Is this Bad sampling frame, Undercoverage, Non-response bias, Response bias, Wording & Order: There are many people in the state who have recently moved here and do not appear on the states list of all licensed drivers.
Bad sampling frame.
Since these names are not on the state lists, they cannot be accurately added to the random selection, meaning this is not a full list to sample from.
Is this Bad sampling frame, Undercoverage, Non-response bias, Response bias, Wording & Order: In one of the questions, the subjects were asked if their commute was made longer or more difficult by buses or bicyclists sharing the roadway.
Wording and Order.
This type of wording sounds like its framed in a way to elicit a response.
Is this Bad sampling frame, Undercoverage, Non-response bias, Response bias, Wording & Order: Some people filled out the survey but refused to answer the questions about their income because they felt it was too personal.
Response Bias.
They did not want to reveal a personal matter, which makes the sample unreliable.
What are the two most common ways to collect data?
Observational studies and experiments.
What is a experiment?
The process of applying some treatment and then observing its effects is called an experiment.
What is an experiment composed of?
Always compares two or more groups. This will usually involve a treatment group and a control group. Control groups do not receive the treatment. You can also compare two treatments without a control group.
What is an observational study?
The process of observing and measuring specific characteristics without attempting to modify the individuals being studied is called an observational study. Observational studies tell what is happening and cannot describe a cause-and-effect relationship.
Is this Observational or Experiment: What is the effect of using a sugar substitute in desserts on the blood sugar level of adults?
Experiment.
This has a cause and effect relationship.
Is this Observational or Experiment: What are the marriage rituals for people living in India today?
Observational.
This is not cause and effect; you are just watching.
What is the response variable in a relationship study?
The response variable measures an outcome of a study.
What is the explanatory variable in a relationship study?
An explanatory variable explains or influences changes in the response variable.
What does the Design of experiments mean?
Plan for collecting the sample
What is the response variable?
Variable measured in the experiment (outcome)
What are the two reasons that there will be Variability?
Treatment effects: This is what we’re looking for- different treatments causing different outcomes
Experimental Error:
Variability among observed values of the response variable for experimental units that receive the same treatment
We want this to be as small as possible
What is a Lurking Variable?
This is another source of variability. A variable that is not among the explanatory variables in a study and yet may influence the interpretation of the relationship among response and explanatory variables.
What is a Confounding Variable?
This is another source of variability. Two variables are confounded when the effects on the response variable cannot be distinguished from each other.
What are the principles of experimental design?
Control: Control the effects of lurking/confounding variables and other sources of variability on the response by carefully planning the study. A control group is one that receives no treatment and is used as a baseline or comparison for the treatment group.
Randomization: Randomly assign experimental units to treatments to reduce or eliminate bias
Replication: Measure the effect of each treatment on many units to reduce the chance variation in the result.
What does randomization refer to in the case of experimental design?
This refers to the methods used to assign those already in a sample to a treatment group.
What are the 3 randomization methods in the case of an experiment?
Completely Randomized Designs: participants are randomly assigned to treatments (including control groups). By randomly assigning subjects to treatments, the experimenter assumes that, on average, lurking variables will affect each treatment group equally; any significant differences between groups can fairly be attributed to the explanatory variable.
Randomized Block Designs: The experimenter divides participants into subgroups called blocks, such that the variability within blocks is less than the variability between blocks.
Matched Pairs Design: A special case of the randomized block design. It is used when the experiment has only two treatment groups, and participants can be grouped into pairs, based on one of more blocking variables.
Give me the Explanatory and response variables, treatment and control groups, and randomization method: A researcher wants to study how much a new drug lowers blood sugar levels in teens and young adults who have abnormally high blood sugar levels. 100 teenagers and 100 young adults who have abnormally high blood sugar levels are randomly selected to be in this study. The blood sugar levels of the subjects are measured. Then, the teenager group will be randomly separated into one group of 50 who will receive the actual drug and 50 who will receive the placebo pill. The same will be done for the young adult group. 8 weeks later, the change in blood sugar levels will be recorded.
Explanatory: Do they receive the drug or placebo
Response: Change in blood sugar after 8 weeks
Treatment: 50 adults, 50 teenagers → receive the drug
Control: 50 adults, 50 teenagers → receive the placebo
Randomization method → Randomized Block Design
What is the placebo effect?
The tendency to react to a drug or treatment regardless of its actual physical function.
What is the bias of the Subjects?
Similar to response bias in sampling, subjects may want to please the researcher or hope for a specific outcome.
What is the Hawthorne effect?
When people behave differently because they know they are being watched.
What is the bias of the researcher?
People subconsciously behave in ways that favor what they believe. Researchers are no different; they may assign subjects to groups or report results in a biased way. They also may treat animals and people different based on the expectations of their treatment.
How do we reduce bias in experimental studies?
Blinding → when the individuals associated with an experiment, as either a subject or experimenter, are not aware of how subjects have been assigned (treatment or control, treatment or placebo).
What is a single blind study?
Those who could influence the results (subjects, administrators, technicians, ect.) are blinded.
What is a double blind study?
Those who evaluate the results (judges, physicians, analysts, ect.) are blinded as well.
Sometimes you cannot reasonably use a blinding method, why?
Say, for example, you have a situation where you have high blood pressure patients and put them on two different diets. Half is low-fat, and the other is low-carb. These individuals can taste what they are eating, so it does not necessarily matter if you blind them, since they could figure out the diet.
What is helpful in organizing large data sets?
Frequency distribution/ Frequency table
What are some common measures of center?
Mean, Median, Mode
What does this stand for? → Σ
This is an uppercase “sigma” → It denotes a sum
What does this stand for? → x
This is a lowercase x → It denotes an individual data value
What does this stand for? → n
This is lowercase n → It denotes the number of values in a sample. Also called the sample size.
What does this stand for? → N
This is uppercase N → It denotes the number of values in a population
What does this stand for? → x̄
It denotes the sample mean.
What does this stand for? → μ
It denotes the population mean.
How is the mean found?
Found by adding all values and dividing by the number of values in the set.
What is the median, and how is it found?
This is the value that is in the middle when listed in ASCENDING order. This is what separates the bottom 50% from the top. If there is not an odd amount of number, take the average of the 2 middle numbers.
What is the mode, and how is it found?
If it exists, it is the value that occurs with the greatest frequency. A data set may have no mode. If the data set has one mode is called unimodal. A data set with multiple modes is called multimodal. A data set with two modes is called bimodal.
Give me the outlines for Mean, Median, and Mode.
Mean → Uses every data value, highly affected by outliers, NOT GOOD for skewed data sets, BUT best for SYMMETRIC data
Median → Not affected by outliers; can be used with ANY data set
Mode → Not necessarily in the center, not affected by outliers, ONLY useful for multimodal or qualitative data
What is the graph of frequency distribution called?
A histogram
Describe a histogram.
Bars of equal width are drawn adjacent to each other (unless there are gaps), a horizontal scale representing quantitative data values, and a vertical scale representing frequency
What is a dotplot?
Shows each value in a dataset as a dot above a number line. There is no y-axis displayed on the dot plot, so you have to count.
What does a normal distribution look like?
It looks like a bell shape; this is unimodal, symmetric, and normal.
What does a right skew look like?
The data is positively skewed; this data comes out towards the left.
What does a left skew look like?
This data is negatively skewed; this data comes out towards the right.
What is a uniform distribution?
Equal spread with no peaks; all the data is the same height and width.
What are the two types of bimodal data you can have, and what is unique?
Symmetric bimodal and Non-symmetric bimodal; these both have two modes.
In a symmetric skew what happens to the mean, median, and mode?
The mean = median = mode; these measures occur under the peak.
In a right-skewed (positively skewed ) distribution, what happens to the mean, median, and mode?
Mode < Median < Mean
The mode will usually occur at the peak, and if there are any outliers, they appear on the right side.
In a left skew (negatively skewed), what happens to the mean, median, and mode?
Mean < Median < Mode
The mode will usually occur at the peak, and if there are any outliers, they appear on the left side.
What are the common measures of variation (measures of spread).
Range and interquartile range, variance, and standard deviation.
What is range, and how do you find it?
The difference between the maximum and the minimum, this will be highly affected by outliers.
Range = maximum data value - Minimum data value
What is the Interquartile range and how do you find it?
Use quartiles to provide a range of values that are not as affected by potential outliers.
Quartiles are values that separate a data set into fourths →
Q1 = the first quartile → Take the median of the numbers BELOW the entire set’s median
Q2 = the second quartile, aka the entire set median
Q3 = the third quartile → Take the median of the numbers ABOVE the entire set’s median
What is the IQR formula?
Difference between the third and first quartile.
IQR = Q3 - Q1
What is a boxplot (or box and whisker plot)?
A visual representation of the 5-number summary that also helps identify outliers.
This can be displayed horizontally or vertically.
An outlier will show up as a circle outside the box plot.
You can also notice skewness based on which ‘whisker’ is longer.
What is variance?
Variance, by definition, is the square of the standard deviation.
Variance = (standard deviation)²
Therefore, standard deviation is the square root of the variance → SD = √variance
True or False: The value of the SD is never negative. It is zero ONLY when all of the data values are the same.
True
What does this stand for → σ²
This is a lowercase sigma THAT IS squared; this stands for Population Variance.
What does this stand for → σ
This is a lower case sigma WITHOUT the square; this is Standard Deviation
What does this stand for → s²
This is a lowercase s THAT IS squared; this stands for Sample Variance