1/151
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
What are the goals of descriptive and inferential statistics? In what ways do scientists depend upon their use in descriptive, correlational, and experimental research designs?
Descriptive statistics summarize and describe what the data looks like (averages, spread, graphs). Inferential statistics use a sample to make educated guesses about a larger population and to decide whether results are likely due to chance. Scientists use descriptive stats to describe patterns in data, correlational stats to describe relationships between variables, and inferential stats in experiments to decide whether an effect is real or just random.
What are the three core elements of “good” research?
Good research has control (limits other explanations), randomization (reduces bias), and replication (results can be repeated).
What types of scales are used in the measurement of independent or dependent variables? How does each scale differ, and what information is added as one increments from nominal to ordinal to interval to ratio scales? Can you think of an example of each?
Nominal: categories only, no order (e.g., major).
Ordinal: ranked order, but gaps aren’t equal (e.g., class year).
Interval: equal spacing, no true zero (e.g., temperature in Fahrenheit).
Ratio: equal spacing and a true zero (e.g., hours studied).
Each step adds more meaningful information and allows more types of math to make sense.
How might one go about graphing a frequency distribution of a set of data? What is the advantage of presenting data in such a graphical format?
Put score values or ranges on the x-axis and how often they occur on the y-axis, then draw bars or a curve. Graphs make it easy to see the shape of the data, where most scores are, and whether there are outliers.
What are the three primary Measures of Central Tendency? How is each affected by extreme scores? In normal vs. skewed distributions, what is the relationship of each of the three measures with each other?
Mean is the average, median is the middle score, and mode is the most common score. Extreme scores strongly affect the mean, slightly affect the median, and usually do not affect the mode. In a normal distribution, mean, median, and mode are about the same. In a skewed distribution, the mean gets pulled toward the tail.
In what ways can we assess the variability of our data? What information does each measure give us regarding how scores are distributed around the mean?
Range shows the distance between the highest and lowest scores. Variance shows how spread out scores are from the mean on average. Standard deviation shows the typical distance of scores from the mean. These tell you how spread out or clustered the data is.
We will be using inferential statistics this semester to ask what particular question about our experimental groups?
We ask whether the differences between groups are likely caused by the treatment or are just due to random chance.
How do bias and random error differ in terms of “systematic” versus “unsystematic” direction?
Bias is a systematic error that pushes results in one direction. Random error is unsystematic noise that adds unpredictable variation but does not push results in one consistent direction.
What makes an estimator unbiased?
An estimator is unbiased if it does not consistently overestimate or underestimate the true population value.
Is the sample mean considered a biased or unbiased estimator of the population mean?
The sample mean is an unbiased estimator of the population mean.
Why is the sample variance formula corrected with (n–1) instead of just (n)?
Using n–1 corrects for the fact that samples tend to underestimate the true population variance, making the estimate more accurate on average.
What are the identifying features of an experimental method in research? In what ways are the goals of experiments different from the goals of descriptive and correlational research?
Experiments manipulate an independent variable and use random assignment. The goal is to test cause and effect. Descriptive and correlational research describe patterns and relationships but do not establish causation.
Within an experiment, how do we identify an INDEPENDENT variable and a DEPENDENT variable?
The independent variable is what the researcher changes. The dependent variable is what the researcher measures as the outcome.
What is internal validity?
Internal validity is how confident we are that the independent variable caused the change in the dependent variable.
What are confounds? How do we identify and control for them?
A confound is anything besides the independent variable that could explain the results. We identify confounds by asking what else differs between groups. We control for them using random assignment, keeping conditions the same, and treating groups equally.
What are the different ways that we can assign subjects to conditions that allow us to provide reasonable control for confounds?
Random assignment, matched pairs, and random sampling help reduce confounds.
What are the sources of variance in an experimental design?
Systematic variance from the independent variable and error variance from random noise and individual differences.
What do we hope contributes to the systematic variance in our experimental designs?
Only the independent variable.
What do we hope does not contribute to systematic variance in our experimental designs?
Confounds and outside factors.
What are the three specific criteria required to determine causality?
The variables must be related, the cause must come before the effect, and there must be no confounds.
What is error variance? Why do we expect it to occur in psychological studies?
Error variance is random noise in the data. It happens because people differ, measurements are imperfect, and situations vary.
How might we reduce error variance from our measures?
By using reliable measures, standardizing procedures, and controlling the environment.
What is the downside to imposing such tight control on our experiments?
The study may become unrealistic and not generalize well to real life.
What is external validity? What could we do to make experiments more externally valid?
External validity is how well results apply to the real world. We can increase it by using more realistic settings, diverse participants, and natural behaviors.
Why does increasing internal validity often decrease external validity?
More control improves cause-and-effect conclusions but makes the situation less realistic, reducing generalizability.
What are the different means by which we can do probability sampling of a population?
Simple random sampling, stratified sampling, cluster sampling, and multistage sampling.
What is the primary advantage of probability sampling?
It produces samples that better represent the population.
If a researcher does not do probability sampling, what are the other options?
Convenience sampling, quota sampling, and volunteer samples.
Why would a researcher consider using a non-probability sample?
Because experiments focus more on testing cause and effect than perfectly representing a population, and non-probability samples are faster and cheaper.
Why should we care about probability for this class?
Because inferential statistics use probability to decide whether results are likely due to chance.
If given frequencies of members of a population, could you calculate the probability of choosing one type?
Yes. Probability is the number of favorable outcomes divided by the total number of outcomes.
What is the difference between a theoretical probability distribution and a frequency distribution from actual sampling? How should they relate?
The theoretical distribution is the ideal curve predicted by math. The empirical distribution is what you actually observe from data. They should look similar if the theory fits the data well.
What are the characteristics of a normal curve?
It is bell-shaped, symmetrical, with most scores near the mean and fewer scores toward the extremes.
What two things does a z-score tell you?
How far a score is from the mean and whether it is above or below the mean.
How does the standard normal curve relate to z-scores?
Z-scores map data onto the standard normal curve, where about 68% of data fall between –1 and +1 standard deviations.
What are the advantages of converting data into z-scores?
It allows comparison across different scales, shows how extreme a score is, and lets you find probabilities.
How do you use z-scores and tables to determine probabilities?
You convert a raw score to a z-score, then use the z-table to find the proportion of data below or above that value.
What are the proper symbols for population and sample means and standard deviations?
Population mean is μ, population standard deviation is σ. Sample mean is x̄, sample standard deviation is s.
How is standard error different from standard deviation?
Standard deviation describes the spread of individual scores. Standard error describes spread of sample means.
What happens to the standard error as sample size increases?
It gets smaller and the curve becomes narrower.
How would you construct a sampling distribution of the mean?
Take many samples from the same population, compute the mean of each sample, and graph those means.
Which sampling distribution would have a smaller standard error: n = 10 or n = 40?
The sampling distribution with n = 40 would have a smaller standard error.
How does the sampling distribution relate to the population mean and standard deviation? What does the central limit theorem say?
The mean of the sampling distribution equals the population mean. The standard deviation of the sampling distribution (standard error) is smaller than the population standard deviation. The central limit theorem says that with large samples, the sampling distribution becomes normal even if the population is not.
How do we compute confidence intervals around the sample mean?
By using the sample mean plus or minus a margin of error based on the standard error and the desired confidence level.
What is the best way to describe a confidence interval? What advantage does it have over a point estimate?
A confidence interval is a range of values that likely contains the true population mean. It is better than a single number because it shows uncertainty and precision instead of pretending the estimate is exact.
In the z-tests & t-tests that we have covered since the first test, the statistic tests the probability of obtaining a group mean based upon the null hypothesis. You should also be able to read and write those in words or in symbolic form.
Why do statistics assume the null hypothesis?
Why do scientists typically test the null hypothesis rather than directly testing the experimental or alternative hypothesis?
Statistics assume the null hypothesis because it represents the idea that no effect or difference exists in the population. By assuming no effect, researchers can calculate the probability of observing the sample results if nothing is actually happening.
Scientists test the null hypothesis rather than the alternative because it is mathematically easier to evaluate probabilities under the assumption of no effect. If the results are very unlikely under the null hypothesis, we reject it and conclude that the alternative hypothesis is likely true.
In one- or two- group tests, we can have directional or nondirectional hypotheses. How do we express these, and how do they differ as we use our statistics to interpret the results of our experiment? When setting up a hypothesis for a study (such as a public health campaign for seat belt use), what is the difference between how you would express the alternative hypothesis and the null hypothesis in symbols?
A directional hypothesis predicts the direction of the effect (greater than or less than). A nondirectional hypothesis predicts that a difference exists but does not specify the direction.
Example:
Directional: H₁: μ₁ > μ₂
Nondirectional: H₁: μ₁ ≠ μ₂
The null hypothesis always states that there is no difference or no effect.
Example: H₀: μ₁ = μ₂
In statistical testing, directional hypotheses use one-tailed tests, while nondirectional hypotheses use two-tailed tests.
In the context of medical testing, explain what a Type I error and a Type II error would represent. Which one is considered a "false positive," and which is a "false negative"?
A Type I error occurs when we reject the null hypothesis even though it is actually true. In medical testing, this means diagnosing a disease when the person is actually healthy. This is called a false positive.
A Type II error occurs when we fail to reject the null hypothesis even though it is false. In medical testing, this means failing to detect a disease that the person actually has. This is called a false negative.
How does choosing a one-tailed test versus a two-tailed test affect the likelihood of rejecting the null hypothesis for a specific z-value (e.g., z = 1.67)?
A one-tailed test places all of the rejection region in one tail of the distribution. This makes it easier to reject the null hypothesis if the result is in the predicted direction.
A two-tailed test splits the rejection region between both tails, making it harder to reject the null hypothesis.
For example, a z-score of 1.67 might be significant in a one-tailed test, but not in a two-tailed test because the cutoff value is larger.
What specific information must you know about a population (the parent population) in order to use a z-test instead of a t-test?
To use a z-test, we must know:
The population mean (μ)
The population standard deviation (σ)
The sample size (n)
If the population standard deviation is unknown, we must use a t-test instead.
Why is it considered scientifically unethical to change a research hypothesis from two-tailed to one-tailed after you have already collected and analyzed your data?
Changing from a two-tailed test to a one-tailed test after seeing the results is unethical because it artificially increases the chances of finding a significant result.
This is considered a form of p-hacking, where researchers manipulate statistical decisions to obtain significant outcomes rather than testing their original hypothesis honestly.
Conceptually, why do we use a t-test instead of a z-test when the population standard deviation is unknown?
When the population standard deviation is unknown, we estimate it using the sample standard deviation. This introduces extra uncertainty.
The t-distribution accounts for this extra uncertainty, especially when the sample size is small, making the t-test more appropriate than the z-test.
How does the sample size (n) relate to the df, and how does this affect the shape of the t-distribution curve? Why is the t-distribution referred to as a "family" of curves, and how does a t-distribution with very low degrees of freedom differ visually from a normal Distribution?
Degrees of freedom are usually related to sample size, typically:
df = n − 1
As sample size increases, degrees of freedom increase and the t-distribution becomes more similar to the normal distribution.
The t-distribution is called a family of curves because its shape changes depending on the degrees of freedom.
With very low df, the curve is wider and has thicker tails than a normal distribution, reflecting more uncertainty.
For each type of test, how is the systematic variability assessed? (hint, it has to do with the means of the treatment group or groups) How is unsystematic variability accounted for? What do the calculations for each of these statistics do with these measures of variability to obtain the z- or t-score?
Systematic variability refers to differences caused by the independent variable or treatment.
Unsystematic variability refers to random error or natural differences among participants.
Statistical tests calculate a ratio of these two types of variability. If systematic variability is large relative to unsystematic variability, the test statistic (z or t) becomes larger and the result is more likely to be significant.
When conducting an independent groups t-test, what is the fundamental question we are asking about the two groups in relation to the null hypothesis population?
The main question is whether the difference between the two sample means is larger than what we would expect by random chance if both groups came from the same population.
In other words, we test whether the two groups differ significantly from each other.
How do the symbolic representations of the alternative hypothesis and null hypothesis differ between a one-tailed (directional) test where you expect the experimental group to perform better, and a two-tailed (non-directional) test?
One-tailed test (directional):
H₀: μ₁ ≤ μ₂
H₁: μ₁ > μ₂
Two-tailed test (nondirectional):
H₀: μ₁ = μ₂
H₁: μ₁ ≠ μ₂
Explain why the degrees of freedom for an independent groups t-test are calculated as df = (n_1 - 1) + (n_2 - 1). What does this reflect about how the mean and variance are estimated for each group?
Each sample loses one degree of freedom when estimating its mean. Because two samples are used, the degrees of freedom are the sum of the two:
This reflects that each group independently estimates its own mean and variance
How does a repeated-measures design help reduce unsystematic (error) variance compared to an independent groups design?
In repeated-measures designs, the same participants are measured multiple times. Because the same individuals are used, differences caused by individual variability are removed.
This reduces unsystematic error variance and makes it easier to detect the effect of the independent variable.
Conceptually, why is the "difference score" (D) the basis for the dependent t-test? How does the test use these scores to evaluate the effect of an independent variable?
The dependent t-test focuses on the difference between two measurements for each participant. These differences are called difference scores (D).
The test determines whether the average difference score is significantly different from zero, which would indicate that the treatment had an effect.
Under what circumstances might a researcher choose a dependent samples t-test over an independent samples t-test? Mention at least one specific advantage related to "individual differences".
researchers use dependent samples t-tests when the same participants are measured twice, such as before-and-after studies.
An advantage is that individual differences are controlled, which reduces random error and increases statistical power.
When is it impossible to use a repeated measures design, and how does "matched random assignment" serve as an alternative to achieve a similar statistical goal?
Repeated measures cannot be used when the treatment would permanently change participants or when participants cannot be tested multiple times.
In these cases, researchers may use matched random assignment, where participants are matched on important characteristics (such as age or IQ) and then randomly assigned to different groups.
Why does matching subjects on a relevant characteristic increase the statistical power of a study? Specifically, what happens to the "random error" in this process?
Matching reduces differences between participants that are unrelated to the treatment. This reduces random error (unsystematic variance), making it easier to detect the true effect of the independent variable.
What is the primary disadvantage of using a matched-groups design if the variable you matched on turns out to be irrelevant to the dependent measure? How does this impact
If participants are matched on a variable that is not actually related to the dependent variable, the matching process does not reduce error variance.
This wastes time and effort and does not improve the statistical power of the study.
We hope that our statistical tests are perfectly accurate, but they are based upon probability theory, which means that sometimes our statistic gives us an answer that isn’t accurate. Understand the difference between Type I and Type II errors, and be able to describe/interpret them.
A Type I error occurs when we reject the null hypothesis even though it is true (false positive).
A Type II error occurs when we fail to reject the null hypothesis even though it is false (false negative).
With regards to Type 1 errors, that error rate is formally set by the level of alpha that we choose (typically p of 0.05). Type 2 errors are just as bad (and maybe worse), but we do have tools that we can do to decrease the likelihood of making a Type 2 error, or, in other words, increasing the POWER of our study to detect a significant difference when there actually is one.
What can we do, during the design of our study, to increase the power of our experimental design? WHY does each of these techniques increase our power, from the point of view of the statistic?
Researchers can increase statistical power by:
Increasing sample size – reduces sampling error.
Increasing the strength of the treatment – increases systematic variance.
Reducing variability within groups – lowers random error.
Using a one-tailed test when appropriate – increases sensitivity.
Each of these increases the ratio of systematic variance to unsystematic variance, making it easier to detect real effects.
If one does not know the standard deviation of the population, how can one use t-values to compute a confidence interval for a mean?
When the population standard deviation is unknown, we estimate a confidence interval using the formula:
Mean ± (t × standard error)
The t-value comes from the t-distribution table based on degrees of freedom and the desired confidence level.
In the context of hypothesis testing, what is the mathematical and conceptual relationship between Power and a Type II Error (Beta)? If a researcher calculates that their study has a power of 0.80, what does this value specifically tell us about their chances of failing to find an effect that actually exists?
Statistical power is defined as:
Power = 1 − β
Where β represents the probability of making a Type II error.
If a study has power = 0.80, this means there is an 80% chance of correctly detecting a real effect and a 20% chance of missing the effect (Type II error).
Suppose you are conducting a study but find that your initial design has very low statistical power. Based on the "general formula" for parametric tests (Systematic Variance / Unsystematic Variance), identify two distinct ways you could modify your experiment to increase power. In your answer, explain how these changes specifically affect the numerator or denominator of the test statistic.
Increasing systematic variance – strengthening the treatment effect.
Reducing unsystematic variance – controlling variables, matching participants, or using repeated measures.
A researcher conducts a study with a massive sample size and finds a "highly significant" result (p < .001), but the actual difference between the groups is extremely small. Using the concepts from the slides, explain the difference between statistical significance and practical significance (effect size). Why can a very large sample size sometimes lead to results that are statistically significant but "unimportant" in the real world?
Statistical significance means that the observed effect is unlikely to have occurred by chance according to the statistical test.
Practical significance refers to whether the effect size is large enough to matter in real life.
With a very large sample size, even tiny differences between groups can become statistically significant because sampling error becomes very small. However, these differences may not be meaningful or important in practical terms.
What is the primary advantage of using a one-way ANOVA over multiple t-tests?
It lets you compare 3 or more groups at once while controlling Type I error better than doing many separate t-tests.
For the one-way ANOVA, how is the variance between groups compared? How is the error term derived (variance within groups)? Can you explain what these are in your own words?
Between-groups variance = how different the group means are from each other.
Within-groups variance = how much scores vary inside each group due to chance/error.
ANOVA compares them with: F = MS between / MS within.
In simple terms: “Are the groups more different from each other than we would expect by random variation alone?”
What does it mean when our F ratio is between zero and 1? Exactly at 1? Greater than 1?
F < 1: within-group variance is bigger than between-group variance, so no evidence of treatment effect.
F = 1: both variances are about the same, suggesting no treatment effect.
F > 1: between-group variance is larger; the bigger it gets, the more likely the treatment had an effect.
How do we express our statistical null hypothesis for the one-way ANOVA design? If we get a significant F-value and reject our null, what can we say about the groups?
Null hypothesis: all population means are equal. Example: μ1 = μ2 = μ3
If F is significant, we say at least one group mean differs from another.
We cannot say exactly which groups differ yet without post-hoc tests.
If we have a study with five groups in a one-way ANOVA and get a significant F-statistic, what might we do next? How might we use a post-hoc analysis to determine which group means are different? Which test is the least conservative? Which is the most conservative? In what cases might we use a less conservative test? In what cases might we feel obligated to use a more conservative test? What are the risks of using a less conservative test? What are the risks of using a more conservative post-hoc test?
Next, do post-hoc tests to compare specific pairs of means.
Least conservative: Fisher’s LSD
More conservative / most conservative of common options: Bonferroni is very conservative; Tukey’s HSD is also conservative and commonly used for all pairwise comparisons.
Use a less conservative test when you had a few planned comparisons and strong predictions.
Use a more conservative test when making many comparisons and you want to reduce false positives.
Risk of less conservative: more Type I errors (false positives).
Risk of more conservative: more Type II errors (missing real differences).
Beyond just “more groups,” explain the mathematical reason why conducting 45 separate t-tests for 10 experimental groups is problematic.
Because each t-test has its own chance of a Type I error, and those chances add up across many tests. With 10 groups, there are 45 pairwise comparisons, so the familywise error rate becomes much higher than .05.
If you use an alpha of .05 for 6 different pairwise comparisons, what is the actual probability that you will commit at least one Type I error?
Use: 1 − (1 − .05)^6 = 1 − (.95)^6 ≈ .265
So the chance is about 26.5%.
If the null hypothesis μ1 = μ2 = μ3 is rejected, list at least three different specific outcomes that could make it false.
Examples:
μ1 ≠ μ2, but μ2 = μ3
μ1 ≠ μ3, but μ1 = μ2
all three means are different
Basically, any pattern where not all means are equal makes the null false.
Why is the F-statistic expected to be approximately 1.00 if the treatment has no effect? What components of variance are being compared in that scenario?
Because if the treatment has no effect, then the variance between groups should be about the same size as the variance within groups.
So MS between ≈ MS within, giving F ≈ 1.
Explain why it is impossible to conduct a one-tailed F-test in an ANOVA.
Because F is a ratio of variances, and variances cannot be negative. So F values are always 0 or positive, and significance is only about whether F is large enough, not whether it is positive or negative.
What is the fundamental difference between a “planned comparison” and an “unplanned comparison”?
Planned comparison: decided before seeing the data, based on your hypotheses.
Unplanned comparison: decided after seeing the data, usually to explore differences.
Which post-hoc test is most likely to commit a Type I error? Which is most likely to commit a Type II error?
Most likely Type I error: Fisher’s LSD
Most likely Type II error: Bonferroni or very conservative post-hoc tests
How does the Bonferroni correction adjust the alpha level to control for multiple comparisons?
It divides the original alpha by the number of comparisons:
new alpha = α / number of tests
Example: .05 / 5 = .01
Be prepared to analyze a set of data for a one-way ANOVA, including specifying the null and interpreting the outcome.
What to remember:
H0: all group means are equal
F = MS between / MS within
If F observed > F critical, reject H0
If significant, conclude at least one mean differs
Then use post-hoc tests to find where the difference is.
Be prepared to calculate Tukey’s HSD post-hoc test on a One-Way ANOVA dataset if asked.
What to remember:
Compare each pair of means
If the mean difference is bigger than Tukey’s HSD value, that pair is significantly different.
What are the advantages of using a factorial design, rather than running multiple experiments focusing on manipulating only one independent variable at a time?
A factorial design lets you test:
the effect of IV1
the effect of IV2
the interaction between them
all in one study. It is more efficient and tells you whether the effect of one IV depends on the other.
What are the hypotheses that a 2-way factorial design tests? If our F-statistic is significant for any one of the hypotheses, how might we interpret the outcome of our study?
A 2-way ANOVA tests 3 null hypotheses:
no main effect of Factor A
no main effect of Factor B
no interaction between A and B
If one is significant:
significant A = Factor A matters overall
significant B = Factor B matters overall
significant interaction = the effect of one factor depends on the other.
If we obtain an interaction effect on a factorial design, how do we understand it?
Graph the cell means. An interaction means the lines are not parallel. The shape of the graph helps show how one IV changes across levels of the other
In the context of a factorial design, what is the difference between a “Factor” and a “Level”?
Factor = independent variable
Level = the categories or conditions within that IV
Example: Factor = dosage; levels = low, medium, high.
How many total experimental conditions (cells) are in a 4x2 factorial design? How many independent variables does this design have?
Cells: 4 × 2 = 8
Independent variables: 2
Compare running two separate independent t-tests versus one 2x2 factorial ANOVA. What specific piece of information does the factorial ANOVA provide that the t-tests cannot?
The factorial ANOVA tells you whether there is an interaction effect. Two separate t-tests cannot test that.
Define a “main effect” in your own words. Can a study have a main effect for Factor A but not for Factor B?
A main effect means one IV has an overall effect on the DV when averaging across the levels of the other IV.
Yes, a study can have a main effect for A but not for B.
Describe a scenario where an interaction exists because the effect of an IV is in the same direction for both levels of the second IV, but is significantly stronger at one level than the other.
This is an interaction where both lines go the same way, but one line is steeper.
Example: caffeine improves memory for both young and old adults, but the improvement is much bigger for young adults.
Describe an interaction where an IV has an effect at only one level of the other IV.
Example: a study method improves test scores only for students in quiet rooms, but not in noisy rooms.
What characterizes an interaction where the effects of one IV go in completely opposite directions depending on the level of the second IV?
This is a crossover interaction. The lines cross.
Example: treatment helps Group 1 but hurts Group 2.
List the four specific sources of variance that the “Total Variance” is broken down into for a two-factor ANOVA.
Variance due to Factor A
Variance due to Factor B
Variance due to A × B interaction
Within-groups variance (error)
In a standard between-subjects factorial ANOVA, what does the “Within-groups variance” actually measure?
It measures random error and individual differences within each cell that are not explained by either IV or their interaction.
Be prepared to analyze a set of data for a two-way ANOVA, including specifying the null hypotheses and interpreting the outcome.
Remember:
3 null hypotheses: no A effect, no B effect, no interaction
A significant interaction is usually the most important result to interpret first
Graph the means to understand the interaction.
What advantages does a repeated measures design give us? Why might we use RM rather than a between-groups design? What concerns do we need to address?
Advantages:
same participants in every condition
fewer participants needed
controls for participant differences
usually more power / less error variance
Use it when you want to compare the same people across conditions or time.
Concerns:
order effects
practice effects
fatigue effects
These are handled with things like counterbalancing.
Provide an example of a research question that requires a repeated measures design rather than a one-way independent groups design.
“How does the same person’s stress level change before, during, and after finals week?”
Same participants are measured multiple times.
Unlike a one-way ANOVA, a repeated measures ANOVA removes a specific type of variance from the error term. What is this variance called, and why is removing it beneficial?
It removes between-subjects variance or participant differences.
This helps because people naturally differ from one another, and removing that from error makes it easier to detect real treatment effects.