(L11/Ch9) T-Tests

0.0(0)
Studied by 0 people
call kaiCall Kai
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/99

flashcard set

Earn XP

Description and Tags

Key Terms

Last updated 1:07 PM on 10/2/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

100 Terms

1
New cards
Alpha-level (α-level)
The probability of making a Type I error (usually .05).
2
New cards
Alternative hypothesis (experimental hypothesis)
The prediction that there will be an effect, for example that your manipulation will have some effect or that variables will relate to each other.
3
New cards
Assumption of normality
It relates to the errors of the model (the residuals), not the observed data. If errors are not normal, estimates stay unbiased, but standard errors, confidence intervals and significance tests need normal errors or a large sample.
4
New cards
Beta-level (β-level)
The probability of making a Type II error. Cohen (1992) suggests a maximum of 0.2.
5
New cards
Binary variable/Dichotomous
A categorical variable with only two mutually exclusive categories.
6
New cards
Bootstrap
A technique for estimating the sampling distribution of a statistic by taking repeated samples, with replacement, from the data. The SD of the resulting distribution estimates the standard error, from which confidence intervals and significance tests are computed.
7
New cards
Boxplot (box-whisker diagram)
Shows the median in the centre of a box that spans the middle 50% of scores (the interquartile range). Whiskers extend to the highest and lowest extreme scores.
8
New cards
Categorical variable
A variable made up of categories of objects or entities.
9
New cards
Central limit theorem
With large samples (above about 30) the sampling distribution is normal whatever the shape of the population. For small samples the t-distribution approximates it better. The standard error of the mean equals s divided by the square root of N.
10
New cards
Cohen's d
An effect size that expresses the difference between two means in standard deviation units, d = (X1 minus X2) divided by s.
11
New cards
Confidence interval
A range around a statistic that is believed to contain the true value in a certain proportion of samples (for example 95%). In the other samples it does not, and you cannot know which case you are in.
12
New cards
Degrees of freedom
Essentially the number of entities free to vary when estimating a parameter. They matter for significance tests and determine the exact form of the distribution of test statistics such as t.
13
New cards
Dummy variables
A way of recoding a categorical variable into dichotomous variables that take only the values 0 or 1. One group is chosen as the baseline (for example a control group) and coded 0.
14
New cards
Effect size
An objective and usually standardized measure of the magnitude of an observed effect. Examples are Cohen's d, Glass's g and Pearson's r.
15
New cards
Error bar chart
Plots the mean with its 95% confidence interval as a line. Error bars can also use the standard error or standard deviation.
16
New cards
Experimental research
Research in which variables are systematically manipulated to see their effect on an outcome variable. The data can support statements about cause and effect.
17
New cards
General linear model
The linear model can cover different designs, such as comparing means of categorical predictors (t-test, ANOVA) and including both categorical and continuous predictors (ANCOVA).
18
New cards
Heterogeneity of variance
The opposite of homogeneity of variance. The variance of one variable is different across levels of another variable.
19
New cards
Homogeneity of variance
The assumption that the variance of one variable is stable (relatively similar) at all levels of another variable.
20
New cards
Hypothesis
A proposed, theory-driven explanation for a narrow phenomenon. It cannot be tested directly, so it is turned into predictions about measurable variables.
21
New cards
Independence of errors/Independent errors
The assumption that one residual (prediction error) does not influence another, so for any two observations the residuals are uncorrelated. For people, one person's error does not influence another person's.
22
New cards
Independent design/Between-groups design/Between-subjects design
Different treatment conditions use different people, so the data are independent.
23
New cards
Independent t-test
A test using the t-statistic that establishes whether two means collected from independent samples differ significantly.
24
New cards
Levene's test
Tests the hypothesis that the variances in different groups are equal. A significant result means the variances differ. With large samples, small differences can be significant. The author does not recommend it.
25
New cards
Linear model
A model of the form outcome = b times predictor + error. The key is that its form is linear, so with a single predictor it is a straight line.
26
New cards
Long format data
Scores on an outcome variable are in a single column, and each row is a combination of attributes of a score, such as the entity or condition. Scores from one entity can appear over several rows.
27
New cards
Mann-Whitney test
A non-parametric test for differences between two independent samples. It is functionally the same as Wilcoxon's rank-sum test, and both are non-parametric equivalents of the independent t-test.
28
New cards
Non-parametric tests
Procedures that do not rely on the restrictive assumptions of parametric tests. In particular, they do not assume a normally distributed sampling distribution.
29
New cards
Normal distribution
A probability distribution that is perfectly symmetrical (skew of 0).
30
New cards
Null hypothesis
The reverse of the experimental hypothesis. It states that your prediction is wrong and the predicted effect does not exist.
31
New cards
One-tailed test
A test of a directional hypothesis. The book generally advises against it because of the temptation to interpret interesting effects in the opposite direction to that predicted.
32
New cards
Outcome variable/Dependent variable
A variable whose values we are trying to predict from one or more predictor variables. It is called dependent in experiments because it is not manipulated, so its value depends on the manipulated variables.
33
New cards
Outlier
An observation very different from most others. Outliers bias statistics such as the mean and their standard errors and confidence intervals.
34
New cards
Paired-samples t-test/Dependent t-test
A test using the t-statistic that establishes whether two means collected from the same sample (or related observations) differ significantly. Also called the matched-pairs t-test.
35
New cards
Parameter
Describes relations between variables in the population. We use sample data to estimate its likely value because we cannot access the population directly.
36
New cards
Parametric test
A test based on a known distribution, usually the normal. It needs four assumptions, a normal sampling distribution, homogeneity of variance, interval or ratio data, and independence.
37
New cards
Population
The collection of units (people, plants, cities and so on) to which we want to generalize a set of findings or a model.
38
New cards
Power
The ability of a test to detect an effect of a particular size (0.8 is a good level to aim for). It equals 1 minus the Type II error rate.
39
New cards
Predictor variable/Independent variable
A variable used to try to predict values of an outcome variable. It is called independent in experiments because the experimenter manipulates it.
40
New cards
Q-Q plot (see IMG6)
Short for quantile-quantile plot. It plots the quantiles of a variable against those of a distribution, often the normal. Points on the diagonal mean the same distribution, and deviations show departures from it.
41
New cards
Raincloud plot (see IMG16)
A visualization that combines a raw data plot, a boxplot and a probability density plot.
42
New cards
Repeated-measures design/Within-subject design
Different treatment conditions use the same people, so the data are related. Also called a related design.
43
New cards
Residual
The difference between the value a model predicts and the value observed in the data. Basically an error. The collection of residuals for all observations are the residuals.
44
New cards
Robust methods
Procedures that give unbiased estimates, confidence intervals and significance tests even when the normal assumptions of the statistic are not met.
45
New cards
Sample
A smaller (hopefully representative) collection of units from a population, used to find out about that population.
46
New cards
Sampling distribution
The probability distribution of a statistic. It is the distribution of values we would get if we took lots of samples from a population and calculated the statistic for each.
47
New cards
Sampling variation
The extent to which a statistic (the mean, median, t, F and so on) varies in samples taken from the same population.
48
New cards
Shapiro-Wilk test
Tests whether a distribution of scores differs significantly from a normal distribution. A significant value means a deviation from normality, but large samples make even small deviations significant.
49
New cards
Standard deviation
An estimate of the average spread of a set of data in the original units of measurement. It is the square root of the variance.
50
New cards
Standard error
The standard deviation of the sampling distribution of a statistic. It shows how much the statistic varies across samples from the same population. Large values mean a sample may not reflect the population accurately.
51
New cards
Standard error of differences
The standard deviation of the sampling distribution of differences between sample means. It measures how variable those differences are.
52
New cards
Standardization
Converting a variable into a standard unit of measurement, typically standard deviation units. It lets us compare data measured in different units.
53
New cards
Systematic variation
Variation due to a genuine effect. It can be explained by the model fitted to the data.
54
New cards
t-statistic
A test statistic with a known distribution (the t-distribution). In the linear model it tests whether a b-value differs from zero, and for two means it tests whether their difference differs from zero.
55
New cards
Test statistic
A statistic for which we know how frequently different values occur. Its observed value is typically used to test hypotheses.
56
New cards
Two-tailed test
A test of a non-directional hypothesis.
57
New cards
Type I error
Believing there is a genuine effect in the population when there is not.
58
New cards
Type II error
Believing there is no effect in the population when there is.
59
New cards
Unsystematic variation
Variation not due to the effect of interest, such as natural differences in intelligence or motivation. It cannot be explained by the model fitted to the data.
60
New cards
Variance
An estimate of the average spread of a set of data. It is the sum of squares divided by the number of values minus 1.
61
New cards
Variance ratio (Hartley's Fmax)
The ratio of the variance of the group with the biggest variance to that of the group with the smallest. It is compared to Hartley's critical values as a test of homogeneity of variance.
62
New cards
Variance sum law
The variance of a sum or of a difference between two independent random variables equals the sum of their variances.
63
New cards
Wide format data (see IMG5)
Scores from a single entity appear in one row, and the levels of the predictor are arranged over different columns.
64
New cards
Wilcoxon signed-rank test
A non-parametric test for differences between two related samples. It is the non-parametric equivalent of the related t-test.
65
New cards
Difference score (lecture and Ch 9)
A person's score in one condition minus their score in the other. The paired-samples t-test tests whether the mean of these scores (D-bar) is 0.
66
New cards
p-value (lecture, see IMG2, IMG3 and IMG4)
The probability of observing this test statistic or one more extreme if the null hypothesis is true. If it is lower than alpha, H0 is rejected.
67
New cards
Pooled standard error (lecture and Ch 9)
The standard error of the difference that uses the pooled variance for both groups, so it assumes equal variances.
68
New cards
Pooled variance (lecture and Ch 9)
A weighted average of the two group variances, weighted by degrees of freedom (n - 1). It is used when group sizes differ, and you do not need to memorise the formula.
69
New cards
Student t-test (Ch 9)
The standard independent t-test. It assumes equal variances and is slightly more powerful.
70
New cards
t-distribution (lecture, see IMG1)
The sampling distribution of t under H0. It is bell-shaped and centred on 0, with fatter tails for small n (few degrees of freedom), and it converges to the standard normal as n grows.
71
New cards
Unpooled standard error (lecture, see IMG6)
The standard error of the difference that uses each group's own variance, the square root of (s1 squared over n1 + s2 squared over n2).
72
New cards
T-test in terms of the linear model
All t-tests use the linear model (outcome = model + error) to compare means. The independent t-test is a linear model with one two-category predictor, and t tests whether b for group (the difference between means) differs from 0. The one-sample t-test tests the mean against a reference value. The paired t-test is a one-sample test on the difference scores.
73
New cards
Three types of t-tests
One-sample compares one group mean to a hypothesized value. Paired-samples compares two means from the same or related entities (within-subjects). Independent-samples compares two means from different entities (between-subjects).
74
New cards
Degrees of freedom for each t-test
One-sample and paired, df = n - 1 (n is the number of pairs for the paired test). Independent (Student), df = n1 + n2 - 2. Welch, lower than that. The df set the shape of the t-distribution used to get p.
75
New cards
T-statistic as a signal-to-noise ratio
t = (observed difference minus the difference expected under H0) divided by the standard error of the difference. The top is the effect (signal, systematic variance) and the bottom is the error (noise, unsystematic variance). A larger absolute t means a larger signal relative to sampling noise.
76
New cards
Standard error - role in the t-test logic
The SE shows how much sample means (or differences) vary by chance alone. A small SE means even a modest difference is unlikely under H0, and a large SE means big differences are plausible by chance. The observed difference is judged relative to the SE.
77
New cards
Two explanations for a larger-than-expected difference
(1) There is no effect and we were unlucky, so sampling variation produced very different means. (2) The samples come from different populations, meaning a genuine effect and H0 is false. The bigger the difference relative to the SE, the more plausible the second.
78
New cards
What b0 and b1 represent
b0 (the intercept) is the mean of the baseline group coded 0. b1 is the difference between the two group means, so the t-test of b1 tests whether that difference is 0.
79
New cards
Median splits - why to avoid them
Dichotomizing a continuous variable at the median (1) distorts information, since similar people end up in opposite groups and dissimilar people in the same group, (2) shrinks effect sizes, and (3) increases the chance of spurious effects. It is only justified with a clear theoretical break point, such as a clinical diagnosis.
80
New cards
Variable types needed for a t-test
The outcome is a quantitative (continuous) variable, because t-tests compare means. In the independent t-test the predictor is categorical with exactly two groups. The paired t-test needs two measurements of the same outcome per participant. The one-sample t-test compares the outcome mean with a numerical reference value.
81
New cards
Postulated value
A value that is assumed or hypothesized rather than observed. H0 uses the postulated value as its model (for example model = 120), whereas HA uses the observed data (model = x-bar).
82
New cards
One-sided p-values
Only one tail counts, in the direction of HA. If t falls in the opposite direction to HA, p is large, always greater than 0.5.
83
New cards
Cohen's benchmarks for d
Small about .3, medium about .5 and large about .8 on the slide. They are only rough guides, so report the value and not just the label.
84
New cards
Paired t-test - what must be normal
The residuals, meaning the difference scores, should be roughly normal, especially with small samples. The raw scores in each condition do not need to be. Compute the differences and check them with a Q-Q plot. Two very non-normal measures can still give normally distributed differences.
85
New cards
Why repeated measures are more powerful
Using the same people removes individual differences such as IQ and motivation. Unsystematic (error) variance drops, so the SE is smaller and confidence intervals are narrower, which makes an effect easier to detect.
86
New cards
Paired design - advantage and trade-off
The advantage is more power, because stable individual differences are removed from the comparison. The trade-off is possible practice and boredom effects from repeating tasks, which the design can address with counterbalancing.
87
New cards
Counterbalancing
Systematically varying the order in which conditions are done. With two conditions, half the participants do A then B and the other half do B then A. It aims to remove systematic bias caused by practice effects or boredom effects.
88
New cards
Practice effect
Participants' performance on a task may be influenced (positively or negatively) if they repeat it, because of familiarity with the experimental situation and the measures used.
89
New cards
Boredom effect
Performance on a task may be influenced (assumed to be negatively) by boredom or lack of concentration when there are many tasks or the task goes on for a long time.
90
New cards
Independent t-test - what you need to calculate it
Only the group means, standard deviations (or variances) and sample sizes. You do not need the raw data.
91
New cards
Assumptions of all t-tests (lecture)
(1) Normally distributed residuals (model errors). (2) Random samples. (3) Independence, meaning observations are independent across participants or pairs and participants do not systematically influence one another. In a paired test the two scores within a person are deliberately dependent, but the pairs should be independent.
92
New cards
Extra assumption of the independent t-test
Equality (homogeneity) of variance, meaning both groups have equal variances. Assess it by comparing the observed SDs or variances, or with Levene's test. If it is violated, use Welch's t-test.
93
New cards
Levene's and Shapiro-Wilk tests - warning
Both are heavily influenced by sample size. In large samples tiny, harmless deviations become significant, and in small samples real problems can be missed. Do not treat them as pass or fail checks. Inspect variances and Q-Q plots, and prefer Welch where appropriate. A variance ratio above 2 is problematic.
94
New cards
Welch's t-test (lecture and Ch 9, see IMG6)
A version of the t-test that is robust to unequal variances, which bias the sampling distribution of t. It (1) uses the unpooled SE, the square root of (s1 squared over n1 + s2 squared over n2), and (2) lowers the df, more so the more unequal the variances are.
95
New cards
Student vs Welch rows
Student assumes equal variances and Welch does not. Field recommends reading the Welch row. The rows differ more when group sizes and SDs differ.
96
New cards
Practical default for an independent-samples test
The Welch t-test, because it is more robust when variances differ. It changes the SE (unpooled) and the df. It can have slightly less power when variances are truly equal, but it gives safer inference when equality of variance is doubtful.
97
New cards
Assessing the normality assumption
Use a Q-Q plot, where points should lie along the diagonal (not exact). A Shapiro-Wilk test is possible, where p below alpha means the assumption is violated, but with the same caution as Levene's test. Normality matters less as n grows (about 30).
98
New cards
Handling assumption violations
Violations distort the sampling distribution and the Type I and Type II error rates. Unequal variances call for Welch. Problematic non-normality calls for a non-parametric test, the Wilcoxon signed-rank test for paired samples or the Mann-Whitney test for independent samples. Bootstrapping is another option.
99
New cards
General procedure for a t-test
Explore the data, check for outliers, normality and homogeneity (boxplots, histograms, descriptives), run the t-test (or a non-parametric test if assumptions fail), and compute an effect size.
100
New cards
Key terms from Chapter 9
Dependent t-test, dummy variables, independent t-test, paired-samples t-test, standard error of differences and variance sum law.