Comprehensive Study Guide on Single Sample and Dependent Means t Tests
Overview of Inferential Tests: Z Test vs. t Test
Scope and Course Context:
Inferential statistics covers both the Z test and the t test (Textbook Chapter 7).
Chapter 6 (Power and Effect Size) is excluded from this course scope.
Review of the Z Test:
The population mean (μ) and population variance (σ2) or standard deviation (σ) are fully known.
An individual raw score (X) or sample mean (M) is converted into a Z score.
The calculated Z score is evaluated against the standard normal distribution (Z distribution).
Introduction to the t Test:
The population variance (σ2) is unknown.
Population variability is estimated directly from sample data.
A t score is calculated and evaluated against a t distribution.
Summary Comparison Matrix
Population Variance Known:
Z Test: Yes
Single Sample t Test: No
Dependent Means t Test: No
Population Mean Known:
Z Test: Yes
Single Sample t Test: Yes (sample mean is compared to a known population mean)
Dependent Means t Test: No (assumes the population mean of difference scores is μD=0)
Original Scores per Participant or Pair:
Z Test: 1 score per person
Single Sample t Test: 1 score per participant
Dependent Means t Test: 2 scores per individual or paired set
Score Analyzed in Procedure:
Z Test: Individual score (X) or sample mean (M)
Single Sample t Test: Individual score (X) or sample mean (M)
Dependent Means t Test: Within-individual or within-pair difference score (D)
Shape of Comparison Distribution:
Z Test: Z distribution (standard normal distribution)
Single Sample t Test: t distribution
Dependent Means t Test: t distribution
Degrees of Freedom (DF):
Z Test: Not applicable
Single Sample t Test: DF=N−1
Dependent Means t Test: DF=N−1 (where N represents the number of participants or pairs)
Test Formulas:
Z Test: Z=σMM−μ
Single Sample t Test: t=SMM−μ
Dependent Means t Test: t=SMDMD−μD
t Test for a Single Sample
Definition and Application:
A single sample t test compares a sample mean (M) to a known population mean (μ) when the population variance (σ2) is unknown.
Example Scenario: Comparing college students' average stress scores to a known national average of 4 on a 7-point rating scale (4out of7).
Estimating Population Variance from Sample Data:
Sample variability provides an estimate of population variability.
Larger population variability yields larger sample variability; smaller population variability yields smaller sample variability.
Tendency to Underestimate: Sample variance calculated with standard division by N tends to be slightly smaller than true population variance, especially in small samples.
Unbiased Population Variance Estimate Formula (S2): S2=N−1SS
Adjustment Mechanism: Dividing by a slightly smaller number (N−1 instead of N) produces a slightly larger variance estimate, correcting for the bias to underestimate.
Degrees of Freedom (DF):
Definition: The amount of independent information available for estimating a population parameter after accounting for estimated parameters.
Formula for Single Sample t Test: DF=N−1
Conceptual Basis: Calculating the sample mean consumes 1 piece of independent information, leaving exactly N−1 scores free to vary independently when estimating variance.
Distribution of Means, t Distribution, and t Table
Estimated Standard Deviation of the Distribution of Means (Standard Error, SM):
Estimated Variance of the Distribution of Means: SM2=NS2
Estimated Standard Error of the Distribution of Means: SM=SM2=NS2
Review of Distribution of Means Formation:
Repeated random samples of size N are drawn from the population.
The mean (M) of each sample is calculated.
The resulting collection of sample means forms the distribution of means.
The standard error (SM) quantifies how much sample means typically vary from one another.
Symbol Notation Reference:
Sample Parameter (Unadjusted): Standard Deviation = Ssample, Variance = Ssample2
True Population Parameter: Standard Deviation = σ, Variance = σ2
Unbiased Population Estimate: Standard Deviation = S, Variance = S2
Distribution of Means Estimate: Standard Error = SM, Variance of Means = SM2
Characteristics of the t Distribution:
A mathematically defined probability curve used as the comparison distribution in t tests.
Why Not Standard Normal? Estimating population variance adds extra uncertainty. To account for this uncertainty, the t distribution is more spread out and has heavier tails (more probability in extreme regions) than a standard normal distribution.
Family of Distributions: A unique t distribution exists for every specific value of degrees of freedom (DF).
Effect of Sample Size (N) and DF:
Smaller sample size / lower DF: Higher uncertainty, heavier tails, more spread out t distribution.
Larger sample size / higher DF: Lower uncertainty, distribution approaches standard normal.
Infinite Degrees of Freedom (DF=∞): The t distribution is identical to the standard normal Z distribution.
Using the t Table (Appendix Table 2):
Provides critical cutoff scores across various degrees of freedom (DF), significance levels (α=0.05, etc.), and directional setups (1-tailed vs. 2-tailed).
Displays positive values. For 2-tailed tests, both positive and negative versions (±tcritical) are used.
Critical values are more extreme for t distributions than for standard normal distributions, particularly when DF is small.
Concept of the t Score:
Measures the number of standard errors a sample mean is located away from the population mean.
Evaluated relative to a t distribution with DF = N - 1$.\n\n# Step-by-Step Example: Single Sample t Test\n\n- Context and Research Question:\n - Do average hours studied per week by students in a dormitory exceed the university average?\n - Known Population Mean (General University Average): \mu = 17\,\text{hours/week}\n - Sample Size: N = 16\,\text{students}\n - Sample Mean: M = 21\,\text{hours/week}\n - Sum of Squared Deviations: SS = 694\n - Significance Level: \alpha = 0.05\n - Test Direction: One-tailed test (predicts study hours exceed general average)\n- Step 1: Restate Question as Hypotheses\n - Population 1: Students living in the dormitory.\n - Population 2: General university students.\n - Research Hypothesis (H_1):Dormitorystudentsstudymorethangeneraluniversitystudents(\mu_1 > \mu_2).\n - Null Hypothesis (H_0):Dormitorystudentsdonotstudymorethangeneraluniversitystudents(\mu_1 \le \mu_2).\n- Step 2: Determine Comparison Distribution Characteristics\n - Distribution Mean (\mu_M):\mu_M = \mu = 17\n - Degrees of Freedom: DF = N - 1 = 16 - 1 = 15\n - Estimated Population Variance (S^2):\n S^2 = \frac{SS}{DF} = \frac{694}{15} = 46.27\n - Variance of Distribution of Means (S_M^2):\n S_M^2 = \frac{S^2}{N} = \frac{46.27}{16} = 2.891875\n - Standard Error (S_M):\n S_M = \sqrt{2.891875} = 1.70\n- Step 3: Determine Critical Cutoff Score\n - Setup: 1−tailedtest,\alpha = 0.05,DF = 15\n - Critical value from ttable:t_{\text{critical}} = +1.753\n - Rejection Region: Top 5\%(t > 1.753)\n- Step 4: Calculate Sample t Score\n t = \frac{M - \mu}{S_M} = \frac{21 - 17}{1.70} = \frac{4}{1.70} = 2.35\n- Step 5: Make Decision and Interpret Results\n - Sample score t = 2.35 > 1.753 falls into the rejection region.\n - Result is statistically significant at \alpha = 0.05.\n - Reject null hypothesis (H_0)andsupportresearchhypothesis(H_1).\n - Conclusion: Dormitory students study significantly more hours per week than general university students.\n- Standard Journal Reporting Format (APA style):\n t(15) = 2.35, p < 0.05, \text{one-tailed}\n - Structure: t(\text{DF}) = \text{calculated score}, p < \text{significance level}, \text{tail type}\n\n# t Test for Dependent Means\n\n- Definition and Overview:\n - Used when two scores are obtained for each individual participant or paired subjects.\n - Neither population mean nor population variance is known.\n - Focuses entirely on difference scores (D) for each participant or pair.\n- Difference Score (D) Calculation:\n D = X_{\text{after}} - X_{\text{before}}\n - Difference scores are also referred to as change scores.\n - Example: Anxiety score before therapy = 30,aftertherapy=20.\n D = 20 - 30 = -10\n The negative score indicates a 10-point reduction in anxiety.\n- Assumption Under Null Hypothesis (H_0):\n - Assumes the population mean of difference scores is zero (\mu_D = 0).\n - Represents a scenario of no mean change or difference in the population.\n- Formation of Comparison Distribution (Distribution of Means of Difference Scores):\n - Repeated random samples of size N are selected from the population.\n - Each individual in a sample provides two paired scores, generating individual difference scores (D).\n - The mean difference score (M_D) is calculated for each sample.\n - The collection of sample mean differences forms a distribution centered around \mu_{M_D} = 0\n\n# Step-by-Step Example: Dependent Means t Test\n\n- Context and Research Question:\n - Do husbands receiving ordinary premarital counseling change in communication quality from before to after marriage?\n - Sample Size: N = 19\,\text{husbands}\n - Raw Difference Scores (D = X_{\text{after}} - X_{\text{before}}):\n - Sum of Difference Scores (\sum D):-229\n - Mean Difference Score (M_D):\n M_D = \frac{\sum D}{N} = \frac{-229}{19} = -12.05\n (On average, communication quality scores decreased by 12.05\,\text{points})\n - Sum of Squared Deviations of Difference Scores (SS_D):2762.90\n - Significance Level: \alpha = 0.05\n - Test Direction: Two-tailed test (tests for change in either direction)\n- Step 1: Restate Question as Hypotheses\n - Population 1: Husbands receiving premarital counseling.\n - Population 2: Husbands whose communication quality does not change from before to after marriage.\n - Research Hypothesis (H_1):\mu_D
eq 0 (Communication quality changes on average).\n - Null Hypothesis (H_0):\mu_D = 0 (Communication quality does not change on average).\n- Step 2: Determine Comparison Distribution Characteristics\n - Mean Difference under H_0(\mu_{M_D}):0\n - Degrees of Freedom: DF = N - 1 = 19 - 1 = 18\n - Estimated Population Variance of Difference Scores (S_D^2):\n S_D^2 = \frac{SS_D}{DF} = \frac{2762.90}{18} = 153.494\n - Variance of Distribution of Mean Differences (S_{M_D}^2):\n S_{M_D}^2 = \frac{S_D^2}{N} = \frac{153.494}{19} = 8.0786 \approx 8.08\n - Standard Error of Mean Differences (S_{M_D}):\n S_{M_D} = \sqrt{8.08} = 2.84\n- Step 3: Determine Critical Cutoff Scores\n - Setup: 2−tailedtest,\alpha = 0.05,DF = 18\n - Rejection region split: 2.5\%(0.025) in each tail\n - Critical values from ttable:t_{\text{critical}} = \pm 2.101\n - Rejection Regions: t > +2.101ort < -2.101\n- Step 4: Calculate Sample t Score\n t = \frac{M_D - \mu_D}{S_{M_D}} = \frac{-12.05 - 0}{2.84} = -4.24\n- Step 5: Make Decision and Interpret Results\n - Sample score t = -4.24 < -2.101 falls into the lower rejection region.\n - Reject null hypothesis (H_0)andsupportresearchhypothesis(H_1).\n - Conclusion: Husbands who received counseling changed significantly in communication quality; specifically, communication quality decreased significantly from before to after marriage.\n\n# Paired Designs in Dependent Means t Tests\n\n- Scope Beyond Repeated Measures:\n - Dependent means t tests do not strictly require two repeated measures from the exact same participant.\n - The procedure applies to any two measurements that are paired or inherently dependent rather than independently sampled.\n- Examples of Paired Designs:\n - Sibling Pairs: "Do younger siblings receive more parental attention than older siblings?" (Each pair consists of one younger and one older sibling; comparisons are evaluated within sibling pairs).\n - Married Couples: "Do husbands and wives differ in relationship satisfaction?" (Each husband is paired with his wife; difference score D = X_{\text{husband}} - X_{\text{wife}}).\n\n# Statistical Assumptions of t Tests\n\n- Definition of Assumption:\n - A required mathematical condition for conducting a hypothesis testing procedure.\n - Forms the foundation for the accuracy of cutoff values in statistical probability tables.\n- The Normality Assumption:\n - Single Sample t Test Assumption: The population from which the sample was selected follows a normal distribution.\n - Dependent Means tTestAssumption:Thepopulationofdifferencescores(D) follows a normal distribution.\n- Robustness of t Tests:\n - Definition of Robustness: A statistical test's ability to maintain accuracy even when mathematical assumptions are violated.\n - t tests are robust to moderate violations of the normality assumption.\n - In practice, populations do not need to be perfectly normal; unless the normality assumption is severely violated, t$$ test results remain valid and reliable.