Lecture 1: Two-Sample t -Test • Scope & Goals
• Compare the means μ < e m > 1 , μ < / e m > 2 \mu<em>1,\,\mu</em>2 μ < e m > 1 , μ < / e m > 2 of two independent populations.
• By the end you should be able to
– derive / recall the test statistic for a two-sample t -test;
– calculate 100 ( 1 − α ) % 100(1-\alpha)\% 100 ( 1 − α ) % CIs for μ < e m > 1 − μ < / e m > 2 \mu<em>1-\mu</em>2 μ < e m > 1 − μ < / e m > 2 .
• Position in the “road-map” of inference
┌── One variable ──┐ ┌──────────── Two variables ───────────┐
│categorical→C I ( p ) CI\ (p) C I ( p ) │ │both categorical→χ 2 \chi^2 χ 2 │
│quantitative→C I ( μ ) CI\ (\mu) C I ( μ ) │ │one cat.+one quant.→ two-sample or paired t │
│ │ │both quantitative→simple regression $t$ │
• Prototype examples
– “Pregnancy & Smoking” (nicotine vs control; guinea-pig maze errors).
– “Do ravens fly to gunshots?” (later turns out to require paired test).
• Terminology & Hypotheses
Population 1: RV X < e m > 1 ∼ ( μ < / e m > 1 , σ < e m > 1 ) X<em>1\sim(\mu</em>1,\sigma<em>1) X < e m > 1 ∼ ( μ < / e m > 1 , σ < e m > 1 ) ; Population 2: X < / e m > 2 ∼ ( μ < e m > 2 , σ < / e m > 2 ) X</em>2\sim(\mu<em>2,\sigma</em>2) X < / e m > 2 ∼ ( μ < e m > 2 , σ < / e m > 2 ) .
Null H < e m > 0 : μ < / e m > 1 = μ < e m > 2 ( ! ⇔ ! μ < / e m > 1 − μ 2 = 0 ) H<em>0: \mu</em>1=\mu<em>2\;(!\Leftrightarrow!\mu</em>1-\mu_2=0) H < e m > 0 : μ < / e m > 1 = μ < e m > 2 ( ! ⇔ ! μ < / e m > 1 − μ 2 = 0 ) . One- or two-sided alternatives chosen a priori :
Ha:\ \mu 1<\mu2,\;\mu 1>\mu2,\;\text{or}\;\mu 1\neq\mu_2.
• Assumptions ("pooled" version)
Two simple random samples; observations inside and across samples independent.
Equal population SDs: σ < e m > 1 = σ < / e m > 2 = σ \sigma<em>1=\sigma</em>2=\sigma σ < e m > 1 = σ < / e m > 2 = σ (checked via similarity of s < e m > 1 , s < / e m > 2 s<em>1,s</em>2 s < e m > 1 , s < / e m > 2 ).
Either both samples are large (CLT) or both parent distributions \( \approx \) Normal.
• Pooled variance estimator & standard error
S < e m > p 2 = ( n < / e m > 1 − 1 ) S < e m > 1 2 + ( n < / e m > 2 − 1 ) S < e m > 2 2 n < / e m > 1 + n < e m > 2 − 2 S<em>p^2=\frac{(n</em>1-1)S<em>1^2+(n</em>2-1)S<em>2^2}{n</em>1+n<em>2-2} S < e m > p 2 = n < / e m > 1 + n < e m > 2 − 2 ( n < / e m > 1 − 1 ) S < e m > 1 2 + ( n < / e m > 2 − 1 ) S < e m > 2 2 S E ( X ˉ < / e m > 1 − X ˉ < e m > 2 ) = S < / e m > p 1 n < e m > 1 + 1 n < / e m > 2 SE(\bar X</em>1-\bar X<em>2)=S</em>p\sqrt{\frac1{n<em>1}+\frac1{n</em>2}} S E ( X ˉ < / e m > 1 − X ˉ < e m > 2 ) = S < / e m > p n < e m > 1 1 + n < / e m > 2 1
• Test statistic
T = X ˉ < e m > 1 − X ˉ < / e m > 2 S < e m > p 1 n < / e m > 1 + 1 n < e m > 2 ∼ H < / e m > 0 t < e m > ( n < / e m > 1 + n 2 − 2 ) T=\frac{\bar X<em>1-\bar X</em>2}{S<em>p\sqrt{\dfrac1{n</em>1}+\dfrac1{n<em>2}}}\;\stackrel{H</em>0}{\sim}t<em>{\,(n</em>1+n_2-2)} T = S < e m > p n < / e m > 1 1 + n < e m > 2 1 X ˉ < e m > 1 − X ˉ < / e m > 2 ∼ H < / e m > 0 t < e m > ( n < / e m > 1 + n 2 − 2 )
• P-value rules
• Right-tail (Ha: \mu 1>\mu2): P ( T ≥ t < / e m > o b s ) P(T\ge t</em>{obs}) P ( T ≥ t < / e m > o b s )
• Left-tail (Ha: \mu 1<\mu2): P ( T ≤ t < / e m > o b s ) P(T\le t</em>{obs}) P ( T ≤ t < / e m > o b s )
• Two-tail (H < e m > a : μ < / e m > 1 ≠ μ < e m > 2 H<em>a: \mu</em>1\neq\mu<em>2 H < e m > a : μ < / e m > 1 = μ < e m > 2 ): 2 P ( T ≤ − ∣ t < / e m > o b s ∣ ) . 2P(T\le -|t</em>{obs}|). 2 P ( T ≤ − ∣ t < / e m > o b s ∣ ) .
Robust to mild non-normality as long as X ˉ < e m > 1 − X ˉ < / e m > 2 \bar X<em>1-\bar X</em>2 X ˉ < e m > 1 − X ˉ < / e m > 2 is \($\approx\) Normal.
• Confidence interval (level C)
CIC(\mu 1-\mu2)=(\bar X 1-\bar X2)\;\pm\;t^{\,df}\,S p\sqrt{\dfrac1{n1}+\dfrac1{n 2}}w i t h with w i t h df=n1+n 2-2a n d and an d t^*t h e t w o − s i d e d ∗ t ∗ − q u a n t i l e c o v e r i n g a r e a C . < / p > < p > • < s t r o n g > W o r k e d e x a m p l e – G u i n e a − p i g n i c o t i n e s t u d y < / s t r o n g > < b r > < / p > < p > D a t a s u m m a r y : < b r > < / p > < p > │ G r o u p │ the two-sided *t*-quantile covering area C.</p><p>• <strong>Worked example – Guinea-pig nicotine study</strong> <br></p><p>Data summary: <br></p><p>│Group│ t h e tw o − s i d e d ∗ t ∗ − q u an t i l eco v er in g a r e a C . < / p >< p > • < s t r o n g > W or k e d e x am pl e – G u in e a − p i g ni co t in es t u d y < / s t r o n g >< b r >< / p >< p > D a t a s u mma r y :< b r >< / p >< p > │ G r o u p │ n│ │ │ \bar x│ │ │ s│ → C o n t r o l : 10 , 23.4 , 12.3 ; T r e a t m e n t : 10 , 44.3 , 21.5. < b r > < / p > < p > – P o o l e d │ → Control: 10, 23.4, 12.3; Treatment: 10, 44.3, 21.5. <br></p><p>– Pooled │ → C o n t r o l : 10 , 23.4 , 12.3 ; T r e a t m e n t : 10 , 44.3 , 21.5. < b r >< / p >< p > – P oo l e d Sp=17.515( R ) . – (R). – ( R ) .– t {obs}=\dfrac{23.4-44.3}{17.515\sqrt{0.1+0.1}}=-2.668
– One-sided P-value \( \approx \) 0.0078 \( \Rightarrow \) strong evidence nicotine increases errors.
– 95 % CI: (-20.4,-4.4) \($ \to \) errors higher by 4–37 on average.
• Unequal variances?
– If \sigma1\neq\sigma 2b u t but b u t n1=n 2, p o o l e d t e s t i s s t i l l f i n e . < b r > < / p > < p > – L a r g e i m b a l a n c e i n b o t h , pooled test is still fine. <br></p><p>– Large imbalance in both , p oo l e d t es t i ss t i l l f in e . < b r >< / p >< p > – L a r g e imba l an ce inb o t h na n d and an d \sigma \($ \to \) transform data or use Welch (not in MATH1041 curriculum).
Lecture 2: Paired t -Test • Motivation – samples are dependent (matching, before/after, twins…). Independence assumption of two-sample test fails.
• Examples
– Ravens before/after gunshot (12 pairs).
– Speed-camera counts before vs after installation (4 pairs).
• Strategy – reduce to a single sample of differences Di=X {1i}-X{2i}. P o p u l a t i o n p a r a m e t e r o f i n t e r e s t : . Population parameter of interest: . P o p u l a t i o n p a r am e t er o f in t er es t : \mu D=E(D). < / p > < p > • < s t r o n g > A s s u m p t i o n s < / s t r o n g > < / p > < o l > < l i > < p > .</p><p>• <strong>Assumptions</strong></p><ol><li><p> . < / p >< p > • < s t r o n g > A ss u m pt i o n s < / s t r o n g >< / p >< o l >< l i >< p > np a i r e d d i f f e r e n c e s a r e i n d e p e n d e n t . < / p > < / l i > < l i > < p > D i s t r i b u t i o n o f paired differences are independent.</p></li><li><p>Distribution of p ai r e dd i f f er e n ces a r e in d e p e n d e n t . < / p >< / l i >< l i >< p > D i s t r ib u t i o n o f D \($ \approx \) Normal (or n large for CLT).
• Test statistic & CI
T{paired}=\frac{\bar D}{S D/\sqrt n}\;\stackrel{H0}{\sim}t {\,n-1}w i t h with w i t h H0:\mu D=0. < b r > < / p > < p > . <br></p><p> . < b r >< / p >< p > CIC(\mu D)=\bar D\;\pm\;t^{*\,\,n-1}\,\dfrac{S_D}{\sqrt n}.
• R implementation
t.test(x_after, x_before, paired=TRUE, alternative="greater")
• Raven example recap
Differences: mean \bar D=1.083, S D , SD , S D SD=1.884, , , n=12. . . t {obs}=1.995, , , p=0.036 \($ \to \) moderate evidence more ravens post-shot.
One-sided 95 % CI: (0.11,\infty)( ( ( \ge0.11 r a v e n i n c r e a s e ) . < / p > < p > • < s t r o n g > S p e e d − c a m e r a e x a m p l e < / s t r o n g > < b r > < / p > < p > D i f f e r e n c e s h u g e ( a l l n e g a t i v e ) , 0.11 raven increase).</p><p>• <strong>Speed-camera example</strong> <br></p><p>Differences huge (all negative), 0.11 r a v e nin cr e a se ) . < / p >< p > • < s t r o n g > S p ee d − c am er a e x am pl e < / s t r o n g >< b r >< / p >< p > D i f f er e n ces h ug e ( a l l n e g a t i v e ) , t_{obs}=4.95, , , p=0.008, r e d u c t i o n , reduction , r e d u c t i o n \ge2222 c a r s / m o n t h ( o n e − s i d e d 95 2222 cars/month (one-sided 95 % CI).</p><p>• <strong>Paired vs two-sample</strong> <br></p><p>– Matched design controls extraneous variation (e.g. "location" in ravens). <br></p><p>– But if data are actually independent, pairing halves degrees-of-freedom \($ \to \) loss of power.</p><p>• <strong>Intro to power</strong> <br></p><p>– Type I error 2222 c a r s / m o n t h ( o n e − s i d e d 95 \alpha, T y p e I I e r r o r , Type II error , T y p e I I er r or \beta, P o w e r , Power , P o w er 1-\beta. < b r > < / p > < p > – B l o o d − p r e s s u r e e x a m p l e : . <br></p><p>– Blood-pressure example: . < b r >< / p >< p > – B l oo d − p r ess u r ee x am pl e : \sigma=12, , , n=100/100 \($ \Rightarrow \) SE=1.7. < b r > < / p > < p > – R e j e c t i o n r e g i o n f o r t w o − s i d e d . <br></p><p>– Rejection region for two-sided . < b r >< / p >< p > – R e j ec t i o n r e g i o n f or tw o − s i d e d \alpha=0.05: \(|\text{diff}|>3.33\).
– If true diff = -3 \( \to \) power \( \approx \) 42 % .
– Formula manipulation shows n\approx251p e r g r o u p n e e d e d f o r 80 per group needed for 80 % power.</p><div data-type="horizontalRule"><hr></div><h4 id="591ad4e5-06ac-4ba6-b97d-81c3f9649962" data-toc-id="591ad4e5-06ac-4ba6-b97d-81c3f9649962" collapsed="false" seolevelmigrated="true">Lecture 3: Data Analysis for Two-Way Tables</h4><p>• <strong>Goal</strong> – Summarise and visualise the joint distribution of two categorical variables; compute marginal & conditional distributions.</p><p>• <strong>Construction of a two-way (contingency) table</strong> <br></p><p>– Choose row and column variables, tally counts for each combination. <br></p><p>– Add row/column totals (margins) for marginal distributions.</p><p>• <strong>Key quantities</strong> <br></p><p>– <em>Marginal</em> distribution of a variable: row/column totals divided by p er g r o u p n ee d e df or 80 n. < / p > < p > – < e m > C o n d i t i o n a l < / e m > d i s t r i b u t i o n : d i s t r i b u t i o n o f o n e v a r i a b l e a t a f i x e d l e v e l o f t h e o t h e r , e . g . .</p><p>– <em>Conditional</em> distribution: distribution of one variable at a fixed level of the other, e.g. . < / p >< p > – < e m > C o n d i t i o na l < / e m > d i s t r ib u t i o n : d i s t r ib u t i o n o f o n e v a r iab l e a t a f i x e d l e v e l o f t h eo t h er , e . g . P(\text{North}|\text{Moderately})=28/84=0.33. < / p > < p > • < s t r o n g > G r a p h i c a l d i s p l a y s < / s t r o n g > < b r > < / p > < p > – S i d e − b y − s i d e b a r c h a r t s ( s p l i t b y l e v e l s ) . < b r > < / p > < p > – C l u s t e r e d b a r c h a r t s ( g r o u p e d ) . < b r > < / p > < p > – B a r c h a r t o f c o n d i t i o n a l p r o p o r t i o n s ( s t a c k e d w i t h 100 .</p><p>• <strong>Graphical displays</strong> <br></p><p>– Side-by-side bar charts (split by levels). <br></p><p>– Clustered bar charts (grouped). <br></p><p>– Bar chart of conditional proportions (stacked with 100 % scale).</p><p>• <strong>Illustrative data sets</strong></p><ol><li><p>Hemisphere (N/S) \( . < / p >< p > • < s t r o n g > G r a p hi c a l d i s pl a y s < / s t r o n g >< b r >< / p >< p > – S i d e − b y − s i d e ba r c ha r t s ( s pl i t b y l e v e l s ) . < b r >< / p >< p > – C l u s t er e d ba r c ha r t s ( g r o u p e d ) . < b r >< / p >< p > – B a r c ha r t o f co n d i t i o na l p r o p or t i o n s ( s t a c k e d w i t h 100 \times \) Stress Level (4 levels).
Airline (Virgin/QANTAS) \( \times \) Lateness category.
Cancer Type \($ \times \) Cancer Site.
• Interpretation – Compare conditional distributions; if patterns differ, variables likely associated (quantified formally in Lecture 4).
Lecture 4: Chi-Square Test of Independence • Purpose – Test H_0o f < e m > n o a s s o c i a t i o n < / e m > ( i n d e p e n d e n c e ) b e t w e e n t w o c a t e g o r i c a l v a r i a b l e s i n a n of <em>no association</em> (independence) between two categorical variables in an o f < e m > n o a ssoc ia t i o n < / e m > ( in d e p e n d e n ce ) b e tw ee n tw oc a t e g or i c a l v a r iab l es inan r\times ct a b l e . < / p > < p > • < s t r o n g > A s s u m p t i o n s < / s t r o n g > < / p > < o l > < l i > < p > O b s e r v a t i o n s i n d e p e n d e n t ( u s u a l l y S R S o r r a n d o m i s e d d e s i g n ) . < / p > < / l i > < l i > < p > E x p e c t e d c o u n t i n e v e r y c e l l table.</p><p>• <strong>Assumptions</strong></p><ol><li><p>Observations independent (usually SRS or randomised design).</p></li><li><p>Expected count in every cell t ab l e . < / p >< p > • < s t r o n g > A ss u m pt i o n s < / s t r o n g >< / p >< o l >< l i >< p > O b ser v a t i o n s in d e p e n d e n t ( u s u a l l y S R S or r an d o mi se dd es i g n ) . < / p >< / l i >< l i >< p > E x p ec t e d co u n t in e v er y ce l l \ge10( r u l e o f t h u m b ) . C o m b i n e l e v e l s i f v i o l a t e d . < / p > < / l i > < / o l > < p > • < s t r o n g > E x p e c t e d c o u n t s f o r m u l a < / s t r o n g > < b r > < / p > < p > (rule of thumb). Combine levels if violated.</p></li></ol><p>• <strong>Expected counts formula</strong> <br></p><p> ( r u l eo f t h u mb ) . C o mbin e l e v e l s i f v i o l a t e d . < / p >< / l i >< / o l >< p > • < s t r o n g > E x p ec t e d co u n t s f or m u l a < / s t r o n g >< b r >< / p >< p > E_{ij}=\frac{(\text{row }i\text{ total})(\text{column }j\text{ total})}{n}< / p > < p > • < s t r o n g > T e s t s t a t i s t i c < / s t r o n g > < b r > < / p > < p > </p><p>• <strong>Test statistic</strong> <br></p><p> < / p >< p > • < s t r o n g > T es t s t a t i s t i c < / s t r o n g >< b r >< / p >< p > X^2=\sum{i=1}^r\sum {j=1}^c\frac{(O{ij}-E {ij})^2}{E{ij}}U n d e r Under U n d er H 0: X^2\sim\chi^2_{df},\;df=(r-1)(c-1).< / p > < p > • < s t r o n g > P − v a l u e < / s t r o n g > – R i g h t − t a i l p r o b a b i l i t y </p><p>• <strong>P-value</strong> – Right-tail probability < / p >< p > • < s t r o n g > P − v a l u e < / s t r o n g > – R i g h t − t ai l p r o babi l i t y P(\chi^2{df}\ge X^2 {obs}). < / p > < p > • < s t r o n g > C h i − s q u a r e d i s t r i b u t i o n f a c t s < / s t r o n g > < b r > < / p > < p > – D o m a i n .</p><p>• <strong>Chi-square distribution facts</strong> <br></p><p>– Domain . < / p >< p > • < s t r o n g > C hi − s q u a r e d i s t r ib u t i o n f a c t s < / s t r o n g >< b r >< / p >< p > – D o main (0,\infty), r i g h t − s k e w . < b r > < / p > < p > – M e a n , right-skew. <br></p><p>– Mean , r i g h t − s k e w . < b r >< / p >< p > – M e an =df, v a r i a n c e , variance , v a r ian ce =2df. < / p > < p > – A s .</p><p>– As . < / p >< p > – A s df\uparrow the curve becomes more symmetric.
• Worked example – Hemisphere \( \times \) Stress
– 4 \( \times \) 2 table, df=3. < b r > < / p > < p > – . <br></p><p>– . < b r >< / p >< p > – X^2_{obs}=6.59, , , p=0.086 \($ \Rightarrow \) weak evidence of association.
– All expected counts \ge10 \( \to \) assumptions OK.
• Cancer Type \( \times \) Site example
– 3\( \times \)3 table, one expected count <10 \\(→ ) c a u t i o n n o t e d . < b r > < / p > < p > – \to \\) caution noted. <br></p><p>– → ) c a u t i o nn o t e d . < b r >< / p >< p > – X^2_{obs}=6.56, , , df=4, , , p=0.161 \( \to \) no evidence of association.
• Airline Lateness example
– Original 3\( \times \)2 table had small counts; merged to 2\($ \times \)2.
– X^2_{obs}=13.17, , , df=1, , , p<0.001 \($ \to \) very strong association.
• Software summary (R)
– Build table: tab<-table(var1,var2) or as.table(cbind(...)).
– chisq.test(tab) returns statistic, df, p-value, expected counts, residuals.
– addmargins(tab) appends row/column totals.
• Quick reference: R functions vs data type
│One categorical│prop.test() │
│One quantitative│t.test() (or z.test if \sigma known) │
│Cat + Quant│t.test(independent or paired) │
│Cat + Cat│chisq.test() │
Bonus: Non-Parametric Sign Test • When? – Very small paired sample, distribution of differences skewed or with outliers \($ \to \) Normal assumption dubious.
• Procedure
– Define X =n u m b e r o f p o s i t i v e d i f f e r e n c e s a m o n g number of positive differences among n u mb er o f p os i t i v e d i f f er e n ces am o n g np a i r s . < b r > < / p > < p > – U n d e r “ n o e f f e c t ” m e d i a n 0 , pairs. <br></p><p>– Under “no effect” median 0, p ai r s . < b r >< / p >< p > – U n d er “ n oe f f ec t ” m e d ian 0 , X\sim Bin(n,0.5). < / p > < p > – C o m p u t e .</p><p>– Compute . < / p >< p > – C o m p u t e P-value from tail(s) of Binomial.
• Speed-camera revisited (n = 4)
– All four sites show reduction \($ \to \) x_{obs}=4. < / p > < p > – .</p><p>– . < / p >< p > – P=0.0625 (best possible with n=4).
– Conclusion: only weak evidence; paired t test was more powerful.
• Trade-off – Non-parametric tests safer but usually less powerful and target medians, not means.
Chapter Wrap-Up & Keywords • You can now:
– Perform and interpret two-sample & paired t tests.
– Build and read two-way tables, compute marginal / conditional distributions.
– Conduct \chi^2 tests for independence and check their assumptions.
– Understand basics of statistical power & sample-size planning.
– Recognise situations needing non-parametric alternatives.
• Core formulas:
– Two-sample T statistic & CI (pooled).
– Paired T statistic & CI.
– Expected counts E_{ij}. < / p > < p > – .</p><p>– . < / p >< p > – \chi^2s t a t i s t i c a n d d e g r e e s − o f − f r e e d o m . < / p > < p > • K e y w o r d s : t w o − s a m p l e t e s t , p o o l e d v a r i a n c e , p a i r e d , t w o − w a y t a b l e , j o i n t / m a r g i n a l / c o n d i t i o n a l d i s t r i b u t i o n , o b s e r v e d v s e x p e c t e d c o u n t s , statistic and degrees-of-freedom.</p><p>• Keywords: two-sample test, pooled variance, paired, two-way table, joint / marginal / conditional distribution, observed vs expected counts, s t a t i s t i c an dd e g r ees − o f − f r ee d o m . < / p >< p > • K ey w or d s : tw o − s am pl e t es t , p oo l e d v a r ian ce , p ai r e d , tw o − w a y t ab l e , j o in t / ma r g ina l / co n d i t i o na l d i s t r ib u t i o n , o b ser v e d v se x p ec t e d co u n t s , \chi^2$$ distribution, sign test, statistical power.