GWAS and Complex Traits

Complex Traits and GWAS

Complex traits are influenced by multiple genetic and environmental factors, unlike Mendelian traits. Genome-wide association studies (GWAS) are used to identify genetic variants associated with these complex traits.

GWAS Applications

GWAS can be applied to study any phenotype. For example, a genetic testing company provided results that showed genetic variants are linked to cilantro preference.

Association Mapping of QTLs

  • Quantitative trait loci (QTLs) are genes contributing to complex traits.

  • Association mapping is useful when controlled crosses are not possible (e.g., in human studies).

  • It relies on recombination events that have occurred over generations.

Association Mapping in Case vs. Control Studies

  • Recombination shuffles variants linked on the same chromosome over many generations.

  • Variants that are physically closer are less likely to undergo recombination between them.

Haplotypes and Linkage

  • Haplotype: A combination of alleles at multiple loci that are transmitted together.

  • Example: Consider two sites:

    • Site 1: A or T

    • Site 2: G or C

    • Possible haplotypes: AG, AC, TG, TC

Linkage Equilibrium vs. Linkage Disequilibrium

  • Linkage Equilibrium: Haplotypes are present at equal frequencies, and knowing the sequence at one site provides no information about the sequence at another site.

  • Linkage Disequilibrium (LD): Variants at two loci are correlated, and the sequence at one site can predict the sequence at another site.

Linkage Disequilibrium Representation

  • LD is represented as a triangle diagram.

  • Block color indicates the strength of the statistical correlation for pairwise comparisons of SNPs.

  • Darker color indicates stronger correlation.

Utility of Linkage Disequilibrium

LD makes it more efficient to identify SNPs associated with complex traits using SNP microarrays.

  • Only tag SNPs need to be surveyed on a microarray to infer the full haplotype.

GWAS to Identify SNPs

GWAS is used to identify SNPs associated with complex traits such as:

  • Type 2 diabetes

  • Autism spectrum disorder

Case-Control Comparison in GWAS

In a GWAS study:

  • Cases: Individuals with the disease.

  • Controls: Individuals without the disease.

SNPs are analyzed for each group.

  • Example data:

    • SNP1: Cases have a higher frequency of C allele (55.4%) compared to controls (47.4%), p-value = 1.2×10−141.2 × 10^{-14}.

    • SNP2: Cases and controls have similar frequencies of C allele (42.8% vs 43.2%), p-value = 0.80.

P-value in GWAS

  • The null hypothesis is that there is no association between a SNP and the disease.

  • The p-value is the probability of observing the extreme bias by random chance if the null hypothesis is true.

  • Typically, the null hypothesis is rejected when p < 0.05.

  • A chi-squared test is commonly used to test this association.

Chi-Squared Test for Independence

The chi-squared test assesses the independence of SNP alleles and case/control status.

  • Formulae:

    • Observed (O) values: OCaG, OCoG, OCaC, OCoC (Cases with G, Controls with G, Cases with C, Controls with C).

    • Expected (E) values: ECaG, ECoG, ECaC, ECoC, calculated based on total allele frequencies.

    • χ2\chi^2 values:

      • χ2CaG=(ECaG−OCaG)2ECaG\chi^2CaG = \frac{(ECaG - OCaG)^2}{ECaG}

      • χ2CoG=(ECoG−OCoG)2ECoG\chi^2CoG = \frac{(ECoG - OCoG)^2}{ECoG}

      • χ2CaC=(ECaC−OCaC)2ECaC\chi^2CaC = \frac{(ECaC - OCaC)^2}{ECaC}

      • χ2CoC=(ECoC−OCoC)2ECoC\chi^2CoC = \frac{(ECoC - OCoC)^2}{ECoC}

  • Total χ2\chi^2 = χ2CaG+χ2CoG+χ2CaC+χ2CoC\chi^2CaG + \chi^2CoG + \chi^2CaC + \chi^2CoC

  • Degrees of freedom (df) = 1.

Degrees of Freedom

With the row and column totals given, calculating one cell value allows you to determine all other values; hence, df = 1.

P-value Determination

The p-value can be obtained from a χ2\chi^2 value and degrees of freedom using:

  • Statistical tables

  • Software like Excel or Google Sheets using the CHIDIST function: CHIDIST(χ2\chi^2 value, df)

Example Chi-Squared Test

  • Given data for SNP1:

    • Cases:

      • O(G) = 1716, E(G) = 1902, χ2\chi^2 = 18.2

      • O(C) = 2132, E(C) = 1946, χ2\chi^2 = 17.8

    • Controls:

      • O(G) = 3089, E(G) = 2903, χ2\chi^2 = 11.9

      • O(C) = 2783, E(C) = 2969, χ2\chi^2 = 11.7

  • Total χ2\chi^2 = 59.6, df = 1, p = 1.2×10−141.2 × 10^{-14}

Manhattan Plot

Manhattan plots are used to visualize the results of GWAS, showing the association of each SNP with the phenotype.

Multiple Testing Correction

  • In GWAS, many SNPs are tested, necessitating correction for multiple testing.

  • To maintain overall significance, the p-value threshold is adjusted:

    • p<0.05Number of SNPs testedp < \frac{0.05}{\text{Number of SNPs tested}}

Calculating Allelic Odds Ratio

  • Odds of C allele given Case status = Freq of C in CaseFreq of G in Case\frac{\text{Freq of C in Case}}{\text{Freq of G in Case}}

  • Odds of C allele given Control status = Freq of C in ControlFreq of G in Control\frac{\text{Freq of C in Control}}{\text{Freq of G in Control}}

  • Odds Ratio = Case C/Case GControl C/Control G=\frac{\text{Case C/Case G}}{\text{Control C/Control G}} =\frac{2132/1716}{2783/3089} = 1.38$$

  • Interpretation: The case group has 1.38 times the odds of having the disease/phenotype compared to the control group.