(L6, 7) IMED2004 - Genetic Variation I-II

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/96

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 6:14 AM on 8/18/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

97 Terms

1
New cards

What are the ten learning objectives for the Genetic Variation lectures?

1. Describe the Human Genome Project.

2. Understand the terms variant and polymorphism and the scale of human variation.

3. Describe SNVs, SNPs, indels and CNVs and their nomenclature.

4. Understand that genetic variation may or may not have functional consequences.

5. Describe repetitive genomic regions, how replication slippage and unequal crossover affect genetic diversity, and why repeats are useful in genetic fingerprinting.

6. Describe structural variation, including balanced versus unbalanced variants.

7. Describe the aims of the HapMap Project.

8. Understand the 1000 Genomes Project.

9. Understand that multiple databases curate human genetic variation and its consequences.

10. Describe the ENCODE Project.

Additional reading:

Chapter 4, Genetics and Genomics in Medicine, 1st Edition, Strachan.

2
New cards

What key pre-1950 discoveries formed the historical background to human genetics? (NOT ASSESSABLE)

1859 — Charles Darwin published On the Origin of Species.

1866 — Gregor Mendel described inheritance of crop traits and phenotype segregation patterns.

1869 — Friedrich Miescher isolated "nuclein" (DNA).

1881 — Albrecht Kossel isolated the basic building blocks of DNA and RNA: A, T, G, C and U.

1882 — Walther Flemming observed chromosome doubling.

Early 1900s — Theodor Boveri and Walter Sutton developed chromosome theory.

1944 — Oswald Avery identified DNA as the transforming principle.

1944-1950 — Erwin Chargaff showed that DNA is involved in heredity and differs between species.

Late 1940s — Barbara McClintock discovered "jumping genes".

Lecturer explanation:

This historical material was included only to provide context and did not need to be remembered.

3
New cards
<p>What was the significance of automated Sanger sequencing? (NOT ASSESSABLE)</p>

What was the significance of automated Sanger sequencing? (NOT ASSESSABLE)

Automated Sanger sequencing was developed in the 1980s and enabled large-scale DNA sequencing.

It became the predominant technology used to generate the first draft of the human genome.

The slide shows:

1. PCR with fluorescent chain-terminating ddNTPs

2. Size separation by capillary gel electrophoresis

3. Laser excitation and fluorescence detection to generate a chromatogram

Lecturer explanation:

The details were presented as background rather than assessable content.

<p>Automated Sanger sequencing was developed in the 1980s and enabled large-scale DNA sequencing.</p><p>It became the predominant technology used to generate the first draft of the human genome.</p><p>The slide shows:</p><p>1. PCR with fluorescent chain-terminating ddNTPs</p><p>2. Size separation by capillary gel electrophoresis</p><p>3. Laser excitation and fluorescence detection to generate a chromatogram</p><p>Lecturer explanation:</p><p>The details were presented as background rather than assessable content.</p>
4
New cards
<p>What was the Human Genome Project?</p>

What was the Human Genome Project?

The Human Genome Project was a large, organised, highly collaborative international effort that generated the first reference sequence of the human genome and several model organisms.

Key dates:

- Concept discussed in 1985

- Formally launched in 1990

- Completed in 2003

.

Major aims:

- Sequence the human genome

- Estimate the number of human genes

- Compare the human genome with model organisms

<p>The Human Genome Project was a large, organised, highly collaborative international effort that generated the first reference sequence of the human genome and several model organisms.</p><p>Key dates:</p><p>- Concept discussed in 1985</p><p>- Formally launched in 1990</p><p>- Completed in 2003</p><p>.</p><p>Major aims:</p><p>- Sequence the human genome</p><p>- Estimate the number of human genes</p><p>- Compare the human genome with model organisms</p>
5
New cards

What major milestones were associated with the Human Genome Project?

1999:

- Human chromosome 22 was sequenced, the first complete human chromosome sequence reported.

2001:

- Initial sequencing and analysis of the human genome was published.

- Initial comparative sequencing of the mouse genome was also undertaken.

2002:

- An updated "complete" human genome sequence was published with additional regions included.

Lecturer explanation:

The term "complete" was qualified because some repetitive regions were still unresolved.

6
New cards
<p>Why was the original human reference genome not the genome of one person?</p>

Why was the original human reference genome not the genome of one person?

The original human reference genome was a patchwork assembled from DNA from a limited number of individuals.

Approximate composition shown:

- ~70% from one individual of blended ancestry

- ~30% from 19 other individuals, mostly of European ancestry

Therefore:

- It was useful as a reference scaffold.

- It did not accurately represent one complete individual genome.

<p>The original human reference genome was a patchwork assembled from DNA from a limited number of individuals.</p><p>Approximate composition shown:</p><p>- ~70% from one individual of blended ancestry</p><p>- ~30% from 19 other individuals, mostly of European ancestry</p><p>Therefore:</p><p>- It was useful as a reference scaffold.</p><p>- It did not accurately represent one complete individual genome.</p>
7
New cards
<p>What were the main components of the human genome estimated from the Human Genome Project?</p>

What were the main components of the human genome estimated from the Human Genome Project?

Approximate genome size:

~3.2 billion base pairs

Major components shown:

- Protein-coding genes: 1.5%

- Introns: 26%

- LINEs: 20%

- SINEs: 13%

- LTR retrotransposons: 8%

- Miscellaneous unique sequences: 12%

- Miscellaneous heterochromatin: 8%

- Segmental duplications: 5%

- Simple sequence repeats: 3%

- DNA transposons: 3%

Lecturer emphasis:

Only a very small fraction of the genome is protein coding, while a large fraction is non-coding and repetitive.

<p>Approximate genome size:</p><p>~3.2 billion base pairs</p><p>Major components shown:</p><p>- Protein-coding genes: 1.5%</p><p>- Introns: 26%</p><p>- LINEs: 20%</p><p>- SINEs: 13%</p><p>- LTR retrotransposons: 8%</p><p>- Miscellaneous unique sequences: 12%</p><p>- Miscellaneous heterochromatin: 8%</p><p>- Segmental duplications: 5%</p><p>- Simple sequence repeats: 3%</p><p>- DNA transposons: 3%</p><p>Lecturer emphasis:</p><p>Only a very small fraction of the genome is protein coding, while a large fraction is non-coding and repetitive.</p>
8
New cards

What are the two broad categories of human genetic variation based on DNA content?

1. Variants that do not alter total DNA content

- Single-nucleotide substitutions

- Inversions

- Translocations

2. Variants that cause a net gain or loss of DNA

- Insertions

- Deletions

- Copy-number changes

- Abnormal chromosome segregation

- Changes ranging from one nucleotide to megabase-scale regions

9
New cards

What is the overall scale and possible phenotypic effect of human genetic variation?

- The most common DNA changes are small-scale changes.

- Some variants alter phenotype.

- Many variants have no detectable phenotypic effect.

- Many remain of unknown significance.

Lecturer emphasis:

Sequencing variants is now much easier than determining what many of those variants actually do.

10
New cards

What is HGVS nomenclature?

HGVS nomenclature is an internationally recognised standard for describing DNA, RNA and protein sequence variants.

It is used in:

- Clinical reports

- Publications

- Databases

HGVS recommends:

- Describe variants at the DNA level whenever possible.

- RNA and protein-level descriptions may be added.

11
New cards

What information is encoded in the HGVS example NC_000023.10:g.33038255C>A?

NC_000023.10:

- Approved reference sequence identifier

g.:

- Linear genomic DNA reference

33038255:

- Base position

C:

- Reference base

>A:

- Substitution to A

The reference sequence should ideally be based on an accepted genome build such as GRCh38/hg38.

DIAGRAM ON SLIDE 8

12
New cards

What HGVS reference-type prefixes were listed?

c. = coding DNA reference sequence

g. = linear genomic reference sequence

m. = mitochondrial DNA reference sequence

n. = non-coding DNA reference sequence

o. = circular genomic reference sequence

p. = protein reference sequence

r. = RNA reference sequence or transcript

13
New cards

What HGVS symbols were used for common variant types?

> indicates substitution at DNA or RNA level

Example: g.123456G>A

Protein substitution:

p.Ser321Arg

del = deletion

Example: c.76del

dup = duplication

Example: c.76dup

ins = insertion

inv = inversion

Example: c.76_83inv

Not used at protein level

fs = frameshift

Example: p.Arg456GlyfsTer17 or p.Arg456Glyfs*17

Ter or * indicates termination.

14
New cards

How is an insertion written in HGVS nomenclature?

Example:

NC_000023.10:g.32862923_32862924insCCT

This means:

- Reference sequence: NC_000023.10

- Genomic reference: g.

- Insertion occurs between positions 32862923 and 32862924

- Inserted sequence is CCT

The flanking positions are separated by an underscore.

DIAGRAM ON SLIDE 10

15
New cards

What are the major possible effects of coding-sequence base substitutions?

A coding-sequence change may be:

- Silent

- Conservative

- Non-conservative

- Nonsense

- Frameshift

The phenotypic effect depends on how the amino-acid sequence and protein function are altered.

16
New cards

What is a silent mutation?

A silent mutation changes the DNA sequence without changing the encoded amino acid.

Example from the slide:

GTG → GTT

Both encode valine.

Therefore:

The protein sequence is unchanged.

17
New cards

What is a conservative amino-acid substitution?

A conservative substitution changes one amino acid to another with similar chemical properties.

Example from the slide:

Valine → leucine

Because the amino acids are similar, the change may have little or no functional effect.

18
New cards

What is a non-conservative amino-acid substitution?

A non-conservative substitution replaces an amino acid with one having substantially different properties.

Example from the slide:

Glutamine → proline

These changes are more likely than conservative substitutions to alter protein function.

19
New cards

What is a nonsense mutation?

A nonsense mutation converts a codon into a stop codon.

Example from the slide:

TCA → TGA

Consequence:

Translation stops prematurely, producing a shortened polypeptide.

20
New cards

What is a frameshift mutation?

A frameshift results when an insertion or deletion changes the reading frame.

Consequence:

- Downstream codons are read differently.

- The amino-acid sequence after the variant changes.

- A premature stop may occur.

Lecturer explanation:

Nonsense and frameshift changes, especially early in the coding sequence, often cause major loss-of-function effects.

DIAGRAM ON SLIDE 11

21
New cards

What is a DNA variant?

A DNA variant is a general term for any difference in DNA sequence observed between individuals or sequences.

Lecturer explanation:

The term "variant" is broader and more neutral than "mutation", which is often reserved for a change-generating event or particular disease-associated changes.

22
New cards

What is an allele?

An allele is an alternative form of a gene sequence found at the same location on a chromosome.

23
New cards

What is a polymorphism?

A polymorphism is a DNA variant that is common in the population.

Typical threshold used in this lecture:

Variant frequency >0.01, or >1%.

24
New cards

What is a rare variant?

A rare variant is a DNA variant with population frequency below 0.01, or below 1%.

Lecturer explanation:

Some publications may use slightly different thresholds, but 1% was the convention used in this lecture.

25
New cards

Why must many individuals be compared to understand human genetic variation?

Humans are diploid:

- Two nuclear genomes, one from each parent

Mitochondrial DNA:

- Is inherited maternally

A single genome cannot define population variation.

Therefore:

- Many individuals must be compared.

- The Human Genome Project reference acts as a scaffold, not a complete description of population variation.

26
New cards

What did sequencing the diploid genomes of Craig Venter and James Watson reveal?

Their genomes were among the first individual diploid genomes compared with the reference.

They showed the extent of sequence variation present within one individual human genome relative to the reference.

Craig Venter:

- 2007

James Watson:

- 2008

27
New cards

How much did the Venter/Watson genomes differ from the reference genome?

Major findings included:

- ~12 million nucleotides differed from reference

- Majority of variants were non-coding

- ~3.2 million SNPs

- 44% of Craig Venter's genes contained a sequence variant

- 17% of those variants encoded an altered protein

28
New cards

What insertion/deletion and structural variation was detected in the early individual genome studies?

Approximately:

- 290,000 heterozygous insertion/deletion variants

- 1-571 bp

- 559,000 homozygous insertion/deletion variants

- 1-82,711 bp

- 90 large inversions

- 62 large copy-number variants

Lecturer explanation:

These studies used relatively early sequencing technology but still revealed extensive individual variation.

29
New cards

What is a single-nucleotide variant, or SNV?

An SNV is a difference at a single nucleotide position.

Single-nucleotide substitution is the most common type of human sequence variation.

30
New cards

What is a single-nucleotide polymorphism, or SNP?

A SNP is an SNV that is common in the population.

Using the threshold in this lecture:

Frequency >0.01, or >1%.

"SNP" is pronounced "snip".

31
New cards

What are major and minor alleles?

At a polymorphic locus:

- Major allele = the more common allele

- Minor allele = the less common allele

Lecturer explanation:

The reference allele is not necessarily the major allele.

32
New cards

Why is the distribution of SNVs across the human genome non-random?

Variation differs across genomic regions because of factors including:

- Regional intolerance or tolerance to genetic variation

- Higher variation in mitochondrial DNA than nuclear DNA

- Excess C→T substitutions

- Evolutionary ancestry

- Mutation hotspots

33
New cards

What does regional intolerance to genetic variation mean?

Some genomic regions tolerate sequence changes poorly.

If variation in a region reduces fitness:

- Variants are less likely to persist.

- The region appears depleted of variation.

Other regions are more tolerant and can accumulate more variants.

Lecturer example:

Odour-receptor genes can be highly variable.

34
New cards

Why are C→T substitutions especially common?

Methylated cytosine can undergo deamination.

5-methylcytosine → thymine

Unlike uracil:

- Thymine is a normal DNA base.

- It may escape repair.

- It can pair with A during subsequent replication.

If this occurs in the germline, the change can become a heritable de novo variant.

35
New cards

What is the approximate de novo SNV mutation rate in humans?

Approximate rate:

1.1-1.4 × 10⁻⁸ per base pair per generation

Equivalent to:

~74 novel SNVs per human genome per generation.

36
New cards

How can SNPs provide information about ancestry?

New variants arise on particular ancestral chromosome segments.

If they persist:

- They remain associated with surrounding ancestral variants.

- Groups of SNPs can therefore mark shared chromosome ancestry.

This principle helps reconstruct population ancestry and ancestral genomic segments.

37
New cards

What is an indel?

An indel is an insertion or deletion of one or more nucleotides.

Modern convention:

- Indels usually refer to relatively small insertion/deletion events.

- Typically about 1-100 nucleotides.

38
New cards

What is a copy-number variant, or CNV?

A CNV is a gain or loss in the number of copies of a DNA sequence.

Modern convention:

- Usually refers to larger sequence gains or losses.

- Often >~100 nucleotides.

Strictly, even a one-base deletion changes copy number, but smaller events are usually called indels.

39
New cards

How common are indels relative to single-nucleotide substitutions?

Indels occur at roughly one-tenth the frequency of single-nucleotide substitutions.

40
New cards

What is the size distribution of indels described in the lecture?

Approximately:

- 90% = 1-10 nucleotides

- 9% = 11-100 nucleotides

- 1% = >100 nucleotides

Key principle:

Short indels are much more common than long indels.

41
New cards

What does the CNV size-distribution graph show?

The number of CNVs falls as CNV size increases.

The graph groups CNVs into:

- 1-10 kb

- 10-50 kb

- 50-100 kb

- 100-200 kb

- 200-500 kb

- 500-1000 kb

- >1000 kb

Small CNVs are much more common than very large CNVs.

DIAGRAM ON SLIDE 20

42
New cards

What major categories of tandemly repetitive DNA were described?

Satellite DNA:

- 20 kb to many hundreds of kb

- Centromeres and heterochromatic regions

Minisatellite DNA:

- 100 bp to 20 kb

- Mainly telomeres and subtelomeric regions

Microsatellite DNA:

- <100 bp

- Widely distributed through euchromatin

Short tandem repeats, STRs:

- Repeat unit 1-6 bp

- A subset of microsatellites

43
New cards

Why are tandem-repeat regions especially variable?

Repeated sequences are unstable during DNA replication and recombination.

Repeat number can change through:

- Replication slippage

- Unequal crossover

As a result:

Individuals may carry different numbers of repeat units at the same locus.

44
New cards

How do microsatellites differ from typical SNPs in allele number?

Typical SNP:

- Usually two alleles

Microsatellite:

- Often many alleles because repeat number can vary

Example on the slide:

(CA)₁₀

(CA)₁₁

(CA)₁₂

These produce fragments of 20, 22 and 24 bp respectively.

DIAGRAM ON SLIDE 22

45
New cards

How does replication slippage produce an insertion in a repetitive sequence?

1. The nascent DNA strand partly dissociates from the template.

2. It reassociates out of register.

3. The nascent strand loops out one or more repeat units.

4. DNA synthesis continues.

5. The new strand contains extra repeat units.

DIAGRAM ON SLIDE 23

46
New cards

How does replication slippage produce a deletion in a repetitive sequence?

1. The nascent strand dissociates and reassociates out of register.

2. The template strand loops out one or more repeat units.

3. DNA polymerase skips the looped-out template repeat.

4. The newly synthesised strand contains fewer repeats.

DIAGRAM ON SLIDE 24

47
New cards

How can microsatellites be genotyped?

PCR primers are designed to flank the repeat region.

Because alleles contain different repeat numbers:

- PCR products differ in length.

- Fragment size reveals the allele carried.

DIAGRAM ON SLIDE 25

48
New cards

Why are microsatellites useful for genetic fingerprinting?

Microsatellites are highly polymorphic and often have many alleles.

Therefore they are very informative for:

- Distinguishing individuals

- Tracking chromosome segments through pedigrees

- DNA fingerprinting

- Confirming cell-line identity

49
New cards

What historical role did microsatellites play in human genetics?

- They became major genetic markers from the 1990s.

- Early Human Genome Project work devoted substantial effort to mapping them.

- ~150,000 were identified.

- They are more informative per locus than SNPs for distinguishing individuals.

- They are harder to automate than SNP genotyping.

Example marker shown:

D13S121 dinucleotide repeat.

DIAGRAM ON SLIDE 26

50
New cards

What is unequal crossing over?

Unequal crossing over is recombination between misaligned repeated sequences.

It can produce:

- A duplication in one chromatid

- A deletion in the other chromatid

It may occur:

- Between homologous chromosomes during meiosis

- Between sister chromatids during mitosis

51
New cards

What key difference between mitosis and meiosis matters for unequal crossover?

Mitosis:

- Duplicated sister chromatids separate during anaphase.

Meiosis:

- Homologous chromosomes pair and can recombine during meiosis I.

- Homologues separate in anaphase I.

- Sister chromatids separate in meiosis II.

Recombination between homologues during meiosis creates opportunities for unequal crossover.

DIAGRAM ON SLIDE 28

52
New cards

How does unequal crossover contribute to minisatellite diversity?

Misaligned recombination between repeat units changes repeat number.

During meiosis:

- Misaligned homologous chromatids can undergo unequal crossover.

During mitosis:

- Misaligned sister chromatids can undergo unequal sister-chromatid exchange during homologous recombination repair.

Outcome:

- One chromatid gains repeat units.

- The other loses repeat units.

Lecturer emphasis:

This is a major mechanism generating minisatellite diversity.

DIAGRAM ON SLIDE 29

53
New cards

Why is the distinction between meiotic and mitotic unequal crossover biologically important?

Meiotic unequal crossover:

- Occurs in germ cells

- Can be transmitted to offspring

Mitotic unequal sister-chromatid exchange:

- Occurs in somatic cells

- Produces somatic variation within the individual

54
New cards

What is structural variation?

Structural variation refers to relatively large-scale changes in DNA organisation.

It includes:

- Balanced structural variants

- Unbalanced structural variants

- Copy-number variation

Moderately large structural variation is common in human genomes.

55
New cards

What is balanced structural variation?

Balanced structural variation changes DNA arrangement without changing total DNA content.

Examples:

- Inversions

- Translocations

Mechanism:

Chromosome fragments break and are rejoined in altered positions or orientations without net DNA gain or loss.

56
New cards

What should you identify in the balanced structural-variation diagram?

Inversion:

- Same DNA segment present in both alleles

- Segment orientation is reversed

Translocation:

- DNA segments are moved to different genomic locations

- Total DNA content is preserved

The labels 1 and 2 represent alternative variants.

DIAGRAM ON SLIDE 31

57
New cards

What is unbalanced structural variation?

Unbalanced structural variation changes DNA content.

Examples include:

- Large deletions

- Large insertions

- Copy-number variation

- Unbalanced chromosome rearrangements

Rare large losses or gains may cause developmental disease.

Some CNVs are pathogenic, whereas others are normal population variants.

58
New cards

What major forms of copy-number variation were illustrated?

1. Insertion/deletion of a sequence element

2. Tandem duplication

3. Interspersed duplication in normal orientation

4. Interspersed duplication in inverted orientation

The diagram uses box A to represent the copied sequence.

DIAGRAM ON SLIDE 32

59
New cards

What does the map of segmental duplications on chromosomes 1, 2 and 3 show?

It maps duplications >10 kb.

Blue connecting lines:

- Intrachromosomal duplications

Red bars:

- Interchromosomal duplications

A and B:

- Recombination hotspots associated with genetic disorders

DIAGRAM ON SLIDE 33

60
New cards

What are the major origins of DNA variation?

DNA variation can arise from:

- DNA replication errors

- Recombination errors

- Chromosome segregation errors

- Copy-number changes caused by crossover errors

- Endogenous DNA damage

- Exogenous DNA damage

61
New cards

Why do most DNA replication errors not become permanent variants?

DNA replication inevitably introduces some errors.

However:

- DNA polymerase has proofreading activity.

- Most misincorporated bases are corrected rapidly.

Only unrepaired errors can persist as sequence variants.

62
New cards

How can chromosome-segregation errors create genetic variation?

Abnormal chromosome segregation can produce gametes with:

- Fewer chromosomes than normal

- More chromosomes than normal

These are large-scale changes in genomic DNA content.

63
New cards

What limitation of the Human Genome Project motivated HapMap?

The Human Genome Project produced a useful consensus/reference scaffold.

However:

- It was a patchwork of several individuals.

- It was not designed to describe common individual differences across populations.

HapMap was developed to catalogue common variants and relationships among them.

64
New cards

What were the main aims of the International HapMap Project?

The goal was to determine:

- Common patterns of DNA sequence variation

- Allele frequencies

- Degree of association between variants

- Haplotype structure across the genome

Project period:

2002-2009

65
New cards

What populations and samples were included in HapMap?

HapMap used cell lines from participants from four population groups:

- CEPH/European ancestry

- Yoruba/African ancestry

- Japanese

- Chinese

Technology:

Mainly microarray-based genotyping

66
New cards

What were the major HapMap Phase I and Phase II datasets?

Phase I:

- ~1 million common SNPs

- Approximately every 5 kb across genome

- 269 DNA samples

Phase II:

- 3.1 million common SNPs

- 270 DNA samples

- Four populations

67
New cards

What earlier SNP-discovery goals were associated with HapMap?

Early goals included:

- Identify 300,000 SNPs

- Determine SNP allele frequencies

- Infer haplotype structure

- Determine correlation among SNPs across the genome

The slide also notes a cost of approximately $45 million.

DIAGRAM ON SLIDE 37

68
New cards

Why was HapMap useful for biomedical research?

HapMap created a public haplotype map describing common patterns of human genetic variation.

It became a resource for finding genes associated with:

- Health

- Disease

- Drug responses

- Environmental responses

69
New cards

What major insight about haplotype structure came from HapMap?

Large genomic regions can contain many SNPs but only a limited number of common haplotypes.

In the example:

- Chromosome 2 region

- 36 SNPs

- Zero obligate recombination events

- Only 7 common haplotypes were observed

This demonstrated that nearby SNPs are often inherited together.

70
New cards

What are tag SNPs and what did the HapMap example show?

A tag SNP is a representative SNP that captures information about a group of correlated SNPs.

In the example:

- SNPs with r² ≥ 0.8 were grouped together.

- Seven tag SNPs captured all SNP variation across the region.

DIAGRAM ON SLIDE 39

71
New cards

What was the 1000 Genomes Project?

The 1000 Genomes Project was a sequencing-based international effort to catalogue human genetic variation across populations.

It used:

- Whole-genome sequencing

- Exome sequencing

- Additional genotyping technologies

72
New cards

What were the main sample sizes and sequencing depths in the 1000 Genomes Project?

Phase 1:

- 1,092 individuals

- 14 populations

- Europe, East Asia, sub-Saharan Africa and the Americas

- Whole-genome sequencing at ~2-6× coverage

- Exome sequencing at ~50-100× coverage

Phase 3:

- Completed 2015

- 2,504 individuals

- 26 populations

- Whole-genome and exome data

73
New cards

What major amounts of human variation were identified by the 1000 Genomes Project?

Across 2,504 individuals from 26 populations:

- 84.7 million SNPs

- Approximately 1 SNP per 100 nucleotides

A typical genome differs from the reference at:

- ~4.1-5.0 million sites

74
New cards

How can most population variants be rare while most variants in one person are common?

Across a large population:

- The vast majority of distinct variants are rare.

Within one individual:

- Most variants carried are common variants.

Therefore:

- At many SNP loci, an individual will be homozygous for common alleles.

75
New cards

How much structural variation does a typical genome contain according to 1000 Genomes?

A typical genome contains approximately:

- 2,100-2,500 structural variants

- Affecting ~20 million bases of sequence

Although structural variants are fewer in number than SNPs, they affect more total bases.

76
New cards

What proportion of common sequence variants are SNPs and short indels?

>99.9% of variants identified in population-based genome data are SNPs and short indels.

Therefore:

Single-nucleotide changes are the most common type of human genetic variation.

77
New cards

How much human genetic variation occurs within versus between populations?

Within populations:

- ~33% of protein-coding loci are polymorphic

- Additional variation occurs in introns, regulatory sequences and flanking sequences

- ~85% of total human genetic variation is found within populations

Between populations:

- ~15% of total variation

- Allele frequencies may differ, especially for morphological traits

78
New cards

What did the 1000 Genomes data show about African genetic diversity? (NOT ASSESSABLE)

African populations showed the highest genetic diversity across multiple variant categories.

This supported an out-of-Africa model in which migration created population bottlenecks.

Lecturer explanation:

This material was presented as FYI/background.

DIAGRAM ON SLIDE 43

79
New cards

What is the serial-founder model of human evolution? (NOT ASSESSABLE)

The serial-founder model proposes that:

1. Human populations originated with high diversity in Africa.

2. Successive migrations out of Africa involved only subsets of the source population.

3. Each migration caused loss of some alleles.

4. Genetic diversity therefore declined through successive founder events.

Lecturer explanation:

This slide was presented as FYI.

DIAGRAM ON SLIDE 44

80
New cards

What major databases were listed for curating human genetic variation?

dbSNP:

- SNP database

dbVar:

- Genomic structural variation

DGV:

- Genomic structural variation

ExAC:

- 60,706 exomes

gnomAD:

- 125,748 exomes and 15,708 genomes

HGV Database:

- Peer-reviewed genome variations

ClinVar:

- Relationships between human variants and phenotypes

ClinGen:

- Dosage-sensitive genes and regions

OMIM:

- Human genes and genetic disorders

GTEx:

- ~1,000 individuals, 54 tissues, RNA-seq plus exome/genome data

ENCODE:

- Diverse functional genomic assays across cell types

DIAGRAM ON SLIDE 45

81
New cards

What is OMIM?

OMIM stands for Online Mendelian Inheritance in Man.

It is a catalogue of:

- Human genes

- Mendelian genetic disorders

- Associated disease-causing variants and functional information

It was first created by Dr Victor McKusick of Johns Hopkins.

82
New cards

What is ClinVar?

ClinVar is a public archive that catalogues relationships between:

- Human genetic variants

and

- Disease or clinical phenotypes

Researchers and clinicians can submit and query variants associated with clinical conditions.

83
New cards

How do OMIM and ClinVar differ?

OMIM:

- Focuses heavily on genes and Mendelian genetic disorders

- Provides detailed gene/disease descriptions

ClinVar:

- Focuses on individual variants and their reported relationships with clinical phenotypes

- Includes both Mendelian and non-Mendelian clinical associations

84
New cards

What general relationship exists between genetic variation and phenotype?

Most human genetic variation is thought to have a neutral effect.

A smaller fraction:

- Is harmful

- May be beneficial in some environments

Even in functional regions:

- Many small DNA changes have no obvious effect.

85
New cards

Why is functional interpretation of non-coding variation difficult?

Protein-coding variants can often be interpreted by examining amino-acid changes.

Non-coding variants may affect:

- Regulatory elements

- Non-coding RNAs

- Chromatin

- Long-range gene regulation

These effects are harder to predict directly from sequence alone.

86
New cards

Why were large functional-genomics consortia developed after the Human Genome Project?

The Human Genome Project showed that only ~1.5% of the genome is protein coding.

Therefore, researchers needed to determine:

- Which non-coding regions are functional

- Which regions regulate gene expression

- Which non-coding variants may affect phenotype

Major efforts included:

- ENCODE

- FANTOM

- GTEx

- Disease-specific sequencing consortia

87
New cards

What is ENCODE?

ENCODE stands for Encyclopedia of DNA Elements.

Its goal is to create a comprehensive catalogue of functional elements in the genome, including:

- Protein-coding elements

- RNA elements

- Promoters

- Long-range regulatory elements

- Transcription-factor binding

- Chromatin features

- Circumstances under which genes are active

88
New cards

What types of assays did ENCODE use to identify functional genomic elements?

The lecture showed examples including:

- DNase-seq

- FAIRE-seq

- ATAC-seq

- ChIP-seq

- Whole-genome bisulphite sequencing

- RNA-seq

- CLIP-seq

- RIP-seq

- Chromosome-conformation methods such as Hi-C

Lecturer explanation:

You do not need to know the technical details of each assay; the important point is that many complementary assays were used.

89
New cards

What functional features can ENCODE identify beyond DNA sequence alone?

ENCODE can identify or infer:

- DNA methylation

- Chromatin modifications

- DNase I hypersensitive sites

- Transcription-factor binding sites

- Promoter architecture

- Long-range regulatory elements

- Protein-coding and non-coding transcripts

- Long-range chromatin interactions

- Other functional genomic elements

DIAGRAM ON SLIDE 51

90
New cards

What were the three main phases of ENCODE?

Phase I, 2003-2007:

- Interrogated ~1% of the human genome

- Mainly array-based technology

Phase II, 2007-2012:

- Introduced sequencing-based assays

- Whole-genome and transcriptome analysis

- Included ChIP-seq and RNA-seq

Phase III, 2012-2020:

- Expanded assay types and production

- Studied RNA binding and 3D genome organisation

- Introduced the Registry of candidate cis-regulatory elements, cCREs

91
New cards

What are ENCODE candidate cis-regulatory elements, or cCREs?

cCREs are candidate genomic regulatory elements identified by integrating evidence of gene-regulatory activity.

The lecture described cCREs as DNase hypersensitive sites supported by:

- H3K4me3

- H3K27ac

or

- CTCF binding

These features suggest regulatory potential.

92
New cards

What was important about ENCODE Phase III for developmental biology?

ENCODE Phase III profiled multiple mouse embryonic tissues across developmental stages.

Examples of tissues shown:

- Forebrain

- Midbrain

- Hindbrain

- Heart

- Liver

- Intestine

- Kidney

- Lung

- Stomach

- Neural tube

- Limb

- Craniofacial prominence

Developmental stages included:

E10.5, E11.5, E12.5, E13.5, E14.5, E15.5, E16.5 and P0.

DIAGRAM ON SLIDE 53

93
New cards

What types of developmental assays were shown in ENCODE Phase III?

Examples shown:

- ATAC-seq

- DNase-seq

- ChIP-seq for histone modifications

- Whole-genome bisulphite sequencing

- mRNA-seq

These assays were applied across many mouse tissues and developmental stages.

Lecturer emphasis:

This generated a major resource for studying how regulatory activity changes during mammalian development.

94
New cards

How could a candidate enhancer identified by ENCODE be experimentally tested? (NOT ASSESSABLE)

Sequence identified as a candidate enhancer could be:

1. PCR amplified

2. Cloned into a reporter construct

3. Injected into a fertilised mouse egg

4. Reimplanted

5. Collected at a developmental stage such as E11.5

6. Assayed by LacZ staining

Reporter expression reveals where the candidate element drives activity.

Lecturer explanation:

The lecturer explicitly said the experimental details did not need to be known.

DIAGRAM ON SLIDE 54

95
New cards

What did ENCODE Phase III report about the number of candidate cis-regulatory elements?

ENCODE Phase III reported:

- 926,535 human cCREs

- 339,815 mouse cCREs

These covered approximately:

- 7.9% of the human genome

- 3.4% of the mouse genome

The SCREEN resource was created to provide access to this registry.

96
New cards

Why is the ENCODE cCRE result important compared with the 1.5% protein-coding fraction?

The Human Genome Project showed:

- ~1.5% of the human genome is protein coding.

ENCODE Phase III assigned regulatory evidence to:

- ~7.9% of the human genome.

Therefore:

Functional genomic activity extends well beyond protein-coding sequence.

97
New cards

What is the IGVF Consortium? (NOT ASSESSABLE)

IGVF stands for Impact of Genomic Variation on Function.

It is a legacy project building on ENCODE and aims to determine how genomic variation affects biological function.

Lecturer explanation:

The lecturer stated that details of this slide did not need to be known.

DIAGRAM ON SLIDE 56