Class 7 - Genome Annotation

0.0(0)
Studied by 0 people
call kaiCall Kai
Locked
learnLearn
examPractice Test
spaced repetitionSpaced Repetition
heart puzzleMatch
flashcardsFlashcards
GameKnowt Play
Card Sorting

1/30

encourage image

There's no tags or description

Looks like no tags are added yet.

Last updated 10:42 PM on 7/12/26
Name
Mastery
Learn
Test
Matching
Spaced
Call with Kai
Chat

No analytics yet

Send a link to your students to track their progress

31 Terms

1
New cards

Bioinformatic analysis

Allowed researchers to put together the sequence of all chromosomes

2
New cards

How does bioinformatic analysis work?

Researchers sequenced thousands of inserts (approximately 1000bp long) and tried to assemble the DNA in chromosomes

3
New cards

Problem with bioinformatic analysis

Many parts of human chromosomes are repetitive (can be mapped to many different regions)

4
New cards

Solution to bioinformatic analysis problem

BAC clones → plasmids which contain bigger DNA fragments were sequenced from both sides (paired-end sequencing). Gave researchers the info of 1000bp on each side and the info that sequences are approximately 200-300kb apart. Allowed researchers to integrate the shorter reads and reconstruct the DNA sequence of all chromosomes

5
New cards

Open Reading Frame (ORF)

DNA can be read in 6 different reading frames. If random distribution of nucleotides is assumed, STOP codon should appear every 21 triplets. If there is a long stretch of nucleotides without interruptions by STOP codons, likely that a gene is in that location.

6
New cards

DNA homology

2 sequences of DNA with similar nucleotide sequences were derived from a common ancestor. DNA sequences conserved throughout evolution.

7
New cards

Exceptions to the central dogma

Non-coding RNAs

  1. rRNAs - (important for 3D structure of ribosomes)

  2. tRNAs - (important for translation)

  3. Long-noncoding RNAs and short-noncoding RNAs are important for gene expression

8
New cards

Theory of finding where different genes are located

If we can map the sequence of the mRNA to the DNA sequences. Only a small part of our DNA is coding for proteins. Finding sequences of genes by looking at transcribed regions.

9
New cards

Why are enzymes from retroviruses widely used in genetics?

  1. Retroviruses use RNA and not DNA as the information carrying material

  2. Retroviruses infect a cell + incorporate information from their building blocks into host DNA

10
New cards

Reverse transcription

RNA transformed into DNA

11
New cards

Reverse transcriptase

Enzyme that can do reverse transcription

12
New cards

mRNA and gene identification

Needs to be reverse transcribed

  1. Magnetic beads carrying oligo-dT single stranded DNA can be used to isolate mRNA molecules

  2. Given RNA has a poly-A tail, it can complementary base pair + be bound with these beads

  3. Using a magnet, mRNA can be separated from other RNAs e.g. tRNA

  4. Reverse transcriptase can then be used to synthesize DNA leading to a DNA-RNA hybrid

13
New cards

Synthesis of double-stranded DNA molecule after reverse transcription

  1. mRNA digested using RNAse

  2. 3’ end of newly synthesized DNA folds back and can serve as free 3’ OH group for DNA polymerase

  3. DNA polymerase synthesizes the second DNA strand using the first strand as a template

  4. Nuclease cuts hairpin resulting in a double stranded cDNA (complementary DNA)

14
New cards

Old way of sequencing fragments

Cloning into a plasmid

  • Method identical to sequencing genomes

1) inserts and plasmids digested with same restriction enzyme

2) cloning using ligase

3) generation of library of bacteria with different fragents

4) sanger sequencing

15
New cards

Problem with cloning into a plasmid

Very labor intensive + expensive

16
New cards

Solution to cloning into a plasmid problem

Next generation sequencing

17
New cards

New methods of sequencing many different fragments efficiently

  1. Sanger sequencing chain termination sequencing of defined fragments (PCR products or cloned DNA)

  • Highly accurate, 600-700bp in typical read

  1. High-throughput sequencing

  • simultaneous sequencing of millions of random fragments of DNA, longer sequences are assembled based on overlaps, using bioinformatics

18
New cards

Library preparation

Required for high throughput sequencing

  • double-stranded DNA isolated from species of interest and randomly fragmented

  • Adapter sequences are ligated to fragment ends that can be used:

1) for hybridization purposes

2) as primers for PCR

3) as primers for sequencing reactions

4) as identifiers in subsequent bioinforatic analyses

19
New cards

Commonly used method for next generation sequencing

Sequencing by synthesis

  • after fragments were PCR amplified, they are now added onto a chip (flowcell)

  • adapter sequences hybridized to chip-bound complementary sequences

  • generates complementary strands by “bridge amplification”

20
New cards

Sequencing by synthesis

  • adds one nucleotide at a time, causes a short fluorescent pulse

  • as different nucleotides are labeled with different dyes, can be analyzed to obtain DNA sequence

21
New cards

Differences between cDNA sequencing and genomic DNA sequencing

Genomic DNA sequencing

  • gives info about whole DNA (including introns)

  • all parts of DNA should (ideally) be equally covered

cDNA sequencing

  • only gives coding regions (5’ untranslated regions (UTR), exons and 3

  • UTR)

  • does not necessarily show every gene expressed in the body

  • derived from turning mRNA into cDNA, so only represents the expressed genes of that tissue

22
New cards

What can cDNA libraries give us information about?

Alternative splicing

23
New cards

Gene-rich regions vs gene deserts

  • gene-rich = chromosomal regions that have more genes than expected

  • gene deserts = regions that have no identifiable genes

biological significance unknown

24
New cards

DNA polymorphisms

sequence differences between individual genomes within a species

25
New cards

Anonymous DNA polymorphisms/DNA markers

DNA polymorphisms that do not affect phenotype but can be used to track specific regions of the genome

26
New cards

SNP (Single nucleotide polymorphism)

  • average density in human genome is 1 per 1000bp, most common polymorphism

  • often just a single nucleotide changed within population

  • may be silent or can cause changes in protein sequence

  • most are in regions not coding for a protein

  • sometimes correlated with the occurrence of specific disease, can also have pharmacological implications

27
New cards

SNPs can be identified using PCR + Sanger Sequencing

in homozygous state → sanger sequencing gives one big peak for the nucleotide

heterozygous individual → signal for 2 different nucleotides at position. Sanger sequencing came to a stop with an A or T at that position.

28
New cards

Deletion-insertion polymorphisms (DIP/InDel)

  • second most common in human genomes

  • once every 10kb

  • size of deletion varies → more bp = less likely they are (most only 2-3bp)

  • caused by problems in 1) DNA replication 2) recombination 3) DNA repair

29
New cards

Simple Sequence Repeats (SSRs)

  • results from differences in number of copies of short DNA sequence that is repeated many times in a chromosome

  • every 30kb, 1-, 2- or 3- bp

  • mostly outside coding regions but might still impact gene expression e.g. fragile X syndrome

  • frequency of new alleles much higher than normal mutation rate - 1/1000

30
New cards

How to analyze SSRs using PCR

  • primers designed to genetic area adjacent to area of SSR

  • PCR amplifies region between forward and reverse primer

  • if there are different numbers of repeats on the 2 homologous chromosomes, PCR fragments will have a different length

  • regular gel electrophoresis can be used to separate the different fragments (based on size)

  • given high frequency of mutations in SSRs, most individuals are heterozygous so should have 2 bands in the gel (each representing 1 allele)

31
New cards

Copy Number Variations (CNVs)

  • tandem sequence that repeats more than 10bp long

  • misalignment during meiosis leads to unequal crossing-over

  • not common, so most CNVs are inherited, rather than being a new mutation