1/30
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
Bioinformatic analysis
Allowed researchers to put together the sequence of all chromosomes
How does bioinformatic analysis work?
Researchers sequenced thousands of inserts (approximately 1000bp long) and tried to assemble the DNA in chromosomes
Problem with bioinformatic analysis
Many parts of human chromosomes are repetitive (can be mapped to many different regions)
Solution to bioinformatic analysis problem
BAC clones → plasmids which contain bigger DNA fragments were sequenced from both sides (paired-end sequencing). Gave researchers the info of 1000bp on each side and the info that sequences are approximately 200-300kb apart. Allowed researchers to integrate the shorter reads and reconstruct the DNA sequence of all chromosomes
Open Reading Frame (ORF)
DNA can be read in 6 different reading frames. If random distribution of nucleotides is assumed, STOP codon should appear every 21 triplets. If there is a long stretch of nucleotides without interruptions by STOP codons, likely that a gene is in that location.
DNA homology
2 sequences of DNA with similar nucleotide sequences were derived from a common ancestor. DNA sequences conserved throughout evolution.
Exceptions to the central dogma
Non-coding RNAs
rRNAs - (important for 3D structure of ribosomes)
tRNAs - (important for translation)
Long-noncoding RNAs and short-noncoding RNAs are important for gene expression
Theory of finding where different genes are located
If we can map the sequence of the mRNA to the DNA sequences. Only a small part of our DNA is coding for proteins. Finding sequences of genes by looking at transcribed regions.
Why are enzymes from retroviruses widely used in genetics?
Retroviruses use RNA and not DNA as the information carrying material
Retroviruses infect a cell + incorporate information from their building blocks into host DNA
Reverse transcription
RNA transformed into DNA
Reverse transcriptase
Enzyme that can do reverse transcription
mRNA and gene identification
Needs to be reverse transcribed
Magnetic beads carrying oligo-dT single stranded DNA can be used to isolate mRNA molecules
Given RNA has a poly-A tail, it can complementary base pair + be bound with these beads
Using a magnet, mRNA can be separated from other RNAs e.g. tRNA
Reverse transcriptase can then be used to synthesize DNA leading to a DNA-RNA hybrid
Synthesis of double-stranded DNA molecule after reverse transcription
mRNA digested using RNAse
3’ end of newly synthesized DNA folds back and can serve as free 3’ OH group for DNA polymerase
DNA polymerase synthesizes the second DNA strand using the first strand as a template
Nuclease cuts hairpin resulting in a double stranded cDNA (complementary DNA)
Old way of sequencing fragments
Cloning into a plasmid
Method identical to sequencing genomes
1) inserts and plasmids digested with same restriction enzyme
2) cloning using ligase
3) generation of library of bacteria with different fragents
4) sanger sequencing
Problem with cloning into a plasmid
Very labor intensive + expensive
Solution to cloning into a plasmid problem
Next generation sequencing
New methods of sequencing many different fragments efficiently
Sanger sequencing chain termination sequencing of defined fragments (PCR products or cloned DNA)
Highly accurate, 600-700bp in typical read
High-throughput sequencing
simultaneous sequencing of millions of random fragments of DNA, longer sequences are assembled based on overlaps, using bioinformatics
Library preparation
Required for high throughput sequencing
double-stranded DNA isolated from species of interest and randomly fragmented
Adapter sequences are ligated to fragment ends that can be used:
1) for hybridization purposes
2) as primers for PCR
3) as primers for sequencing reactions
4) as identifiers in subsequent bioinforatic analyses
Commonly used method for next generation sequencing
Sequencing by synthesis
after fragments were PCR amplified, they are now added onto a chip (flowcell)
adapter sequences hybridized to chip-bound complementary sequences
generates complementary strands by “bridge amplification”
Sequencing by synthesis
adds one nucleotide at a time, causes a short fluorescent pulse
as different nucleotides are labeled with different dyes, can be analyzed to obtain DNA sequence
Differences between cDNA sequencing and genomic DNA sequencing
Genomic DNA sequencing
gives info about whole DNA (including introns)
all parts of DNA should (ideally) be equally covered
cDNA sequencing
only gives coding regions (5’ untranslated regions (UTR), exons and 3
UTR)
does not necessarily show every gene expressed in the body
derived from turning mRNA into cDNA, so only represents the expressed genes of that tissue
What can cDNA libraries give us information about?
Alternative splicing
Gene-rich regions vs gene deserts
gene-rich = chromosomal regions that have more genes than expected
gene deserts = regions that have no identifiable genes
biological significance unknown
DNA polymorphisms
sequence differences between individual genomes within a species
Anonymous DNA polymorphisms/DNA markers
DNA polymorphisms that do not affect phenotype but can be used to track specific regions of the genome
SNP (Single nucleotide polymorphism)
average density in human genome is 1 per 1000bp, most common polymorphism
often just a single nucleotide changed within population
may be silent or can cause changes in protein sequence
most are in regions not coding for a protein
sometimes correlated with the occurrence of specific disease, can also have pharmacological implications
SNPs can be identified using PCR + Sanger Sequencing
in homozygous state → sanger sequencing gives one big peak for the nucleotide
heterozygous individual → signal for 2 different nucleotides at position. Sanger sequencing came to a stop with an A or T at that position.
Deletion-insertion polymorphisms (DIP/InDel)
second most common in human genomes
once every 10kb
size of deletion varies → more bp = less likely they are (most only 2-3bp)
caused by problems in 1) DNA replication 2) recombination 3) DNA repair
Simple Sequence Repeats (SSRs)
results from differences in number of copies of short DNA sequence that is repeated many times in a chromosome
every 30kb, 1-, 2- or 3- bp
mostly outside coding regions but might still impact gene expression e.g. fragile X syndrome
frequency of new alleles much higher than normal mutation rate - 1/1000
How to analyze SSRs using PCR
primers designed to genetic area adjacent to area of SSR
PCR amplifies region between forward and reverse primer
if there are different numbers of repeats on the 2 homologous chromosomes, PCR fragments will have a different length
regular gel electrophoresis can be used to separate the different fragments (based on size)
given high frequency of mutations in SSRs, most individuals are heterozygous so should have 2 bands in the gel (each representing 1 allele)
Copy Number Variations (CNVs)
tandem sequence that repeats more than 10bp long
misalignment during meiosis leads to unequal crossing-over
not common, so most CNVs are inherited, rather than being a new mutation