1/157
Looks like no tags are added yet.
Name | Mastery | Learn | Test | Matching | Spaced | Call with Kai | Chat |
|---|
No analytics yet
Send a link to your students to track their progress
genetic variation in a population
generated by mutation
variants
passed to offspring by vertical transmission or between individuals by horizontal gene flow
selection
favourable variants shapes which members of the population survive
organisms occupying the same niche
they are in necessarily in competition
physical niches i..e gut or soil
geographic/demographic niches i.e australian soils or guts of children
nutritional niches i.e lactose fermenters or phototrophs
competition means two organisms cannot occupy the same niche for long — have to evolve different niches or one will out compete the other and drive it to extinction
vertical transmission
direct replication of the chromosome during cell division
if replicaiton is faithful → both daughter cells are genetically identical
if errors in replication → new genetic variants will rise and be in competition with other variants of their species
random mutation
random changes in gnetic code due to DNA damage or replication errors.
increased diversity (relatively slow)
horizontal gene transfer
Movement of genetic material from one organism to another independent of cell division
increased diversity (rapid)
genetic drift
random changes to the frequency of variants through mutation or HGT
may increase/decrease (tends to impact loci under weak selective pressure more strongly)
selection
increase survival of most fit variants
decreased diversity (impact depends on strength of selection)
migration
Movement of subset of the population into a new isolated region (or host population), generates a genetic bottleneck
decreased diversity (rapid, often heavily random)
allele
each variant (gene/promoter/other element) differing by one or more nucleotides — normally alleles all share the same function
bacterial species
collection of strains with a conserved core of genes and phenotypes
strain
subvariant of a bacterial species with a common ancestor
may be defined by genetic content and/or phenotype
strains of the same species may have very different phenotypes
isolates
individual pure cultures from different sources
represent a single snapshot of an evolving lineage (and can themselves evolve in the lab)
clade
group consisting of all organisms descended from a single common ancestor
may include multiple species or genera but is typically used in bacteriology to refer to groupings within a species
clone
group of genetically homogenous cells which have arisen from a single parent cell
clonality
tendency of a population towards forming clonal groups
more clonal organisms tend to mostly evolve slowly through mutations and binary cell division rather than rapidly evolving and acquiring new genes
horizontal gene transfer
results in acquistion of genetic material from outside the cell
may occur through recombination of foreign DNA into the chromosome
by movement of mobile genetic elements i.e plasmids, phages, integrative and conjugative elements
“two roads to rome” of bacterial evolution
spotaneous DNA damage
error in DNA replication is the main cause
rate of 1 in 108 to 1011 nucleotides is copied incorrectly by the DNA polymerase which uses proof reading activity to correctly copy the template
induced DNA damage by chemical alteration
alkylation
UV-induced thymine dimers
oxygen radicals
initial damage, if not repaired, will result in muation of the DNA which will be in herited by daughter cells via cell division
alkylation
electrophiles add alkyl groups to phosphates, stalls replication
carcinogens, ethylmethane sulphonate
UV-induced thymine dimers
DNA absorbs UV at 260nM
forms intra-strand pyrimidine dimers, main T-T
distortion of double helix prevents DNA replication → lethal
oxygen radicals
cause single and double stranded breaks
gamma radiation and x-ray
direct repair mechanisms
restoration to original undamaged state
photoreactivation
nucleotide excision repair (short match repair)
mismatch repair
indirect repair
damage bypass system using DNA replication (not necessarily restoring the undamaged state)
recombination repair
SOS repair
silent mutation
no effect on protein sequence
due to degeneracy of genetic code
missense mutation
one amino acid in the protein is replaced by another
replacement with amino acid of similar biochemical profile
replacement with an amino accid with a differnt biochemical profile
can result in complete or partial loss of function, change in function or temperature sensitive mutants
nonsense mutation
mutation gives rise to stop codon
stop codons TAG (amber), TGA (opal), TAA (ochre)
usually results in complete loss of function as rpemature termination of polypeptide chain gives truncated protein
In-frame indel
loss or gain of a multiple of 3 bases results in insertion or deletion of amino acids
often does not result in complete loss of function
frame shift mutation
insertion shift DNA sequence out of frame
usually results in complete loss of function — either premature stop codon or entirely new amino acid sequence
impact depnds on location within the gene
mutation rate is impacted by
genetic drift
population size — larger populations have slower rates of drift
ecology — some lifestyles select for rapid mutation e.g. host restricted bacteria/viruses undergoing rapid transmission face frequent selection bottlenecks
dN/dS ratio
degree of selection acting on a gene is given by this
dN = non-synonymous nucleotide substitutions per non-synonymous site
dS = synonymous nucleotide substitution per synonymous site
intepreting dN/dS
non-synonymous mutations are presumed to impact fitness — expect them less frequently in residues under greater selective pressure
dN/dS=1 there is no selective pressure (neutral selection)
dN/dS<1 deleterious alleles are rapidly removed from the population (purifying sleection, the locos is conserved)
dN/dS>1 polymorphisms are frequent and are retained in the population (positive selection, the locus is flexible)
homologs
pair of two sequences which are predicted to have a common ancestor
can arise through gene duplication (makes paralogs) or speciation (produces orthologs)
assumptions of models of homology
identical residue pairs are aligned
different amino acids with similar physiochemical properties will form a pair
scores are attributed to each possible matches, mismatch and gap
gaps represent indels in an evolutionary context
highest total score is most likely relationship
measuring homology of nucleotides
either a base matches or it doesn’t
can measure the percentage of nucleotide matching in an alignment to measure the homology — % identity
can be interpreted as evolutionary distance between sequences
epidemiology
study of distribution and determinants of disease and health related events and its application in control and prevention
epidemiology
deals with one defined population at risk
risks lead directly to cases
identifies previously unknown causes
infectious disease epidemiology
two or more populations (human, pathogen, possible vectors or alternate hosts)
a case is itself risk factor
the cause often known
two or more populations
humans — sometimes the only reservoir of pathogens
infectious agents — helminths, bacteria, fungi, protozoa, virusese, prions
vectors — mosquito, snails, blackfly
animals — sheep, goats, mice and ticks, bats
what is infectious disease epidemiology used for?
identification of causes of new, emerging infections
surveillance of infectious disease
indentification of source of outbreaks (environmental or human)
studies of routes of transmission and natural history of infections
identification of new interventions
R
the basic reproductive number — R0
the mean number of individuals directly infected by an infectious case through the total infectious period, when introduced to a susceptible population
Endemic (R=1) — transmission occurs but numbers of cases remains constant
Epidemic (R>1) — the number of cases increases
Pandemic (R>1) — when epidemics occur at several continents (global epidemic)
identification of the infectious agent is required
to administer the treatment
for prognosis
to initiate appropriate infectious disease control measures
to take suitable preventive steps
to understand epidemiology
to know the disease history
phenotypic microbial typing methods
based on the assumption that an organims genotype affects its phenotype — therefore the phenotype tells us about the relatedness of organisms
rapid and inexpensive
i.e biotyping, antibiotyping, serotyping, phage-typing
API strips
grow bacteria from patient sample (overnight), inoculate API strip and incubate
capsules contain various tests for metabolic activities which give colorimetric readings
15 identification systems covering more than 600 bacteria species
time drawbacks of phenotyping
does not detect emergence of noval variants
culture dependent which makes it slow for some organisms
requires more infrastructure than molecular typing
sensitivity drawbacks of phenotyping
antimicrobial treatment iff often commenced before sampling for diagnosis — microbe dies and samples are not useful for phenotyping (almost 50% of diagnostic samples are in this category)
some diseases have very low infectious doses and are difficult to detect
accuracy drawbacks of phenotyping
seropositivity may be confounded by cross reaction of antibodies to commensals or other pathogens
antibiotyping may be confounded if lab conditions do not result in the expression of antimicrobial resistance (AMR) determinants
phage typing can be confounded by evolution of phage-resistant variants
biotyping may be confounded when species are heterogenous for biochemical pathways
molecular typing
based on detection of characteristic nucleotide sequences or proteins
strenths of molecular typing
culture independent and therefore typically faster
more sensitive since amplification allows detection of very low numbers of organisms
weaknesses of molecular typing
antibiotyping remains important and requires culture (presence/absence of an AMR gene is not always confirmation of resistance/sensitivity)
need to be aware that detection of an organism does not always mean it is responsible for disease
change in staff training and education
initial cost of set up i.e equipment, reagants, cryogenic storage etc
identification of a causative agent of an outbreak is confounded by
presence of other strains in the environment that do not cause disease
asymptomatic carriage of disease-causing agents by hosts
need to discriminate between strains genotypically to determine which strains are endemic or epidemic
strains
subvariant of a species with a common ancestor
have permanent genetic changes to their DNA content through mutation or HGT (mobile element/recombination)
emergence of a new strain may result in
commensal becoming pathogen
pathogen becoming able to colonise and cause disease in a new host
pathogen become able to cause more sever disease in current host
advantages of PFGE
good resolution
whole genome is visualised
gold standard method widely used
disadvantages of PFGE
expensive equipment and training required
long run times
requires expert interpretation (incomplete digestion, skewed lanes, some species produce nuclease that degrades DNA, other strains modify restriction sites and block digestion)
advantages of ribotyping
very reproducible
has been automated
detects long term evolutionary trends
disadvantages of ribotyping
less discriminatory than PFGE (fewer bands, not whole genome)
does not discriminate on short evolutionary timescales such as an outbreak
advantages of RAPD
Inexpensive, efficient and sensitive — most useful for rapid typing and distinction between unrelated isolates
disadvantages of RAPD
poor reproducibility
hard to standardise between labs
variation in band intensity
advantanges of REP-PCR
quick and more cost effective than PFGE
disadvantages of REP-PCR
requires enough Reps in close proximity to generate enough PCR products for discrimination of strains
advantages of AFLP
no need for known sequences in the genome
high reproducibility
many loci are simultaneously analysed
by changing the selective primers different loci can be analysed
whole genome analysis is (theoretically) possible
disadvantage of AFLP
high protocol complexity
advantages of MLVA
highly reproducible/standardisable
disadvantages of MLVA
species specific design
advantages of MLST analysis
cheaper in the early days of sequencing
data is unambiguous, reproducible between labs, easily transferred and compared between labs, scalable and automated using high throughput sequencing
disadvantages of MLST
cost is no longer better than whole-genome sequencing
databases are only available for some pathogens
significantly impacted by recombination affecting the MLST
advantages of SNP-based phylogeny
high accuracy and resolution over long timescales for clonal species
relatively straightforward analysis pipeline
potentialy impacted by SNPs accrued during lab storage
disadvantages of SNP-based phylogeny
requires a reference genome sufficiently closely relatd to your set of isolates — especially difficult for novel or non-model organisms
reference choice can impact the number of SNPs called, limiting reproducibility across studies
over long timescales removing recombination from panmictic species can essentially result in chucking out the entire dataset
advatanges of K-MER approaches
by far the highest discriminatory ower f the methods discussed here
reference-agnostic approach avoids bias
disadvantages of K-MER approaches
most complicated analysis pipeline requiring technical expertise and computing time
scientific community still disagrees how many SNPs should be considered a single strain (varies by organism)
phylogeny
model of the relationships between organisms, genes, proteins, or other structures based on common ancestry
only makes sense when the characters being compared have a high degree of homology
four common uses of phylogeny
classification
grouping of genes, proteins, and other molecular sequences including non-coding sequences.
epidemiological investigations
analysis of parallel evolution between host and parasite
redial trees
all branch lengths indicate sequence diversity
lenths are proportional to that diversity
dendrogram trees
only horizontal branches show sequence diversity
vertival lines show how the horizontal branches are related
may be shown as rectangular or circular, both are dendrograms
cladogram
branching tree diagram assumed to be an estimate of a phylogeny where the branches are of equal lengths.
show common ancestry but do not indicate the amount of evolutionary time seperating taxa.
usually emphasizes that the diagram represent a hypothesis about the actual evolutionary history of a group
phylogram
branching tree diagram that is assumed to be an estimate of a phylogeny
branch lengths are proportional to the amount of inferred evolutionary change based on differences in the sequences being compared
ultrametric tree
phylogram showing inferred evolutionary change over time
“molecular clock” is inferred based on the number of sequence changes and the expected rate of change to conver the branch lengths to real-world time
evolutionary time
the passage of time as measured by the number of mutations introduced into a lineage during replication.
often assumed to be constant within a phylogeny although we know this is not always the case
root
most distant branche of a tree
represent the most recent common ancestor of all taxa at the tips of the tree
outlier
typically the most divergent sequence from your dataset
often an ortholog from another closely related species is chosen
midpoint rooting
find the longest path between two taxa and use the midpoint as the root
useful when no outgroup is available or when ingroup-outgroup status of taxa is uncertain
clade
group that includes a common ancestor and all its descendants (living and extinct)
closely related branches may not represent functionally different taxa: often these are part of a single clone or represent functionally identical proteins
monophyletic
contains only one continuous evolutionary lineage and no other taxa
not missing any members of that lineage
polyphyletic
group with no recent common ancestor
can think of it as a group where one or more common ancestors are missing
paraphyletic
containing an ancester and only some of its descendants
almost like an incomplete monophyletic group
distance matrices
generate a tree the alignment is converted to a distance matrix comparing each of the sequences in a pairwise fashion
generated using the probability of a subsitition occuring
DNA — equal probability for exchange of nucleotides
proteins — probability based on redundancy of genetic code, use ‘point accepted mutation’ matrix for these calculations
disadvantage of neighbour joining method
conversion to a matrix results in loss of information which is sometimes undesirable. compute time required grows exponentially whick makes this impractical for large datasets
maximum parsimony
makes use of the optimality criterion that the best tree requires the fewest residue substitiutions to explain the data
advantage of maximum parsimony
final tree will include all the informative positions in the alignment
disadvantage of maximum parsimony
several trees may be equally persimonous
maximum likelihood
most common technique used
uses probability of substitutions at each position as the optimality criteria
advantage of maximum likelyhood
final tree will include all the informative positions in the alignment
disadvantage of maximum likelihood
for large datasets the mostl likely tree will never be found
bayesian methods
for phylogenies no resolved by maximum likelihood
involves simulating trees in a markov chain monte carlo simulation and keeping the trees with the hightest probabilities
posterior probabilities of each branch are written into the final tree
advantages of bayesian methods
includes all informative positions
disadvantages of bayesian method
very complex phylogenies may not be fully resolved
bootstrapping
resampling-based method for giving a value to tree branches
if a branch is highly likely to be real, it will appear many times in the bootstrap analsis and will have a bootstrap value close to 100
boostrap value close to 0 — most likely to be present by chance alone