mol gen unit 3

4/11:

question of the unit: how to turn a gene on?

ability of pol to bind to promoter

understand it best for organisms that have it driven by sigma 70

some GTFs bind directly to DNA and recognize seq associated with promoters. some bind to the proteins that bind to DNA and link them to the constitutive components of the RNA pol II complex. others modify that complex in way necessary for initiation of xcription.

can tell this story different ways: focus on biology of promoter sequences in euk genes

far more contacts for proteins in promoters in euks than proks

emphasize sequences found upstream of xcriptional start: TATA box and BRE. there are additional sequences downstream of xcriptional initiation site. identified in various colors. up above we see the names of some GTFs that will be important here, especially TBP (TATA binding protein). bind in seq specific ways.

the way that we should think about gene expression in euks: all of these GTFs are req for expression of every protein-encoding gene.

start to call the GTFs by name and their roles

TATA binding proteins and TFIID sit down as complex, flagging the promoter for others and eventually RNA pol II to sit down

brings in TFIIB which binds to DNA directly in a seq specific way, and TFIIA which doesn’t — scaffolding protein whose presence is necessary for RNA pol II to interact with the other GTFs. this is the complex to which pol II binds. sits down pre-bound to another GTF called TFIIF that draws in TFIIE and TFIIH. TFIIH is more important to us bc it’s also an enzyme — a protein kinase that phosphorylates the tail that extends from the RNA pol II holoenzyme complex. it is the tail of the holoenzyme that is responsible for getting the pol to bind to promoter in seq specific way

what are phosphate groups? small relative to protein, but always CHARGED. if we phosphorylate the tail of RNA pol II at various positions, we change its charge and conformation. → allows pol to let go of the complex that formed there. we need it to bind more tightly to promoters so it doesn’t express the entire genome equally, but it has to be able to let go — we let go via phosphorylation

these are essentially ALL REQUIRED for euk xcription in essentially ALL instances, but never enough in cells.

  1. missing the impact of chromatin structure

  2. possibility that other proteins will bind to genome and regulate rate at which these complexes form in both proks and euks


e coli can metabolize glucose and lactose and others. look at expression of a lactase and measure its expression in culture

in sources where there is no galactose and only glucose, expression of lactase is low. if we take this culture and spike it with lactose, we eventually see expression of lactase once the glucose is depleted! only then do they turn on the cells necessary to digest lactose. has genetic circuitry to detect what sugars are present and to prioritize using glucose until it’s gone.

how does it work? works at level of xcriptional initiation. there are proteins in addition to the RNA pol and the sigma 70 subunit, proteins that are sensitive to presence of glucose and lactose that affect the rate at which rna pol binds to promoters. there are going to be additional proteins involved for which their presence is not necessary in some, but not all genes. conversely, we also have proteins that will be inhibitory at some but not all genes. repressors and activators — transcriptional regulators! technically are TFs, but they are NOT our GTFs! just don’t call them TFs. important in both proks and euks

activators act positively, repressors act negatively.

in most instances in e coli, they do their job by binding to the genome at sites very near promoters (potentially even overlap with promoters!)

what about euks? these transcriptional regulators often bind at great distance from promoters. concept of loops in chromatin structure. for xcriptional regs, we req loops where the loops may be enormous.

enhancers. story we may have been told before: activators enhance gene expression by binding to RNA pol and stabilizing it at promoter sites — not untrue, just oversimplified. there are very few (basically NONE) in which activators interact directly with the RNA pol or the GTFs to get it to sit down. instead, interact w complex of proteins called mediator. in vitro, not req for xcription but IS req in all euk cells.

what does mediator look like? a bunch of proteins all called MEDs. does it look the same in all cells? do they even all look the same in a single cell? no good answer. don’t know, but probably not biggest way we distinguish expression between different cells/tissues.

rna pol is huge, especially with all these other proteins tacked on. for a promoter to be used, it has to be accessible by entire complex. space between 2 nucleosomes rarely allows complex to bind in its entirety — nucleosome organization really important! thus it’s ALSO really important at enhancer sites, not just initiation sites! here we have a clear mechanism by which we can drive differential expression of genes bc we can change degree to which can (genuinely cannot hear in the recording at all. 29:00). at promoter sites but also activator sites.


elongation: looks like rna pol adding nucleotides to 3’ ends of the xcript using watson crick base pairing to determine what nucleotide gets added. in euks it’s more complicated bc processing occurs.

RNA pol will interact w nucleosomes (not always, but sometimes — we can and do see gene expression near the nucleosomes even if the ends of the loops see expression a lot more). RNA pol needs to xscribe regions of DNA wrapped around nucleosomes — do the nucleosomes get displaced?

(at least in euks) in DNA replication, part of the nucleosome is partially disassembled and part of it remains. what abt xcription — does it disrupt nucleosomal structure? NO! it certainly “affects” it, but does it require that the histones are displaced? no. RNA pol can transcribe around nucleosomes!

chromatin structure involved in determining when transcriptional start sites are used + rate at which it goes to completion. rna pol can also fall off during xcription like DNA pol can during synthesis. how does the degree to which a region of DNA is packed affect how RNA pol will fall off? we don’t know.

termination:

lost the drawing of prok gene. promoter on left, sequence of YFG. eventually xcription has to stop. in the YFG there is a stop codon, but does it have anything to do with transcriptional termination? NO! those are associated with translation, not xcription. there is no clear arrangement of nucleotides that in all situations triggers termination in proks and euks. understand best in proks, some details in euks.

xcriptional termination in proks: two distinct ways

  1. Rho-dependent

  2. Rho-independent

Rho is a protein that binds in seq specific way to some, but not all, e coli xcripts. binds to rut sites. Rho isn’t a small protein — average sized alone, but bind in complexes of SIX, so it’s kind of massive! matters bc it ultimately interacts w message but moves from rut binding site to a seq motif that allows it to interact w rna pol making it fall off. xcription terminated

independent: no Rho. instead, form a 3D stem and loop RNA structure where we force DS character. interacts w rna pol, it pauses, decr processivity (decr likelihood of next nucleotide being added, incr odds of rna pol falling off)

in euks: look at ends of RNA and compare to in proks. it has a methylated G at 5’ end! at 3’ end, poly-A tail. the two ends very commonly interact with each other, so usually our xcripts are circularized!

how does the poly-A tail happen? an enzyme adds it on to end of all messages — polyadenylation A enzyme. how do we know where to do it? there are seqs in the DNA encoding for RNA that is responsible for binding a complex of enzymes responsible for processing 3’ ends of transcripts. that seq is broadly AAUAAA, to which a protein known as CPSF (cleavage and polyadenylation stimulating factor) binds and brings in other players involved in processing 3’ ends: include other accessory proteins called CFI/II, which bring in cleavage stimulating factor to cleave the xcript at a particular seq. transcript gets cleaved as its being made → if we cleave at a certain position, we still have xcription occurring via RNA pol downstream of the AAUUAAA (?). occurs with transcription. once it cleaves, poly-A. several hundred A residues!

how does xcriptional termination occur? don’t rly know. maybe just falls off, or maybe cleavage affects the RNA pol biology so it falls off more readily, but we don’t know.

splicing. we are left with exons. xcript being made in nucleus transiently contains introns that have to be removed to produce a xcript with only exons. we have to talk abt the mechanism by which this works bc it accounts for huge amt of diversity in proteins expressed in our cells. remember we have ~20,250 different genes in our genome. we don’t make ~20k proteins, we make between 500k-1M proteins! thanks to splicing — next time


4/14:

splicing is important bc we can get more than 1 protein from a gene

the way that xcripts get spliced (esp in complicated organisms) is dependent on the enviro. splicing is in euks, not proks.

**image

euks have nuclei. in humans, splicing is REQUIRED before transport out of nucleus (nuclear export), not true for all euks. essentially no genes in our genome that don’t have multiple BIG introns

**image

introns are physically removed from the xscript. splicing is necessary for complete processing of the pre-mRNA. what machine splices? …splicing can occur without any protein at all. self-splicing introns essentially remove themselves, but don’t make up most of the introns in a complicated organism

most exons are ~100-200 base pairs in length. introns vary more… in humans, most are 100-2000 but some (not many) are greater than 30k bps in length!

**image

start to tease out mechanisms: there’s several, but we’ll look at one in particular. we can identify where introns are bc there are stereotypical patterns at junctions between it and its surrounding exons. will call this either the 5’ junction or the exon-intron junction, and the 3’ end of the intron or the intron-exon junction. we see classical seq in the exons itself and in the intron that will be spliced. in the image, Y11 means 11 pyrimidines. consensus site. there’s also an internal site involved in the biochem responsible for removing this seq: the branch site. found internally and spacing can vary from one to the next. has its own characteristics. the A highlighted is the “actual” branch site.

**image

focus on left part of image. the A residue attacks the phosphodiester bond linking the intron to its upstream exon (exon-intron junction). remember that RNAs can fold in on themselves! when the A attacks and breaks that bond, it becomes site-lysed. right side of image is what it looks like biochemically after it has attacked. the A has made a 2’ junction, so the A has 3 phosphodiester bonds instead of the regular 2. have now interrupted the message so that intron makes a loop that will be called the lariat. now the exon has an exposed reactive OH group that will attack the intron-exon junction and the intron pops off (think of biochem!) now we have the lariat, which will go away.

no enzyme involved in this whole process, but it’s enzyme-like… we say it’s ribozyme-catalyzed bc it’s catalyzed by a molecule made up of RNA, not proteins.

**image (above)— this is what an actual self-splicing intron looks like. everything in black is the intron; the two red lines are the upstream and downstream exons!

**image: more 3D structure of an actual self-splicing intron. looks like a protein… RNA can have discreet 3D structures that are associated with their activities

none of our introns are self-splicing, but the mechanism isn’t so different… some of our introns have the same basic structure/arrangement and same mechanism of removal as the self-splicing introns.

**image

we see characteristic seqs are the two junctions, we see a central A that acts as a branch site. in our cells, splicing reqs accessory molecules that bind to xscript to be spliced — we call the large group of molecules important for splicing the spliceosome. GIGANTIC— made of dozens of molecules, some are proteins, some are RNA. we call the complexes of protein and RNA that are important for forming the spliceosome “small nuclear ribonucleoprotein” complexes— snRNPs. bring splice sites close together so that the same mechanism we see in self-splicing will occur. snrnps will bind, brings the branch site to the 5’ end of intron, similar biochem rxn where we see cyclized A residue, exposed end of exon attacks the other junction, form lariat.

**whiteboard drawing

snrnps open us up to the possibility of regulation — can splice out the way we’d expect, or could even use the blue binding sites and cut out two introns and an intervening exon! this forms a new, similar but different protein. by cutting out an exon, we are losing ~30-60 amino acids… and it’s okay. all hell does not break loose. we create a protein that we can expect will have a similar, but likely not identical, function to one produced by the original transcript. there are ~8 introns in a typical human gene —> # of combinations possible becomes difficult to sort out, and not all work. not random, bc not all possible combos will be used.

**image

look at a xcript for which there are 3 exons and 2 introns. top example: can splice as described before to either take out only the introns or also cut out the intervening exon — skipping

second example: four exons, three introns. can produce 124 product, or 134. use one intervening exon or another, but never both

third and fourth: the sequence in yellow can be either an intron or an exon — cut out or left in! is an exon in name, as it can be retained, but can also be cut out like an intron.

**image

we can have proteins that direct snrnps to/away from certain sequences so that they are or are not spliced in certain ways


4/16:

how can splicing be regulated?

SR proteins — best understood that bind to mRNA and regulate binding of SNRNPs. really high conc of Ser and Arg residues. bind in exons and regulate the assembly of the spliceosome at corresponding seqs. thus how spliceosome binding varies from cell to cell.

alt splicing is highly regulated.

splicing useful bc allows expression of dif versions of same protein → greater degree of complexity.

generally, we see a lot more introns in multicellular organisms than in unicellular ones like yeast. splicing is likely to be important for having multiple versions of proteins that are important to development. main difference between us and other animals (like wildebeests) is not the genome, but the order in which the genes are expressed in early development.

when you have neurons that are activating their neighbors to make APs, the way that they interact w each other is physical and biochemical. there are proteins found on upstream and downstream neurons, and those proteins interact w one another to allow the neurons to bind with one another to allow for synaptic transmission. these are subject to differential splicing events — dif versions of these proteins expressed in dif cells, or even in dif regions of the same cell.

***image

from a single gene, we can express one version in one region and another version in another region. affects how our nervous system works

**image

*whiteboard drawing

evo perspective: introns give chance for recombination events that wouldn’t be possible in organisms without them. if recombination in coding regions was random, there’s a low likelihood we’ll form a functional protein. exons correspond (more generally) to protein domains. protein structure “rhymes” — very few structural domains relative to the overall diversity of life on the planet. include kinase, ATPase, etc domains. enzymes that hydrolyze ATP all have ATPase domains, and all of those domains look very similar to each other. domains correspond roughly to the exons we see in our genes. so if instead of seeing recombination under the conditions he has drawn, we make the second recombination event, we have a much higher possibility of generating a functional protein product. we have disrupted no coding sequence. we have generated a protein in which we still have functional folding domains. benefit of introns: we build in recomb events that generate diversity.

**image of proteins involved in chromatin structure in euks.

we see that they’re all made up of same basic domains. they differ in the number of domains and their sequences, but the domains themselves are very very similar.

exon shuffling: the differences between us and our neighbors (evolutionarily speaking) is less important when we think about individual nucleotides and amino acids. big differences between species is instead accounted for bythe way the exons have been recombined to form new genes.

**image

cell bio perspective: messages are not exported from nucleoplasm until they are spliced — sent thru nuclear pores only when processed. has a unique sequence seen only after splicing — at the junction between exons. proteins (EJC — exon junctional complex) bind in seq specific way at the exon junctions. hooks message up to machinery necessary for export.


why did we learn the LAC operon? it is the genetic system to which we will compare all others. LAC operon is our first example of understanding how a gene can be “on”.

**whiteboard drawing

active beta galactosidase is measured in cultures. if you take bacteria and grow them in glycerol (can use for energy, but not well), and in absence of lactose, cells do not make active b-gal. if spiked with lactose, they’ll turn it on essentially immediately. if instead of glycerol they’re grown in glucose, then spike with lactose, they won’t turn it on until the glucose is gone. tells us that they “know” what food sources are available to them and only make machinery necessary to metabolize lactose when it’s present, but it’s obv not the only thing they sense. how can they possibly know to wait until glucose is gone?

originally thought there was a biochem explanation, but it is in fact genetical. it’s not that they always make b-gal that they turn on/off biochemically, but instead when we look at b-gal, we can measure its concentration for not just the active b-gal and perform the same experiment and we get the same results.

**whiteboard drawing

made random mutations. saw three kinds of mutations routinely:

  1. blue: class I mutations. cells always produce b-gal. identified regions in genome where these mutations land — which we will call I, O, P.

  2. green: class II mutations. in genes we call Z, Y, A.

  3. red: class III mutations. found in genes O and P again.

O and P mutations can result in overexpression or underexpression.

we can map.

**whiteboard drawing

somewhere in here there must be a gene to express b-gal, and somewhere there must be seqs responsible for regulation of that gene. how do we know which is which? they couldn’t target each to mutate… didn’t have the tech. did partial diploid analysis. see if any of them could act from far away — on an entirely different DNA molecule together. if one of these is your favorite lac operon gene (YFLOG) and it could act on a gene other than this one, it strongly suggests that it codes for a protein that is regulating that gene. any gene that can exert its effects from a different chromosome works in ‘trans’. if it can’t, and YFLOG had to be on the same chromosome as the other genes, it had to act in ‘cis’. we’ll call these assays the cis/trans test to see what the different genes each do.

**whiteboard drawing

could add a gene to e coli using a plasmid. they put together dif combos of these versions of the gene they had — I O P Z Y A. wild type and mutant versions put in combinations together to see what happened.

bacterial DNA in black, exogenous (plasmid) in red. we will look at pair-wise combos but remember that all of the others are still there. look at O and Z together. bacterial are both wild type; use O+ and Z+ for the plasmid too. phenotype would be normal, meaning that if we grow in glycerol, expression of LACZ only when we add lactose.

now let’s try O+ and Z+ bacterial, O- and Z- plasmid. still behaves normally… bc we still have a normal copy of each behaving in a dominant way in the diploids (or so it SEEMS)

now if we separate the O- and Z- and put them on separate chromosomes? those cells do not behave normally. we have to modify our past conclusion — they have to be on the same chromosome. thus O and Z are cis — the normal copies HAVE to be on the same chromosome! phenotypically, looks like b-gal expression is on all the time.


different genes in combo: instead of O and Z we look at I and O? the two strains drawn on board behave the same as each other — we need a normal copy of each for regular behavior to occur, but don’t have to be on same chromosome… trans. why? because I codes for a protein that binds to O…. and proteins diffuse!

however if we have no normal copies of either, then b-gal is expressed all the time regardless of lactose presence when grown on glycerol


4/18:

why did they do the partial diploid analyses? what problem trying to solve? they know they’re looking at genes important for regulating glucose metabolism at the level of gene expression — they asked what it means for a gene to be on/off, and how do you turn them on/off?

*whiteboard drawing

P is the promoter. promoters are seq to which pol binds. has other protein binding sequences — the operator (O)

the mRNA produced from this xcript has 3 coding sequences in it. e coli has coding machinery for 3 dif genes. will be translated separately in 3 different forms. Z encodes b-gal. Y encodes a permease allowing lactose to cross membranes better (transporter protein). we don’t know what A does.

lac I has its own promoter — codes for lac repressor (lac R). if no lactose around, it binds @ operator → RNA pol can’t bind to promoter. ground state: always off due to this repressor.

**image

when lactose is around, the lac R binds to it and changes shape and can no longer bind to the operator → RNA pol now has potential to bind. this is called an inducible system bc presence of a certain compound induces expression of genes necessary to metabolize that compound.

example of repressable system: AA synthesis. don’t need to make all of them all the time — once we’ve got enough of a particular AA floating around, the system that makes them is inactivated.


the previous observations about O/Z, I/O, etc are different when grown in glucose instead of glycerol

how do we build a genetic circuit that is sensitive to two compounds @ the same time (lactose and glucose)?

CAP protein, unlike LacR, is an xcriptional activator and will drive expression of genes by binding to DNA and stabilizing RNA pol binding @ promoter sites. in order to drive max expression of lac operon, must not only remove Lac R but stabilize RNA pol binding to promoter site. have CAP bind to cis-acting site upstream of the binding site (the CAP-binding site). CAP is dependent on glucose presence, not lactose.

CAP binds only when there is NO glucose! binds to a catabolite that acts as a sign that glucose is depleted — cAMP.

cAMP binds to and regulates CAP. when bound, now CAP binds to the CAP binding site. thus cAMP and glucose are inversely related — when conc of one is high, conc of the other is low.

lac Z is dependent on conc of both lactose and glucose; glucose regulates more than just lac Z

combinatorial control — combo of signals determine if a gene is expressed or not.


*2 images

not the only way genes can be regulated in e coli. sigma 70 story, but there are others. here’s another inducible system — involved in nitrogen metabolism. story we think in broad terms is gonna be the same: see expression of genes only if RNA pol can bind, see sigma 54 instead of 70, see xcriptional activator not near but downstream (a few hundred bps down). but that’s enough because they’re far enough apart that DNA can bend to bring the promoter and the place where the xcriptional activator binds together.

**image

can make it work by regulating degree to which helix is twisted in DNA molecule. e coli have scriptional regulator that binds to promoter sites. when binds to mercury, it undergoes conformational change that affects degree to which DNA is twisted, bringing sequences in major groove into right orientation to maximize binding of RNA pol.

**image

in bacteria, regulatory sequences are somewhat close to the promoter (even in above case where they’re a few hundred bp down). in euks, those cis acting seqs that drive expression can be thousands, tens of thousands, hundreds of thousands of base pairs. we call groups of those seqs to which TAs bind enhancers. usually have many enhancers for each gene.

some fundamental questions: how do we get an enhancer that is so far away into the right conformation for it to work. if there’s multiple enhancers, do we need binding to all/some/certain combos for it to work? do some cells use different enhancers? do different button-pushing combos lead to differences in cells, development, etc over course of life? what matters here beside cis-acting seqs and their proteins? are there things in cells that change DNA such that it can/cannot bind to regulatory proteins? inherited? inherited from one cell to another?


*image

lymphocytes and neurons are functionally identical in terms of accessible genome (ignoring antibody junk in lymphocytes). note differences in shape and size — what tells a neuron to do that? partially genes, partially environment. link between nature v nurture — are they truly separable from each other, if enviro affects the genes and their expression?

how do we do this? if the dif between 2 cells isn’t in the genes, then what? and how might gene expression affect the difference between the two?

Sir John Gerden. is the difference between and embryo and an adult the idea that the embryo has all the genes necessary to create any body part, and the cells ultimately only retain the genes that are relevant to whatever organ/whatever they differentiated into? this was the leading idea for a while, but we know this isn’t true… you can cut a branch off a plant and stick it in some dirt, and it’ll grow an entire new plant. that wouldn’t be possible if the branch retained only branch-genes and got rid of all the root genes, flower genes, etc. tougher to prove in animals… cut off a limb, see what happens. depending on the organism, you might regrow the limb…

take a salamander. cut off leg — can we grow a whole new salamander from that leg? sure you can! compare to identical twins — 1 egg, 1 sperm. embryo splits early on, and both halves regrow and create two identical (clone!) offspring.

vanishing twins — two embryos fuse — can create conjoined twins, or have no visible impact at all. some of us in class maybe chimeric thru this process.

can take a frog egg and replace the nucleus from an adult frog’s skin cells (or any other of its cells) and it will grow to be a regular tadpole.


4/21:

looking at complicated multicellular euks, we see LOTS of enhancers — there might be dozens. also see insulators, which tend to squash gene expression when proteins bind to insulators far from the promoter (can be hundreds to hundreds of thousands of base pairs away!). much more complicated system for a typical human gene than in the typical e coli gene. combinatorial control — enormously important for understanding gene expression in proks and euks; whether a gene is expressed depends on the combination of ? experiences at any given time (can’t make out what he says in the recording). in euks there are just a lot more. what combinations allow for gene expression? what is the outcome? what does it mean for a gene to be on?

in e coli we can think of gene expression as near-binary, where genes are off or on, and we can flip that switch back and forth. this might be an oversimplification, but it’s good enough for us rn.

euks: because there’s more combinations, it’s less like an on/off switch, and more like a rheostat/dimmer switch! it’s tunable to a degree that most prokaryotic genes cannot — “turning it up to 11”

**image for cow cloning

if we can use any nucleus to clone an entire animal, then what events make it so that cells differentiate into specific tissues?

Shinya Yamanaka — a japanese scientist in a race to isolate stem cells. isolated from sheep, cows, primates, and now human embryos. lost that race to pull individual stem cells from human embryos and grow them in a lab to differentiate into the 3 major tissue types. he won the next race though — found that you can make them yourself. pulled out an epithelial cell, had to “de-differentiate” it. at some level, this happens at least partially in all tumors, particularly regarding metastasis — must de-differentiate so that they might gain characteristics of different tissues where they land. think of teratomas (tumors that can grow hair/teeth/etc).

can we do this in a regulated way? we want to not only produce a stem cell, but a stem cell that is pleuripotent (can become the 3 tissue types). all we have to do is drive the expression of 3 transcriptional activators (he thought it was 4, but we know now it’s just 3 necessary) and suddenly we have grown stem cells that will grow forever and have induced pleuripotency.

means: you can take a cell from yourself and turn it into a stem cell that can grow any tissue type. can use to study a genetic disorder. can use to turn them into a cell type you’re missing, and transplant it back into yourself. call these human induced pleuripotent cells (HIPs). there is an institute in aurora dedicated to STEM cell research. they decided early on to go all-in on HIPs.

it’s the reverse of all the cloning slides we’ve seen — we are taking a cell that already expresses a specific set of genes, and making it so that it is capable of expressing any set of genes!

what does the chromatin structure look like in a pleuripotency induced stem cell compared to a differentiated cell? — the 3 TAs that induce pleuripotency initiate chromatin reprogramming — all chromatin is now available for gene expression! new question: how do we get them to re-differentiate? what is the fear of using this tech therapeutically? … teratomas could form.

what is the degree to which we see differential gene expression?

**image — called a heat map. looks at 10k genes simultaneously. y axis has thin lines stacked, and each corresponds to one gene. moving left to right, we see cells derived from tumors from a variety of different tissues. colors tell us the ratio of how much of a particular gene is expressed in the tumor vs a non-tumor in the same organ. if it expresses more than the normal, we see green. if the tumors express less of a gene than we see in a normal cell, we see red. in between, no color at all. not all identical to each other… but we can compare. i.e. we see a chunk of green genes all together, and we can compare to lung tumors where those same genes are underexpressed (red) vs breast tumors that don’t show much difference in expression compared to the regular cells.

**image

story of various cancers looks a lot like the story for induced pleuripotent stem cells

if we take a region of the genome and pack it differently, gene expression can be changed in big ways.

although much gene expression occurs at distal ends of chromatin loops, many are expressed in euchromatin for which the stretches of DNA that are expressed are wrapped into nucleosomal structure. can imagine a scenario in which a promoter’s ability to bind RNA pol depends on the particular arrangement in that nucleosomal structure. can imagine that TAs can do their job when they bind to enhancers by just changing chromatin structure. very different story from the story of the lac operon bc there was no chromatin structure to be changed. imagine now that we have a transcriptional activator binds to an enhancer. brings in a HAT (histone acetyltransferase) that will add acetyls to histone tails → generally leads to unpacking and relaxation! tend not to act just at one nucleosome, as the HATS tend to sit on acetylated nucleosomes so that they will activate their neighbors.

not the only way that transcriptional activators must act. not only can they bind to HATs, they can bind to chromatin remodeling enzymes (learned about ages ago…tend to act like a pulley/spool to pull DNA off of the nucleosome). just have to pull on one end to spool off a particular stretch of DNA. chromatin remodeling enzymes tend to have those kinds of activities.

if transcriptional activators can do this, transcriptional repressors can do similar/opposite activities. we can activate a gene by bringing in a histone acetyltransferase or by bringing in an activating chromatin remodeling complex, we can turn a gene’s expression off by removing acetyl groups or by reversing the actions of the chromatin remodeling complex. a typical gene might have a dozen enhancer sites to which these kinds of activators can bind. you can imagine that in some cells to express YFG, you need to change chromatin structure, and in some you won’t need to because the structure is already available.

HIPs — all genes necessary for differentiation into the three tissue types are ALL available for expression.

remember that this is all just part of the story of expressing any given gene… very rarely will this be the only step involved in expressing YFG.

**image

can they act like CAP? can they bind in trans and directly affect RNA pol’s ability to bind to promoter site? yes — thus our transcriptional activators can act similarly to LAC, and in other ways. this slide emphasizes repression. a) removing the activator will repress expression. b) bind at distinct sites and directly interact with the activator to affect its functionality. c) direct repression preventing RNA pol or activators from binding; squelching gene expression by removing acetyl groups (discussed earlier today)

number of combos we have by which we can both positively and negatively affect gene expression in euks is far more complicated than in proks.

**image

insulators — mysterious behaviors. have a promoter and enhancer. enhancer binds to transcriptional activators that act on promoter to drive expression even at a distance. here’s the issue: clearly are stretches of DNA between enh and promoters that regulate availability of a promoter by that particular enhancer — these are insulators. sometimes will interfere w ability of a trans-acting TA to act on its promoter. insulates a trans-acting binding event and prevents a trans-acting TA and its enhancer from having its typical effect. if we have multiple promoters, and an enhancer that can act on multiple genes (scenario C) then the insulator state nearby determines which one is picked — a flip-flop switch where we express one or the other, not both. HOWEVER, there are several enhancers for each gene on average in the human genome… so it IS possible to express both at the same time if another enhancer is activated for the one that is not picked by the insulator!

looking at D, it’s like an OR switch. specific combo of enhancers occupied vs unoccupied


4/23:

paper discussion

think abt and characterize xcription of groups of genes together. emphasizes xcriptional networks

CAP activates ~100 genes at least, not just LAC operon. expecting many genes to be regulated. Lac is downstream of CAP and the Lac repressor Lac I, which regulates Lac Z, Y, A.

this paper examines yeast (saccharomyces cerevisiae). bc they are haploid, we can modify them really easily

asked: what are all of the genes that encode for xcriptional regulators — in yeast, appears to be 5%; and 2% overall encode for transcriptional activators. thus we have 144 TA genes and thus 144 CAP-like proteins in yeast.

for these 144, how many genes does each regulate on average? for each individual gene, how many TAs regulate it? can we start to build networks and understand how one TA might regulate its neighbors, and what do they look like?

researches tried to tag these TAs 1-144. the tag they added comes from a TF known as c-myc. just a few AA residues added to the end — like a flag. way to pull this protein out from the cells that express it.

now we have in principle 144 strains of yeast that are all identical except for the TA tag they received. these cells bust open and we break apart their genomes. can do really carefully to find the regions where these TFs bind. can also assume that each TA could bind in multiple spots.

106 strains, tagged TAs, pulled them out of a cracked cell in a way that DNA remains bound, see what sequences are bound.

first question: how many promoters does each activator bind? — varies enormously, but broadly we see the following: if we look at it in a way based in an promoter-based view, ~1/3 bind 1-20 promoters (out of 8k total). ~1/3 bind 21-40. ~1/3 bind 41-180.

experiment performed under standard lab conditions for optimal growth. thus we can only say under normal conditions, this is how genes are regulated.


for the 144 genes that encode for TAs, what TFs regulate their expression?

figure 3 shows a few examples of how they are regulated.

Ste12 is a TF involved in production of “gametes” in yeast. regulates expression of itself — autocatalytic loop — positive feedback loop! good for changing cellular behaviors quickly and irreversibly. antihomeostatic. anaphylaxis is another human example of non-self-limiting positive feedback loop.

single-input motifs: Leu3 regulates 3 genes under these conditions. push one button; generate lots of responses. one TF when activated can regulate a set of genes.

multi-input motifs: 4 genes are regulated by 3 different TAs. possible that we can turn on certain genes by pressing different buttons and don’t need all TAs for the desired genes to be activated, or possible that we need all three to be active and thus limits the conditions under which this pathway could occur. both are possible.

multi-component loops: another form of a positive feedback loop. A activates b activates C activates d activates A. a square!

feed-forward loop: A regulates b which regulates C which regulated d; but A also regulates d.

chain: A regulates b regulates C regulates d regulates E regulates f and g. specific order of operations, like making a sandwich. this is used in mitosis! specific order of genes expressed in particular order, where each is expressed one at a time. same goes for the cell cycle as a whole


what happens when we look at these ~106 TFs in groups? organized the TFs into groups based on known function

  • the developmental genes interregulate each other to a significant degree, and regulate others (like a few from metabolism).

  • metabolism genes tends to regulate each other extensively, as well as other systems. but are very tightly linked to each other

  • cell cycle genes: regulate each other and organized into sequential pattern we discussed previously. again, also regulates TFs from other groups

  • did not discuss others


in humans, ~15% of our genes are dedicated to regulating expression of your genome. true for all multicellular organism. difference between us and yeast? we have brain, lungs, appendages, head, etc all in specific spots.


4/25:

how to think about complexity of gene regulatory networks in multicellular organisms like us. how is it that in development we reproducibly create an organism with our degree of polarity (head/tail, proximal/distal, dorsal/ventral, etc). do it through gene networks where TFs control expression of other TFs which in concert regulate expression of genes to create particular body parts in particular locations.

Turing model — describes polarity. didn’t know abt TFs, but knew about morphogens (hormone-like molecules that regulate production of morphological structures) that turn on TFs— close in function to TFs themselves.

**image

morphogens/TFs: one can be activated at one end, one at the other end. because we have combinatorial control, we can have TFs dependent on conc of both morphogens — blue. we could also have one that’s dependent on just red and blue — black line. etc etc etc. same story can be true across proximal/distal axis, ventral/dorsal, etc. can start to see that if we mapped D-V, we’d have another set of peaks **image.

thus we’d have even more specific location-based gene expression. this part is what Crick figured out

another way: Turing’s original idea. body is a set of stripes — vertebrae, spine, ribs, hands/fingers, etc. diffusion model.

**image

have an activator: positive regulates self (+ feedback loop) and activates another thing/partner. the partner is inhibitory and inhibits the activator.

imagine now that they diffuse at different rates, for one reason or another. (i think he said the inhibitor diffuses faster). what would we see? see a zone of activation where conc of activator is high enough, but as we move further and further away we have less expression of that activator. have a halo of inactivation around the central zone. this is how cheetahs get their spots, how tigers and zebras get their stripes, rare sheep with black/white stripes. we don’t really have stripes like that on our skin, but we have them in our spinal column and feet and whatever else.

gradient of molecules that determine back/front, up/down, left/right, etc.

work done mostly in insects because they’re easy.

even before fertilization, the fly egg is morphologically polar — one end looks different than the other; one becomes head and one becomes tail.

**image

if you take material from anterior end of egg and put it in the posterior end/middle, at least for a while it’ll appear to form two heads.

bicoid — what makes this work. perhaps this is how A/P patterning occurs!

other possibilities: let’s say we have an egg that is uniform throughout. can imagine that at the point at which the sperm fertilizes the egg becomes a specific structure and determines the orientation of the rest of the organism… not how it works in us or in flies, but in some simple organisms!

think of bicoid as a transcriptional activator. will bind to enhancers and stabilize RNA pol binding at sites where genes rely on bicoid.

**image

mRNA is found at one end of the egg and not the other. mRNA interact with motor proteins that bring them to one end and not the other. if mRNA is conc on one side, then the proteins that they encode for are free to diffuse. shortly after fertilization, we have syncitial development where we have mitosis without cytokinesis — 1 cell, 1k nuclei. bicoid is a TF that does its job in the nucleus and is free to diffuse.

why important? bicoid is the kind of TF that exhibits thresholding behavior.

if we incr temp, will see incr rates of diffusion and will see more bicoid on opposite end of egg. but we don’t end up with flies with bigger heads — we must be missing part of the story. if there’s another factor working from opposite end, then its diffusion would also incr with incr temps, and the center of the overlapping point would be the same.

**image

we do change things if we change dosing of bicoid.

**image

another experiment: old-fashioned genetic screen. random mutations, see which ones lead to developmental defects. mutagenized only mothers; looked for eggs that are mutated. look for embryonic lethals (don’t develop at all). can bend them into groups

**image

gap genes — large segments where the middle of the embryo is gone — i.e. kruppel.

pair-rule genes — when missing, results in deletion of every-other body segment. i.e. even-skipped (we will talk about IN DEPTH) — missing even segments 2, 4, 6, 8, etc. other examples are odd-skipped and fushi terazu.

segment polarity genes —he segments each are polar themselves — have a front half and back half. segment polarity genes only affect one half of each segment. i.e. gooseberry

think of all these as encoding TFs.


**image

**whiteboard drawing

bicoid is a maternal gene — expressed solely by the mother and laid down in the egg before fertilization. leads to gap genes — kruppel and hunchback. these genes lead to regulation of pair rule genes which lead to regulation of segment polarity genes. all of these together regulate homeotic selector (HOX) genes. HOX genes encode for more transcription factors


fineness increases the further down we move in these developmental programs

**image

kruppel fluoresces over a wide chunk of the middle, but is pretty specific — expression in one can be high, while its neighbor’s is low.

hairy: kruppel is still active, but we just don’t see it in some — get stripes that can be different from each other. one can be hairy + krupple; one can be hairy + not kruppel

engrailed: 2-3 nuclei. even more specific, even more combos

**image

exquisite degree of fineness between cells — layer is essentially 1 nucleus thick! cephalic furrow shows the EXACT location where the head develops — not in the stripe to the left, not in the stripe to the right, ONLY there. what’s the sequence of transcriptional events that drives the expression of this gene here? specific combination of genes in play for expression of one stripe vs another??

will look @ a particular gene: eve (codes for even-skipped protein).

**image

expressed at this stage in development in 7 stripes. what’s the combo of TAs thats necessary for expression in stripe 1, and is it different from the combo of TAs necessary for expression in stripe 2?

**image — structure of eve in drosophila

metallic blue sequence is the actual gene — coding sequence. promoter is just slightly upstream of the green arrow. have enhancers for which binding TAs is necessary for driving expression of this gene. some downstream, some upstream of coding. can get rid of each enhancer to see which stripes go away, if any.

easy to say one enh is necessary for expression of stripe 3, but maybe it works collaboratively with the other enhancers at the other stripes. how can we tease that out? confused, missed info, ASK.

there’s an enh for each stripe, and they can overlap a little bit. there is a distinct mechanism to drive expression in different locations. likely true for a lot more complicated stories as well.

first understood this @ eve stripe 2. **image

if we are gonna express this gene at this location, then we need a mechanism to drive expression here and inhibit expression elsewhere in order to have fine order of control.

**image

a maternal gene regulates expression of 3 gap genes which regulate expression of even-skipped, and particular order allows its expression in the stripe, and a different order will NOT allow expression elsewhere.


4/28:

**image

how do you reproducibly make a fly with all the right parts in all the right spots? — gene expression patterns (bicoid, gap, pair-rule, hox, etc). divide into finer and finer segments


spatial and temporal expression of stripe 2 is determined by all gene expression patterns before it (gap, maternal effect, whatever).

**image

**whiteboard drawing

be able to draw these FROM SCRATCH, WITHOUT NOTES — that should allow you to answer most questions in-depth!

also think about how you would try to explain to someone else who has taken genetics, but not any dev bio or mol gen.

bicoid turns on hunchback. very strong thresholding effect. get expression of hunchback only when we have enough bicoid — threshold is relatively low though, don’t need a ton of bicoid to drive hunchback.

bicoid similarly turns on giant, a xcriptional regulator. thresholding effect, but threshold is much higher.

hunchback inhibits kruppel.

even skipped is regulated by all four others: kruppel inhibits, hunchback activates, giant inhibits, bicoid activates.

**note that hunchback can be both an activator and inhibitor — FOR EXAM, think about how a protein could be both, based on context it’s in?

will see even-skipped where we see bicoid and hunchback, but not where we see giant and kruppel. (generally, see expression of a gene wherever its activators are and its inhibitors are not). perfect example of combinatorial control

this is just the story for one of the seven stripes… the only one we’ll focus on, but know that by definition they must use different patterns/combinations of inputs!

**image — what the enhancer looks like

more than one copy of each protein binds at the enhancer, and the repressors that bind will bind also to the DNA. repressors can work in lots of ways — can repress even-skipped to avoid expression out of stripe two by binding to DNA and keep RNA pol II from binding (but probably not this). could bind and modify histones and convert to heterochromatin. could bind at sites that other transcriptional activators would normally bind to. many more possibilities. this is the inherent challenge in understanding gene expression in euks, esp complicated euks.


moving on to epigenetics ("above genetics”)

classical examples: crocodilian sex ratio is determined by temperature at which the eggs are incubated. clownfish gender is plastic.

epigenetic effects have some permanence — get passed down minimally from cell to cell, and maximally from individual to individual

what is the mechanism + what are the ramifications of this occurring?

how do we manipulate this?


**image

can change activity of a TF, and can usually change it back, but even if we can’t it’ll probably die soon anyway. very little cell memory for TF changes.

what about chromatin structure? some TF regulate enzymes that control chromatin form, and the chromatin structure is passed down from cell to cell

describe an experiment that suggests this is true? the “peppermint candy” yeast experiment where we put a gene into a region of heterochromatin, and occasionally it turns into euchromatin and can be expressed (and stays that way).

when DNA is replicated, the histones are displaced. half of the histone proteins are left behind and new proteins have to come and bind. the ones left behind bear modifications from the previous cell. those modifications remain and provide signals to recapitulate machinery to recreate the modifications seen before the S phase.

if we can take the cause of some environmental insult and activate a TA that converts heterochromatin to euchromatin, we can do the opposite. only need to activate it once. can be passed from cell to cell.

epigenetic memory might also be found in DNA methylation story. DNA in euk cells gets methylated, specifically CG/GC dinucleotides. the Cs get methylated, if they get deaminated, they turn into T and struggle to be repaired. so why do we methylate them? what good is it? → the Cs that are methylated regulate protein binding.

**whiteboard drawing

TAs that bind to the enhancers stabilize RNA pol II binding at the promoter — only one possible story. but let’s imagine there’s lots of CPG islands in the enhancer — dozens of CG dinucleotides. there are cytosine methyltransferase proteins that methylate the Cs — generally if one gets methylated, most will be methylated. will be added in such a way that C can continue to bind to G. an enzyme will sneak into the major groove and add the methyl off to the side of the C so that it’s actually hanging off into the major groove.

when we methylate the Cs, proteins bind more poorly. that TA is less likely to bind when that region (enhancer) is methylated. so we might turn on/activate/whatever the TA we want, but it doesn’t matter if that section is methylated.

methylation of these Cs gets passed down from generation to generation.

dependent on whats happening in environment and whether the template DNA was methylated here or not.

what happens even as we take that signal away — likely remains methylated. the DNA has memory. will look the way its predecessor looked. still possible to turn that back off again, but would need addition of a secondary signal rather than just the removal of the original signal.

**image

there are opportunities for methylation. proteins that bind to methylated C. tend to be histone deacetylases.

histone deacetylases remove acetyls from histones, pushing towards heterochromatin (packed more tightly) → another way that expression is restricted/regulated

thus methylation is broadly associated with decreased gene expression

next we need to talk about places where it matters — talk abt mech on wednesday. gene called IGF2 (insulin-like growth factor 2). promotes muscle growth. highly regulated by methylation. express only 1 copy of this gene. copy is from paternal chromosome. maternal and paternal chromosomes are differentially methylated. get this from the dad because he wants baby to be big and strong, while the mom wants baby to be small and easy to actually give birth to.


4/30:

gene modifications can be reversed, i.e. thru introduction of a hormone or creating a negative feedback loop

if this hormone acts bc it regulates gene expression via an epigenetic mechanism (changes dna and its modificaiton, or changes chromatin in an inheritable way) the change can persist over time and through cell replication. epigenetic changes can be heritable! looks kinda like lemarck’s giraffe

what are these mechanisms and how do they work?

revisiting 1st unit here: talked abt histone acetyltransferases, how other enzymes do the opposite, introduced idea that DNA can get methylated (but we introduced it as a PROBLEM… but it has some biological function — regulates protein binding! in broad sense, the transcriptional regulators that bind to DNA by making contacts in atoms available in the major groove, they bind less well when the cytosine residues in the major groove are methylated.. when we talk abt RNA pol or TAs being necessary to express a gene, they bind less well in regions that are methylated. steric hindrance and methylation has effect on chromatin structure. methylated C —> fewer DNA regulators, plus packed more tightly so less is available for binding

**image

two genes in same cell. no poly-A, has 5’ and 3’, ATG indicates start, TGA indicates stop. that’s the coding sequence. hell of a lot of CG islands. image on left has essentially all methylated, on right is all unmethylated — general rule that when we methylate a region of the genome, we methylate all of that region pretty quickly, and other regions won’t get methylated at all.

for a couple hundred genes in the genome (~1%), we have imprinting — inherit a copy of the gene from each parent, but only one or the other is expressed.

Igf2 story

insulin is a potent anabolic hormone. made when glucose levels are high to shuttle the glucose into tissues and start anabolic cascade to store it

Igf1 is produced downstream of growth hormone, which is made majorly in puberty. GH → Igf1 in liver → → → effects everywhere — in bone and muscle especially. hence why during pubescent kids tend to get taller and put on muscle more easily. even after puberty, this pathway still exists and will be activated by exercise (in which case it will become dependent on catecholamines like epi/norepi/dopamine).

Igf2 looks largely the same, as it promotes bone and muscle growth and is a potent anabolic hormone, but we only make a little of it. maximal expression was during gestation. responsible for rapid growth that occurs in utero. this is the system depicted in the image above. kind of a classical and important way to think abt epigenetic combinatory control. scenario where we have 2 genes (one for Igf2, one for H19). nobody knows what H19 does, other than the fact that it codes for an RNA molecule instead of a protein. have essentially an either/or switch. express one or the other, not both Igf2 and H19. if we get rid of H19, the pups born are small. what does it look like? we have an enh for these genes. will act either on the promoter for Igf2 OR the promoter for H19, not both. the one it acts on is very chromosome-dependent. will express one and only one from one chromosome, and on the other we will express the other gene. the maternal chromosome expresses H19 and not Igf2. the paternal chromosome expresses Igf2 and not H19. talking abt expression primarily in utero. what’s the mechanism? in maternal, this region is not highly methylated, so proteins can bind in a seq specific way, and we tend to see more euchromatin. an insulator can bind to its partners (CTCF — affects DNA looping in chromatin structures). aggregates the possibility that these TAs will work on this promoter, and H19 is activated instead — limits growth

paternal chromosome: this region gets methylated, CTCF can’t bind, that enhancer has access for the Igf2 promoter because the chromatin is arranged different than it was before.

inherited the copy that expresses H19 from mom, Igf2 from dad, and did not express both.

how might such a system get set up? sexual reproduction incr genetic diversity, but what’s the cost? if you couple together two perfect greyhounds, you get a really good child. however that’s not typical of nature — one will be fitter than the other, and the offspring will be intermediate in fitness. dilutes out the genetic information. some possibility that there might be value in competition — different sperm cells compete for example. game theory: male wants offspring to be as big and strong as possible at birth because it’s strongly associated with quality of life and ability to pass on genes. female wants offspring to be big and strong enough, but too big and strong reduces the possibility for her to have more babies as it can take a toll on her body or kill her. we are seeing some interplay between the competition of paternal vs maternal interests.

what is the benefit of imprinting? what does it mean in terms of gamete production?

**whiteboard drawing

diploid individuals with the maternal and paternal chromosomes. can they pass on both chromosomes to the next generation? meiosis dictates that we randomly select one. if Igf2 were partitioned into sperm cell, doesn’t need methylation, but if the maternal one is, it must be methylated (whichever chromosome gets put into the sperm will be methylated if it was not already)! complementary story; if partitioning into egg, must have epigenetic reprogramming of paternal genes to be demethylated, but maternal genes are already demethylated. still 50/50 despite how methylation is energetically expensive. happens to ~1% of our genes during each meiotic cycle.

**image

talk now less abt imprinting and more maintenance of epigenome.

just like the way that DNA structure can be inherited, so can the pattern at which the way the CPG islands are methylated. how? let’s reverse engineer it

**whiteboard image

big long portion of genome. have a region that is highly methylated (if it’s methylated on one strand, it’ll be methylated on the other strand too). on other end of genome, we don’t have much methylation. DNA rep is semi-conservative. means if we copy this DNA, the two template strands will be methylated, at least in the regions where they were methylated previously. we’ve done nothing to the template strand. generate a new strand using the whole mechanism we dicussed in this class, which looks like the e coli story we discussed in unit 1 — hemimethylated DNA. e coli used this to its advantage. regulation of reinitiation of DNA synthesis. repair where they need to recognize the newly synthesized strand. if DNA methlation looks like it did in the template, then all we need to do is to invoke a machine to make it so. need to search for every site with hemimethylation, which will have a C residue, and nearby there will be another spot nearby that can be methylated. machine not only recognizes the methylated C, but will add a methyl to the residue that is roughly across from it in the double helix (CG dinucleotide). maintenance methylases/methyltransferases. now have 2 classes of enzymes that methylate your genome. one does it in a pretty precise way. can’t be too precise, but precise enough that is brought to regions of genome that are methylated, in a way that is dependent on conditions in which the cell exists. ultimately these methylated C groups are preserved even through DNA rep.

where else does it matter for us?

**image — water fleas

water fleas are clones of one another —genetically identical, but look totally different. when do we start to see spiky ones emerge? when we expose them to predators. it is NOT that a predator comes about and a water flea magically sprouts a horn…. but its offspring do, and will pass on that trait consistently. even if we take the predators away, the offspring of the offspring will retain the spike. genetic memory passed not only from cell to cell, but from organism to organism. occurs in mammals too, like mice

**image

mice to which BPA was added to diet. commonly used industrial chemical still found in canned goods, but used to be found in plastic a TON esp nalgene bottles. what happens to mice who are fed extremely high concentrations of BPA? pups/adolecent mice shown are all littermates and are all genetically identical as they are produced basically thru cloning. coat color is dependent on amount of BPA that they saw while in gestation in the mouse. effect is relatively mild, but real. why do we care? if it affects coat color, it almost certainly affects smth else. in mice, coat color is highly associated with metabolism… yellow ones gain weight at far higher rate. color and body mass are linked. some things like this might matter, some may not.


5/2:

talking about regulatory RNAs. large number of RNA molecules produced in cells that don’t code for proteins, don’t make up ribosome/spliceosomes, sole function is to regulate expression of other genes.

circumstantial evidence that RNAs matter for regulation. use petunias. wanted to make deep purple petunias — attempted to increase gene dosage. ID’d gene responsible for particular pigment, and added more copies of that gene. worked in a really general sense. as you add more of that gene, they get darker and darker until they reach an inflection point… the plants started getting lighter in color. the mRNA molecules get processed in a way that ultimately restricts expression of that gene

Fire Mello: interested in antisense technology. takes advantage of fact that you can knock down (reduce) expression by adding a seq of RNA that is complementary to the message that encodes for that gene. you have a mRNA molecule that encodes for a protein. GKD (gene knockdown) — want to reduce expression as much as you can to start to determine what that gene does. antisense works like this: take a seq of mRNA and generate an RNA molecule that was complementary to it. if done at right place + time, xlation rates are reduced, produce less of that protein. seemed random at times, nobody knew how it worked. Fire Mello did this work in c. elegans to figure out how development worked. tried using GKD for their favorite gene, didn’t work. control group was to add a “sense” strand — have a sequence that matches (is complementary to the antisense). looks like the sequence you’re trying to target. sense group also had no effect. however, they also added a third control where they mixed the antisense + sense → saw an effect that was highly specific… a knockdown! lowered expression of their favorite gene! this is known as RNA interference/RNAi. need to have it as this DS sense/nonsense strand so that it can bind the machinery for GKD to actually occur.

**whiteboard drawing

this mechanism by which RNAi works takes advantage of other mechanisms that preexist in these organisms.

**image (first one taken on 5/2; took pic out of order)

worm cell — feed DS RNA to these cells and make it such that the DS RNA is made up of seqs that correspond to a gene in the genome of that cell. DS RNA finds its way into cytoplasm and is processed by group of enzymes known as slicer and dicer. dicer cuts up into small chunks of ~20 nucleotides. name we use is dependent on where that RNA molecule came from — can be exogenous. when it is, it’s called siRNA (small interfering RNA). internal source of other DS RNA in this process is that cells make them in variety of different ways. i.e. mRNA with introns which have stem-loop structures processed by same machinery that exogenous RNA gets processed by. chopped up 20 bp RNA is called miRNA (microRNA). RISC — RNA induced silencing complex — another name for GKD/RNAi. is the complex that allows for silencing. RISC separates the strands. one is now a guide RNA (similar process to CRISPR/cas9). capacity to bind to mRNA in a seq specific way. in this case, the guide RNA is the antisense! has to bind using complementary base pairing with the message. if the match between the guide and messenger is near perfect, we activate pathway of degradation. when the guide binds, the message (mRNA) gets degraded, so it can’t get translated and expression is reduced. this is what happened in the petunias, and in the Fire Mello experiments.

other outcome: antisense oligo interacts with the xcript well, but not perfectly. well enough that it binds, but not so well that it attracts the machinery that initiates degradation. here we have a relatively stable complex, where the mRNA is now bound to guide RNA. also see knockdown and silencing, because the ribosomes can’t read through it, so it won’t produce a protein.

regulation of gene expression that occurs post-transcription. regulating rate at which proteins are produced from those transcription products.

this whole complex works not only in the cytoplasm, but in the nucleoplasm as well. back in nucleoplasm, seeks out the strand it’s complementary to — one of the DNA strands that coded for that mRNA that made it. FINAL EXAM QUESTION (JUST KIDDING HE MESSED IT UP SO HE WON’T PUT IT ON THE EXAM. I GOT THE INFO CORRECTED BUT DON’T WORRY ABOUT IT)— is the guide seq that RISC is using here gonna be complementary to the coding or noncoding strand (template/nontemplate)? if it’s binding to DS DNA at the gene that encoded for the RNA that it silences, then it must bind complementary to the mRNA and must be complementary to … template or nontemplate? let’s draw

**whiteboard drawing — INSERT THE ONE FROM THE END OF THE DAY; HE MUCKED UP THE FIRST TIME.

the template strand is used to make the mRNA. so the mRNA is complementary to the template. if the oligo is complementary to mRNA, then the oligo must be complementary to the nontemplate strand and the sequence would match the template strand.

why does it matter that it’s brought to particular region in genome? additional complex of proteins involved too

**Image

additional activity/machine separate from the mRNA degradation… when those RNA molecules bind to genome in seq specific way, series of processing events bring in chromatin remodeling enzymes. leads to differentially modifying nucleosomes. euchromatin → heterochromatin. that heterochromatin region grows in such a way that the gene is silenced and neighboring genes are commonly silenced as well. why does this matter for us? we make a surprising # of small RNA molecules processed by this machinery (at least a few thousand). most commonly produce these miRNAs from introns. the way they’re spliced and the DS miRNAs are processed ultimately feeds back into regulation of the genes from which they were produced.

**image

zebrafish. have been processed so we can identify where specific miRNAs can be found — expression patterns are distinct from each other. now we have an additional level of complexity to combinatory control present in euks and not in proks.

also matters enormously for cancers.

**image

his favorite protein is Ras. the protein at the center of the signalling cascade that leads to DNA replication. almost all solid tumors make more Ras than they should. Ras has a small miRNA that regulates its activity. mucking up the miRNA contributes to this overexpression, at least in some cases. can thus potentially use RNAs to attempt to treat conditions like this.


5/5:

miRNA is endogenous, siRNA is exogenous.

there is an antiviral component to this whole thing — DS RNA are not common in euks. but commonly produced in life cycles of RNA viruses. siRNA effect on pathways — at least in part an antiviral defense. very clear in plants! animals have a robust and complex immune system while plants have a lesser immune system. no matter the origin, we can now process DS RNA. take advantage of that when we do RNAi to knock down expression of genes. evolutionary viral defense has been co-opted.

hard to know how many of these we’ve found. probably at least a couple thousand. how many mRNA molecules are regulated by small RNA molecules that are produced endogenously — at least 60%.

3 or 4 fundamental questions for which we are not sure we have complete answers

Q1: how is sequence-specific protein binding achieved?

the (incomplete) story he told us in unit 1: thru interactions between the protein and atoms that are available in the major/minor grooves, where there is potential to distinguish base pair identity and conformation. would expect for seq spec binding that binding pocket would be 18-24 bps in length. would require more than a full turn of the DNA. that’s not what we actually see though… closer to 12-14 or fewer! they make far fewer contacts than you’d expect to be the minimum necessary to bind in a sequence-specific way. seems to not be enough interactions for specificity… maybe there’s another piece to the story. can you imagine scenarios by which this may be achieved?

1 reasonable idea is to use accessory proteins… but perhaps a better one. maybe the DNA structure we discuss throughout our school career isn’t complete… maybe it can be unzipped and we don’t have to worry about major/minor grooves. maybe when we look at DNA, it interconverts between the unzipped and double helical structures — “breathing” — allows for seq spec binding to the SS DNA — have strutures to which proteins can bind, for example, in a sequence nonspecific way but can also bring in other proteins like the accessory protein theory

nitrogenous bases potentially exposed — allows for specificity as well.


Q2: where does transcription occur? true that it occurs in the cytoplasm of proks and nuclei of euks. told a complicated chromatin looping story— DNA less tightly wound at the ends of the loops, and that’s true, but there’s also transcriptional hotspots. what is the nature of these? what do we find in them? GTFs, RNA pol II, mediator complex, and BRD4. BRD4 binds to acetylated histones. we occasionally find these loops. promoter sequences, when they find their way into transcriptional hotspot, that’s when they get activated. if we get transcription at this promoter seq, modern microscopy allows us to follow the life of that RNA molecule. appears as though even though transcription starts here, it ends/continues in part outside of those transcriptional hotspots. thus, the promoter needs to be in the hotspot, but doesn’t need to remain there. how did it get here to begin with? is this droplet always there and the DNA goes to + from it? or does the droplet form bc DNA binds to some set of proteins necessary for the formation of the hotspot? enhancers bind to their TAs, which have large region of intrinsically disordered protein — doesn’t form traditional secondary structures. tends to interact in multivalent way — forms globs with others with intrinsically disordered domains. all proteins here have intrinsically disordered domains. current model: TAs bind enhancers, and maybe interact with each other — leads to nucleating sites in which the hotspots form. when hotspot forms, promoter finds its way into the hotspot and that is what triggers initiation of transcription, but then it can quickly leave. will continuously go back and forth. switch between hotspot and the nucleolus — dynamic kissing. processing of ends and splicing occurs simultaneously, and the spliceosome components are located away from the hotspot.

does this droplet look the same in brain cells and liver cells? we do not know at all!! could throw out any hypothesis and can’t really refute them. how do they differ in STEM and non-STEM cells? between diseased and healthy cells? can we control this process?


Q3: how does the environment interact w genes to determine phenotypes? EXAM

epigenetics… but we are in our infancy of understanding epigenetics. fundamental truths: there are enzymes that modify DNA and chromatin. often call these writers. i.e. histone acetylases, histone methylases, cytosine methylases, etc. set of enzymes that recognize writers — readers. can be chromatin remodelling enzymes, TAs. if there are dozens of writers, there are hundreds of readers. info from writers and readers can be passed from cell to cell. writers change in a way that is preserved in DNA replication. must be a way to reset things and erase DNA + chromatin mods — erasers. plays role in deprogramming cells.

cancer — 1000s of genes are misregulated in cancer. if we have a tumor with thousands of misregulated genes… which of them are causing the cancer, and which of them are symptoms of the cancer? there must be master regulators that are misregulated that are critically important and seem to contribute to development of cancer. protooncogenes — when overexpressed, leads to tumors. tumor suppressor gene — should restrict development of tumors, and downregulating these genes leads to tumors. these two genes are good targets for treatment. could we use drugs that affect writers, readers, and erasers as anti-tumor drugs? can we use drugs that encourage expression of tumor suppressor genes? could we use drugs to cause protooncogenes to be packaged into heterochromatin to slow tumor development? there are currently 12 of these drugs prescribed on the market. 5x as many are in development. the oldest was first put on the market in 2015. if we can know how to manipulate structure of the genome, we can use that to hopefully treat patients

common question — how do they work? poorly. extend life expectancy by months at best.

wouldn’t you be nervous taking a drug affecting DNA structure/packaging? …everything you do affects gene expression and chromatin remodelling. taking flonase, for example


behavior is at least in some ways connected to genes. breed animals specifically for behaviors. could we use these kinds of drugs to affect behavior? should we? some mood modifying drugs do affect DNA/chromatin structure!


to prep for exam: focus on unit 3, relate back to unit 1 and 2. discussed these things to talk about today bc they required integration of all previous units. start to think in detail about lessons we learned in unit 1 that are applicable to unit 3.

prepare for big ideas about units 1 and 2? old exams. USE THEM. make study guide. WRITE YOUR OWN QUESTIONS. DO THIS. HIGHLY HIGHLY EMPHASIZED. TELLS YOU WHAT YOU THINK IS IMPORTANT — GIVES SENSE OF WHAT HE THINKS IS IMPORTANT.

will read a paper abt epigenetic behavior in rats for wednesday. MUST READ BEFORE CLASS. A L O T THERE. do what you can, but try to understand the question they are asking + broadly how they are asking it.


5/7:

started recording super late soz.

pup licking incr = stress decr = decr cortisol = incr GR → NGFI-A

pup licking decr = stress incr = incr cortisol = decr GR

stress response: makes it so that lots of glucose is available to fuel that stress response.

glucocorticoid receptor (GR) is a steroid hormone receptor in nucleus of cells, when cortisol binds it, it binds to DNA — GR is a TA for which its activity is dependent on ability to bind to cortisol

when we produce cortisol, lots of negative feedback events in both thalamus and pituitary, but also at level of GR expression. GR expression is turned off when cortisol is present. the more stress + duration of stress, the greater the negative response.

looking at GR expression as measure of stress in these rats

we know that in low stress state induced by maternal licking of pups, we turn on expression of gene coding for GR — NGFI-A. what’s the mechanism in stressed pups that accounts for shutting down expression of GR by altering NGFI-A binding or activity? is this epigenetic, and if so, how long does it last and can we manipulate it?

looking at CpG seqs. in these pups that weren’t taken care of mothers that exhibit high stress, what does promoter for GR look like? what is its methylation state?

epigenetic mechs we’ve discussed — direct mod of genome by methylating cytosines or regulation of modification of histone proteins (chromatin remodelling).

bisulfide mapping — if you take DNA from cell and treat w bisulfate, if the C is not methylated and is deaminated, becomes U; but if it IS methylated and get deaminated by bisulfate, it becomes T.

pups to non-licking mothers see methylation. consistent w story — those pups are stressed, don’t express GR, that region of genome is methylated and NGFI-A can no longer bind to drive its expression.

pups born to licking mothers are not methylated.

possible that methylation state in pups isn’t dependent on maternal behavior — how do we determine if this difference truly is bc of the maternal licking behavior.

take pups born to the licking moms. give them surrogate mothers. see if we get incr methylation when the surrogate is non-licking. found that in the earliest days of life, the findings align with what we’d expect — stress incr when surrogate is non-licking. thus it is maternal care that matters, not predetermined in genome!

methylation pattern established during week 1. once established pattern is generated, it lasts as long as they measure. thus, maternal care in very beginning of life has lasting effects.

look next at chromatin structure. use CHIP process — cut genome into chunks, pull out specific fragments, see if it is hetero/euchromatin — essentially, are the histones acetylated or not. we know when we methylated Cs, we reduce protein binding overall but call in proteins that shift the genome in that region to a heterochromatin state.

pups for which GR is low (no maternal care), the region that encodes for that GR is less acetylated, meaning it is more tightly packed, meaning gene expression is reduced. the histone proteins aren’t acetylated to the levels seen in pups that are taken care of by their mothers.

this is both nature and nurture — differential nurturing of their young, but that has a genetic basis and a molecular and epigenetic effect in the offspring. thus nature and nurture are linked to each other!

next: does methylation of genome and subsequent deacetylation of histones associated w lack of maternal care block NGFI-A binding? yes! those pups w/o maternal care don’t have NGFI-A binding to the promoter seq. threefold difference — unambiguously enormous difference in NGFI-A binding between pups depending on maternal care.

can we fix it? pharmacological intervention? can we use drugs/small organic compounds to reverse this effect? can you take these pups for which their DNA is methylated and for which their chromatin is tightly packed due to lack of pup licking, and can we add drugs that affect the chromatin? they used drug called TSA which inhibits the deacetylases responsible for taking acetyls off of histones. if we add these drugs, can we take this region that was densely packed, and can we open it back up for gene expression? yes, we see reversal in gene expression! they change so they have incr GR expression! thus this condition might be drugable (treatable with drugs); this is reversible! even for behaviors influenced by epigentic effects, behaviors that affected gene expression early in life (or even in utero), might be able to prescribe drugs that treat these! however, these drugs have enormous side effects. current specificity is not that high, so they have off-target effects and affect the packing of many other regions you don’t want to affect. only really a consideration in cancer patients —adjuvant therapy. don’t do much on their own, but when added to other drugs, have synergistic effects to make the other drugs in the cocktail work better!

considering if these drugs can be used not only for cancer, but also for muscular dystrophy, autoimmune diseases, social effects on epigenome (how does early trauma affect genome and long-term biology and success on the planet — mostly in social/pack animals like hyenas).


5/9:

office hours 12-2 monday and tuesday

exam wednesday 10am

alternative splicing — system that allows 1 gene to encode for more than 1 protein. can occur in a single cell where might splice same message in different ways — produce 2+ versions of protein. can also produce dif versions of proteins in dif cells. important for dev bio. what’s dif between proks and euks? complicated, but part of the story is that in euks the dif between single and multi celled organisms lies in differences in splicing, chromatin structure, etc for more variety. allows for specialization not possible in simpler organisms. describe one mech by which splice site is regulated? SR proteins bind to xcript, bind in exons, and act like TFs for splicing (do NOT CALL THEM SPLICING FACTORS) — determine where spliceosomal components bind or don’t bind and push splicing one direction or another. sequence specific!


lac operon: extremely important for gene expression. should know its parts and the things that bind and what they do. be comfortable with doing a question like #2 — might involve other pieces though!

note that CAP and lacR are not competitors for the same binding site. when high lactose low glucose, lacR is not bound to operator — bound to lactose and conformation changed instead. CAP is bound to cAMP and binds to its enhancer. we get transcription.

low lactose low glucose: lacR bound to operator; no transcription because RNA pol cannot bind. CAP is bound, but doesn’t matter bc we can’t stabilize the binding of RNA pol if RNA pol doesn’t bind in first place

high lactose high glucose: lacR not bound, CAP not bound. we see “leaky transcription” — we do get transcription, but minimally. nowhere near max efficiency

lac operon is a base model for other gene regulatory models in proks and euks. when we understand it in euks so poorly, we have to reference something — proks, even tho we know it works in fundamentally different ways.


fibroblasts: likely that there has been some progression towards differentiation. we are starting with a differentiated cell, it’s already determined as a fibroblast, but if it only takes a few myogenic factors to change them to muscle cells, it’s reasonable to expect that they have something within them that helps them “switch over”


mutation in lacO/I: how can we mutate lacO to block expression or allow expression all the time? make operator so that it can never bind repressor → expression all the time. to get expression never (or “delayed”), could make it work for lac repressor to bind all the time even when bound to lactose. what about lacI? to get expression all the time, either a mutation in lacI that either makes it so it can no longer bind to DNA OR so that it can no longer bind to lactose. how can you say smth abt the relationship between O and I? use cis-trans test AKA partial diploid AKA merodiploid analysis. for trans operating genes, doesn’t need to be on same chromosome. trans acting factors are PROTEINS — can DIFFUSE.


even skipped: **image

start by knowing the 4 TRs that feed into even skipped @ stripe 2 and how they combine in network to drive expression in the right time and place in development. bicoid, hunchback, giant, kruppel, eve. more complicated story than the broad ones discussed in class bc we have TRs that are sometimes activators, sometimes repressors. bicoid is upstream of them all bc it is maternal gene; regulated the gap genes hunchback, giant, and kruppel. they cooperate together to determine fine segmentation of the embryo, such as the pair-rule gene eve. what do each TRs do? bicoid → hunchback. bicoid → giant. hunchback —| kruppel. bicoid → eve. kruppel —| eve. giant —| eve. hunchback → eve. at the core of this question, thinking about combinatorial control where the map looks “just like” lac operon, but instead of 2 TRs that feed into particular circuit, we have 4 which regulate each other in a far more complicated circuit/network. how many copies of bicoid bind? some number other than 1 → raises possibility of important dosage effects. would not be asked number of TFs needed to bind for it to be expressed in a particular way, but know how these work and interact. thresholding effect — so long as we have more bicoid than the threshold of hunchback and giant, we get expression… but the thresholds for either are different. how might we achieve that? can imagine it’s a dosage effect. one of those genes is sensitive to smaller amts of bicoid, one is sensitive to greater amts of bicoid. what’s important abt spacial arrangement with which a TR binds to DNA could be important for its ability to regulate genes — yeah, why? most simple way to think abt it: have a TF that acts in negative way. how does it work? does it likely bind to promoter and keep RNA pol from binding bc of competitive inhibition? unlikely, not in euks. most of these TFs bind away from the promoter. if true, then a TF that slows gene expression may not interact w promoter in any way. instead, might bind to TAs and block their activity. thus their spacing and orientation on genome would be critically important.


TR models and bridge: fundamental question: how does such a distant enh affect the promoter? 1: binds to enhancer that acts as landing site, then zips down. model 2: does not req that it zip along, rather it invokes bending. how to determine? exp1: take enh and promoter on separate DNA molecules, see gene expression. exp2, link enh and promoter using avidin (protein). where do we see expression? in the second model, not first. does it support idea of zipping or looping? note that in the diagram, it is not to scale. DNA is WAY smaller than any protein. getting specific binding at a site in the DNA, relies on interactions with atoms in major/minor grooves. if zipping along, less likely to make interactions with the grooves; so it makes more interactions with backbone. thus it can’t scan through avidin because it doesn’t look like DNA. could invoke hopscotch answer where it jumps over. more likely though that the protein brings the enh and promoter closer together. otherwise, when not bridged, the two segments are diffusing around the cell. if you bring them closer, you increase local concentration, they move TOGETHER through cell! no longer have to come together in solution, just have to bend — LOOPING! remember that this experiment doesn’t necessarily prove or disprove anything — the hopscotch theory could very well be possible, and would align with the same results, as it’s easier to hop when the segments are closer. so really you could argue both, but the looping theory is perhaps more strongly supported.



side question — drawing lac operon interacts like the bicoid network. **image

CAP really refers to cAMP concentration (low glucose = high cAMP = CAP binds = activation). lacI really refers to lactose levels (low = lacI binds = inhibition).