DNA Replication and Repair

DNA Replication

Each strand of a DNA double helix contains a sequence of nucleotides that is exactly complementary to the nucleotide sequence of its partner strand. Each strand can therefore serve as a template, or mold, for the synthesis of a new complementary strand. In other words, if we designate the two DNA strands as S and Sʹ, strand S can serve as a template for making a new strand Sʹ, while strand Sʹ can serve as a template for making a new strand S.

Thus, the genetic information in DNA can be accurately copied by the beautifully simple process in which strand S separates from strand Sʹ, and each separated strand then serves as a template for the production of a new complementary partner strand that is identical to its former partner. The ability of each strand of a DNA molecule to act as a template for producing a complementary strand enables a cell to copy, or replicate, its genes before passing them on to its descendants. Although simple in principle, the process is awe-inspiring, as it can involve the copying of billions of nucleotide pairs with incredible speed and accuracy: a human cell undergoing division will copy the equivalent of 1000 books like this one in about 8 hours and, on average, get no more than a few letters wrong. This impressive feat is performed by a cluster of proteins that together form a replication machine.

DNA replication produces two complete double helices from the original DNA molecule, with each new DNA helix being identical in nucleotide sequence (except for rare copying errors) to the original DNA double helix. Because each parental strand serves as the template for one new strand, each of the daughter DNA double helices ends up with one of the original (old) strands plus one strand that is completely new; this style of replication is said to be semiconservative.

The DNA double helix is normally very stable: the two DNA strands are locked together firmly by the large numbers of hydrogen bonds between the bases on both strands. As a result, only temperatures approaching those of boiling water provide enough thermal energy to seperate the two strands. To be used as a template, however, the double helix must first be opened up and the two strands separated to expose the nucleotide bases. The process of DNA synthesis is begun by intiator proteins that bind to specific DNA sequences called replication origins. Here the initiator proteins pry the two DNA strands apart, breaking the hydrogen bonds between the bases.

Although the hydrogen bonds collectively make the DNA helix very stable, individually each hydrogen bond is weak. Separating a short length of DNA a few base pairs at a time therefore does not require a large energy input, and the initiator proteins can readily unzip short regions of the double helix at normal temperatures. In simple cells such as bacteria or yeast, replication origins span approximately 100 nucleotide pairs. They are composed of DNA sequences that attract the initiator proteins and are especially easy to open. A-T base pairs are held together by fewer hydrogen bonds than G-C base pairs. Therefore DNA rich in A-T base pairs is easier to pull apart, and A-T rich stretches of DNA are typically found at replication origins.

A bacterial genome which is typically contained in a circular DNA molecule of several million nucleotide pairs has a single replication origin. The human genome which is very much larger has approximately 10,000 such origins an average of 220 origins per chromosome. Begining DNA replication at many places at once greatly shortens the time a cell needs to copy its entire genome. Once an initiator protein binds to DNA at a replication origin and locally opens up the double helix it attracts a group of proteins that carry out DNA replication. These proteins form a replication machine, in which each protein carries our a specific function.

Dna molecules in the process of being replicated contain Y-shaped junctions called replication forks. Two replication forks are formed at each replication origin.

At each form a replication machine moves along the DNA opening up the two strands of the double helix and using each strand as a template to make a new daughter strand. The two forks move away from the origin in opposite directions, unzipping the DNA double helix and copying the DNA a they go.

DNA replication in both bacterial and eukaryotic chromosome is therefore termed bidirectional. The forks move very rapidly at about 1000 nucleotides pairs per second in bacteria and 100 nucleotide pairs per second in humans. The slower rate of fork movement in humans, may be due to the difficulties in replicating DNA through the more complex chromatic structure.of eukaryotic chromosomes.

The movement of a replication fork is driven by the action of the replication machine at the heart of which is an enzyme called DNA polymerase. This enzyme catalyses the addition of nucleotides to the 3’ end of a growing DNA strand, using one of the original, parental DNA strands as a template. Base pairing between an incoming nucleotide and the template strand determines which of the four nucleotides (A, G, T or C) will be selected. The final product is a new strand of DNA that is complementary in a nucleotide sequence to the template.

The polymerization reaction involves the formation of phosphodiester bond between the 3’ end of the growing DNA chain and the 5’-phosphate group of the incoming nucleotide, which enters the reaction as a deoxy-ribonucleoside triphosphate. The eenergy for polymerization is provided by the incoming deoxyribonucleoside triphosphate itself: hydrolysis of one of its high energy phosphate bonds fuels the reaction that links the nucleotide monomer to the chain, releasing pyrophosphate is further hydrolyzed to inorganic phosphate which makes the polymerization reaction effectively irreversible.

The 5’ to 3’ direction of the DNA polymerization reaction poses a problem at the replication fork. The sugar phosphate backbone of each strand of a DNA double helix has a unique chemical direction or polarity determined by the way each sugar residue is linked to the nexxt, and the two strands in the double helix are antiparallel; that is they run in opposite directions. As a consequence at each replication fork one new DNA strans is being made on a template that runs in one direction (3’ to 5’), whereas the other new strans is being made on a template that runs in the opposite direction (5’ to 3’). The replication fork is therefore asymmetrical.

All DNA polymerases add new subunits only to the 3’ end of a DNA strand. As a result a new DNA chain can be synthesized only in a 5’ to 3’ direction. This can easily account for the synthesis of one of the two strands of DNA at the replication fork, but what happens on the other? This conundrum is solved by the use of a backstitching maneuver. The DNA strans that appears to grow in the incorrect 3’ to 5’ direction is actually made discontinuosly, in successive separate small pieces-with the DNA polymerase moving backward with respect to the direction of replication fork movememnt so that each new DNA fragment can be polymerized in the 5’ to 3’ direction. The resulting okazaki fragments are the pair of biochemists who discovered them are later joined together to form a continuous new strans. The DNA strans that is made discontinuosly in this way is called the lagging strand, because the cumbersome backstitching mechanism imparts a slight delay to its synthesis; the other strans which is synthesized continuously is called the leading strand.

Although they differ in subtle details the replication forks of all cells, prokaryotic and eukaryotic have leading and lagging strands. This common feature arises from the fact that all DNA polymerases work only in the 5’ to 3’ direction a restriction that allows DNA polymerase to check its work.

DNA polymerase is so accurate that it makes only one error in every 10^9 nucleotide pairs it copies. This error rate is much lower than can be explained simply by the accuracy of complementary base pairs, other less stable base pairs for example, G-T and C-A can also be formed. Such incorrect base pairs are formed much less frequently than correct ones, but if allowed to remain they would result in accumulation of mutations. This disaster is avoided because DNA polymerase has two special qualities that greatly increase the accuracy of DNA replication. First the enzyme carefully monitors the base-pairing between each incoming nucleoside triphosphate and the template strand. Only when the match is correct does DNA polymerase undergo a small structural rearrangement that allows it to catalyze the nucleotide-addition reaction. Second, when DNA polymerase does make a rare mistake and adds the wrong nucleotide, it can correct the error through an activity called proofreading. Proofreading takes place at the same time as DNA synthesis. Before the enzyme adds the next nucleotide to a growing DNA strand, it checks whether the previously added nucleotide is correctly base-paired to the template strand. If so, the polymerase adds the next nucleotide; if not, the polymerase clips the mispaired nucleotide and tries again.

Polymerization and proofreading are tightly coordinated and the two reactions are carried out by different catalytic domains in the same polymerase molecule.

This proofreading mechanism is possible only for DNA polymerase that synthesize DNA exclusively in the 5’ to 3’ direction. If a DNA polymerase were able to synthesize in the 3’ to 5’ direction(circumventing the need for backstitching on the lagging strand), it would be unable to proofread. That’s because if this backward polymerase were to remove an incorrectly paired nucleotide from the 5’ end, it would create a chemical deadend- a strand that could no longer be elongated.

Thus for a DNA polymerase to function as a self-correcting enzyme that removes its own polymerization errors as it moves along the DNA , it must proceed only in the 5’ to 3’ direction. The cumbersome backstitching mechanism on the lagging strand can be seen asa necessary consequence of maintaining this crucial proofreading activity.

We have seen that the accuracy of DNA replication depen ds on the requirement of the DNA polymerase for a correctly base-paired 3’ end before it can add more nucleotides to a growing DNA strand. How then can the polymerase begin a completely new DNA strand? To ger the process started, a different enzyme is needed—one that can begin a new polynucleotide strand simply by joining two nucleotides together without the need for a base-paired end. This enzyme does not, however, synthesize DNA. It makes a short length of a closely related type of nuclei acid(RNA)-using the DNA strand as a template. This short length of RNA, about 10 nucleotides long, is base-paired to the template strand and provides a base-paired 3’ end as a starting point for DNA polymerase.

An RNA fragment thus serves as a primer for DNA synthesis and the enzyme that synthesizes the RNA primer is known as primase. Primase is an example of an RNA polymerase an enzyme that synthesizes RNA using DNA as a template. A strand of RNA is very similar chemically to a single strand of DNA except that it is made of ribonucleotide subunits, in which the sugar is ribose, not deoxyribose, RNA also differs from DNA in that it contains the base uracil instead of thymine. However because U can form a base pair with A the RNA primer is synthesized on the DNA strand by complementary base-pairing in exactly the same way as is DNA. For the leading strand, an RNA primer is needed only to start replication at a replication origin'; at that point the DNA polymerase simply takes over, extending this primer with DNA synthesized in the 5’ to 3’ direction. But on the lagging strand, where DNA aynthesis is discontinuous, new primers are continuously needed to keep polymerization going. The movement of the replication fork continually exposes unpaired bases on the lagging strand template, and new RNA primers must be laid down at intervals along the newly exposes, single stranded stretched. DNA polymerase then adds a deoxyribonucleotide to the 3’ end of each new primer to produce another Okazaki fragment, and it will continue to elongate this fragment until it runs into the previously synthesized RNA primer.

To produce a continuous new DNA strand from the many separate pieces of nucleic acid made on the lagging strand, three additional enzymes are needed. These act quickly to remove the RNA primer, replace it with DNA, and join the remaining DNA fragments together. A nuclease degrades the RNA primer, a DNA polymerase called a repair polymerase replaces the RNA primers with DNA using the end of the adjacent Okazaki fragment as its primer and the enzyme DNA ligase joins the 5’-phosphate end of one DNA fragment to the adjacent 3’-hydroxyl end of the next.

Because it was discovered first, the repair polymerase involved in this process is often called DNA polymerase I; the polymerase that carries out the bulk of DNA replication at the forks is known as DNA polymerase III. Unlike DNA polymerases I and III, primase does not proofread its work. As a result, primers frequently contain mistakes. But because p;rimers are made of RNA instead of DNA, they stand out as a suspect copy to be automatically removed and replaced by DNA. The repair polymerase that makes this DNA like the replicative polymerase proofreads as it synthesizes. In this way, the cells replication machinery is able to begin new DNA strands and at the same time ensure that all of the DNA is copied faithfully.

DNA replication requires the cooperation of a large number of proteins that act in concert to synthesize new DNA. These proteins form part of a remarkably complex replication machine. The first problem faced by the replication machine is accessing the nucleotides that lie ahead of the replication fork and are thus buried within the double helix. For DNA replication to occur the double helix must be continuously pried apart so that the incoming nucleoside triphosphates can form base pairs with each template strand. Two types of replication proteins- DNA helicases and single strand DNA binding proteins-cooperate to carry out this task. A helicase sits at the very front of the replication machine where it uses the energy of ATP hydrolysis to propel itself forward, prying apart the double helix as it speeds along the DNA.

Single strand DNA binding proteins then latch onto the single stranded DNA exposed by the helicase, preventing the strands from re-forming base pairs and keeping them in an elongated form so that they can serve as efficient templates. This localized unwinding of the DNA double helix itself presents a problem. As the helicase moves forward, prying open the double helix, the DNA ahead of the fork gets wound more tightly. This excess twisting in front of the replication fork creates tension in the DNA that—if allowed to build makes unwinding the double helix increasingly difficult and ultimately impedes the forward movement of the replication machinery.

Enzymes called DNA topoisomerases, relieve this tension. A DNA topoisomerase produces a transient, single strand nick in the DNA backbone, which temporarily releases the built up tension; the enzyme then reseals the nick before falling off the DNA.

Back at the replication fork an additional protein called a sliding clamp, keeps DNA polymerase firmly attached to the template while it is synthesizing new strands of DNA. Left on their own, most DNA polymerase molecules will synthesize only a short string of nucleotides before falling off the DNA remplate strand. The sliding clamp forms a ring around the newly formed DNA double helix and by tightly gripping the polymerase allows the enzyme to move along the template strand without falling off as it synthesizes new DNA.
Assembly of the clamp around DNA requires the activity of another replication protein, the clamp loader, which hydrolyzes ATP each time it locks a sliding clamp around a newly formed DNA double helix. This loading needs to occur only once per replication cycle on the leading strand; on the lagging strand, however, the clamp is removed and then reattached each time a new Okazaki fragment is made. In bacteria, this happens approximately once per second. Most of the proteins involved in DNA replication are held together in a large multienzyme complex that moves as a unit along the parental DNA double helix, enabling DNA to be synthesized on both strands in a coordinated manner. This complex can be likened to a miniature sewing machine composed of protein parts and powered by nucleoside triphosphate hydrolysis.

Because DNA replication proceeds only in the 5’ to 3’ direction, the lagging strands of the replication fork must be synthesized in the form of discontinuous DNA fragments, each of which is initiated from an RNA primer laid down by a primase. A serious problem arises, however, as the replication fork approaches the end of a chromosome: although the leading strand can be replicated all the way to the chromosome tip, the lagging strand cannot.When the final RNA primer on the lagging strands is removed there is no enzyme that can replace it with DNA.

Without a strategy to deal with this problem, the lagging strand would become shorter with each round of DNA replication and after repeated cell divisions, the chromsomes themselves would shrink eventually losing valuable genetic information. Bacteria avoid this end replication problem by having circular DNA molecules as chromsomes. Eukaryotes get around it by adding long, repetitive nucleotide sequences to the ends of every chromosome. These sequences, which are incorporated into structures called telomeres, attract an enzyme called telomerase to the chromosome ends. Telomerase carries its own RNA template, which it uses to add multiple copies of the same repetitive DNA sequence to the lagging-strand template. In many dividing cells, telomeres are continuosly replenished and the resulting extended templates can then be copied by conventional DNA replication, ensuring that no peripheral chromosomal sequences are lost.

In addition to allowing replication of chromsome ends, telomeres form structures that mark the true ends of a chromosome. These structures allow the cell to distinguish unambiguously between the natural ends of chromosomes and the double strand DNA breaks that sometimes occur accidentally in the middle of chromosomes. These breaks are dangerous and must be immediately repaired as we will see shortly. In addition to attracting telomerase, the repetitive DNA sequences found within telomeres attrach other telomere-binding proteins that not only physically protect chromosome ends, but help mantain telomere length. Cells that divide at a rapid rate throughout the life of the organism those that line the gut or generate blood cells in the bone marrow, for example- keep their telomerase fully active. many other cell types, however gradually turn down their telomerase activity. After many rounds of cell division, the telomeres in these descendent cells will shrink, until they essentially disappear. At this ppint, these cells will cease dividing. In theory such a mechanism could provide a safeguard against the uncontrolled proliferation of cells including abnormal cells that have accumulated mutations that could promote the development of cancer.

DNA Repair

Just like any other molecule in the cell, DNA is continually undergoing thermal collisions with other molecules, often resulting in major chemical changes in the DNA. For example, in the time it takes to read this sentence, a total of about a trillion (1012) purine bases (A and G) will be lost from DNA in the cells of your body by a spontaneous reaction called depurination.

Depurination does not break the DNA phosphodiester backbone but instead removes a purine base from a nucleotide, giving rise to lesions that resemble missing teeth. Another common reaction is the spontaneous loss of an amino group (deamination) from a cytosine in DNA to produce the base uracil.

The ultraviolet radiation in sunlight is also damaging to DNA; it promotes covalent linkage between two adjacent pyrimidine bases, forming, for example, the thymine dimer shown.

It is the failure to repair thymine dimers that spells trouble for individuals with the disease xeroderma pigmentosum. These are only a few of many chemical changes that can occur in our DNA. Others are caused by reactive chemicals produced as a normal part of cell metabolism. If left unrepaired, DNA damage leads either to the substitution of one nucleotide pair for another as a result of incorrect base pairing during replication or to deletion of one or more nucleotide pairs in the daughter DNA strand after DNA replication. Some types of DNA damage(thymine dimers) can stall the DNA replication machinery at the site of the damage.

In addition to this chemical damage, DNA can also be altered by replication itself. The replication machinery that copies the DNA can-albeit rearely-incorporate an incorrect nucleotide that it fails to correct via proofreading. For each form of DNA damage, cells possess a mechanism for repair.

The thousands of random chemical changes that occur every day in the DNA of a human cell-through thermal collisions or exposure to reactoive metabolic by products, DNA-damaging chemicals, or radiation are repaired by a variety of mechanisms, each catalyzed by a different set of enzymes. Nearly all these repair mechanisms depend on the double-helical structure of DNA, which provides two copies of the genetic information one in each strand of the double helix. Thus if the sequence in one strand is accidentaly damaged, information is not lost irretrievably, because a backup version of the altered strand remains in the complem,entary sequence of nucleotides in the other undamaged strand. Most DNA damage creates structures that are never encountered in an undamaged DNA strand; thus the good strand is easily distinguished from the bad.

The basic pathway for repairing damage to DNA involves three basic steps:

  1. The damaged DNA is recognized and removed by one of a variety of mechanisms. These involve nucleases, which cleave the covalent bonds that join the damaged nucleotides to the rest of the DNA strand, leaving a small gap on one strand of the DNA double helix.

  2. A repair DNA polymerase binds to the 3’ hydroxyl end of the cut DNA strand. The enzyme, then fills in the gap by making a complementary copy of the information present in the undamaged strand. Although they differ from the DNA polymerase that replicates DNA, repair DNA polymerases synthesize DNA strands in the same way. For example, they elongate chains in the 5’ to 3’ direction and have the same type of proofreading activity to ensure that the template strand is copied accurately. In many cells, the repair polymerase is the same enzyme that fills in the gaps left after the RNA primers are removed during the normal DNA replication process.

  3. When the repair DNA polymerase has filled in the gap, a break remains in the sugar phosphate backbone of the repaired strand. This nich in the helix is sealed by DNA ligase, the same enzyme that joins the Okazaki fragments during replication of the lagging DNA strand.

Although the high fidelity and proofreading abilities of the cells replication machinery generally prevent replication errors from occuring, rare mistakes do happen,. Fortunately the cell has a backup system called mismatch repair that is dedicated to correcting these errors. The replication machine makes approximately one mistake per 10^7 nucleotides synthesized; DNA mismatch repair corrects 99% of these replication errors, increasing the overall accuracy to one mistake in 10^9 nucleotides synthesized. This level of accuracy is much, much higher than that generally encountered in our day-to-day lives.

Whenever the replication machinery makes a copying mistake, it leaves behind a mispaired nucleotide. If left uncorrected, the mismatch will result in a permanent mutation in the next round of DNA replication.

In most cases, however a complex of mismatch repair proteins will detect the DNA mismatch remove a portion of the DNA strand containing the error, and then resynthesize the missing DNA. This repair mechanism restores the correct sequence.

To be effective, the mismatch repair system must be able to recognize which of the DNA strands contains the error. Removing a segment from the strand that contains the correct sequence would only compound the mistake. The way the mismatch system solves this problem is by recognizing and removing only the newly made DNA. In bacteria, newly synthesized DNA lacks a type of chemical modification (a methyl group added to certain adenines) that is present on the preexisting parent DNA. Newly synthesized DNA is unmethylated for a short time, during which the new and template strands can be easily distinguished. Other cells use different strategies for distinguishing their parent DNA from a newly replicated strand. In humans, mismatch repair plays an important role in preventing cancer. An inherited predisposition to certain cancers(especially some types of colon cancer is caused by mutations in genes that encode mismatch repair proteins. Human cells have two copies of these genes(one from each parent) and individuals who inherit one damaged mismatch repair gene are unaffected until the undamaged copy of the same gene is randomly mutated in a somatic cell. This mutant cell and all of its progenty are then deficient in mismatch repair; they therefore accumulate mutations more rapidly than do normal cells. Because cancers arise from cells that have accumulated multiple mutations, a cell deficient in mismatch repair has a greatly enhanced chance of becoming cancerous. Thus, inheriting a single damaged mismatch repair gene strongly predisposes an individual to cancer.

The repair mechanisms rely on the genetic redundancy built into every DNA double helix. If nucleotides on one strand are damaged, they can be repaired using the information present in the complementary strand. This feature makes the DNA double helix especially well-suited for stably carrying genetic information from one generation to the next. When both strands of the double helix are damaged at the same time, mishaps at the replication fork, radiation and various chemical assaults can all fracture DNA, creating a double-strand break. Such lesions are particularly dangerous, because they can lead to the fragmentation of chromosomes and the subsequent loss of genes. This type of damage is especially difficult to repair. Every chromosome contains unique information; if a chromosome experiences a double-strand break, and the broken pieces become separated, the cell has no spare copy it can use to reconstruct the information that is now missing. To handle this potentially disastrous type of DNA damage, cells have evolved two basic strategies. The first involves hurriedly sticking the broken ends back together, before the DNA fragments drift apart and get lost. This repair mechanism, called nonhomologous end joining, occurs in many cell types and is carried out by a specialized group of enzymes that ‘clean’ the broken ends and rejoin them by DNA ligation. This ‘quick and dirty’ mechanism rapidly seals the break, but it comes with a price: in ‘cleaning’ the break to make it ready for ligation, nucleotides are often lost at the site of repair. If this imperfect repair disrupts the activity of a gene, the cell could suffer serious consequences. Thus, nonhomologous end joining can be a risky strategy for fixing broken chromosomes. Fortunately, cells have an alternative error-free strategy for repairing double strand breaks called homologous recombination.

The challenge in a repairing a double strand break as mentioned previously is finding an intact template to guide the repair. However if a double strand break occurs in a double helix shortly after that stretch of DNA has been replicated, the undamaged copy can serve as a template to guide the repair of both broken strands of DNA. This information on the undamaged strands of the intact double helix can be used to repair the complementary strands in the broken DNA. Because the two DNA molecules are homologous they have identical or nearly identical nucleotide sequences outside the broken region-this mechanism is known as homologous recombination. It results in a flawless repair of the double strand break, with no loss of genetic information.

Homologous recombination most often occurs shortly after a cell’s DNA has been replicated before cell division, when the duplicated helices are still physically close to each other.

To initiate the repair, a recombination-specific nuclease chews back the 5’ ends of the two broken strands at the break.

Then, with the help of specialized enzymes (called recA in bacteria and Rad52 in eukaryotes), one of the broken 3’ ends ‘invades’ the unbroken homologous DNA duplex and searches for a complementary sequence through base-pairing.

Once an extensive, accurate match is made, the invading strand is elongated by a repair DNA polymerase, using the complementary undamaged strand as a template.

After the repair polymerase has passed the point where the break occured, the newly newly elongated strand rejoins its original partner, forming base pairs that hold the two strands of the broken double helix together.

Repair is then completed by additional DNA synthesis at the 3’ ends of both strands of the broken double helix.

Followed by DNA ligation.

The net result is two intact DNA helices, for which the genetic information from one was used as a template to repair the other. Homologous recombination can also be used to repair many other types of DNA damage, making it perhaps the most handy DNA repair mechanism available to the cell: all that is needed is an intact homologous chromosome to use as a partner—a situation that occurs transiently each time a chromosome is duplicated. The all purpose nature of homologous recombinational repair probably explains why this mechanism and the proteins that carry it out, have been conserved in virtually all cells on Earth. Homologous recombination is versatile, and it also has a crucial role in the exchange of genetic information that occurs during the formation of the gametes—sperm and eggs. This change, during the specialized form of cell division called meiosis, enhances the generation of genetic diversity within species during sexual reproduction.

On occasion, the cell’s DNA replication and repair processes fail and allow a mutation to arise. This permanent change in the DNA sequence can have profound consequences. If the change occurs in a particular position in the DNA sequence, it could alter the amino acid sequence of a protein in a way that reduces or eliminates that proteins ability to function. For example, mutation of a single nucleotide in the human hemoglobin gene can cause the disease sickle-cell anemia. The hemoglobin protein is used to transport oxygen in the blood. Mutations in the hemoglobin gene can produce a protein that is less soluble than normal hemoglobin and forms fibrous intracelullar precipitates, which produce the characteristic sickle shape of affected red blood cells.

Because these cells are more fragile and frequently tear as they travel through the bloodstream, patients with this potentially life-threatening disease have fewer red blood cells than usual—that is they are anemic. Moreover, the abnormal red blood cells that remain can aggregate and block small vessels, causing pain and organ failure. We know about sickle cell hemoglobin because individuals with the mutation survive; the mutation even provides a benefit an increased resistance to malaria.

The example of sickle-cell anemia, which is an inherited disease, illustrates the consequences of mutations arising in the reproductive germ-line cells. A mutation in a germ-line cell will be passed on to all the cells in the body of the multicellular organism that develop from it, including the gametes responsible for the production of the next generation.

The many other cells in a multicellular organism must also be protected against mutation-in this case, against mutations that arise during the life of the individual. Nucleotide changes that occur in somatic cells can give rise to variant cells, some of which grow and divide in an uncontrolled fashion at the expense of the other cells in the organism. In the extreme case, an unchecked cell proliferation known as cancer results. Cancers are responsible for about 30% of the deaths that occur in Europe and North America, and they are caused primarily by a gradual accumulation of random mutations in a somatic cell and its descendants.

Increasing the mutation frequency even two or threefold could cause a disastrous increase in the incidence of cancer by accelerating the rate at which such somatic cell variants arise.

Thus, the high fidelity with which DNA sequences are replicated and maintained is important both for germ line cells, which transmit the genes to the next generation, and for somatic cells, which normally function as carefully regulated members of the complex community of cells in a multicellular organism. We should therefore not be surprised to find that all cells possess a very sophisticated set of mechanisms to reduce the number of mutations that occur in their DNA devoting hundreds of genes to these repair processes.

Although the majority of mutations do neither harm nor good to an organism, those that have severely harmful consequences are usually eliminated through natural selection; individuals carrying the altered DNA may die or experience decreased fertility, in which case these changes will be gradually lost from the population. By contrast, favorable changes will tend to persist and spread. But even where no selection operates—at the many sites in the DNA where a change of nucleotide has no effect on the fitness of the organism—the genetic message has been faithfully preserved over tens of millions of years. Thus humans and chimpanzees, after about 5 million years of divergent evolution, still have DNA sequences that are at least 98% identical. Even humans and whales, after 10 or 20 times this amount of time, have chromosomes that are unmistakably similar in their DNA sequence.

Thus our genome—and those of our relatives— contains a message from the distant past. Thanks to the faithfulness of DNA replication and repair, 100 million years of evolution have scarcely changed its essential content.