Comprehensive Guide to Human Genome Organization and Chromosome Chromosomes and Genetic Variation
Fundamental Components and Definition of the Genome
The organization of the genome involves a hierarchical structure starting from the basic unit of the cell, leading to the chromosome, which contains genes composed of DNA. In eukaryotic organisms, this DNA is organized specifically into chromosomes characterized by distinct features such as a centromere, telomeres, and two distinct arms known as the p arm and the q arm. This distinguishes eukaryotic DNA from prokaryotic DNA, which lacks these specific chromosomal structures. Each species possesses a characteristic number of chromosomes () relative to their genome size. For instance, humans possess chromosomes, potatoes have , rice has , monkeys have , and mice have .
Genetic Units and Chromosomal Terminology
A gene is physically defined as a specific sequence of DNA that encodes for a functional product, which can be a protein or a non-translated RNA species such as tRNA, rRNA, or snRNA. Every gene is located at a specific physical coordinate on a chromosome referred to as its locus. Variations or mutations within the DNA sequence of a gene at a particular locus result in alleles. While many genes exhibit multiple alleles, this variation is not limited to coding regions; non-coding DNA can also possess alleles of specific sequences. When a chromosomal site possesses multiple alleles within a population, it is described as being polymorphic. Structurally, genes are composed of coding regions called exons and non-coding regions called introns. A polynucleotide chain forms the DNA, which then packages into the chromosome.
Anatomy and Morphological Classification of Chromosomes
Chromosomes are defined by their structural regions, including the centromere, which is not necessarily the physical center of the structure but rather a rounded, constrained, and narrow region that holds sister chromatids together. The arms of the chromosome consist of a complex network of proteins and DNA where genes are situated, divided into heterochromatin and euchromatin regions. Telomeres are found at the ends of the chromosomes, serving as protective caps for gene-rich regions. They consist of repetitive, non-coding DNA sequences, specifically categorized as microsatellites and minisatellites. Based on the centromere's position, chromosomes are categorized into four main types. Metacentric chromosomes have the centromere located exactly or nearly exactly in the center, making the p and q arms equal in length. Submetacentric chromosomes have the centromere slightly shifted toward one arm, resulting in arms of similar but unequal length. Acrocentric chromosomes feature a centromere located very close to the p arm, making it much shorter than the q arm. Telocentric chromosomes have the centromere located at the very end of the p arm, making the p arm barely visible or entirely absent, though the q arm remains distinguishable. Some references also include a subtype called subtelocentric.
Karyotyping Procedures and Visualization
Karyotyping is the systematic process of pairing and ordering all chromosomes of an organism to provide a genome-wide snapshot. It describes the chromosome count and physical appearance under a light microscope, focusing on length, centromere position, banding patterns, sex chromosome differences, and other physical traits. The preparation of a karyotype involves several standardized staining steps. First, lymphocytes are collected and treated with fitohemaglutinina (phytohemagglutinin) to stimulate cell division. The culture is incubated for days. Then, a chemical such as Colchicina (colchicine) is added to arrest the cells in metaphase. The harvested lymphocytes are treated with a hypotonic solution to cause the cells to swell. Finally, these swollen cells are fixed, dropped onto glass slides, dried, and stained. A cariotipo (karyotype) is a real image of the chromosomes, whereas an idiograma (idiogram) is a diagrammatic representation showing the specific banding patterns.
Classification of the Human Karyotype
The 23 pairs of human chromosomes are classified into seven groups (A through G) based on size and centromere location. Group A includes chromosomes , , and , which are large and metacentric. Group B includes chromosomes and , which are large and submetacentric. Group C includes chromosomes , , , , , , , and the X chromosome; these are medium-sized and submetacentric. Group D contains chromosomes , , and , which are medium and acrocentric. Group E consists of chromosomes , , and , categorized as relatively short and submetacentric. Group F comprises chromosomes and , which are short and may be metacentric or submetacentric. Finally, Group G includes chromosomes , , and the Y chromosome, which are short and acrocentric. Notably, telocentric chromosomes do not naturally occur in humans.
Chromosomal Homology and Structural Alterations
Chromosomes are classified as homologous if they are identical in terms of gene sequences and loci, typically representing one chromosome from the mother and one from the father. Heterologous or non-homologous chromosomes, such as the X and Y pair, differ in sequence and loci. Structural alterations can occur within these chromosomes. A deletion involves the removal of a chromosomal segment (e.g., segment is lost from a sequence ). A duplication occurs when a segment is repeated (e.g., is doubled). An inversion happens when a segment is reversed within the chromosome (e.g., becomes ). A translocation involves moving a segment from one chromosome to a nonhomologous one. Lastly, an insertion involves the placement of a genetic area into a new location on a different chromosome, such as a segment of chromosome being inserted into chromosome .
Composition and Statistics of the Human Genome
The human genome is constituted by a sequence of approximately nucleotides, organized into pairs of somatic chromosomes and sex chromosomes. It contains roughly coding genes. These genes are composed of exons and introns in a ratio of . Non-repetitive sequences contain protein-coding genes, while repetitive regions contain genes for rRNA, tRNA, and snRNA. Approximately of the genome is comprised of exons ( exons total, averaging per gene), which are the coding regions. Introns make up and are generally ten times larger than exons. On average, there are to genes located on each chromosome, with an individual gene spanning approximately . Repetitive DNA makes up about of the genome and includes transposons, intergenic regions, pseudogenes, tandem repeats, introns, UTRs (untranslated regions), and non-coding RNA genes.
Types of Repetitive DNA and Polymorphisms
Repetitive DNA is categorized by the length and location of cycles. Satellite DNA consists of tandem repeats located in telomeres and centromeres, playing a structural role in protein binding; it is typically found in heterochromatin with units of . Minisatellites, or Variable Number of Tandem Repeats (VNTRs), range from nucleotides (or ) and are found in subtelomeric and intergenic regions, often in euchromatin. Microsatellites, also known as Short Tandem Repeats (STRs) or STRPs, consist of very simple combinations of base pairs (e.g., ). These are used for DNA fingerprinting because they vary rapidly between individuals. They are distributed throughout chromosomes, often in non-coding regions or introns. A polymorphism is strictly defined as a variation in the DNA sequence where the least common variant is present in at least of the population (1 in 100 people). This distinguishes it from rare variants. Polymorphisms can range from Single Nucleotide Polymorphisms (SNPs) to complex copy number alterations.
Analytical Techniques for Genetic Variation
Restriction Fragment Length Polymorphisms (RFLPs) use restriction endonucleases to recognize specific palindromic sequences. The presence or absence of a site produces varying fragment lengths, visualized via a Southern blot. This technique is used to detect mutations, such as one that might destroy a restriction site in a disease state like the MstI site. The RFLP process involves DNA extraction, digestion with restriction enzymes, separation by agarose gel electrophoresis, transfer to a nitrocellulose or nylon membrane, hybridization with a marked probe, and autoradiography. STRPs are amplified using PCR with primers flanking the repeat block, with results visualized on an agarose gel. VNTRs are used less frequently today as they cluster near chromosome ends. DNA Fingerprinting utilizes these various repeats (VNTRs, STRs) to create unique profiles for individuals, useful in paternity testing and forensics.
Transposons and Pseudogenes
Transposons represent the largest component of the human genome, accounting for . They are mobile elements (jumping genes) that can replicate and insert copies of themselves elsewhere. Class 1 (Retrotransposons) utilize an RNA intermediate and enzymes like reverse transcriptase and integrase. Class 2 (DNA transposons) are rarer and use a "cut and paste" mechanism via the enzyme transposase without an RNA intermediate. Transposition can be conservative (moving without replication) or replicative (producing a new copy at a target site while the original remains). Transposons can cause mutations, chromosomal rearrangements, or gene silencing via methylation. Pseudogenes are "fossil genes" or ancestral sequences that have lost their coding capacity due to the accumulation of multiple mutations and remain silenced within the genome.