Comprehensive Guide to Human Genome Organization and Chromosome Chromosomes and Genetic Variation

Fundamental Components and Definition of the Genome

The organization of the genome involves a hierarchical structure starting from the basic unit of the cell, leading to the chromosome, which contains genes composed of DNA. In eukaryotic organisms, this DNA is organized specifically into chromosomes characterized by distinct features such as a centromere, telomeres, and two distinct arms known as the p arm and the q arm. This distinguishes eukaryotic DNA from prokaryotic DNA, which lacks these specific chromosomal structures. Each species possesses a characteristic number of chromosomes (2n2n) relative to their genome size. For instance, humans possess 4646 chromosomes, potatoes have 4848, rice has 2424, monkeys have 4242, and mice have 4040.

Genetic Units and Chromosomal Terminology

A gene is physically defined as a specific sequence of DNA that encodes for a functional product, which can be a protein or a non-translated RNA species such as tRNA, rRNA, or snRNA. Every gene is located at a specific physical coordinate on a chromosome referred to as its locus. Variations or mutations within the DNA sequence of a gene at a particular locus result in alleles. While many genes exhibit multiple alleles, this variation is not limited to coding regions; non-coding DNA can also possess alleles of specific sequences. When a chromosomal site possesses multiple alleles within a population, it is described as being polymorphic. Structurally, genes are composed of coding regions called exons and non-coding regions called introns. A polynucleotide chain forms the DNA, which then packages into the chromosome.

Anatomy and Morphological Classification of Chromosomes

Chromosomes are defined by their structural regions, including the centromere, which is not necessarily the physical center of the structure but rather a rounded, constrained, and narrow region that holds sister chromatids together. The arms of the chromosome consist of a complex network of proteins and DNA where genes are situated, divided into heterochromatin and euchromatin regions. Telomeres are found at the ends of the chromosomes, serving as protective caps for gene-rich regions. They consist of repetitive, non-coding DNA sequences, specifically categorized as microsatellites and minisatellites. Based on the centromere's position, chromosomes are categorized into four main types. Metacentric chromosomes have the centromere located exactly or nearly exactly in the center, making the p and q arms equal in length. Submetacentric chromosomes have the centromere slightly shifted toward one arm, resulting in arms of similar but unequal length. Acrocentric chromosomes feature a centromere located very close to the p arm, making it much shorter than the q arm. Telocentric chromosomes have the centromere located at the very end of the p arm, making the p arm barely visible or entirely absent, though the q arm remains distinguishable. Some references also include a subtype called subtelocentric.

Karyotyping Procedures and Visualization

Karyotyping is the systematic process of pairing and ordering all chromosomes of an organism to provide a genome-wide snapshot. It describes the chromosome count and physical appearance under a light microscope, focusing on length, centromere position, banding patterns, sex chromosome differences, and other physical traits. The preparation of a karyotype involves several standardized staining steps. First, lymphocytes are collected and treated with fitohemaglutinina (phytohemagglutinin) to stimulate cell division. The culture is incubated for 343-4 days. Then, a chemical such as Colchicina (colchicine) is added to arrest the cells in metaphase. The harvested lymphocytes are treated with a hypotonic solution to cause the cells to swell. Finally, these swollen cells are fixed, dropped onto glass slides, dried, and stained. A cariotipo (karyotype) is a real image of the chromosomes, whereas an idiograma (idiogram) is a diagrammatic representation showing the specific banding patterns.

Classification of the Human Karyotype

The 23 pairs of human chromosomes are classified into seven groups (A through G) based on size and centromere location. Group A includes chromosomes 11, 22, and 33, which are large and metacentric. Group B includes chromosomes 44 and 55, which are large and submetacentric. Group C includes chromosomes 66, 77, 88, 99, 1010, 1111, 1212, and the X chromosome; these are medium-sized and submetacentric. Group D contains chromosomes 1313, 1414, and 1515, which are medium and acrocentric. Group E consists of chromosomes 1616, 1717, and 1818, categorized as relatively short and submetacentric. Group F comprises chromosomes 1919 and 2020, which are short and may be metacentric or submetacentric. Finally, Group G includes chromosomes 2121, 2222, and the Y chromosome, which are short and acrocentric. Notably, telocentric chromosomes do not naturally occur in humans.

Chromosomal Homology and Structural Alterations

Chromosomes are classified as homologous if they are identical in terms of gene sequences and loci, typically representing one chromosome from the mother and one from the father. Heterologous or non-homologous chromosomes, such as the X and Y pair, differ in sequence and loci. Structural alterations can occur within these chromosomes. A deletion involves the removal of a chromosomal segment (e.g., segment DD is lost from a sequence ABCDEFGHABCDEFGH). A duplication occurs when a segment is repeated (e.g., BCBC is doubled). An inversion happens when a segment is reversed within the chromosome (e.g., BCDEBCDE becomes EDCBEDCB). A translocation involves moving a segment from one chromosome to a nonhomologous one. Lastly, an insertion involves the placement of a genetic area into a new location on a different chromosome, such as a segment of chromosome 44 being inserted into chromosome 2020.

Composition and Statistics of the Human Genome

The human genome is constituted by a sequence of approximately 3.2×1093.2 \times 10^{9} nucleotides, organized into 2222 pairs of somatic chromosomes and 22 sex chromosomes. It contains roughly 20,00020,000 coding genes. These genes are composed of exons and introns in a ratio of 1:241:24. Non-repetitive sequences contain protein-coding genes, while repetitive regions contain genes for rRNA, tRNA, and snRNA. Approximately 11.5%1-1.5\% of the genome is comprised of exons (180,000180,000 exons total, averaging 1010 per gene), which are the coding regions. Introns make up 24%24\% and are generally ten times larger than exons. On average, there are 500500 to 1,1001,100 genes located on each chromosome, with an individual gene spanning approximately 1050kb10-50\,kb. Repetitive DNA makes up about 98%98\% of the genome and includes transposons, intergenic regions, pseudogenes, tandem repeats, introns, UTRs (untranslated regions), and non-coding RNA genes.

Types of Repetitive DNA and Polymorphisms

Repetitive DNA is categorized by the length and location of cycles. Satellite DNA consists of tandem repeats located in telomeres and centromeres, playing a structural role in protein binding; it is typically found in heterochromatin with units of 1001000pb100-1000\,pb. Minisatellites, or Variable Number of Tandem Repeats (VNTRs), range from 1010010-100 nucleotides (or 2070bases20-70\,bases) and are found in subtelomeric and intergenic regions, often in euchromatin. Microsatellites, also known as Short Tandem Repeats (STRs) or STRPs, consist of very simple combinations of 262-6 base pairs (e.g., CACACACACACA). These are used for DNA fingerprinting because they vary rapidly between individuals. They are distributed throughout chromosomes, often in non-coding regions or introns. A polymorphism is strictly defined as a variation in the DNA sequence where the least common variant is present in at least 1%1\% of the population (1 in 100 people). This distinguishes it from rare variants. Polymorphisms can range from Single Nucleotide Polymorphisms (SNPs) to complex copy number alterations.

Analytical Techniques for Genetic Variation

Restriction Fragment Length Polymorphisms (RFLPs) use restriction endonucleases to recognize specific palindromic sequences. The presence or absence of a site produces varying fragment lengths, visualized via a Southern blot. This technique is used to detect mutations, such as one that might destroy a restriction site in a disease state like the MstI site. The RFLP process involves DNA extraction, digestion with restriction enzymes, separation by agarose gel electrophoresis, transfer to a nitrocellulose or nylon membrane, hybridization with a marked probe, and autoradiography. STRPs are amplified using PCR with primers flanking the repeat block, with results visualized on an agarose gel. VNTRs are used less frequently today as they cluster near chromosome ends. DNA Fingerprinting utilizes these various repeats (VNTRs, STRs) to create unique profiles for individuals, useful in paternity testing and forensics.

Transposons and Pseudogenes

Transposons represent the largest component of the human genome, accounting for 4550%45-50\%. They are mobile elements (jumping genes) that can replicate and insert copies of themselves elsewhere. Class 1 (Retrotransposons) utilize an RNA intermediate and enzymes like reverse transcriptase and integrase. Class 2 (DNA transposons) are rarer and use a "cut and paste" mechanism via the enzyme transposase without an RNA intermediate. Transposition can be conservative (moving without replication) or replicative (producing a new copy at a target site while the original remains). Transposons can cause mutations, chromosomal rearrangements, or gene silencing via methylation. Pseudogenes are "fossil genes" or ancestral sequences that have lost their coding capacity due to the accumulation of multiple mutations and remain silenced within the genome.