Lecture 6: Genetic Algorithms - Principles, Components, and Applications

Course Learning Objectives

The Artificial Intelligence course (AI 106) at Horus University in Egypt (HUE), specifically within the Innovative Technologies curriculum, aims to equip students with several core competencies. Students are expected to learn how to formulate complex problems into AI state-space and agent-based models efficiently. The curriculum emphasizes the application of modern AI frameworks and software tools for the development of intelligent systems. Furthermore, students must master the design of original heuristics, fitness functions, and tactics to facilitate optimized problem-solving. A critical component of the course involves analyzing the performance of AI algorithms through computational modeling and various complexity metrics. Finally, students are tasked with implementing Machine Learning (ML), Fuzzy Logic, and Neural models to successfully navigate and solve engineering challenges.

Principal Heuristic Algorithms

Heuristic algorithms serve as essential tools for searching for global maxima or minima by combining randomness with specific rules. Genetic Algorithms were introduced by Holland in 1975, finding their inspiration in the fields of genetics and natural selection. Simulated Annealing was developed by Kirkpatrick in 1983, drawing its operational principles from statistical mechanics. Particle Swarm Optimization was introduced by Eberhart and Kennedy in 1995, inspired by the social behavior observed in swarms of insects or flocks of birds. These techniques utilize heuristic rules to guide search processes that would otherwise be computationally prohibitive.

Introduction to Genetic Algorithms (GA)

Genetic Algorithm is defined as a search-based optimization technique deeply rooted in the principles of Genetics and Natural Selection. It is frequently employed to identify optimal or near-optimal solutions for difficult problems that might otherwise require a lifetime to solve through exhaustive search. GAs are widely utilized for solving optimization problems across research fields and within machine learning. Optimization is defined as the process of making something better; in any given process, there is a set of inputs and a set of outputs designed to achieve an improved state.

In GAs, researchers maintain a pool or population of possible solutions to a specific problem. These solutions undergo processes of recombination and mutation, modeled after natural genetics, to produce new children. This iterative process repeats over various generations. Every individual, or candidate solution, receives an assigned fitness value based on its objective function value. Individuals with higher fitness values are granted a higher probability of mating and yielding more "fitter" individuals. This evolutionary process continues, evolving better solutions over successive generations until a specific stopping criterion is reached.

Evolutionary Strategies and Population Dynamics

The propagation of individuals to the next generation depends on their fitness. For instance, in a simple two-member evolutionary strategy, a parent (P) and a child (C) are compared. If the child member is more fit than the parent, it becomes the parent for the next generation. Consider a sequence where in Generation 1, the Parent has fitness F=5.2F = 5.2 and the Child has F=9.2F = 9.2. Because the child is fitter, it becomes the parent in Generation 2. If the Child in Generation 2 has F=8.2F = 8.2 and the Parent has F=9.2F = 9.2, the Parent remains the parent in Generation 3. If in Generation 3, the parent with F=9.2F = 9.2 produces a child with F=11.1F = 11.1, the child takes over as the parent in Generation 4, and so on.

An individual represents a single solution, while the population is the complete set of individuals involved in the search at any given time. A chromosome is the specific representation of a solution for the given problem. A chromosome is further subdivided into genes, where a gene represents a GA's representation of a single control factor. The specific value that a gene takes for a particular chromosome is known as an Allele.

Computational and Real-World Representations

Genes are the basic instructions for building Genetic Algorithms, and a chromosome is a sequence of these genes. There is a distinction between the Genotype and the Phenotype. The Genotype represents the population in the computation space, where solutions are formatted in a way that is easily understood and manipulated by a computing system (e.g., bit strings or sequences of numbers). The Phenotype represents the population in the actual real-world solution space, showing solutions as they are represented in real-world situations.

In the 8 Queens Problem, the representation follows an mapping between genotype and phenotype. The genotype might be a permutation of the numbers 1 through 8 (e.g., 1,3,5,2,6,4,7,81, 3, 5, 2, 6, 4, 7, 8), while the phenotype is the actual configuration of queens on an 8×88 \times 8 chessboard. This mapping allows the computational algorithm to manipulate simple numerical strings while searching for a valid physical board state.

Core Components and Population Data Structures

The Fitness Function is a critical objective measure that takes a solution as input and produces the suitability of that solution as output. While the fitness function and the objective function are often the same, they can differ depending on the specific problem. Genetic Operators, such as crossover, mutation, and selection, are used to change the genetic composition of offspring.

A population consists of a specific number of individuals being tested, defined by phenotype parameters and search space information. The main data structures in GA include chromosomes, phenotypes, objective function values, and fitness values. Two vital aspects of managing a population are the initial population generation (which can be random or seeded) and the population size. Examples of population data structures include collections of chromosomes such as Chromosome 1 (111100010111100010), Chromosome 2 (0111101101111011), Chromosome 3 (1010101010101010), and Chromosome 4 (1100110011001100).

Basic Structure and Components of GA

The basic structure of a GA begins with an initial population. Parents are selected for mating, and crossover and mutation operators are applied to these parents to generate new offspring. These children are then inserted into the population, and the process repeats. The overall components required for a GA include a problem definition as input, encoding principles (defining genes and chromosomes), an initialization procedure, parent selection for reproduction, recombination (crossover), mutation, evaluation via the fitness function, and a termination condition.

Encoding is the process of representing individual genes. Binary Encoding is the most common method, representing chromosomes as bit strings where integers are encoded exactly (e.g., 110100011010110100011010). Integer Encoding is used for discrete valued genes where binary is insufficient, such as encoding directions (North, South, East, West) as {0,1,2,3}\{0, 1, 2, 3\}. Real value Encoding is used for problems defining genes with continuous variables (e.g., 0.2,0.4,0.6,0.40.2, 0.4, 0.6, 0.4). Permutation Encoding represents chromosomes as a string of numbers in sequence, which is essential for ordering problems like the Traveling Salesman Problem (TSP).

Selection and Recombination Operators

Selection is the process of choosing parents for crossing. Roulette Wheel Selection involves a circular wheel divided into nn segments, where nn is the number of individuals. Each individual receives a portion of the circle proportional to its fitness value. A fixed point is chosen, the wheel is rotated, and the region in front of the fixed point determines the selected parent. This is repeated for the second parent.

Crossover, or recombination, is the process of producing a child from parent solutions. It involves selecting a random pair of individuals, choosing a random cross site along the string length, and swapping the position values between the two strings following that site. Single Point Crossover cuts chromosomes at one point and exchanges sections. Two-Point Crossover selects two points and exchanges the contents between them. Uniform Crossover uses a randomly generated binary crossover mask of the same length as the chromosomes; each gene in the offspring is copied from one or the other parent based on the mask's value (0 or 1). The Crossover Probability (PcP_c) describes how often crossover is performed; if no crossover occurs, offspring are exact copies of parents.

Mutation and Replacement Policies

Following crossover, strings are subjected to mutation based on a Mutation Probability (PmP_m). For binary representations, Bit Flipping converts 0 to 1 and vice versa based on a mutation chromosome. Interchanging (Swap) involves choosing two random positions and interchanging their bits. Reversing involves choosing a random position and reversing the bits next to it. If mutation is not performed, offspring remain unchanged after crossover. If performed, one or more parts of the chromosome are altered.

Replacement, or survivor selection, determines which individuals are retained for the next generation. Random Replacement involves the children replacing two randomly chosen individuals. Weak Parent Replacement replaces a weaker parent with a stronger child. Both Parents replacement implies the children simply replace their parents.

Termination and Application Areas

The GA search terminates when convergence criteria are met. Common stopping conditions include reaching a maximum number of generations, the elapsing of a specified time period, or the observation of no change in the population's best fitness over a set number of generations. It is noted that the maximum number of generations typically takes priority over other termination conditions.

Genetic Algorithms are applied in numerous areas, including optimization, Neural Networks, economics, scheduling, image processing, Machine Learning, DNA analysis, robot trajectory generation, vehicle routing, the parametric design of aircraft, multimodal optimization, and solving the Traveling Salesman Problem along with its various applications.

Detailed Numerical Example: Maximizing f(x)=x2f(x) = x^2

Consider the problem of maximizing the function f(x)=x2f(x) = x^2 where xx varies between 0 and 31. This requires a 5-bit binary string representation.

In the Initial Population phase, four strings are randomly selected: String 1 (0110001100, x=12x=12, f(x)=144f(x)=144), String 2 (1100111001, x=25x=25, f(x)=625f(x)=625), String 3 (0010100101, x=5x=5, f(x)=25f(x)=25), and String 4 (1001110011, x=19x=19, f(x)=361f(x)=361). The sum of fitness values is 1155, with an average of 288.75 and a maximum of 625.

The probability for each string is calculated as: Probi=f(x)i∑i=1nf(x)iProb_i = \frac{f(x)_i}{\sum_{i=1}^n f(x)_i} Resulting in: String 1 (12.47%), String 2 (54.11%), String 3 (2.16%), and String 4 (31.26%).

The Expected Count is calculated as: Expected Count=f(x)iAvg f(x)\text{Expected Count} = \frac{f(x)_i}{\text{Avg } f(x)} Yielding: String 1 (0.4987), String 2 (2.1645), String 3 (0.0866), and String 4 (1.2502). The Actual Counts used for the mating pool are 1, 2, 0, and 1 respectively.

In the Crossover phase, the mating pool consists of String 1 (0110001100), String 2 (1100111001), String 2 again (1100111001), and String 4 (1001110011). Applying crossover at points 4 and 3 produces offspring: offspring 1 (0110101101, x=13x=13, f(x)=169f(x)=169), offspring 2 (1100011000, x=24x=24, f(x)=576f(x)=576), offspring 3 (1101111011, x=27x=27, f(x)=729f(x)=729), and offspring 4 (1000110001, x=17x=17, f(x)=289f(x)=289). The fitness sum increases to 1763, with an average of 440.75.

Finally, the Mutation phase uses mutation chromosomes to flip bits. For offspring 1 (0110101101) with mutation chromosome (1000010000), the result is (1110111101, x=29x=29, f(x)=841f(x)=841). Other offspring remain similar if the mutation chromosome is (0000000000), except for offspring 4 (1000110001) which, with mutation chromosome (0010100101), becomes (1010010100, x=20x=20, f(x)=400f(x)=400). After mutation, the population sum grows to 2546, with an average fitness of 636.5 and a maximum of 841.