CompBio practical file _241023_145953

Index of Experiments

Experiment 1: NCBI GenBank and FASTA

Aim

Introduction to NCBI GenBank and FASTA.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

The National Center for Biotechnology Information (NCBI) is part of the US National Library of Medicine (NLM) and plays a crucial role in the field of bioinformatics. It provides access to various biological data, including genetic sequences, which are essential for research in genomics, proteomics, and molecular biology. The extensive databases hosted by NCBI, like GenBank, PubMed, and BLAST, serve as platforms for researchers and healthcare professionals to share, access, and analyze biological data effectively. The NCBI databases collectively support crucial tasks, including genetic research, medical studies, and biological discoveries that have profound implications on human health and disease.

Procedure

  1. Open the NCBI website.

  2. For Nucleotide:

    • Select "nucleotide" from the drop-down menu.

    • Enter desired gene name (e.g., AR).

    • Optional: Filter by taxon.

    • Click "GenBank" to view nucleotide sequence (formatted into header, feature, origin).

    • Click "FASTA" for FASTA format (starts with '>').

  3. For Protein:

    • Select "Protein" from the menu.

    • Enter desired protein name.

    • Filter as needed.

    • Access complete protein information and sequences.

    • Click "GenPept" for the complete protein sequence, "FASTA" for FASTA format.

Results

NCBI, GenPept, GenBank, and FASTA formats for protein and nucleotides can be downloaded.

Experiment 2: dbSNP

Aim

Introduction to dbSNP.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

dbSNP (Single Nucleotide Polymorphism Database) serves as a comprehensive database of genetic variations across various species, focusing particularly on human SNPs, microsatellites, and small-scale insertions/deletions (indels). It provides crucial data concerning genetic variation, which is instrumental in understanding population genetics, evolutionary biology, and the genetic basis of diseases. Researchers leverage dbSNP for association studies, genotype-phenotype correlations, and to investigate the genetic factors contributing to complex traits and conditions.

Procedure

  1. Open the NCBI website.

  2. Select "SNP" from the drop-down menu.

  3. Enter desired gene name (e.g., BRCA2) and search.

  4. Click the "rs" number in search results for more details on the dbSNP page.

Results

The genetic variation within and across different species can be determined using dbSNP.

Experiment 3: OMIM and ClinVar

Aim

Introduction to OMIM and ClinVar.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

ClinVar is a public archive that catalogs the relationships between genetic variations and phenotypes, paving the way for understanding the implications of genetic variants in human health and disease. It acts as a vital tool for genetic counselors, clinicians, and researchers to assess the clinical significance of variants. OMIM (Online Mendelian Inheritance in Man) is a rich resource that compiles detailed information on human genes and genetic phenotypes, frequently updated to reflect the latest research findings and clinical data. OMIM is particularly useful for those investigating Mendelian disorders and aids in linking genotype to phenotype.

Procedure

  1. Open NCBI website.

  2. For ClinVar:

    • Select "ClinVar" from menu, enter desired gene name (e.g., ADCY5).

  3. For OMIM:

    • Select "OMIM" from menu, enter desired gene name (e.g., ATF2).

    • Select "Gene summaries" for detailed information.

Results

ClinVar helps in understanding genotype-phenotype relationships while OMIM provides data on Mendelian disorders.

Experiment 4: Protein Data Bank (PDB)

Aim

Introduction to PDB.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

The Protein Data Bank (PDB) serves as a central repository for three-dimensional structural data of biological macromolecules, including proteins, nucleic acids, and complex assemblies. This data is critical for understanding the molecular mechanisms underpinning biological functions and provides insight into the interactions and structures of biomolecules. The information in PDB is obtained through experimental techniques such as X-ray crystallography, NMR spectroscopy, and cryo-electron microscopy, enabling researchers to visualize and study the structure-function relationships in proteins and other macromolecules.

Procedure

  1. Open PDB website.

  2. Select "3D structures" to analyze protein structure.

Results

Visualization of protein 3D structures and experimental information is possible.

Experiment 5: UniProt/SwissProt

Aim

Introduction to UniProt/SwissProt.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

The UniProt Knowledgebase (UniProtKB) is a comprehensive and curated resource for protein sequence and functional information. SwissProt, a component of UniProt, specifically focuses on providing high-quality, manually annotated protein entries. This database includes details about protein functions, interactions, structures, post-translational modifications, and other relevant data that support biological research and application in fields like drug discovery and molecular biology. Researchers often refer to UniProt/SwissProt for reliable information on protein data to explore molecular mechanisms and functions systematically.

Procedure

  1. Open the UniProt website.

  2. Select "UniProt KB" from the search menu and search for desired protein (e.g., Filaggrin).

Results

Information regarding protein expression, interactions, and structure can be obtained from SwissProt.

Experiment 6: BLAST

Aim

Introduction to BLAST.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

BLAST (Basic Local Alignment Search Tool) is a widely used algorithm for comparing nucleotide or protein sequences against a database to identify regions of similarity. This tool plays an essential role in bioinformatics, allowing researchers to find homologous sequences, interpret genetic relationships, and identify functions for unknown genes based on sequence similarity. BLAST can be tailored to different types of sequences and offers multiple alignment options, making it highly adaptable for various research needs.

Procedure

  1. Go to the BLAST website.

  2. Choose the type of BLAST to perform (nucleotide to nucleotide, protein to protein, etc.).

  3. Enter sequence in FASTA format or upload a file.

  4. Click "BLAST" to view results.

Results

Users can select optional alignments between sequences using BLAST.

Experiment 7: Dotmatcher

Aim

Introduction to Dotmatcher.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

Dotmatcher is a bioinformatics tool used to visualize similarities between two sequences by generating a graphical representation known as a dot plot. This method is effective for identifying conserved regions, repeated sequences, and structural elements in nucleic acid or protein sequences. Dotplots allow for a straightforward visual comparison and can help elucidate functional and evolutionary relationships between sequences, thus aiding in multiple areas of genomic research and sequence analysis.

Procedure

  1. Open the Dotmatcher website.

  2. Upload sequences in FASTA format.

  3. Adjust window size/threshhold as needed.

  4. Run Dotmatcher for the plot.

Results

Visual representation of sequence similarity is provided through the dot plot.

Experiment 8: EMBOSS Needle and Water

Aim

Introduction to EMBOSS Needle and Water.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

EMBOSS is a software package for bioinformatics, and within it, Needle and Water are used for sequence alignment. Needle implements the Needleman-Wunsch algorithm for global alignment, aligning all residues in the sequences from end to end, while Water employs the Smith-Waterman algorithm for local alignment, focusing on aligning the most similar subsequences. These tools are essential for researchers conducting comparisons of DNA, RNA, and protein sequences to understand structural and functional attributes more thoroughly.

Procedure

  1. Open the EMBOSS website.

  2. Select DNA or protein according to needs, enter sequences.

Results

Optimal alignments between sequences can be determined using both methods.

Experiment 9: Clustal Omega

Aim

Introduction to Clustal Omega.

Requirements

  • Laptop/PC

  • Web browser

  • Stable internet connection

Theory

Clustal Omega is a powerful bioinformatics tool used for multiple sequence alignment, often utilized to reveal evolutionary relationships among sequences. By aligning three or more sequences simultaneously, it identifies conserved regions and helps understand the genetic relationships among different species. This insight is crucial for phylogenetic studies, functional genomics, and comparative genomics, leading to better predictions about molecular characteristics and biological functions.

Procedure

  1. Open Clustal Omega website.

  2. Select the type of sequences (Protein/DNA/RNA), input FASTA sequences.

Results

Identification of conserved regions among multiple sequences is possible, indicated by '*.'