CompBio practical file _241023_145953
Index of Experiments
Experiment 1: NCBI GenBank and FASTA
Aim
Introduction to NCBI GenBank and FASTA.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
The National Center for Biotechnology Information (NCBI) is part of the US National Library of Medicine (NLM) and plays a crucial role in the field of bioinformatics. It provides access to various biological data, including genetic sequences, which are essential for research in genomics, proteomics, and molecular biology. The extensive databases hosted by NCBI, like GenBank, PubMed, and BLAST, serve as platforms for researchers and healthcare professionals to share, access, and analyze biological data effectively. The NCBI databases collectively support crucial tasks, including genetic research, medical studies, and biological discoveries that have profound implications on human health and disease.
Procedure
Open the NCBI website.
For Nucleotide:
Select "nucleotide" from the drop-down menu.
Enter desired gene name (e.g., AR).
Optional: Filter by taxon.
Click "GenBank" to view nucleotide sequence (formatted into header, feature, origin).
Click "FASTA" for FASTA format (starts with '>').
For Protein:
Select "Protein" from the menu.
Enter desired protein name.
Filter as needed.
Access complete protein information and sequences.
Click "GenPept" for the complete protein sequence, "FASTA" for FASTA format.
Results
NCBI, GenPept, GenBank, and FASTA formats for protein and nucleotides can be downloaded.
Experiment 2: dbSNP
Aim
Introduction to dbSNP.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
dbSNP (Single Nucleotide Polymorphism Database) serves as a comprehensive database of genetic variations across various species, focusing particularly on human SNPs, microsatellites, and small-scale insertions/deletions (indels). It provides crucial data concerning genetic variation, which is instrumental in understanding population genetics, evolutionary biology, and the genetic basis of diseases. Researchers leverage dbSNP for association studies, genotype-phenotype correlations, and to investigate the genetic factors contributing to complex traits and conditions.
Procedure
Open the NCBI website.
Select "SNP" from the drop-down menu.
Enter desired gene name (e.g., BRCA2) and search.
Click the "rs" number in search results for more details on the dbSNP page.
Results
The genetic variation within and across different species can be determined using dbSNP.
Experiment 3: OMIM and ClinVar
Aim
Introduction to OMIM and ClinVar.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
ClinVar is a public archive that catalogs the relationships between genetic variations and phenotypes, paving the way for understanding the implications of genetic variants in human health and disease. It acts as a vital tool for genetic counselors, clinicians, and researchers to assess the clinical significance of variants. OMIM (Online Mendelian Inheritance in Man) is a rich resource that compiles detailed information on human genes and genetic phenotypes, frequently updated to reflect the latest research findings and clinical data. OMIM is particularly useful for those investigating Mendelian disorders and aids in linking genotype to phenotype.
Procedure
Open NCBI website.
For ClinVar:
Select "ClinVar" from menu, enter desired gene name (e.g., ADCY5).
For OMIM:
Select "OMIM" from menu, enter desired gene name (e.g., ATF2).
Select "Gene summaries" for detailed information.
Results
ClinVar helps in understanding genotype-phenotype relationships while OMIM provides data on Mendelian disorders.
Experiment 4: Protein Data Bank (PDB)
Aim
Introduction to PDB.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
The Protein Data Bank (PDB) serves as a central repository for three-dimensional structural data of biological macromolecules, including proteins, nucleic acids, and complex assemblies. This data is critical for understanding the molecular mechanisms underpinning biological functions and provides insight into the interactions and structures of biomolecules. The information in PDB is obtained through experimental techniques such as X-ray crystallography, NMR spectroscopy, and cryo-electron microscopy, enabling researchers to visualize and study the structure-function relationships in proteins and other macromolecules.
Procedure
Open PDB website.
Select "3D structures" to analyze protein structure.
Results
Visualization of protein 3D structures and experimental information is possible.
Experiment 5: UniProt/SwissProt
Aim
Introduction to UniProt/SwissProt.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
The UniProt Knowledgebase (UniProtKB) is a comprehensive and curated resource for protein sequence and functional information. SwissProt, a component of UniProt, specifically focuses on providing high-quality, manually annotated protein entries. This database includes details about protein functions, interactions, structures, post-translational modifications, and other relevant data that support biological research and application in fields like drug discovery and molecular biology. Researchers often refer to UniProt/SwissProt for reliable information on protein data to explore molecular mechanisms and functions systematically.
Procedure
Open the UniProt website.
Select "UniProt KB" from the search menu and search for desired protein (e.g., Filaggrin).
Results
Information regarding protein expression, interactions, and structure can be obtained from SwissProt.
Experiment 6: BLAST
Aim
Introduction to BLAST.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
BLAST (Basic Local Alignment Search Tool) is a widely used algorithm for comparing nucleotide or protein sequences against a database to identify regions of similarity. This tool plays an essential role in bioinformatics, allowing researchers to find homologous sequences, interpret genetic relationships, and identify functions for unknown genes based on sequence similarity. BLAST can be tailored to different types of sequences and offers multiple alignment options, making it highly adaptable for various research needs.
Procedure
Go to the BLAST website.
Choose the type of BLAST to perform (nucleotide to nucleotide, protein to protein, etc.).
Enter sequence in FASTA format or upload a file.
Click "BLAST" to view results.
Results
Users can select optional alignments between sequences using BLAST.
Experiment 7: Dotmatcher
Aim
Introduction to Dotmatcher.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
Dotmatcher is a bioinformatics tool used to visualize similarities between two sequences by generating a graphical representation known as a dot plot. This method is effective for identifying conserved regions, repeated sequences, and structural elements in nucleic acid or protein sequences. Dotplots allow for a straightforward visual comparison and can help elucidate functional and evolutionary relationships between sequences, thus aiding in multiple areas of genomic research and sequence analysis.
Procedure
Open the Dotmatcher website.
Upload sequences in FASTA format.
Adjust window size/threshhold as needed.
Run Dotmatcher for the plot.
Results
Visual representation of sequence similarity is provided through the dot plot.
Experiment 8: EMBOSS Needle and Water
Aim
Introduction to EMBOSS Needle and Water.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
EMBOSS is a software package for bioinformatics, and within it, Needle and Water are used for sequence alignment. Needle implements the Needleman-Wunsch algorithm for global alignment, aligning all residues in the sequences from end to end, while Water employs the Smith-Waterman algorithm for local alignment, focusing on aligning the most similar subsequences. These tools are essential for researchers conducting comparisons of DNA, RNA, and protein sequences to understand structural and functional attributes more thoroughly.
Procedure
Open the EMBOSS website.
Select DNA or protein according to needs, enter sequences.
Results
Optimal alignments between sequences can be determined using both methods.
Experiment 9: Clustal Omega
Aim
Introduction to Clustal Omega.
Requirements
Laptop/PC
Web browser
Stable internet connection
Theory
Clustal Omega is a powerful bioinformatics tool used for multiple sequence alignment, often utilized to reveal evolutionary relationships among sequences. By aligning three or more sequences simultaneously, it identifies conserved regions and helps understand the genetic relationships among different species. This insight is crucial for phylogenetic studies, functional genomics, and comparative genomics, leading to better predictions about molecular characteristics and biological functions.
Procedure
Open Clustal Omega website.
Select the type of sequences (Protein/DNA/RNA), input FASTA sequences.
Results
Identification of conserved regions among multiple sequences is possible, indicated by '*.'