Study Notes on Bioinformatics
Introduction to Bioinformatics
Definition of Bioinformatics
Genomics Education Programme
Bioinformatics is a relatively new and evolving discipline that combines skills and technologies from computer science and biology to help us better understand and interpret biological data.
Wikipedia Definition
Bioinformatics is an interdisciplinary field of science that develops methods and software tools for understanding biological data, especially when the data sets are large and complex.
Key Components of Bioinformatics
Design and usage of tools and databases
Essential for life scientists managing large biological and biomedical datasets.
Disciplines Involved
Computer Science
Mathematics
Statistics
Biology
Processes Involved
Acquisition
Storage
Distribution
Analysis
Visualization
Comparison
Prediction
Biological Data Types and Sources
Sequence Data
DNA
RNA
Protein
Sources of Biological Data
DNA
Obtained through Sanger Sequencing or Next Generation Sequencing experiments.
RNA
Sourced from RNAseq experiments.
Protein
Derived from Edman Degradation or Mass Spectrometry experiments.
Example of Sequence Data
Given a sample transcript:
MPYAL AUGCCGUAUGCUCUU
MPYA AUGCCGUAUGCU
MPY AUGCCGUAU
MP AUG
RNA sequences and their respective bases are crucial for further biological interpretations.
Model Organisms in Bioinformatics
Arabidopsis Thaliana
Commonly known as mouse-ear cress, it is a small plant from the mustard family (Brassicaceae).
It is a popular model organism in plant biology and genetics.
Genome Complexity: Small genome of approximately 135 Mbp.
Notable for being the first plant to have its genome sequenced, aiding studies in flower development and light sensing.
Biochemical Characteristics of A. thaliana
Cold Hardy Plant: At low temperatures, cellular membranes expand due to phospholipid biosynthesis.
Involvement of enzymes such as CTP:Phosphocholine Cytidylyltransferase.
Storing and Managing Sequence Data
Databases
A structured collection of data designed for efficient storage, retrieval, and manipulation.
Raw Data: Includes DNA, RNA, and protein sequences, meaning no additional metadata is included regarding these sequences.
Overview of NCBI Resources
National Center for Biotechnology Information (NCBI)
NCBI advances science and health by providing access to biomedical and genomic information. It includes relevant databases and tools:
All databases relating to chemicals, bioassays, genes, proteins, sequence analysis, and more.
Research Opportunities: Participation in NCBI’s research initiatives or submission of data or manuscripts.
Introduction to GenBank
GenBank is the NIH genetic sequence database which collects all publicly available DNA sequences.
Part of the International Nucleotide Sequence Database Collaboration (DDBJ, ENA, GenBank).
Regular releases occur every two months, containing detailed notes about changes and growth statistics.
Analyzing Sequence Data with GenBank
GenBank provides several functionalities:
Search sequences using identifiers or annotations with Entrez Nucleotide.
Sequence alignment via BLAST (Basic Local Alignment Search Tool).
Access to sites and tools including ASN.1 and flat file formats.
Example Submission and Access Information for GenBank Records
Example of Arabidopsis thaliana's CTP:phosphocholine cytidylyltransferase gene:
>AF165912.1 Arabidopsis thaliana CTP:phosphocholine cytidylyltransferase (CCT) gene
Submission includes details about length, molecule type, and organism information, often linked to scientific publications.
Aquiring Sequence Data for Analysis
Sequences can be analyzed through bioinformatics tools available in various settings, including software installed on personal systems or accessed via educational resources such as Moodle.
Learning resources available include videos on sequencing experiments and a guide on the BLAST suite tools.
Transcription & Translation in Bioinformatics
Transcription: The process of converting a DNA sequence into RNA. RNA serves as a temporary copy of the genetic message coded by DNA.
Translation: The conversion of mRNA into a polypeptide (protein), crucial for cellular functions.
Proteins consist of amino acids linked by peptide bonds.
The process involves detecting start and stop codons within sequences to ensure correct translation.
Tools and Capabilities for Biological Analysis
BLAST: The Basic Local Alignment Search Tool, utilized for discovering regions of similarity between biological sequences and calculating statistical significance.
Functions include processed comparisons among nucleotide or protein sequences.
Embedded Tools in Bioinformatics
Comprehensive tools for sequence conversion, comparison against other species, and genome analysis.
Application of tools includes assessing biological relevance, evolutionary connections, and functional annotations.
Educational Resources
Access to various tools and software through Moodle including guide videos on sequencing experiments for DNA, RNA, and protein sequences alongside amino acid codon tables.
Short videos and guides highlight methodologies and functionalities for users wanting to delve deeper into molecular biology studies using bioinformatics.
Summary of Potential Bioinformatics Applications
Applications range from comparative genomics to specific research projects, enhancing understanding of genetic functions and interactions. The growing field opens doors for future advancements in biotechnology and genetic research.