Study Notes on Bioinformatics

Introduction to Bioinformatics

Definition of Bioinformatics

Genomics Education Programme
  • Bioinformatics is a relatively new and evolving discipline that combines skills and technologies from computer science and biology to help us better understand and interpret biological data.

Wikipedia Definition
  • Bioinformatics is an interdisciplinary field of science that develops methods and software tools for understanding biological data, especially when the data sets are large and complex.

Key Components of Bioinformatics

  • Design and usage of tools and databases

    • Essential for life scientists managing large biological and biomedical datasets.

  • Disciplines Involved

    • Computer Science

    • Mathematics

    • Statistics

    • Biology

  • Processes Involved

    • Acquisition

    • Storage

    • Distribution

    • Analysis

      • Visualization

      • Comparison

      • Prediction

Biological Data Types and Sources

Sequence Data
  • DNA

  • RNA

  • Protein

Sources of Biological Data
  • DNA

    • Obtained through Sanger Sequencing or Next Generation Sequencing experiments.

  • RNA

    • Sourced from RNAseq experiments.

  • Protein

    • Derived from Edman Degradation or Mass Spectrometry experiments.

Example of Sequence Data

  • Given a sample transcript:

MPYAL AUGCCGUAUGCUCUU
MPYA AUGCCGUAUGCU
MPY AUGCCGUAU
MP AUG
  • RNA sequences and their respective bases are crucial for further biological interpretations.

Model Organisms in Bioinformatics

Arabidopsis Thaliana
  • Commonly known as mouse-ear cress, it is a small plant from the mustard family (Brassicaceae).

  • It is a popular model organism in plant biology and genetics.

  • Genome Complexity: Small genome of approximately 135 Mbp.

  • Notable for being the first plant to have its genome sequenced, aiding studies in flower development and light sensing.

Biochemical Characteristics of A. thaliana

  • Cold Hardy Plant: At low temperatures, cellular membranes expand due to phospholipid biosynthesis.

  • Involvement of enzymes such as CTP:Phosphocholine Cytidylyltransferase.

Storing and Managing Sequence Data

Databases
  • A structured collection of data designed for efficient storage, retrieval, and manipulation.

  • Raw Data: Includes DNA, RNA, and protein sequences, meaning no additional metadata is included regarding these sequences.

Overview of NCBI Resources

National Center for Biotechnology Information (NCBI)
  • NCBI advances science and health by providing access to biomedical and genomic information. It includes relevant databases and tools:

    • All databases relating to chemicals, bioassays, genes, proteins, sequence analysis, and more.

  • Research Opportunities: Participation in NCBI’s research initiatives or submission of data or manuscripts.

Introduction to GenBank

  • GenBank is the NIH genetic sequence database which collects all publicly available DNA sequences.

  • Part of the International Nucleotide Sequence Database Collaboration (DDBJ, ENA, GenBank).

  • Regular releases occur every two months, containing detailed notes about changes and growth statistics.

Analyzing Sequence Data with GenBank

  • GenBank provides several functionalities:

    • Search sequences using identifiers or annotations with Entrez Nucleotide.

    • Sequence alignment via BLAST (Basic Local Alignment Search Tool).

    • Access to sites and tools including ASN.1 and flat file formats.

Example Submission and Access Information for GenBank Records

  • Example of Arabidopsis thaliana's CTP:phosphocholine cytidylyltransferase gene:

>AF165912.1 Arabidopsis thaliana CTP:phosphocholine cytidylyltransferase (CCT) gene
  • Submission includes details about length, molecule type, and organism information, often linked to scientific publications.

Aquiring Sequence Data for Analysis

  • Sequences can be analyzed through bioinformatics tools available in various settings, including software installed on personal systems or accessed via educational resources such as Moodle.

  • Learning resources available include videos on sequencing experiments and a guide on the BLAST suite tools.

Transcription & Translation in Bioinformatics

  • Transcription: The process of converting a DNA sequence into RNA. RNA serves as a temporary copy of the genetic message coded by DNA.

  • Translation: The conversion of mRNA into a polypeptide (protein), crucial for cellular functions.

  • Proteins consist of amino acids linked by peptide bonds.

  • The process involves detecting start and stop codons within sequences to ensure correct translation.

Tools and Capabilities for Biological Analysis

  • BLAST: The Basic Local Alignment Search Tool, utilized for discovering regions of similarity between biological sequences and calculating statistical significance.

  • Functions include processed comparisons among nucleotide or protein sequences.

Embedded Tools in Bioinformatics

  • Comprehensive tools for sequence conversion, comparison against other species, and genome analysis.

  • Application of tools includes assessing biological relevance, evolutionary connections, and functional annotations.

Educational Resources

  • Access to various tools and software through Moodle including guide videos on sequencing experiments for DNA, RNA, and protein sequences alongside amino acid codon tables.

  • Short videos and guides highlight methodologies and functionalities for users wanting to delve deeper into molecular biology studies using bioinformatics.

Summary of Potential Bioinformatics Applications

  • Applications range from comparative genomics to specific research projects, enhancing understanding of genetic functions and interactions. The growing field opens doors for future advancements in biotechnology and genetic research.