Skip to main content

Biological Databases – Types of Data and DatabasesNucleotide Sequence Databases (EMBL, GenBank, DDBJ)


Biological Databases – Types of Data and Databases
Nucleotide Sequence Databases (EMBL, GenBank, DDBJ)

1. Introduction

Biological databases are systematic, computerized collections of biological information that allow efficient storage, retrieval, updating, and analysis of large volumes of biological data. With the advent of genome sequencing, molecular biology, and bioinformatics, biological databases have become essential tools in biological research.
These databases support studies in genomics, proteomics, evolutionary biology, taxonomy, medicine, agriculture, and biotechnology.

2. Types of Data Stored in Biological Databases
Biological databases store diverse types of biological information, including:

1. Sequence Data
DNA sequences
RNA sequences
Protein sequences

2. Structural Data

Three-dimensional structures of proteins
Nucleic acid structures

3. Functional Data

Gene functions
Enzyme activity
Regulatory elements

4. Genomic Annotation Data

Gene location
Exons, introns
Promoters and regulatory regions

5. Expression Data

Transcriptome data
Gene expression profiles

3. Classification of Biological Databases
Based on content and level of data processing, biological databases are classified into:

A. Primary Databases

Contain raw experimental data
Direct submissions from researchers
Minimal annotation
Examples:
GenBank, EMBL, DDBJ, Protein Data Bank (PDB)

B. Secondary Databases

Data derived from primary databases
Highly curated and analyzed
Provide functional and structural annotations
Examples:
UniProt, PROSITE, Pfam, SCOP


C. Composite (Integrated) Databases


Combine information from multiple databases
Reduce redundancy
Provide non-overlapping datasets
Examples:
RefSeq, UniGene, Ensembl

4. Nucleotide Sequence Databases

Nucleotide sequence databases store DNA and RNA sequences obtained through sequencing experiments. They are essential for gene discovery, genome analysis, comparative genomics, and evolutionary studies.

The three major global nucleotide sequence databases are:
GenBank (USA)
EMBL-ENA (Europe)
DDBJ (Japan)

These databases function under the International Nucleotide Sequence Database Collaboration (INSDC).


5. International Nucleotide Sequence Database Collaboration (INSDC)


INSDC is a global consortium that ensures:
Free and open access to nucleotide sequence data
Daily exchange of data among databases
Uniform data formats and annotation standards.


Members of INSDC:
GenBank – NCBI (USA)
EMBL-ENA – EMBL-EBI (Europe)
DDBJ – National Institute of Genetics (Japan)

6. GenBank
Overview
GenBank is a comprehensive nucleotide sequence database maintained by the National Center for Biotechnology Information (NCBI), USA. It is one of the largest and most widely used biological databases.
Types of Data Stored

Genomic DNA
cDNA and mRNA sequences
ESTs (Expressed Sequence Tags)
Whole genome sequences
Organelle genomes


7. EMBL (European Molecular Biology Laboratory Database)

Overview
The EMBL nucleotide database is maintained by the European Bioinformatics Institute (EMBL-EBI) and is now part of the European Nucleotide Archive (ENA).



8. DDBJ (DNA Data Bank of Japan)

Overview
DDBJ is maintained by the National Institute of Genetics (NIG), Japan. It mainly accepts sequence submissions from Asian countries but is globally accessible.
Data Stored
DNA and RNA sequences
Whole genome sequences
Environmental and metagenomic data
Special Features
Uses data formats similar to GenBank and EMBL
Exchanges data daily with other INSDC members
Provides online submission tools
9. Comparison of GenBank, EMBL, and DDBJ




➡ All three contain identical data but differ in access portals and management.


10. Importance of Nucleotide Sequence Databases
Preserve genetic information
Support genome sequencing projects
Enable gene identification and annotation
Facilitate evolutionary and phylogenetic studies
Assist in medical, agricultural, and environmental research
11. Applications
Comparative genomics
Molecular taxonomy
Gene cloning and primer design
Mutation analysis
Crop improvement and breeding programmes
12. Conclusion
Biological databases play a central role in modern biological research. Among them, nucleotide sequence databases such as GenBank, EMBL, and DDBJ are primary repositories that store DNA and RNA sequences. Through the INSDC collaboration, these databases ensure global data sharing, accuracy, and accessibility, making them indispensable resources for genomics, bioinformatics, and biotechnology.




Comments

Popular Posts

𓆞 Western Blotting Notes

Western Blotting (Immunoblotting) ❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥  Introduction Western blotting, also known as immunoblotting, is a widely used analytical technique for the detection, identification, and quantification of specific proteins in a complex biological sample. The technique combines protein separation by gel electrophoresis with specific antigen–antibody interaction. The method was developed by Towbin et al. (1979) (Burnette 1981---its group work) and is called “Western” in analogy to Southern blotting (DNA) and Northern blotting (RNA). Principle The principle of Western blotting involves: Separation of proteins based on molecular weight using SDS-PAGE Transfer (blotting) of separated proteins onto a membrane Specific detection of the target protein using primary and secondary antibodies Visualization using enzymatic or fluorescent detection systems 👉 Antigen–antibody specificity is the core principle of Western blotting. Steps Involved in Western Blotting 1. Sa...

✩‧₊ Plaque Blotting Technique

Plaque Blotting Technique *ੈ✩‧₊˚༺☆༻*ੈ✩‧₊˚*ੈ✩‧₊˚༺☆༻*ੈ✩‧₊˚ Introduction Plaque blotting is a molecular biology screening technique used to identify specific DNA or RNA sequences present in bacteriophage plaques formed on a bacterial lawn. It is especially useful in the screening of recombinant phage libraries such as λ (lambda) phage genomic or cDNA libraries. This technique combines: Plaque assay (to isolate individual phage clones) Blotting technique (to transfer nucleic acids onto a membrane) Hybridization (to detect specific sequences using labeled probes) Principle of Plaque Blotting The principle of plaque blotting is based on nucleic acid hybridization. Each plaque represents a clone of phage particles containing identical DNA. DNA from phage particles in plaques is: Released Denatured into single strands Transferred onto a nitrocellulose or nylon membrane The membrane is incubated with a labeled DNA/RNA probe complementary to the target sequence. Hybridization between probe and t...

Genetically modified microbes - biodegradation, biopesticides, bioremediation, mineral leaching and biofertilizers.

 Genetically Modified Microbes (GMMs) covering biodegradation, biopesticides, bioremediation, mineral leaching and biofertilizers.  Genetically Modified Microbes (GMMs) Introduction Genetically Modified Microbes (GMMs) are microorganisms such as bacteria, fungi, yeast or algae whose genetic material has been altered using recombinant DNA technology to enhance or introduce desirable traits. These microbes are engineered to improve efficiency, specificity and speed of biological processes useful in agriculture, industry and environmental management. GMMs play a vital role in sustainable development by reducing dependence on chemical fertilizers, pesticides and polluting industrial processes. 1. Genetically Modified Microbes in Biodegradation Definition Biodegradation is the microbial breakdown of complex organic pollutants into simpler, non-toxic substances. Role of GMMs Natural microbes often degrade pollutants slowly. Genetic modification enhances: Enzyme activity Substrate sp...

Protein Sequence DatabasesPIR, SWISS-PROT and TREMBEL

Protein Sequence Databases PIR, SWISS-PROT and TREMBEL 1. Introduction Protein sequence databases are biological databases that store information about amino acid sequences of proteins, along with their functional, structural, and biochemical characteristics. Since proteins are the functional molecules of the cell, protein databases are essential for understanding gene expression, metabolism, enzymatic activity, signaling pathways, and evolution. Protein sequence databases mainly contain data derived from translated nucleotide sequences and experimental protein studies. 2. Types of Protein Sequence Databases Protein sequence databases are broadly classified into: A. Primary Protein Databases Contain original protein sequence data Minimal or no manual annotation B. Secondary Protein Databases Derived from primary databases Provide curated functional and structural information C. Composite Protein Databases Combine protein data from multiple sources Reduce redundancy 3. Protein Informati...

❃LC-MS (LIQUID CHROMATOGRAPHY – MASS SPECTROMETRY)

LC-MS (LIQUID CHROMATOGRAPHY – MASS SPECTROMETRY)  ┏━━━━━ •❃°•°❀°•°❃•━━━━•━━━┓ 1. INTRODUCTION LC-MS is a hyphenated analytical technique combining Liquid Chromatography (LC) and Mass Spectrometry (MS). It is used for separation, identification, and quantification of compounds in complex mixtures. LC separates analytes based on polarity, size, or charge, while MS detects molecules based on mass-to-charge ratio (m/z). Developed in the 1970s–1980s, LC-MS is now widely used in pharmaceutical, clinical, environmental, and food analysis. Importance : Detects trace levels of compounds (ng–pg range) Analyzes non-volatile, thermally labile compounds that cannot be analyzed by GC-MS Provides structural information through mass fragmentation Example: Detection of drugs in plasma, protein identification in proteomics, pesticide residue analysis in food. 2. COMPONENTS OF LC-MS The LC-MS system has three main parts: A. Liquid Chromatograph (LC) Function: Separates components of a mixture befor...

protoplast fusion

Protoplast Fusion – Detailed Notes 1. Definition Protoplast fusion is a technique in which two or more protoplasts (cells without cell walls) are fused to form a single hybrid cell. It is widely used in plant biotechnology for hybridization, gene transfer, and somatic hybrid production. Also called somatic hybridization or somatic cell fusion. 2. Principle Cell wall removal: Plant cells are treated with cell wall-degrading enzymes (cellulase, pectinase) to generate protoplasts. Fusion of protoplasts: The naked cells are induced to fuse physically or chemically. Hybrid cell formation: Nuclei from different protoplasts combine to form a heterokaryon. Regeneration: The hybrid cell regenerates a new cell wall and divides, eventually forming a somatic hybrid plant. Key Concept: Protoplast fusion bypasses sexual incompatibility barriers, allowing hybridization between distant species or genera. 3. Steps in Protoplast Fusion Step 1: Isolation of Protoplasts Plant tissues (leaves, callus, cell...

Micropropagation for Large-Scale Production of Medicinal Plants, Tree Species and Ornamentals –

Micropropagation for Large-Scale Production of Medicinal Plants, Tree Species and Ornamentals –  1. Introduction Micropropagation is an in-vitro clonal propagation technique used for rapid multiplication of plants under aseptic and controlled laboratory conditions. It enables the production of a large number of genetically uniform, disease-free plants from a small amount of starting material (explant). This technique is especially important for medicinal plants, forest tree species and ornamental plants, where conventional propagation is slow, seasonal or inefficient. 2. Principle of Micropropagation Micropropagation is based on totipotency, the inherent ability of a single plant cell to regenerate into a complete plant when provided with: Suitable nutrient medium Proper plant growth regulators Controlled light, temperature and humidity Sterile conditions. 3. Stages of Micropropagation Micropropagation generally involves five stages : Stage I – Selection and Sterilization of Expla...

Secondary Databases (PROSITE, PRINTS, BLOCKS)

Secondary Databases (PROSITE, PRINTS, BLOCKS  Secondary Databases Introduction Biological databases are broadly classified into primary and secondary databases. Primary databases store raw experimental data (e.g., nucleotide or protein sequences), whereas secondary databases contain derived information obtained by analyzing primary sequence data. Secondary databases are mainly used to: Identify protein families Detect conserved motifs, patterns, and domains Predict protein function Study structure–function relationships Examples of secondary databases include PROSITE, PRINTS, BLOCKS, Pfam, etc. 1. PROSITE Database Definition PROSITE is a secondary database that documents protein domains, families, and functional sites in the form of patterns and profiles. Developed by Swiss Institute of Bioinformatics (SIB) Maintained along with UniProt Principle PROSITE is based on the idea that functionally important regions of proteins are conserved during evolution. These conserved regions can ...

𓆉 INDEX PAGE -NOTETHEPOINT43

INDEX PAGE   MAIN    CONTENT 1.   HSST BOTANY SYLLABUS, DETAILED NOTES, MCQ 2.  SET GENERAL PAPER SYLLABUS, DETAILED NOTES, 50MCQ 3.  SET BOTANY SYLLABUS, DETAILED NOTES, MCQ 4. MSC BOTANY THIRD SEMESTER SYLLABUS, NOTES (KERALA UNIVERSITY ) 5. MSC BOTANY THIRD SEMESTER QUESTION PAPER (KERALA UNIVERSITY ) 6. MSC BOTANY FOURTH SEMESTER SYLLABUS &NOTES (KERALA UNIVERSITY ) 7. FOURTH SEMESTER MSC BOTANY PREVIOUS QUESTION PAPER  (KERALA UNIVERSITY )