Skip to main content

Biological Databases – Types of Data and DatabasesNucleotide Sequence Databases (EMBL, GenBank, DDBJ)


Biological Databases – Types of Data and Databases
Nucleotide Sequence Databases (EMBL, GenBank, DDBJ)

1. Introduction

Biological databases are systematic, computerized collections of biological information that allow efficient storage, retrieval, updating, and analysis of large volumes of biological data. With the advent of genome sequencing, molecular biology, and bioinformatics, biological databases have become essential tools in biological research.
These databases support studies in genomics, proteomics, evolutionary biology, taxonomy, medicine, agriculture, and biotechnology.

2. Types of Data Stored in Biological Databases
Biological databases store diverse types of biological information, including:

1. Sequence Data
DNA sequences
RNA sequences
Protein sequences

2. Structural Data

Three-dimensional structures of proteins
Nucleic acid structures

3. Functional Data

Gene functions
Enzyme activity
Regulatory elements

4. Genomic Annotation Data

Gene location
Exons, introns
Promoters and regulatory regions

5. Expression Data

Transcriptome data
Gene expression profiles

3. Classification of Biological Databases
Based on content and level of data processing, biological databases are classified into:

A. Primary Databases

Contain raw experimental data
Direct submissions from researchers
Minimal annotation
Examples:
GenBank, EMBL, DDBJ, Protein Data Bank (PDB)

B. Secondary Databases

Data derived from primary databases
Highly curated and analyzed
Provide functional and structural annotations
Examples:
UniProt, PROSITE, Pfam, SCOP


C. Composite (Integrated) Databases


Combine information from multiple databases
Reduce redundancy
Provide non-overlapping datasets
Examples:
RefSeq, UniGene, Ensembl

4. Nucleotide Sequence Databases

Nucleotide sequence databases store DNA and RNA sequences obtained through sequencing experiments. They are essential for gene discovery, genome analysis, comparative genomics, and evolutionary studies.

The three major global nucleotide sequence databases are:
GenBank (USA)
EMBL-ENA (Europe)
DDBJ (Japan)

These databases function under the International Nucleotide Sequence Database Collaboration (INSDC).


5. International Nucleotide Sequence Database Collaboration (INSDC)


INSDC is a global consortium that ensures:
Free and open access to nucleotide sequence data
Daily exchange of data among databases
Uniform data formats and annotation standards.


Members of INSDC:
GenBank – NCBI (USA)
EMBL-ENA – EMBL-EBI (Europe)
DDBJ – National Institute of Genetics (Japan)

6. GenBank
Overview
GenBank is a comprehensive nucleotide sequence database maintained by the National Center for Biotechnology Information (NCBI), USA. It is one of the largest and most widely used biological databases.
Types of Data Stored

Genomic DNA
cDNA and mRNA sequences
ESTs (Expressed Sequence Tags)
Whole genome sequences
Organelle genomes


7. EMBL (European Molecular Biology Laboratory Database)

Overview
The EMBL nucleotide database is maintained by the European Bioinformatics Institute (EMBL-EBI) and is now part of the European Nucleotide Archive (ENA).



8. DDBJ (DNA Data Bank of Japan)

Overview
DDBJ is maintained by the National Institute of Genetics (NIG), Japan. It mainly accepts sequence submissions from Asian countries but is globally accessible.
Data Stored
DNA and RNA sequences
Whole genome sequences
Environmental and metagenomic data
Special Features
Uses data formats similar to GenBank and EMBL
Exchanges data daily with other INSDC members
Provides online submission tools
9. Comparison of GenBank, EMBL, and DDBJ




➡ All three contain identical data but differ in access portals and management.


10. Importance of Nucleotide Sequence Databases
Preserve genetic information
Support genome sequencing projects
Enable gene identification and annotation
Facilitate evolutionary and phylogenetic studies
Assist in medical, agricultural, and environmental research
11. Applications
Comparative genomics
Molecular taxonomy
Gene cloning and primer design
Mutation analysis
Crop improvement and breeding programmes
12. Conclusion
Biological databases play a central role in modern biological research. Among them, nucleotide sequence databases such as GenBank, EMBL, and DDBJ are primary repositories that store DNA and RNA sequences. Through the INSDC collaboration, these databases ensure global data sharing, accuracy, and accessibility, making them indispensable resources for genomics, bioinformatics, and biotechnology.




Comments

Popular Posts

Electroporation – Detailed Notes

Electroporation – Detailed Notes Definition : Electroporation is a physical method of gene transfer in which cells are exposed to a brief, high-voltage electric pulse, creating temporary pores in the cell membrane. This allows DNA, RNA, proteins, or other molecules to enter the cytoplasm. It is widely used in bacteria, yeast, plant protoplasts, and mammalian cells. Key Concept: The electric field destabilizes the membrane, making it permeable to macromolecules. 1. Principle Cells are suspended in a conductive medium. A brief electrical pulse induces transient pores in the plasma membrane. DNA or other molecules present in the medium enter the cell through these pores. Membrane reseals after the pulse, and the molecule is retained inside the cell. Advantages of Principle: Direct and rapid. Works in many cell types. Does not require chemical carriers or viral vectors. 2. Materials Required Cells – bacterial, yeast, plant protoplasts, mammalian cells. DNA/RNA/other macromolecule – purifie...

❃GC-MS (GAS CHROMATOGRAPHY – MASS SPECTROMETRY) DETAILED NOTES

GC-MS (GAS CHROMATOGRAPHY – MASS SPECTROMETRY) DETAILED NOTES ┏━━━━━ •❃°•°❀°•°❃•━━━━•━━━ 1. INTRODUCTION GC-MS is a hyphenated analytical technique combining: Gas Chromatography (GC) : Separates volatile compounds in a mixture. Mass Spectrometry (MS) : Identifies and quantifies compounds based on mass-to-charge ratio (m/z). Developed in the 1950s–1970s, GC-MS is now widely used in forensic, pharmaceutical, environmental, and food analysis. Importance : Identifies unknown compounds Detects trace contaminants Provides both qualitative and quantitative information Example: Detection of pesticides in water, drug analysis in biological fluids, environmental pollutant detection. 2. PRINCIPLE GC-MS principle is based on two steps: A. Separation (GC) Sample is vaporized and carried by an inert gas (helium, nitrogen) through a capillary column. Compounds separate based on: Volatility Boiling point Interaction with column stationary phase Result: Different compounds elute at different retention ...

Micropropagation for Large-Scale Production of Medicinal Plants, Tree Species and Ornamentals –

Micropropagation for Large-Scale Production of Medicinal Plants, Tree Species and Ornamentals –  1. Introduction Micropropagation is an in-vitro clonal propagation technique used for rapid multiplication of plants under aseptic and controlled laboratory conditions. It enables the production of a large number of genetically uniform, disease-free plants from a small amount of starting material (explant). This technique is especially important for medicinal plants, forest tree species and ornamental plants, where conventional propagation is slow, seasonal or inefficient. 2. Principle of Micropropagation Micropropagation is based on totipotency, the inherent ability of a single plant cell to regenerate into a complete plant when provided with: Suitable nutrient medium Proper plant growth regulators Controlled light, temperature and humidity Sterile conditions. 3. Stages of Micropropagation Micropropagation generally involves five stages : Stage I – Selection and Sterilization of Expla...

DNA FOOTPRINTING

DNA FOOTPRINTING Introduction DNA footprinting is a molecular biology technique used to identify the specific site(s) on DNA where proteins (such as transcription factors) bind. It reveals the exact nucleotide sequences protected by bound proteins against cleavage by nucleases or chemical agents. It is widely used to study DNA-protein interactions, transcription regulation, and gene expression control. Definition DNA footprinting: A technique used to locate the binding site of DNA-binding proteins on DNA by detecting protected regions that are resistant to enzymatic or chemical cleavage. Principle DNA-binding proteins protect the DNA segment they occupy. DNA exposed to nucleases (DNase I) or chemical cleavage agents is cut at accessible regions. Regions bound by protein remain unaffected, leaving a “footprint”. When fragments are separated on a denaturing polyacrylamide gel, the missing bands correspond to protein-binding sites. Key idea: Cleavage occurs everywhere except where the pro...

RESTRICTION MAPPING

RESTRICTION MAPPING Introduction Restriction mapping is a molecular biology technique used to determine the relative positions of restriction enzyme recognition sites on a DNA molecule. It involves digestion of DNA with one or more restriction endonucleases followed by analysis of fragment sizes using agarose gel electrophoresis. Restriction mapping is essential for DNA characterization, cloning strategies, gene localization, and genome analysis. Definition Restriction mapping is the process of identifying the number, order, and distances between restriction enzyme cleavage sites within a DNA fragment by analyzing the pattern of fragments generated after enzymatic digestion. Principle Restriction enzymes cut DNA at specific palindromic nucleotide sequences. When DNA is digested with: Single restriction enzyme → produces fragments based on its recognition sites Multiple restriction enzymes → produces fragments whose sizes reveal the relative positions of sites By comparing fragment size...

Third Semester M.Sc. Degree Examination, February 2024 231: PLANT BREEDING, HORTICULTURE AND BIOSTATISTICS

Third Semester M.Sc. Degree Examination, February 2024                 Botany BO 231: PLANT BREEDING, HORTICULTURE AND BIOSTATISTICS (2019 Admission onwards) Time: 3 Hours I.Answer the following questions. 1.What is atomic gardening? 2.Name the cardamom research institute in Kerala. 3.Explain advantages of distant hybridisation. 4.Describe plant variety rights. 5.Write short notes on arboriculture. 6.What is vermicomposting? 7.Give short notes on cut flower industry. 8.What is ANOVA? 9.Describe the properties of binomial distribution. 10. Explain the use of LSD. Max. Marks: 75 (10 x 1 = 10 Marks) II.Answer the following questions in not more that 50 words. 11. (a) What do you mean by genetic modification techniques? OR (b) What is center of diversity of a species? 12. (a) Compare auto and allopolyploidy. OR (b) What are requirements of back cross breeding? 13. (a) Describe ideotype breeding and its significance. OR (b) What is the role of seed cer...

••CLASSIFICATION OF ALGAE - FRITSCH

      MODULE -1       PHYCOLOGY  CLASSIFICATION OF ALGAE - FRITSCH  ❖F.E. Fritsch (1935, 1945) in his book“The Structure and  Reproduction of the Algae”proposed a system of classification of  algae. He treated algae giving rank of division and divided it into 11  classes. His classification of algae is mainly based upon characters of  pigments, flagella and reserve food material.     Classification of Fritsch was based on the following criteria o Pigmentation. o Types of flagella  o Assimilatory products  o Thallus structure  o Method of reproduction          Fritsch divided algae into the following 11 classes  1. Chlorophyceae  2. Xanthophyceae  3. Chrysophyceae  4. Bacillariophyceae  5. Cryptophyceae  6. Dinophyceae  7. Chloromonadineae  8. Euglenineae    9. Phaeophyceae  10. Rhodophyceae  11. Myxophyce...

Fourth Semester M.Sc. Degree Examination, September 2019BotanySpecial Paper II - ElectiveBO 242 a: BIOTECHNOLOGY(2013 Admission onwards)

Reg. No.......  Name......... G-5263 Fourth Semester M.Sc. Degree Examination, September 2019 Botany Special Paper II - Elective BO 242 a: BIOTECHNOLOGY (2013 Admission onwards) Max. Marks: 75 1. Answer the following questions: 1. Humulin 2. YAC 3. Cybrids 4. Hybridomas 5. IPR 6. Gene therapy 7. C DNA library 8. AFLP 9. Hairy root culture 10. Somacional variation (10 x 1=10 Marks) II. Answer the following questions in not more than 50 words : 11. (a) What are immobilized enzymes? What is its advantage? OR (b) Write a short note on molecular farming. 12. (a) Give an account of bioprocess technology for the production of secondary metabolites. OR (b) What are bioreactors? How it operates? 13. (a) What are probiotics?. How do they work? OR (b) Discuss the methodology and application of western blotting. 14. (a) Briefly explain the application of protoplast culture OR (b) Write a short note on gene therapy 15. (a) What are reporter genes? Discuss its utility in transformation studies O...

Information retrieval from databases - search concepts, Tools for searching, homology searching, finding Domain and Functional site homologies

Information retrieval from databases - search concepts, Tools for searching, homology searching, finding Domain and Functional site homologies Information Retrieval from Databases 1. Introduction Information retrieval in bioinformatics refers to the process of extracting relevant biological data (DNA, RNA, protein sequences, structures, or functional information) from databases. Aim : Identify sequences, functions, or structural features for analysis, comparison, and annotation. Databases can be primary (raw sequence data) or secondary/derived (annotated, processed data). 2. Search Concepts in Biological Databases 2.1 Types of Searches Exact Match Search Returns results only if the query exactly matches database entries. Useful for known accession numbers or IDs. Pattern/Keyword Search Searches based on specific motifs, keywords, or annotations. Example: “kinase domain,” “signal peptide.” Similarity/Homology Search Detects sequences similar to the query based on sequence alignment. Use...

Protein Sequence DatabasesPIR, SWISS-PROT and TREMBEL

Protein Sequence Databases PIR, SWISS-PROT and TREMBEL 1. Introduction Protein sequence databases are biological databases that store information about amino acid sequences of proteins, along with their functional, structural, and biochemical characteristics. Since proteins are the functional molecules of the cell, protein databases are essential for understanding gene expression, metabolism, enzymatic activity, signaling pathways, and evolution. Protein sequence databases mainly contain data derived from translated nucleotide sequences and experimental protein studies. 2. Types of Protein Sequence Databases Protein sequence databases are broadly classified into: A. Primary Protein Databases Contain original protein sequence data Minimal or no manual annotation B. Secondary Protein Databases Derived from primary databases Provide curated functional and structural information C. Composite Protein Databases Combine protein data from multiple sources Reduce redundancy 3. Protein Informati...