Skip to main content

Protein Sequence DatabasesPIR, SWISS-PROT and TREMBEL


Protein Sequence Databases
PIR, SWISS-PROT and TREMBEL

1. Introduction

Protein sequence databases are biological databases that store information about amino acid sequences of proteins, along with their functional, structural, and biochemical characteristics. Since proteins are the functional molecules of the cell, protein databases are essential for understanding gene expression, metabolism, enzymatic activity, signaling pathways, and evolution.
Protein sequence databases mainly contain data derived from translated nucleotide sequences and experimental protein studies.

2. Types of Protein Sequence Databases

Protein sequence databases are broadly classified into:

A. Primary Protein Databases

Contain original protein sequence data
Minimal or no manual annotation

B. Secondary Protein Databases
Derived from primary databases
Provide curated functional and structural information

C. Composite Protein Databases
Combine protein data from multiple sources
Reduce redundancy
3. Protein Information Resource (PIR)

Overview
Protein Information Resource (PIR) is one of the earliest protein sequence databases, developed to store and analyze protein sequences.

Maintained by

Georgetown University (USA)
In collaboration with NBRF (National Biomedical Research Foundation)


Data Content

Protein sequences
Functional information
Evolutionary relationships
Classification into protein families

Unique Features
Organized into protein superfamilies
Emphasis on evolutionary and functional classification
Non-redundant dataset

Advantages
High-quality annotations
Useful for comparative protein studies

Limitations
Smaller than newer databases
Less frequently updated compared to UniProt


4. SWISS-PROT Database

Overview
SWISS-PROT is a manually curated, high-quality protein sequence database known for its accuracy and reliability.

Maintained by
Swiss Institute of Bioinformatics (SIB)
European Bioinformatics Institute (EMBL-EBI)

Data Content

Amino acid sequences
Protein function
Enzyme activity
Post-translational modifications
Domain structure
Subcellular localization


Key Features

Manual curation by experts
Minimal redundancy
High annotation accuracy
Extensive cross-references


SWISS-PROT Entry Includes : 
Accession number
Protein name
Organism
Function
Sequence length
Amino acid sequence

Advantages
Highly reliable
Preferred for functional studies
Limitations
Slow growth due to manual annotation

5. TrEMBL (Translated EMBL)

Overview
TrEMBL is a computer-annotated protein database that contains protein sequences translated from nucleotide sequence databases.

Maintained by
EMBL-EBI
Swiss Institute of Bioinformatics

Data Source
Translations of coding sequences from:
EMBL
GenBank
DDBJ
Key Features
Automatically annotated
Large and rapidly growing database
Supplement to SWISS-PROT

Advantages
Covers newly discovered proteins
Fast data availability

Limitations

Annotation may contain errors
Less reliable than SWISS-PROT

6. UniProt Knowledgebase (UniProtKB)

SWISS-PROT and TrEMBL together form the UniProt Knowledgebase (UniProtKB).
Components
UniProtKB/Swiss-Prot – reviewed, manually curated
UniProtKB/TrEMBL – unreviewed, automatically annotated

Purpose
Provide comprehensive protein sequence and functional information
Serve as a central protein knowledge hub


7. Comparison of PIR, SWISS-PROT, and TrEMBL


8. Applications of Protein Sequence Databases

Protein function prediction
Identification of conserved domains
Comparative protein analysis
Phylogenetic studies
Drug target identification
Enzyme characterization

9. Importance of Protein Sequence Databases
Link genes to protein function
Support proteomics research
Assist in metabolic pathway analysis
Aid in molecular evolution studies
Help in crop improvement and biotechnology

10. Conclusion
Protein sequence databases such as PIR, SWISS-PROT, and TrEMBL play a vital role in modern bioinformatics. While SWISS-PROT provides high-quality, manually curated protein data, TrEMBL ensures rapid availability of newly sequenced proteins. PIR contributes valuable evolutionary and functional classifications. Together, these databases support comprehensive protein research and biological discovery.

Comments

Popular Posts

Genetically modified microbes - biodegradation, biopesticides, bioremediation, mineral leaching and biofertilizers.

 Genetically Modified Microbes (GMMs) covering biodegradation, biopesticides, bioremediation, mineral leaching and biofertilizers.  Genetically Modified Microbes (GMMs) Introduction Genetically Modified Microbes (GMMs) are microorganisms such as bacteria, fungi, yeast or algae whose genetic material has been altered using recombinant DNA technology to enhance or introduce desirable traits. These microbes are engineered to improve efficiency, specificity and speed of biological processes useful in agriculture, industry and environmental management. GMMs play a vital role in sustainable development by reducing dependence on chemical fertilizers, pesticides and polluting industrial processes. 1. Genetically Modified Microbes in Biodegradation Definition Biodegradation is the microbial breakdown of complex organic pollutants into simpler, non-toxic substances. Role of GMMs Natural microbes often degrade pollutants slowly. Genetic modification enhances: Enzyme activity Substrate sp...

Micropropagation for Large-Scale Production of Medicinal Plants, Tree Species and Ornamentals –

Micropropagation for Large-Scale Production of Medicinal Plants, Tree Species and Ornamentals –  1. Introduction Micropropagation is an in-vitro clonal propagation technique used for rapid multiplication of plants under aseptic and controlled laboratory conditions. It enables the production of a large number of genetically uniform, disease-free plants from a small amount of starting material (explant). This technique is especially important for medicinal plants, forest tree species and ornamental plants, where conventional propagation is slow, seasonal or inefficient. 2. Principle of Micropropagation Micropropagation is based on totipotency, the inherent ability of a single plant cell to regenerate into a complete plant when provided with: Suitable nutrient medium Proper plant growth regulators Controlled light, temperature and humidity Sterile conditions. 3. Stages of Micropropagation Micropropagation generally involves five stages : Stage I – Selection and Sterilization of Expla...

❃LC-MS (LIQUID CHROMATOGRAPHY – MASS SPECTROMETRY)

LC-MS (LIQUID CHROMATOGRAPHY – MASS SPECTROMETRY)  ┏━━━━━ •❃°•°❀°•°❃•━━━━•━━━┓ 1. INTRODUCTION LC-MS is a hyphenated analytical technique combining Liquid Chromatography (LC) and Mass Spectrometry (MS). It is used for separation, identification, and quantification of compounds in complex mixtures. LC separates analytes based on polarity, size, or charge, while MS detects molecules based on mass-to-charge ratio (m/z). Developed in the 1970s–1980s, LC-MS is now widely used in pharmaceutical, clinical, environmental, and food analysis. Importance : Detects trace levels of compounds (ng–pg range) Analyzes non-volatile, thermally labile compounds that cannot be analyzed by GC-MS Provides structural information through mass fragmentation Example: Detection of drugs in plasma, protein identification in proteomics, pesticide residue analysis in food. 2. COMPONENTS OF LC-MS The LC-MS system has three main parts: A. Liquid Chromatograph (LC) Function: Separates components of a mixture befor...

✩‧₊ Plaque Blotting Technique

Plaque Blotting Technique *ੈ✩‧₊˚༺☆༻*ੈ✩‧₊˚*ੈ✩‧₊˚༺☆༻*ੈ✩‧₊˚ Introduction Plaque blotting is a molecular biology screening technique used to identify specific DNA or RNA sequences present in bacteriophage plaques formed on a bacterial lawn. It is especially useful in the screening of recombinant phage libraries such as λ (lambda) phage genomic or cDNA libraries. This technique combines: Plaque assay (to isolate individual phage clones) Blotting technique (to transfer nucleic acids onto a membrane) Hybridization (to detect specific sequences using labeled probes) Principle of Plaque Blotting The principle of plaque blotting is based on nucleic acid hybridization. Each plaque represents a clone of phage particles containing identical DNA. DNA from phage particles in plaques is: Released Denatured into single strands Transferred onto a nitrocellulose or nylon membrane The membrane is incubated with a labeled DNA/RNA probe complementary to the target sequence. Hybridization between probe and t...

Biological Databases – Types of Data and DatabasesNucleotide Sequence Databases (EMBL, GenBank, DDBJ)

Biological Databases – Types of Data and Databases Nucleotide Sequence Databases (EMBL, GenBank, DDBJ) 1. Introduction Biological databases are systematic, computerized collections of biological information that allow efficient storage, retrieval, updating, and analysis of large volumes of biological data. With the advent of genome sequencing, molecular biology, and bioinformatics, biological databases have become essential tools in biological research. These databases support studies in genomics, proteomics, evolutionary biology, taxonomy, medicine, agriculture, and biotechnology. 2. Types of Data Stored in Biological Databases Biological databases store diverse types of biological information, including: 1. Sequence Data DNA sequences RNA sequences Protein sequences 2. Structural Data Three-dimensional structures of proteins Nucleic acid structures 3. Functional Data Gene functions Enzyme activity Regulatory elements 4. Genomic Annotation Data Gene location Exons, introns Promoters a...

𓆞 Western Blotting Notes

Western Blotting (Immunoblotting) ❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥  Introduction Western blotting, also known as immunoblotting, is a widely used analytical technique for the detection, identification, and quantification of specific proteins in a complex biological sample. The technique combines protein separation by gel electrophoresis with specific antigen–antibody interaction. The method was developed by Towbin et al. (1979) (Burnette 1981---its group work) and is called “Western” in analogy to Southern blotting (DNA) and Northern blotting (RNA). Principle The principle of Western blotting involves: Separation of proteins based on molecular weight using SDS-PAGE Transfer (blotting) of separated proteins onto a membrane Specific detection of the target protein using primary and secondary antibodies Visualization using enzymatic or fluorescent detection systems 👉 Antigen–antibody specificity is the core principle of Western blotting. Steps Involved in Western Blotting 1. Sa...

Fourth Semester M.Sc. Degree Examination, May 2020BotanyBO 241 BIOINFORMATICS(2013 Admission Onwards)

Reg. No.:....... Name:......... J-4881 Fourth Semester M.Sc. Degree Examination, May 2020 Botany BO 241 BIOINFORMATICS (2013 Admission Onwards) Max. Marks: 75 I. Answer the following questions. 1. What are Secondary biological databases? 2. What is a Locus? 3. State the importance of E-value in sequence alignment? 4. Write the expansion of PHYLIP. 5. Distinguish proteome and proteomics. 6. Describe optimal alignment. 7. Define clade in a phylogenetic tree. 8. What is PIR? 9. List out any two tool used for molecular docking. 10. Write the name of submission tool for NCBI. (10 x 1=10 Marks) II. Answer the following questions in not more than 50 words. 11. (a) Give a short note on GenBank format. OR (b) Write the difference between scaled and unscaled phylogenetic trees. 12. (a) What are the two classes of data of UniProt? OR (b) State the difference between Orthologous and Xenologous sequences 13. (a) Write a brief note on character based phylogenetic analysis. OR (b) What is the role of...

Secondary Databases (PROSITE, PRINTS, BLOCKS)

Secondary Databases (PROSITE, PRINTS, BLOCKS  Secondary Databases Introduction Biological databases are broadly classified into primary and secondary databases. Primary databases store raw experimental data (e.g., nucleotide or protein sequences), whereas secondary databases contain derived information obtained by analyzing primary sequence data. Secondary databases are mainly used to: Identify protein families Detect conserved motifs, patterns, and domains Predict protein function Study structure–function relationships Examples of secondary databases include PROSITE, PRINTS, BLOCKS, Pfam, etc. 1. PROSITE Database Definition PROSITE is a secondary database that documents protein domains, families, and functional sites in the form of patterns and profiles. Developed by Swiss Institute of Bioinformatics (SIB) Maintained along with UniProt Principle PROSITE is based on the idea that functionally important regions of proteins are conserved during evolution. These conserved regions can ...

❥ Southern Blotting Notes

Southern Blotting  ❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥  Introduction Southern blotting is a molecular biology technique used for the detection of specific DNA sequences in a complex mixture of DNA. It was developed by Edwin M. Southern in 1975. The method involves restriction digestion of DNA, separation by gel electrophoresis, transfer (blotting) onto a membrane, and hybridization with a labeled DNA probe. Principle of Southern Blotting The technique is based on the principle of complementary base pairing. A single-stranded labeled DNA probe hybridizes specifically with its complementary DNA sequence immobilized on a membrane. Detection of the label confirms the presence and size of the target DNA fragment. Steps Involved in Southern Blotting. 1. Isolation of DNA Genomic DNA is extracted from cells or tissues. DNA must be pure and intact to ensure accurate results. 2. Restriction Enzyme  Digestion DNA is digested using specific restriction endonucleases. Produces DNA f...