Skip to main content

Information retrieval from databases - search concepts, Tools for searching, homology searching, finding Domain and Functional site homologies

Information retrieval from databases - search concepts, Tools for searching, homology searching, finding Domain and Functional site homologies

Information Retrieval from Databases

1. Introduction

Information retrieval in bioinformatics refers to the process of extracting relevant biological data (DNA, RNA, protein sequences, structures, or functional information) from databases.
Aim: Identify sequences, functions, or structural features for analysis, comparison, and annotation.
Databases can be primary (raw sequence data) or secondary/derived (annotated, processed data).
2. Search Concepts in Biological Databases

2.1 Types of Searches

Exact Match Search

Returns results only if the query exactly matches database entries.
Useful for known accession numbers or IDs.

Pattern/Keyword Search

Searches based on specific motifs, keywords, or annotations.
Example: “kinase domain,” “signal peptide.”

Similarity/Homology Search
Detects sequences similar to the query based on sequence alignment.
Uses scoring matrices to assess similarity (e.g., BLOSUM, PAM).
Useful for identifying homologous genes or proteins.


Complex Query Search

Combines Boolean operators (AND, OR, NOT) to refine results.
Example: “kinase AND human NOT viral.”


2.2 Search Parameters

Query sequence or keyword
Database selection (nucleotide, protein, structural, functional)
Algorithm choice (BLAST, FASTA, PSI-BLAST)
Threshold or cut-off (E-value, score, % identity)
Filters (organism, date, length, sequence type)


3. Tools for Searching Biological Databases


3.1 Nucleotide Sequence Databases


GenBank (NCBI)
EMBL (European Nucleotide Archive)
DDBJ (DNA Data Bank of Japan)
Search Tools:
BLASTN – nucleotide vs nucleotide
FASTA – nucleotide similarity search


3.2 Protein Sequence Databases

SWISS-PROT / UniProtKB – curated protein sequences
PIR / TrEMBL – unreviewed protein sequences
Search Tools:
BLASTP – protein vs protein
PSI-BLAST – iterative search for distant homologs
HMMER – profile-based search using hidden Markov models

3.3 Structural Databases

Protein Data Bank (PDB) – 3D protein structures
SCOP / CATH – structural classification of proteins
Search Tools:
BLAST 3D – structure-based sequence search
DALI – structural alignment


3.4 Specialized Databases
Pfam – protein families
PROSITE – protein motifs
InterPro – integrated database of protein domains


4. Homology Searching

Homology searching identifies evolutionarily related sequences based on similarity.


4.1 Concept

Homologous sequences: share a common ancestor.
Types:

Orthologs – homologs in different species
Paralogs – homologs in the same species
Homology suggests similar structure or function.


4.2 Methods

1. Pairwise Sequence Alignment

Tools: BLAST, FASTA
Measures similarity (% identity) and E-value


2. Multiple Sequence Alignment (MSA)
Tools: Clustal Omega, MUSCLE
Identifies conserved residues and motifs


3. Profile-based Searching

Uses Position-Specific Scoring Matrices (PSSM)
Tool: PSI-BLAST, HMMER

4. Structural Homology
Comparing 3D structures for similarity
Tools: DALI, CATH, SCOP


5. Finding Domain and Functional Site Homologies


5.1 Protein Domains

Definition: Conserved part of protein with specific function/structure.
Examples: kinase domain, zinc finger, SH2 domain.
Domains often determine protein function.


5.2 Domain Databases and Tools

Pfam – HMM-based domain identification
SMART – domains in signaling and extracellular proteins
InterPro – integrates multiple domain databases
PROSITE – motifs and functional sites

5.3 Functional Site Prediction

Active sites, binding sites, or motifs are predicted based on:
Conserved residues across homologs
3D structure information
Known motifs (PROSITE patterns)
Tools:
ScanProsite – motif scanning
MotifScan – identifies functional motifs
CDD (Conserved Domain Database) – identifies domains and key residues


5.4 Steps to Identify Domain/Functional Homology
Input protein sequence
Perform sequence similarity search (BLASTP/PSI-BLAST)
Check conserved domains (Pfam, SMART, InterPro)
Predict functional motifs (PROSITE, ScanProsite)
Validate with structure-based tools if available

6. Summary / Workflow for Information Retrieval
1.Define the query (sequence, accession, or keyword)
2.Select the appropriate database (nucleotide, protein, structural)
3.Choose the search algorithm (BLAST, FASTA, HMMER)
4. Adjust parameters (E-value, filters)
5. Analyze results:
          Sequence similarity
          Homology inference
          Domain identification
           Functional site prediction
            Validate and annotate sequences
            Optional: Structural or evolutionary analysis


7. Key Points

Homology searches are more reliable than keyword searches for function prediction.
Iterative profile-based methods (PSI-BLAST, HMMER) detect distant homologs.
Domain and motif identification is essential for functional annotation.
Integrating sequence, domain, and structure information gives robust predictions.

Comments

Popular Posts

Protein Structure Database (PDB)

Protein Structure Database (PDB) Introduction The Protein Structure Database (PDB) is the primary global repository for the three-dimensional (3D) structures of biological macromolecules such as proteins, nucleic acids, and protein–ligand complexes. These structures are determined experimentally using techniques like X-ray crystallography, Nuclear Magnetic Resonance (NMR) spectroscopy, and Cryo-Electron Microscopy (Cryo-EM). PDB plays a vital role in understanding: Protein structure and function Molecular interactions Drug discovery and design Structural biology and bioinformatics History and Development Established in 1971 Founded by Brookhaven National Laboratory (USA) Initially contained only 7 protein structures Now maintained by the Worldwide Protein Data Bank (wwPDB) Members of wwPDB RCSB PDB (USA) PDBe (Europe) PDBj (Japan) BMRB (Biological Magnetic Resonance Data Bank) Objectives of PDB To collect, store, and distribute 3D structural data of biomolecules To provide free and ope...

❥ Southern Blotting Notes

Southern Blotting  ❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥  Introduction Southern blotting is a molecular biology technique used for the detection of specific DNA sequences in a complex mixture of DNA. It was developed by Edwin M. Southern in 1975. The method involves restriction digestion of DNA, separation by gel electrophoresis, transfer (blotting) onto a membrane, and hybridization with a labeled DNA probe. Principle of Southern Blotting The technique is based on the principle of complementary base pairing. A single-stranded labeled DNA probe hybridizes specifically with its complementary DNA sequence immobilized on a membrane. Detection of the label confirms the presence and size of the target DNA fragment. Steps Involved in Southern Blotting. 1. Isolation of DNA Genomic DNA is extracted from cells or tissues. DNA must be pure and intact to ensure accurate results. 2. Restriction Enzyme  Digestion DNA is digested using specific restriction endonucleases. Produces DNA f...

𓆞 Western Blotting Notes

Western Blotting (Immunoblotting) ❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥ 𓆞❥  Introduction Western blotting, also known as immunoblotting, is a widely used analytical technique for the detection, identification, and quantification of specific proteins in a complex biological sample. The technique combines protein separation by gel electrophoresis with specific antigen–antibody interaction. The method was developed by Towbin et al. (1979) (Burnette 1981---its group work) and is called “Western” in analogy to Southern blotting (DNA) and Northern blotting (RNA). Principle The principle of Western blotting involves: Separation of proteins based on molecular weight using SDS-PAGE Transfer (blotting) of separated proteins onto a membrane Specific detection of the target protein using primary and secondary antibodies Visualization using enzymatic or fluorescent detection systems 👉 Antigen–antibody specificity is the core principle of Western blotting. Steps Involved in Western Blotting 1. Sa...

✩‧₊ Plaque Blotting Technique

Plaque Blotting Technique *ੈ✩‧₊˚༺☆༻*ੈ✩‧₊˚*ੈ✩‧₊˚༺☆༻*ੈ✩‧₊˚ Introduction Plaque blotting is a molecular biology screening technique used to identify specific DNA or RNA sequences present in bacteriophage plaques formed on a bacterial lawn. It is especially useful in the screening of recombinant phage libraries such as λ (lambda) phage genomic or cDNA libraries. This technique combines: Plaque assay (to isolate individual phage clones) Blotting technique (to transfer nucleic acids onto a membrane) Hybridization (to detect specific sequences using labeled probes) Principle of Plaque Blotting The principle of plaque blotting is based on nucleic acid hybridization. Each plaque represents a clone of phage particles containing identical DNA. DNA from phage particles in plaques is: Released Denatured into single strands Transferred onto a nitrocellulose or nylon membrane The membrane is incubated with a labeled DNA/RNA probe complementary to the target sequence. Hybridization between probe and t...

DNA FOOTPRINTING

DNA FOOTPRINTING Introduction DNA footprinting is a molecular biology technique used to identify the specific site(s) on DNA where proteins (such as transcription factors) bind. It reveals the exact nucleotide sequences protected by bound proteins against cleavage by nucleases or chemical agents. It is widely used to study DNA-protein interactions, transcription regulation, and gene expression control. Definition DNA footprinting: A technique used to locate the binding site of DNA-binding proteins on DNA by detecting protected regions that are resistant to enzymatic or chemical cleavage. Principle DNA-binding proteins protect the DNA segment they occupy. DNA exposed to nucleases (DNase I) or chemical cleavage agents is cut at accessible regions. Regions bound by protein remain unaffected, leaving a “footprint”. When fragments are separated on a denaturing polyacrylamide gel, the missing bands correspond to protein-binding sites. Key idea: Cleavage occurs everywhere except where the pro...

RESTRICTION MAPPING

RESTRICTION MAPPING Introduction Restriction mapping is a molecular biology technique used to determine the relative positions of restriction enzyme recognition sites on a DNA molecule. It involves digestion of DNA with one or more restriction endonucleases followed by analysis of fragment sizes using agarose gel electrophoresis. Restriction mapping is essential for DNA characterization, cloning strategies, gene localization, and genome analysis. Definition Restriction mapping is the process of identifying the number, order, and distances between restriction enzyme cleavage sites within a DNA fragment by analyzing the pattern of fragments generated after enzymatic digestion. Principle Restriction enzymes cut DNA at specific palindromic nucleotide sequences. When DNA is digested with: Single restriction enzyme → produces fragments based on its recognition sites Multiple restriction enzymes → produces fragments whose sizes reveal the relative positions of sites By comparing fragment size...

••CLASSIFICATION OF ALGAE - FRITSCH

      MODULE -1       PHYCOLOGY  CLASSIFICATION OF ALGAE - FRITSCH  ❖F.E. Fritsch (1935, 1945) in his book“The Structure and  Reproduction of the Algae”proposed a system of classification of  algae. He treated algae giving rank of division and divided it into 11  classes. His classification of algae is mainly based upon characters of  pigments, flagella and reserve food material.     Classification of Fritsch was based on the following criteria o Pigmentation. o Types of flagella  o Assimilatory products  o Thallus structure  o Method of reproduction          Fritsch divided algae into the following 11 classes  1. Chlorophyceae  2. Xanthophyceae  3. Chrysophyceae  4. Bacillariophyceae  5. Cryptophyceae  6. Dinophyceae  7. Chloromonadineae  8. Euglenineae    9. Phaeophyceae  10. Rhodophyceae  11. Myxophyce...

Electroporation – Detailed Notes

Electroporation – Detailed Notes Definition : Electroporation is a physical method of gene transfer in which cells are exposed to a brief, high-voltage electric pulse, creating temporary pores in the cell membrane. This allows DNA, RNA, proteins, or other molecules to enter the cytoplasm. It is widely used in bacteria, yeast, plant protoplasts, and mammalian cells. Key Concept: The electric field destabilizes the membrane, making it permeable to macromolecules. 1. Principle Cells are suspended in a conductive medium. A brief electrical pulse induces transient pores in the plasma membrane. DNA or other molecules present in the medium enter the cell through these pores. Membrane reseals after the pulse, and the molecule is retained inside the cell. Advantages of Principle: Direct and rapid. Works in many cell types. Does not require chemical carriers or viral vectors. 2. Materials Required Cells – bacterial, yeast, plant protoplasts, mammalian cells. DNA/RNA/other macromolecule – purifie...

Fourth Semester M.Sc. Degree Examination, September 2019BotanySpecial Paper II - ElectiveBO 242 a: BIOTECHNOLOGY(2013 Admission onwards)

Reg. No.......  Name......... G-5263 Fourth Semester M.Sc. Degree Examination, September 2019 Botany Special Paper II - Elective BO 242 a: BIOTECHNOLOGY (2013 Admission onwards) Max. Marks: 75 1. Answer the following questions: 1. Humulin 2. YAC 3. Cybrids 4. Hybridomas 5. IPR 6. Gene therapy 7. C DNA library 8. AFLP 9. Hairy root culture 10. Somacional variation (10 x 1=10 Marks) II. Answer the following questions in not more than 50 words : 11. (a) What are immobilized enzymes? What is its advantage? OR (b) Write a short note on molecular farming. 12. (a) Give an account of bioprocess technology for the production of secondary metabolites. OR (b) What are bioreactors? How it operates? 13. (a) What are probiotics?. How do they work? OR (b) Discuss the methodology and application of western blotting. 14. (a) Briefly explain the application of protoplast culture OR (b) Write a short note on gene therapy 15. (a) What are reporter genes? Discuss its utility in transformation studies O...