Skip to main content

Information retrieval from databases - search concepts, Tools for searching, homology searching, finding Domain and Functional site homologies

Information retrieval from databases - search concepts, Tools for searching, homology searching, finding Domain and Functional site homologies

Information Retrieval from Databases

1. Introduction

Information retrieval in bioinformatics refers to the process of extracting relevant biological data (DNA, RNA, protein sequences, structures, or functional information) from databases.
Aim: Identify sequences, functions, or structural features for analysis, comparison, and annotation.
Databases can be primary (raw sequence data) or secondary/derived (annotated, processed data).
2. Search Concepts in Biological Databases

2.1 Types of Searches

Exact Match Search

Returns results only if the query exactly matches database entries.
Useful for known accession numbers or IDs.

Pattern/Keyword Search

Searches based on specific motifs, keywords, or annotations.
Example: “kinase domain,” “signal peptide.”

Similarity/Homology Search
Detects sequences similar to the query based on sequence alignment.
Uses scoring matrices to assess similarity (e.g., BLOSUM, PAM).
Useful for identifying homologous genes or proteins.


Complex Query Search

Combines Boolean operators (AND, OR, NOT) to refine results.
Example: “kinase AND human NOT viral.”


2.2 Search Parameters

Query sequence or keyword
Database selection (nucleotide, protein, structural, functional)
Algorithm choice (BLAST, FASTA, PSI-BLAST)
Threshold or cut-off (E-value, score, % identity)
Filters (organism, date, length, sequence type)


3. Tools for Searching Biological Databases


3.1 Nucleotide Sequence Databases


GenBank (NCBI)
EMBL (European Nucleotide Archive)
DDBJ (DNA Data Bank of Japan)
Search Tools:
BLASTN – nucleotide vs nucleotide
FASTA – nucleotide similarity search


3.2 Protein Sequence Databases

SWISS-PROT / UniProtKB – curated protein sequences
PIR / TrEMBL – unreviewed protein sequences
Search Tools:
BLASTP – protein vs protein
PSI-BLAST – iterative search for distant homologs
HMMER – profile-based search using hidden Markov models

3.3 Structural Databases

Protein Data Bank (PDB) – 3D protein structures
SCOP / CATH – structural classification of proteins
Search Tools:
BLAST 3D – structure-based sequence search
DALI – structural alignment


3.4 Specialized Databases
Pfam – protein families
PROSITE – protein motifs
InterPro – integrated database of protein domains


4. Homology Searching

Homology searching identifies evolutionarily related sequences based on similarity.


4.1 Concept

Homologous sequences: share a common ancestor.
Types:

Orthologs – homologs in different species
Paralogs – homologs in the same species
Homology suggests similar structure or function.


4.2 Methods

1. Pairwise Sequence Alignment

Tools: BLAST, FASTA
Measures similarity (% identity) and E-value


2. Multiple Sequence Alignment (MSA)
Tools: Clustal Omega, MUSCLE
Identifies conserved residues and motifs


3. Profile-based Searching

Uses Position-Specific Scoring Matrices (PSSM)
Tool: PSI-BLAST, HMMER

4. Structural Homology
Comparing 3D structures for similarity
Tools: DALI, CATH, SCOP


5. Finding Domain and Functional Site Homologies


5.1 Protein Domains

Definition: Conserved part of protein with specific function/structure.
Examples: kinase domain, zinc finger, SH2 domain.
Domains often determine protein function.


5.2 Domain Databases and Tools

Pfam – HMM-based domain identification
SMART – domains in signaling and extracellular proteins
InterPro – integrates multiple domain databases
PROSITE – motifs and functional sites

5.3 Functional Site Prediction

Active sites, binding sites, or motifs are predicted based on:
Conserved residues across homologs
3D structure information
Known motifs (PROSITE patterns)
Tools:
ScanProsite – motif scanning
MotifScan – identifies functional motifs
CDD (Conserved Domain Database) – identifies domains and key residues


5.4 Steps to Identify Domain/Functional Homology
Input protein sequence
Perform sequence similarity search (BLASTP/PSI-BLAST)
Check conserved domains (Pfam, SMART, InterPro)
Predict functional motifs (PROSITE, ScanProsite)
Validate with structure-based tools if available

6. Summary / Workflow for Information Retrieval
1.Define the query (sequence, accession, or keyword)
2.Select the appropriate database (nucleotide, protein, structural)
3.Choose the search algorithm (BLAST, FASTA, HMMER)
4. Adjust parameters (E-value, filters)
5. Analyze results:
          Sequence similarity
          Homology inference
          Domain identification
           Functional site prediction
            Validate and annotate sequences
            Optional: Structural or evolutionary analysis


7. Key Points

Homology searches are more reliable than keyword searches for function prediction.
Iterative profile-based methods (PSI-BLAST, HMMER) detect distant homologs.
Domain and motif identification is essential for functional annotation.
Integrating sequence, domain, and structure information gives robust predictions.

Comments

Popular Posts

❃HPLC – High Performance Liquid Chromatography

HPLC – High Performance Liquid Chromatography ┏━━━━━ •❃°•°❀°•°❃•━━━━•━━━┓  1. Introduction High Performance Liquid Chromatography (HPLC) is an advanced analytical technique used for the separation, identification, and quantification of components present in a mixture. It is based on the differential distribution of analytes between a stationary phase and a liquid mobile phase under high pressure. HPLC is widely used in biochemistry, biotechnology, pharmaceuticals, food analysis, environmental studies, and clinical diagnostics. 2. Principle of HPLC The principle of HPLC is based on partition, adsorption, ion-exchange, or size-exclusion mechanisms, depending on the type of column used. A liquid mobile phase is pumped at high pressure through a column packed with fine stationary phase particles Sample components interact differently with the stationary phase Components with stronger interaction elute slower Components with weaker interaction elute faster Separated components are detec...

Microbial Production of PharmaceuticalsSomatostatin, Humulin and Interferons

Microbial Production of Pharmaceuticals Somatostatin, Humulin and Interferons 1. Introduction Advances in recombinant DNA technology have enabled microorganisms to produce human therapeutic proteins safely, economically and in large quantities. Microbial systems such as Escherichia coli and yeast (Saccharomyces cerevisiae) are widely used for the production of pharmaceuticals that were earlier isolated from human or animal tissues. Important microbial-derived pharmaceuticals include somatostatin, human insulin (Humulin) and interferons. 2. Advantages of Microbial Production of Pharmaceuticals High yield and rapid production Cost-effective and scalable Free from animal pathogens Consistent product quality Easy genetic manipulation 3. General Steps in Microbial Production of Recombinant Pharmaceuticals Isolation of target gene Construction of recombinant DNA Insertion into suitable vector Transformation into host microorganism Expression of protein Downstream processing and purification ...

••CLASSIFICATION OF ALGAE - FRITSCH

      MODULE -1       PHYCOLOGY  CLASSIFICATION OF ALGAE - FRITSCH  ❖F.E. Fritsch (1935, 1945) in his book“The Structure and  Reproduction of the Algae”proposed a system of classification of  algae. He treated algae giving rank of division and divided it into 11  classes. His classification of algae is mainly based upon characters of  pigments, flagella and reserve food material.     Classification of Fritsch was based on the following criteria o Pigmentation. o Types of flagella  o Assimilatory products  o Thallus structure  o Method of reproduction          Fritsch divided algae into the following 11 classes  1. Chlorophyceae  2. Xanthophyceae  3. Chrysophyceae  4. Bacillariophyceae  5. Cryptophyceae  6. Dinophyceae  7. Chloromonadineae  8. Euglenineae    9. Phaeophyceae  10. Rhodophyceae  11. Myxophyce...

Single Nucleotide Polymorphisms (SNPs) – Detailed Notes

Single Nucleotide Polymorphisms (SNPs) – Detailed Notes 1. Definition SNPs are single base-pair variations in the DNA sequence that occur at a specific position in the genome among individuals of a species. Example: At a specific locus, one individual may have A while another has G: Copy code Individual 1: …A T C G A T…   Individual 2: …A T C G G T… SNPs are the most common type of genetic variation in most organisms. 2. Characteristics of SNPs Single base change: Involves substitution of one nucleotide for another (A↔G, C↔T). Biallelic nature: Most SNPs have only two alleles in a population. Widespread in the genome: Found in coding regions (exons), non-coding regions (introns, promoters, intergenic regions). Stable inheritance: Passed from generation to generation like other genetic markers. Frequency: Occur approximately every 100–300 bp in the human genome. 3 . Types of SNPs SNPs are categorized based on location or effect on gene function: A. Based on genomic location Cod...

Intellectual Property Rights (IPR) – Detailed Notes

Intellectual Property Rights (IPR) – Detailed Notes 1. Introduction Intellectual Property Rights (IPR) are legal rights granted to creators and inventors over their creations or inventions. They protect innovation and creativity, providing the owner exclusive rights to use, sell, or license their creation. IPR encourages research, development, and economic growth by rewarding creativity. 2. Importance of IPR Protects inventions, designs, and creative work. Prevents unauthorized use, copying, or commercialization. Encourages innovation and research. Provides financial benefits to inventors through licensing or royalties. Supports economic growth and competitiveness. Safeguards traditional knowledge and biodiversity. 3. Types of Intellectual Property Rights A. Patents Definition: Exclusive right granted to an inventor for a new invention for a limited period (usually 20 years). Requirements: Novelty – must be new and not published. Inventive step – non-obvious to someone skilled in the f...

SCAR (Sequence Characterized Amplified Region) Markers

SCAR (Sequence Characterized Amplified Region) Markers   Introduction SCAR markers are PCR-based DNA markers derived from RAPD, AFLP, or other random markers. Developed by Paran and Michelmore in 1993 to convert dominant, less reproducible markers into specific, reproducible, co-dominant markers. SCAR markers are locus-specific, reproducible, and sequence-characterized, making them ideal for marker-assisted selection (MAS). Principle SCAR markers are designed based on known DNA sequences obtained from cloned RAPD/AFLP fragments. Specific primers (18–24 bp) are synthesized to amplify a single, defined locus. The PCR amplification of this region generates a distinct band, which is highly reproducible and can distinguish homozygotes from heterozygotes if designed as co-dominant. Key idea: Random marker (e.g., RAPD) → Cloning & sequencing → Design specific primers → PCR → SCAR marker Materials Required Genomic DNA from the organism Specific primers (18–24 bp) designed from sequence...

❥NORTHERN BLOTTING

NORTHERN BLOTTING – 30 MARK DETAILED NOTES  π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž ❥ π“†ž❥ π“†ž❥  Northern blotting is a molecular biology technique used to detect specific RNA molecules in a complex mixture. It provides information about gene expression, RNA size, and transcript abundance by hybridizing RNA with a labeled complementary DNA or RNA probe. πŸ“Œ Named by analogy to Southern blotting (DNA detection). 2. Principle The principle of Northern blotting is based on: Separation of RNA molecules by size using denaturing agarose gel electrophoresis Transfer (blotting) of separated RNA onto a nylon or nitrocellulose membrane Hybridization of membrane-bound RNA with a labeled complementary probe Detection of RNA–probe hybrids by autoradiography or chemiluminescence ✔ Only RNA sequences complementary to the probe will be detected. 3. Types of RNA Analyzed mRNA (most common) rRNA tRNA miRNA and siRNA (with modified protocols) 4. Requirements / Materials Total RNA or poly(A)+ RNA Denaturing agarose ...

𓆉 INDEX PAGE -NOTETHEPOINT43

INDEX PAGE   MAIN    CONTENT 1.   HSST BOTANY SYLLABUS, DETAILED NOTES, MCQ 2.  SET GENERAL PAPER SYLLABUS, DETAILED NOTES, 50MCQ 3.  SET BOTANY SYLLABUS, DETAILED NOTES, MCQ 4. MSC BOTANY THIRD SEMESTER SYLLABUS, NOTES (KERALA UNIVERSITY ) 5. MSC BOTANY THIRD SEMESTER QUESTION PAPER (KERALA UNIVERSITY ) 6. MSC BOTANY FOURTH SEMESTER SYLLABUS &NOTES (KERALA UNIVERSITY ) 7. FOURTH SEMESTER MSC BOTANY PREVIOUS QUESTION PAPER  (KERALA UNIVERSITY )

Suspension culture and development - methodology, kinetics of growth and production formation, elicitation methods, hairy root culture. Detailed notes

Suspension culture and development - methodology, kinetics of growth and production formation, elicitation methods, hairy root culture. Detailed notes 1. Introduction Suspension culture is a type of plant tissue culture in which single cells or small cell aggregates are grown in liquid nutrient medium under continuous agitation. It is mainly used for: Large-scale biomass production Secondary metabolite production Cell physiology and biochemical studies Genetic manipulation and selection. 2. Methodology of Suspension Culture 2.1 Source of Explant Usually initiated from friable callus Callus derived from: Leaf Stem Root Hypocotyl Friable callus is preferred as it disintegrates easily into single cells. 2.2 Preparation of Cell Suspension Friable callus is transferred into liquid MS medium Medium contains: Carbon source (usually sucrose) Auxins (2,4-D commonly used) Culture maintained in: Conical flasks Orbital shaker (100–150 rpm) 2.3 Culture Conditions Parameter Requirement Temperature ...

❥ Southern Blotting Notes

Southern Blotting  ❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥ π“†ž❥  Introduction Southern blotting is a molecular biology technique used for the detection of specific DNA sequences in a complex mixture of DNA. It was developed by Edwin M. Southern in 1975. The method involves restriction digestion of DNA, separation by gel electrophoresis, transfer (blotting) onto a membrane, and hybridization with a labeled DNA probe. Principle of Southern Blotting The technique is based on the principle of complementary base pairing. A single-stranded labeled DNA probe hybridizes specifically with its complementary DNA sequence immobilized on a membrane. Detection of the label confirms the presence and size of the target DNA fragment. Steps Involved in Southern Blotting. 1. Isolation of DNA Genomic DNA is extracted from cells or tissues. DNA must be pure and intact to ensure accurate results. 2. Restriction Enzyme  Digestion DNA is digested using specific restriction endonucleases. Produces DNA f...