335 resultados para PROTEIN FAMILIES
em Indian Institute of Science - Bangalore - Índia
Resumo:
Background: Development of sensitive sequence search procedures for the detection of distant relationships between proteins at superfamily/fold level is still a big challenge. The intermediate sequence search approach is the most frequently employed manner of identifying remote homologues effectively. In this study, examination of serine proteases of prolyl oligopeptidase, rhomboid and subtilisin protein families were carried out using plant serine proteases as queries from two genomes including A. thaliana and O. sativa and 13 other families of unrelated folds to identify the distant homologues which could not be obtained using PSI-BLAST. Methodology/Principal Findings: We have proposed to start with multiple queries of classical serine protease members to identify remote homologues in families, using a rigorous approach like Cascade PSI-BLAST. We found that classical sequence based approaches, like PSI-BLAST, showed very low sequence coverage in identifying plant serine proteases. The algorithm was applied on enriched sequence database of homologous domains and we obtained overall average coverage of 88% at family, 77% at superfamily or fold level along with specificity of similar to 100% and Mathew's correlation coefficient of 0.91. Similar approach was also implemented on 13 other protein families representing every structural class in SCOP database. Further investigation with statistical tests, like jackknifing, helped us to better understand the influence of neighbouring protein families. Conclusions/Significance: Our study suggests that employment of multiple queries of a family for the Cascade PSI-BLAST searches is useful for predicting distant relationships effectively even at superfamily level. We have proposed a generalized strategy to cover all the distant members of a particular family using multiple query sequences. Our findings reveal that prior selection of sequences as query and the presence of neighbouring families can be important for covering the search space effectively in minimal computational time. This study also provides an understanding of the `bridging' role of related families.
Resumo:
Background: Disulphide bridges are well known to play key roles in stability, folding and functions of proteins. Introduction or deletion of disulphides by site-directed mutagenesis have produced varying effects on stability and folding depending upon the protein and location of disulphide in the 3-D structure. Given the lack of complete understanding it is worthwhile to learn from an analysis of extent of conservation of disulphides in homologous proteins. We have also addressed the question of what structural interactions replaces a disulphide in a homologue in another homologue. Results: Using a dataset involving 34,752 pairwise comparisons of homologous protein domains corresponding to 300 protein domain families of known 3-D structures, we provide a comprehensive analysis of extent of conservation of disulphide bridges and their structural features. We report that only 54% of all the disulphide bonds compared between the homologous pairs are conserved, even if, a small fraction of the non-conserved disulphides do include cytoplasmic proteins. Also, only about one fourth of the distinct disulphides are conserved in all the members in protein families. We note that while conservation of disulphide is common in many families, disulphide bond mutations are quite prevalent. Interestingly, we note that there is no clear relationship between sequence identity between two homologous proteins and disulphide bond conservation. Our analysis on structural features at the sites where cysteines forming disulphide in one homologue are replaced by non-Cys residues show that the elimination of a disulphide in a homologue need not always result in stabilizing interactions between equivalent residues. Conclusion: We observe that in the homologous proteins, disulphide bonds are conserved only to a modest extent. Very interestingly, we note that extent of conservation of disulphide in homologous proteins is unrelated to the overall sequence identity between homologues. The non-conserved disulphides are often associated with variable structural features that were recruited to be associated with differentiation or specialisation of protein function.
Resumo:
We explore the fuse of information on co-occurrence of domains in multi-domain proteins in predicting protein-protein interactions. The basic premise of our work is the assumption that domains co-occurring in a polypeptide chain undergo either structural or functional interactions among themselves. In this study we use a template dataset of domains in multidomain proteins and predict protein-protein interactions in a target organism. We note that maximum number of correct predictions of interacting protein domain families (158) is made in S. cerevisiae when the dataset of closely related organisms is used as the template followed by the more diverse dataset of bacterial proteins (48) and a dataset of randomly chosen proteins (23). We conclude that use of multi-domain information from organisms closely-related to the target can aid prediction of interacting protein families.
Resumo:
Protein functional annotation relies on the identification of accurate relationships, sequence divergence being a key factor. This is especially evident when distant protein relationships are demonstrated only with three-dimensional structures. To address this challenge, we describe a computational approach to purposefully bridge gaps between related protein families through directed design of protein-like ``linker'' sequences. For this, we represented SCOP domain families, integrated with sequence homologues, as multiple profiles and performed HMM-HMM alignments between related domain families. Where convincing alignments were achieved, we applied a roulette wheel-based method to design 3,611,010 protein-like sequences corresponding to 374 SCOP folds. To analyze their ability to link proteins in homology searches, we used 3024 queries to search two databases, one containing only natural sequences and another one additionally containing designed sequences. Our results showed that augmented database searches showed up to 30% improvement in fold coverage for over 74% of the folds, with 52 folds achieving all theoretically possible connections. Although sequences could not be designed between some families, the availability of designed sequences between other families within the fold established the sequence continuum to demonstrate 373 difficult relationships. Ultimately, as a practical and realistic extension, we demonstrate that such protein-like sequences can be ``plugged-into'' routine and generic sequence database searches to empower not only remote homology detection but also fold recognition. Our richly statistically supported findings show that complementary searches in both databases will increase the effectiveness of sequence-based searches in recognizing all homologues sharing a common fold. (C) 2013 Elsevier Ltd. All rights reserved.
Resumo:
NrichD
Resumo:
Sequence motifs occurring in a particular order in proteins or DNA have been proved to be of biological interest. In this paper, a new method to locate the occurrences of up to five user-defined motifs in a specified order in large proteins and in nucleotide sequence databases is proposed. It has been designed using the concept of quantifiers in regular expressions and linked lists for data storage. The application of this method includes the extraction of relevant consensus regions from biological sequences. This might be useful in clustering of protein families as well as to study the correlation between positions of motifs and their functional sites in DNA sequences.
Resumo:
Protein structure comparison is essential for understanding various aspects of protein structure, function and evolution. It can be used to explore the structural diversity and evolutionary patterns of protein families. In view of the above, a new algorithm is proposed which performs faster protein structure comparison using the peptide backbone torsional angles. It is fast, robust, computationally less expensive and efficient in finding structural similarities between two different protein structures and is also capable of identifying structural repeats within the same protein molecule.
Resumo:
The availability of the genome sequence of Mycobacterium tuberculosis H37Rv has encouraged determination of large numbers of protein structures and detailed definition of the biological information encoded therein; yet, the functions of many proteins in M. tuberculosis remain unknown. The emergence of multidrug resistant strains makes it a priority to exploit recent advances in homology recognition and structure prediction to re-analyse its gene products. Here we report the structural and functional characterization of gene products encoded in the M. tuberculosis genome, with the help of sensitive profile-based remote homology search and fold recognition algorithms resulting in an enhanced annotation of the proteome where 95% of the M. tuberculosis proteins were identified wholly or partly with information on structure or function. New information includes association of 244 proteins with 205 domain families and a separate set of new association of folds to 64 proteins. Extending structural information across uncharacterized protein families represented in the M. tuberculosis proteome, by determining superfamily relationships between families of known and unknown structures, has contributed to an enhancement in the knowledge of structural content. In retrospect, such superfamily relationships have facilitated recognition of probable structure and/or function for several uncharacterized protein families, eventually aiding recognition of probable functions for homologous proteins corresponding to such families. Gene products unique to mycobacteria for which no functions could be identified are 183. Of these 18 were determined to be M. tuberculosis specific. Such pathogen-specific proteins are speculated to harbour virulence factors required for pathogenesis. A re-annotated proteome of M. tuberculosis, with greater completeness of annotated proteins and domain assigned regions, provides a valuable basis for experimental endeavours designed to obtain a better understanding of pathogenesis and to accelerate the process of drug target discovery. (C) 2014 Elsevier Ltd. All rights reserved.
Resumo:
Background: In the post-genomic era where sequences are being determined at a rapid rate, we are highly reliant on computational methods for their tentative biochemical characterization. The Pfam database currently contains 3,786 families corresponding to ``Domains of Unknown Function'' (DUF) or ``Uncharacterized Protein Family'' (UPF), of which 3,087 families have no reported three-dimensional structure, constituting almost one-fourth of the known protein families in search for both structure and function. Results: We applied a `computational structural genomics' approach using five state-of-the-art remote similarity detection methods to detect the relationship between uncharacterized DUFs and domain families of known structures. The association with a structural domain family could serve as a start point in elucidating the function of a DUF. Amongst these five methods, searches in SCOP-NrichD database have been applied for the first time. Predictions were classified into high, medium and low-confidence based on the consensus of results from various approaches and also annotated with enzyme and Gene ontology terms. 614 uncharacterized DUFs could be associated with a known structural domain, of which high confidence predictions, involving at least four methods, were made for 54 families. These structure-function relationships for the 614 DUF families can be accessed on-line at http://proline.biochem.iisc.ernet.in/RHD_DUFS/. For potential enzymes in this set, we assessed their compatibility with the associated fold and performed detailed structural and functional annotation by examining alignments and extent of conservation of functional residues. Detailed discussion is provided for interesting assignments for DUF3050, DUF1636, DUF1572, DUF2092 and DUF659. Conclusions: This study provides insights into the structure and potential function for nearly 20 % of the DUFs. Use of different computational approaches enables us to reliably recognize distant relationships, especially when they converge to a common assignment because the methods are often complementary. We observe that while pointers to the structural domain can offer the right clues to the function of a protein, recognition of its precise functional role is still `non-trivial' with many DUF domains conserving only some of the critical residues. It is not clear whether these are functional vestiges or instances involving alternate substrates and interacting partners. Reviewers: This article was reviewed by Drs Eugene Koonin, Frank Eisenhaber and Srikrishna Subramanian.
Resumo:
Lipocalins constitute a superfamily of extracellular proteins that are found in all three kingdoms of life. Although very divergent in their sequences and functions, they show remarkable similarity in 3-D structures. Lipocalins bind and transport small hydrophobic molecules. Earlier sequence-based phylogenetic studies of lipocalins highlighted that they have a long evolutionary history. However the molecular and structural basis of their functional diversity is not completely understood. The main objective of the present study is to understand functional diversity of the lipocalins using a structure-based phylogenetic approach. The present study with 39 protein domains from the lipocalin superfamily suggests that the clusters of lipocalins obtained by structure-based phylogeny correspond well with the functional diversity. The detailed analysis on each of the clusters and sub-clusters reveals that the 39 lipocalin domains cluster based on their mode of ligand binding though the clustering was performed on the basis of gross domain structure. The outliers in the phylogenetic tree are often from single member families. Also structure-based phylogenetic approach has provided pointers to assign putative function for the domains of unknown function in lipocalin family. The approach employed in the present study can be used in the future for the functional identification of new lipocalin proteins and may be extended to other protein families where members show poor sequence similarity but high structural similarity.
Resumo:
Primary microcephaly is an autosomal recessive disorder characterized by smaller than normal brain size and mental retardation. It is genetically heterogeneous with seven loci: MCPH1-MCPH7. We have previously reported genetic analysis of 35 families, including the identification of the MCPH7 gene STIL. Of the 35 families, three families showed linkage to the MCPH2 locus. Recent whole-exome sequencing studies have shown that the WDR62 gene, located in the MCPH2 candidate region, is mutated in patients with severe brain malformations. We therefore sequenced the WDR62 gene in our MCPH2 families and identified two novel homozygous protein truncating mutations in two families. Affected individuals in the two families had pachygyria, microlissencephaly, band heterotopias, gyral thickening, and dysplastic cortex. Using immunofluorescence study, we showed that, as with other MCPH proteins, WDR62 localizes to centrosomes in A549, HepG2, and HaCaT cells. In addition, WDR62 was also localized to nucleoli. Bioinformatics analysis predicted two overlapping nuclear localization signals and multiple WD-40 repeats in WDR62. Two other groups have also recently identified WDR62 mutations in MCPH2 families. Our results therefore add further evidence that WDR62 is the MCPH2 gene. The present findings will be helpful in genetic diagnosis of patients linked to the MCPH2 locus.
Resumo:
In peptide and protein structures, occurrence of (phi,psi.) angles in the disallowed region of the Ramachandran map almost always suggests local regions of error or poor accuracy. However, very rarely genuine disallowed conformations occur as noted in the current study in proteins of known structure available at ultra-high resolution (<= 1.2 (A) over circle). In the current work, extent of conservation of genuine disallowed conformations in homologous proteins of known structures has been analyzed. From a dataset of 124 protein domain families, with structure of at least one constituent member in each family available at a resolution of 1.2 (A) over circle or better, we have analyzed the conservation of 221 disallowed conformations. It is observed that the disallowed conformation is only moderately conservedin protein domain families. In the gross dataset no particular residue type adopting disallowed conformation elicit high conservation of residue type though there are alignment positions in the dataset with complete conservation of both the residue type and the disallowed conformation. Conserved disallowed conformation in protein domain families play biologically significant role in roughly 50% of the cases. The residues with the disallowed conformation or its flanking residues are often located within or around the functional site of the protein. (C) 2013 Elsevier B.V. All rights reserved.
Resumo:
Primary microcephaly (MCPH) is an autosomal-recessive congenital disorder characterized by smaller-than-normal brain size and mental retardation. MCPH is genetically heterogeneous with six known loci: MCPH1-MCPH6. We report mapping of a novel locus, MCPH7, to chromosome 1p32.3-p33 between markers D1S2797 and D1S417, corresponding to a physical distance of 8.39 Mb. Heterogeneity analysis of 24 families previously excluded from linkage to the six known MCPH loci suggested linkage of five families (20.83%) to the MCPH7 locus. In addition, four families were excluded from linkage to the MCPH7 locus as well as all of the six previously known loci, whereas the remaining 15 families could not be conclusively excluded or included. The combined maximum two-point LOD score for the linked families was 5.96 at marker D1S386 at theta = 0.0. The combined multipoint LOD score was 6.97 between markers D1S2797 and D1S417. Previously, mutations in four genes, MCPH1, CDK5RAP2, ASPM, and CENPJ, that code for centrosomal proteins have been shown to cause this disorder. Three different homozygous mutations in STIL, which codes for a pericentriolar and centrosomal protein, were identified in patients from three of the five families linked to the MCPH7 locus; all are predicted to truncate the STIL protein. Further, another recently ascertained family was homozygous for the same mutation as one of the original families. There was no evidence for a common haplotype. These results suggest that the centrosome and its associated structures are important in the control of neurogenesis in the developing human brain.
Genome-wide analysis and experimentation of plant serine/threonine/tyrosine-specific protein kinases
Resumo:
Protein tyrosine phosphorylation plays an important role in cell growth, development and oncogenesis. No classical protein tyrosine kinase has hitherto been cloned from plants. Does protein tyrosine kinase exist in plants? To address this, we have performed a genomic survey of protein tyrosine kinase motifs in plants using the delineated tyrosine phosphorylation motifs from the animal system. The Arabidopsis thaliana genome encodes 57 different protein kinases that have tyrosine kinase motifs. Animal non-receptor tyrosine kinases, SRC, ABL, LYN, FES, SEK, KIN and RAS have structural relationship with putative plant tyrosine kinases. In an extended analysis, animal receptor and non-receptor kinases, Raf and Ras kinases, mixed lineage kinases and plant serine/threonine/tyrosine (STY) protein kinases, form a well-supported group sharing a common origin within the superfamily of STY kinases. We report that plants lack bona fide tyrosine kinases, which raise an intriguing possibility that tyrosine phosphorylation is carried out by dual-specificity STY protein kinases in plants. The distribution pattern of STY protein kinase families on Arabidopsis chromosomes indicates that this gene family is partly a consequence of duplication and reshuffling of the Arabidopsis genome and of the generation of tandem repeats. Genome-wide analysis is supported by the functional expression and characterization of At2g24360 and phosphoproteomics of Arabidopsis. Evidence for tyrosine phosphorylated proteins is provided by alkaline hydrolysis, anti-phosphotyrosine immunoblotting, phosphoamino acid analysis and peptide mass fingerprinting. These results report the first comprehensive survey of genome-wide and tyrosine phosphoproteome analysis of plant STY protein kinases.