993 resultados para Databases, Protein
Resumo:
Protein-coding genes evolve at different rates, and the influence of different parameters, from gene size to expression level, has been extensively studied. While in yeast gene expression level is the major causal factor of gene evolutionary rate, the situation is more complex in animals. Here we investigate these relations further, especially taking in account gene expression in different organs as well as indirect correlations between parameters. We used RNA-seq data from two large datasets, covering 22 mouse tissues and 27 human tissues. Over all tissues, evolutionary rate only correlates weakly with levels and breadth of expression. The strongest explanatory factors of purifying selection are GC content, expression in many developmental stages, and expression in brain tissues. While the main component of evolutionary rate is purifying selection, we also find tissue-specific patterns for sites under neutral evolution and for positive selection. We observe fast evolution of genes expressed in testis, but also in other tissues, notably liver, which are explained by weak purifying selection rather than by positive selection.
Resumo:
Les anomalies du tube neural (ATN) sont des anomalies développementales où le tube neural reste ouvert (1-2/1000 naissances). Afin de prévenir cette maladie, une connaissance accrue des processus moléculaires est nécessaire. L’étiologie des ATN est complexe et implique des facteurs génétiques et environnementaux. La supplémentation en acide folique est reconnue pour diminuer les risques de développer une ATN de 50-70% et cette diminution varie en fonction du début de la supplémentation et de l’origine démographique. Les gènes impliqués dans les ATN sont largement inconnus. Les études génétiques sur les ATN chez l’humain se sont concentrées sur les gènes de la voie métabolique des folates du à leur rôle protecteur dans les ATN et les gènes candidats inférés des souris modèles. Ces derniers ont montré une forte association entre la voie non-canonique Wnt/polarité cellulaire planaire (PCP) et les ATN. Le gène Protein Tyrosine Kinase 7 est un membre de cette voie qui cause l’ATN sévère de la craniorachischisis chez les souris mutantes. Ptk7 interagit génétiquement avec Vangl2 (un autre gène de la voie PCP), où les doubles hétérozygotes montrent une spina bifida. Ces données font de PTK7 comme un excellent candidat pour les ATN chez l’humain. Nous avons re-séquencé la région codante et les jonctions intron-exon de ce gène dans une cohorte de 473 patients atteints de plusieurs types d’ATN. Nous avons identifié 6 mutations rares (fréquence allélique <1%) faux-sens présentes chez 1.1% de notre cohorte, dont 3 sont absentes dans les bases de données publiques. Une variante, p.Gly348Ser, a agi comme un allèle hypermorphique lorsqu'elle est surexprimée dans le modèle de poisson zèbre. Nos résultats impliquent la mutation de PTK7 comme un facteur de risque pour les ATN et supporte l'idée d'un rôle pathogène de la signalisation PCP dans ces malformations.
Resumo:
Protein–ligand binding site prediction methods aim to predict, from amino acid sequence, protein–ligand interactions, putative ligands, and ligand binding site residues using either sequence information, structural information, or a combination of both. In silico characterization of protein–ligand interactions has become extremely important to help determine a protein’s functionality, as in vivo-based functional elucidation is unable to keep pace with the current growth of sequence databases. Additionally, in vitro biochemical functional elucidation is time-consuming, costly, and may not be feasible for large-scale analysis, such as drug discovery. Thus, in silico prediction of protein–ligand interactions must be utilized to aid in functional elucidation. Here, we briefly discuss protein function prediction, prediction of protein–ligand interactions, the Critical Assessment of Techniques for Protein Structure Prediction (CASP) and the Continuous Automated EvaluatiOn (CAMEO) competitions, along with their role in shaping the field. We also discuss, in detail, our cutting-edge web-server method, FunFOLD for the structurally informed prediction of protein–ligand interactions. Furthermore, we provide a step-by-step guide on using the FunFOLD web server and FunFOLD3 downloadable application, along with some real world examples, where the FunFOLD methods have been used to aid functional elucidation.
Resumo:
Homology-driven proteomics is a major tool to characterize proteomes of organisms with unsequenced genomes. This paper addresses practical aspects of automated homology-driven protein identifications by LC-MS/MS on a hybrid LTQ orbitrap mass spectrometer. All essential software elements supporting the presented pipeline are either hosted at the publicly accessible web server, or are available for free download. (C) 2008 Elsevier B.V. All rights reserved.
Resumo:
Motivation: DNA assembly programs classically perform an all-against-all comparison of reads to identify overlaps, followed by a multiple sequence alignment and generation of a consensus sequence. If the aim is to assemble a particular segment, instead of a whole genome or transcriptome, a target-specific assembly is a more sensible approach. GenSeed is a Perl program that implements a seed-driven recursive assembly consisting of cycles comprising a similarity search, read selection and assembly. The iterative process results in a progressive extension of the original seed sequence. GenSeed was tested and validated on many applications, including the reconstruction of nuclear genes or segments, full-length transcripts, and extrachromosomal genomes. The robustness of the method was confirmed through the use of a variety of DNA and protein seeds, including short sequences derived from SAGE and proteome projects.
Resumo:
The initiation of glycogen synthesis requires the protein glycogenin, which incorporates glucose residues through a self-glucosylation reaction, and then acts as substrate for chain elongation by glycogen synthase and branching enzyme. Numerous sequences of glycogenin-like proteins are available in the databases but the enzymes from mammalian skeletal muscle and from Saccharomyces cerevisiae are the best characterized. We report the isolation of a cDNA from the fungus Neurospora crassa, which encodes a protein, GNN, which has properties characteristic of glycogenin. The protein is one of the largest glycogenins but shares several conserved domains common to other family members. Recombinant GNN produced in Escherichia coli was able to incorporate glucose in a self-glucosylation reaction, to trans-glucosylate exogenous substrates, and to act as substrate for chain elongation by glycogen synthase. Recombinant protein was sensitive to C-terminal proteolysis, leading to stable species of around 31 kDa, which maintained all functional properties. The role of GNN as an initiator of glycogen metabolism was confirmed by its ability to complement the glycogen deficiency of a S. cerevisiae strain (glg1 glg2) lacking glycogenin and unable to accumulate glycogen. Disruption of the gnn gene of N. crassa by repeat induced point mutation (RIP) resulted in a strain that was unable to synthesize glycogen, even though the glycogen synthase activity was unchanged. Northern blot analysis showed that the gnn gene was induced during vegetative growth and was repressed upon carbon starvation. (C) 2004 Elsevier B.V. All rights reserved.
Resumo:
Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES)
Resumo:
Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP)
Resumo:
Cowpea aphid-borne mosaic virus (CABMV) causes major diseases in cowpea and passion flower plants in Brazil and also in other countries. CABMV has also been isolated from leguminous species including, Cassia hoffmannseggii, Canavalia rosea, Crotalaria juncea and Arachis hypogaea in Brazil. The virus seems to be adapted to two distinct families, the Passifloraceae and Fabaceae. Aiming to identify CABMV and elucidate a possible host adaptation of this virus species, isolates from cowpea, passion flower and C.hoffmannseggii collected in the states of Pernambuco and Rio Grande do Norte were analysed by sequencing the complete coat protein genes. A phylogenetic tree was constructed based on the obtained sequences and those available in public databases. Major Brazilian isolates from passion flower, independently of the geographical distances among them, were grouped in three different clusters. The possible host adaptation was also observed in fabaceous-infecting CABMV Brazilian isolates. These host adaptations possibly occurred independently within Brazil, so all these clusters belong to a bigger Brazilian cluster. Nevertheless, African passion flower or cowpea-infecting isolates formed totally different clusters. These results showed that host adaptation could be one factor for CABMV evolution, although geographical isolation is a stronger factor.
Resumo:
Background: In protein sequence classification, identification of the sequence motifs or n-grams that can precisely discriminate between classes is a more interesting scientific question than the classification itself. A number of classification methods aim at accurate classification but fail to explain which sequence features indeed contribute to the accuracy. We hypothesize that sequences in lower denominations (n-grams) can be used to explore the sequence landscape and to identify class-specific motifs that discriminate between classes during classification. Discriminative n-grams are short peptide sequences that are highly frequent in one class but are either minimally present or absent in other classes. In this study, we present a new substitution-based scoring function for identifying discriminative n-grams that are highly specific to a class. Results: We present a scoring function based on discriminative n-grams that can effectively discriminate between classes. The scoring function, initially, harvests the entire set of 4- to 8-grams from the protein sequences of different classes in the dataset. Similar n-grams of the same size are combined to form new n-grams, where the similarity is defined by positive amino acid substitution scores in the BLOSUM62 matrix. Substitution has resulted in a large increase in the number of discriminatory n-grams harvested. Due to the unbalanced nature of the dataset, the frequencies of the n-grams are normalized using a dampening factor, which gives more weightage to the n-grams that appear in fewer classes and vice-versa. After the n-grams are normalized, the scoring function identifies discriminative 4- to 8-grams for each class that are frequent enough to be above a selection threshold. By mapping these discriminative n-grams back to the protein sequences, we obtained contiguous n-grams that represent short class-specific motifs in protein sequences. Our method fared well compared to an existing motif finding method known as Wordspy. We have validated our enriched set of class-specific motifs against the functionally important motifs obtained from the NLSdb, Prosite and ELM databases. We demonstrate that this method is very generic; thus can be widely applied to detect class-specific motifs in many protein sequence classification tasks. Conclusion: The proposed scoring function and methodology is able to identify class-specific motifs using discriminative n-grams derived from the protein sequences. The implementation of amino acid substitution scores for similarity detection, and the dampening factor to normalize the unbalanced datasets have significant effect on the performance of the scoring function. Our multipronged validation tests demonstrate that this method can detect class-specific motifs from a wide variety of protein sequence classes with a potential application to detecting proteome-specific motifs of different organisms.
Resumo:
Advances in novel molecular biological diagnostic methods are changing the way of diagnosis and study of metabolic disorders like growth hormone deficiency. Faster sequencing and genotyping methods require strong bioinformatics tools to make sense of the vast amount of data generated by modern laboratories. Advances in genome sequencing and computational power to analyze the whole genome sequences will guide the diagnostics of future. In this chapter, an overview of some basic bioinformatics resources that are needed to study metabolic disorders are reviewed and some examples of bioinformatics analysis of human growth hormone gene, protein and structure are provided.
Resumo:
Retinitis pigmentosa (RP) is a name given to a group of inherited retinal dystrophies that lead to progressive photoreceptor degeneration, and thus, visual impairment. It is evident at both the clinical and the molecular level that these are heterogeneous disorders, with wide variation in severity, mode of inheritance, and phenotype. The genetics of RP are not simple; the disease can be inherited in dominant, recessive, X-linked, and digenic modes. Autosomal dominant RP (adRP) results from mutations in at least ten mapped loci, but there may be dozens of genetic loci where mutations can cause RP. To date, there are over a hundred genes known to cause retinal degenerative diseases, and less than half of these have been cloned (RetNet). Among the dozens of retinitis pigmentosa loci known to exist, only a few have been identified and the remainders are inferred from linkage studies. Today, the genes for seven of the twelve-adRP loci have been identified, and these are rhodopsin, peripherin/RDS, NRL, ROM1, CRX, RP13 and RP1. My research projects involved a combination of the continued search for genes involved in retinal dystrophies, as well the investigation into the role of peripherin/RDS and RP1 in the disease etiology of autosomal dominant RP. ^ Most of the mutations leading to inherited retinal disorders have been identified in predominately retina expressed genes like rhodopsin, peripherin/RDS, and RP1. Expressed sequence tags (ESTs) that were retina-specific were culled from sequence databases and, together with laboratory analysis, were analyzed as potential candidate genes for retinal dystrophies. Thirteen of the fifty-five identified retina-specific ESTs mapped to within candidate regions for inherited retinopathies. One of these is RP1L1, a homologue of RP1 and a potential cause of adRP. ^ Once a disease-associated gene has been identified, elucidating the role of that gene in the visual process is essential for understanding what happens when the process is defective as it is in adRP. My next projects involved investigating the role of a novel 5′ donor +3 splice site mutation on the mRNA of peripherin/RDS in adRP affected individuals, and comparative sequencing in RP1 to define conserved regions of the protein. Comparative sequencing is a powerful way to delineate critical regions of a sequence because different regions of a gene have different functions, and each region is subject to different levels of functional or structural constraints. Establishing a framework of conserved domains is beneficial not only for structural or functional studies, but can also aid in determining the potential effects of mutations. With the completion of sequencing of human genome, and other organisms such as Saccharomyces cerevisiae, Caenorhabditis elegans , and Drosophila, the facility of comparative sequencing will only increase in the future. Comparative sequencing has already become an established procedure for pinpointing conserved regions of a protein, and is an efficient way to target regions of a protein for experimental and/or evolutionary analysis. ^
Resumo:
Objective. To determine the accuracy of the urine protein:creatinine ratio (pr:cr) in predicting 300 mg of protein in 24-hour urine collection in pregnant patients with suspected preeclampsia. ^ Methods. A systematic review was performed. Articles were identified through electronic databases and the relevant citations were hand searching of textbooks and review articles. Included studies evaluated patients for suspected preeclampsia with a 24-hour urine sample and a pr:cr. Only English language articles were included. The studies that had patients with chronic illness such as chronic hypertension, diabetes mellitus or renal impairment were excluded from the review. Two researchers extracted accuracy data for pr:cr relative to a gold standard of 300 mg of protein in 24-hour sample as well as population and study characteristics. The data was analyzed and summarized in tabular and graphical form. ^ Results. Sixteen studies were identified and only three studies met our inclusion criteria with 510 total patients. The studies evaluated different cut-points for positivity of pr:cr from 130 mg/g to 700 mg/g. Sensitivities and specificities for pr:cr of 130mg/g -150 mg/g were 90-93% and 33-65%, respectively; for a pr:cr of 300 mg/g were 81-95% and 52-80%, respectively; for a pr:cr of 600-700mg/g were 85-87% and 96-97%, respectively. ^ Conclusion. The value of a random pr:cr to exclude pre-eclampsia is limited because even low levels of pr:cr (130-150 mg/g) may miss up to 10% of patients with significant proteinuria. A pr:cr of more than 600 mg/g may obviate a 24-hour collection.^
Resumo:
Protein-coding gene families are sets of similar genes with a shared evolutionary origin and, generally, with similar biological functions. In plants, the size and role of gene families has been only partially addressed. However, suitable bioinformatics tools are being developed to cluster the enormous number of sequences currently available in databases. Specifically, comparative genomic databases promise to become powerful tools for gene family annotation in plant clades. In this review, I evaluate the data retrieved from various gene family databases, the ease with which they can be extracted and how useful the extracted information is.
Resumo:
The Drosophila retinal degeneration C (rdgC) gene encodes an unusual protein serine/threonine phosphatase in that it contains at least two EF-hand motifs at its carboxy terminus. By a combination of large-scale sequencing of human retina cDNA clones and searches of expressed sequence tag and genomic DNA databases, we have identified two sequences in mammals [Protein Phosphatase with EF-hands-1 and 2 (PPEF-1 and PPEF-2)] and one in Caenorhabditis elegans (PPEF) that closely resemble rdgC. In the adult, PPEF-2 is expressed specifically in retinal rod photoreceptors and the pineal. In the retina, several isoforms of PPEF-2 are predicted to arise from differential splicing. The isoform that most closely resembles rdgC is localized to rod inner segments. Together with the recently described localization of PPEF-1 transcripts to primary somatosensory neurons and inner ear cells in the developing mouse, these data suggest that the PPEF family of protein serine/threonine phosphatases plays a specific and conserved role in diverse sensory neurons.