832 resultados para genome mining


Relevância:

60.00% 60.00%

Publicador:

Resumo:

Aspartyl proteases are a class of enzymes that include the yeast aspartyl proteases and secreted aspartyl protease (Sap) superfamilies. Several Sap superfamily members have been demonstrated or suggested as virulence factors in opportunistic pathogens of the genus Candida. Candida albicans, Candida tropicalis, Candida dubliniensis and Candida parapsilosis harbour 10, four, eight and three SAP genes, respectively. In this work, genome mining and phylogenetic analyses revealed the presence of new members of the Sap superfamily in C. tropicalis (8), Candida guilliermondii (8), C. parapsilosis(11) and Candida lusitaniae (3). A total of 12 Sap families, containing proteins with at least 50% similarity, were discovered in opportunistic, pathogenic Candida spp. In several Sap families, at least two subfamilies or orthologous groups were identified, each defined by > 90% sequence similitude, functional similarity and synteny among its members. No new members of previously described Sap families were found in a Candida spp. clinical strain collection; however, the universality of SAPT gene distribution among C. tropicalis strains was demonstrated. In addition, several features of opportunistic pathogenic Candida species, such as gene duplications and inversions, similitude, synteny, putative transcription factor binding sites and genome traits of SAP gene superfamily were described in a molecular evolutionary context.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

The chemistry of natural products has been remarkably growing in the past few decades in Brazil. Aspects related to the isolation and identification of new natural products, as well as their biological activities, have been achieved in different laboratories working on this subject in the country. More recently, the introduction of new molecular biology tools has strongly influenced the research on natural products, mainly those produced by microorganisms, creating new possibilities to assess the chemical diversity of secondary metabolites. This paper describes some ideas on how the research on natural products can have a considerable input from molecular biology in the generation of chemical diversity. We also explore the role of microbial natural products in mediating interspecific interactions and their relevance to ecological studies. Examples of the generation of chemical diversity are highlighted by using genome mining, mutasynthesis, combinatorial biosynthesis, metagenomics, and synthetic biology, while some aspects of microbial ecology are also discussed. The idea to bring up this topic is linked to the remarkable development of molecular biology techniques to generate useful chemicals from different organisms. Here, we focus mainly on microorganisms, even though similar approaches have also been applied to the study of plants and other organisms. Investigations in the frontier of chemistry and biology require interactions between different areas, characterizing the interdisciplinarity of this research field. The necessity of a real integration of chemistry and biology is pivotal to finding correct answers to a number of biological phenomena. The use of molecular biology tools to generate chemical diversity and control biosynthetic pathways is largely explored in the production of important biologically active compounds. Finally, we briefly comment on the Brazilian organization of research in this area, the necessity of new strategies for the graduation programs, and the establishment of networks as a way of organization to overcome some of the problems faced in the area of natural products.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Polyketides and non-ribosomal peptides are natural products widely found in bacteria, fungi and plants. The biological activities associated with these metabolites have attracted special attention in biopharmaceutical studies. Polyketide synthases act similarly to fatty acids synthetases and the whole multi-enzymatic set coordinating precursor and extending unit selection and reduction levels during chain growth. Acting in a similarly orchestrated model, non-ribosomal peptide synthetases biosynthesize NRPs. PKSs-I and NRPSs enzymatic modules and domains are collinearly organized with the parent gene sequence. This arrangement allows the use of degenerated PCR primers to amplify targeted regions in the genes corresponding to specific enzymatic domains such as ketosynthases and acyltransferases in PKSs and adenilation domains in NRPSs. Careful analysis of these short regions allows the classifying of a set of organisms according to their potential to biosynthesize PKs and NRPs. In this work, the biosynthetic potential of a set of 13 endophytic actinobacteria from Citrus reticulata for producing PKs and NRP metabolites was evaluated. The biosynthetic profile was compared to antimicrobial activity. Based on the inhibition promoted, 4 strains were considered for cluster analysis. A PKS/NRPS phylogeny was generated in order to classify some of the representative sequences throughout comparison with homologous genes. Using this approach, a molecular fingerprint was generated to help guide future studies on the most promising strains.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

The chemical ecology and biotechnological potential of metabolites from endophytic and rhizosphere fungi are receiving much attention. A collection of 17 sugarcane-derived fungi were identified and assessed by PCR for the presence of polyketide synthase (PKS) genes. The fungi were all various genera of ascomycetes, the genomes of which encoded 36 putative PKS sequences, 26 shared sequence homology with beta-ketoacyl synthase domains, while 10 sequences showed homology to known fungal C-methyltransferase domains. A neighbour-joining phylogenetic analysis of the translated sequences could group the domains into previously established chemistry-based clades that represented non-reducing, partially reducing and highly reducing fungal PKSs. We observed that, in many cases, the membership of each clade also reflected the taxonomy of the fungal isolates. The functional assignment of the domains was further confirmed by in silico secondary and tertiary protein structure predictions. This genome mining study reveals, for the first time, the genetic potential of specific taxonomic groups of sugarcane-derived fungi to produce specific types of polyketides. Future work will focus on isolating these compounds with a view to understanding their chemical ecology and likely biotechnological potential.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Here, we report the draft genome sequences of three actinobacterial isolates, Micromonospora sp. RV43, Rubrobacter sp. RV113, and Nocardiopsis sp. RV163 that had previously been isolated from Mediterranean sponges. The draft genomes were analyzed for the presence of gene clusters indicative of secondary metabolism using antiSMASH 3.0 and NapDos pipelines. Our findings demonstrated the chemical richness of sponge-associated actinomycetes and the efficacy of genome mining in exploring the genomic potential of sponge-derived actinomycetes.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

We are indebted with Marnix Medema, Paul Straight and Sean Rovito, for useful discussions and critical reading of the manuscript, as well as with Alicia Chagolla and Yolanda Rodriguez of the MS Service of Unidad Irapuato, Cinvestav, and Araceli Fernandez for technical support in high-performance computing. This work was funded by Conacyt Mexico (grants No. 179290 and 177568) and FINNOVA Mexico (grant No. 214716) to FBG. PCM was funded by Conacyt scholarship (No. 28830) and a Cinvestav posdoctoral fellowship. JF and JFK acknowledge funding from the College of Physical Sciences, University of Aberdeen, UK.

Relevância:

40.00% 40.00%

Publicador:

Resumo:

We have sequenced the genome of Desulfosporosinus sp. OT, a Gram-positive, acidophilic sulfate-reducing Firmicute isolated from copper tailing sediment in the Norilsk mining-smelting area in Northern Siberia, Russia. This represents the first sequenced genome of a Desulfosporosinus species. The genome has a size of 5.7 Mb and encodes 6,222 putative proteins.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Somatic copy number aberrations (CNA) represent a mutation type encountered in the majority of cancer genomes. Here, we present the 2014 edition of arrayMap (http://www.arraymap.org), a publicly accessible collection of pre-processed oncogenomic array data sets and CNA profiles, representing a vast range of human malignancies. Since the initial release, we have enhanced this resource both in content and especially with regard to data mining support. The 2014 release of arrayMap contains more than 64,000 genomic array data sets, representing about 250 tumor diagnoses. Data sets included in arrayMap have been assembled from public repositories as well as additional resources, and integrated by applying custom processing pipelines. Online tools have been upgraded for a more flexible array data visualization, including options for processing user provided, non-public data sets. Data integration has been improved by mapping to multiple editions of the human reference genome, with the majority of the data now being available for the UCSC hg18 as well as GRCh37 versions. The large amount of tumor CNA data in arrayMap can be freely downloaded by users to promote data mining projects, and to explore special events such as chromothripsis-like genome patterns.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Ultra-high-throughput sequencing (UHTS) techniques are evolving rapidly and may soon become an affordable and routine tool for sequencing plant DNA, even in smaller plant biology labs. Here we review recent insights into intraspecific genome variation gained from UHTS, which offers a glimpse of the rather unexpected levels of structural variability among Arabidopsis thaliana accessions. The challenges that will need to be addressed to efficiently assemble and exploit this information are also discussed.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

In conducting genome-wide association studies (GWAS), analytical approaches leveraging biological information may further understanding of the pathophysiology of clinical traits. To discover novel associations with estimated glomerular filtration rate (eGFR), a measure of kidney function, we developed a strategy for integrating prior biological knowledge into the existing GWAS data for eGFR from the CKDGen Consortium. Our strategy focuses on single nucleotide polymorphism (SNPs) in genes that are connected by functional evidence, determined by literature mining and gene ontology (GO) hierarchies, to genes near previously validated eGFR associations. It then requires association thresholds consistent with multiple testing, and finally evaluates novel candidates by independent replication. Among the samples of European ancestry, we identified a genome-wide significant SNP in FBXL20 (P = 5.6 × 10(-9)) in meta-analysis of all available data, and additional SNPs at the INHBC, LRP2, PLEKHA1, SLC3A2 and SLC7A6 genes meeting multiple-testing corrected significance for replication and overall P-values of 4.5 × 10(-4)-2.2 × 10(-7). Neither the novel PLEKHA1 nor FBXL20 associations, both further supported by association with eGFR among African Americans and with transcript abundance, would have been implicated by eGFR candidate gene approaches. LRP2, encoding the megalin receptor, was identified through connection with the previously known eGFR gene DAB2 and extends understanding of the megalin system in kidney function. These findings highlight integration of existing genome-wide association data with independent biological knowledge to uncover novel candidate eGFR associations, including candidates lacking known connections to kidney-specific pathways. The strategy may also be applicable to other clinical phenotypes, although more testing will be needed to assess its potential for discovery in general.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

It is common practice in genome-wide association studies (GWAS) to focus on the relationship between disease risk and genetic variants one marker at a time. When relevant genes are identified it is often possible to implicate biological intermediates and pathways likely to be involved in disease aetiology. However, single genetic variants typically explain small amounts of disease risk. Our idea is to construct allelic scores that explain greater proportions of the variance in biological intermediates, and subsequently use these scores to data mine GWAS. To investigate the approach's properties, we indexed three biological intermediates where the results of large GWAS meta-analyses were available: body mass index, C-reactive protein and low density lipoprotein levels. We generated allelic scores in the Avon Longitudinal Study of Parents and Children, and in publicly available data from the first Wellcome Trust Case Control Consortium. We compared the explanatory ability of allelic scores in terms of their capacity to proxy for the intermediate of interest, and the extent to which they associated with disease. We found that allelic scores derived from known variants and allelic scores derived from hundreds of thousands of genetic markers explained significant portions of the variance in biological intermediates of interest, and many of these scores showed expected correlations with disease. Genome-wide allelic scores however tended to lack specificity suggesting that they should be used with caution and perhaps only to proxy biological intermediates for which there are no known individual variants. Power calculations confirm the feasibility of extending our strategy to the analysis of tens of thousands of molecular phenotypes in large genome-wide meta-analyses. We conclude that our method represents a simple way in which potentially tens of thousands of molecular phenotypes could be screened for causal relationships with disease without having to expensively measure these variables in individual disease collections.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Cultivated peanut (Arachis hypogaea) is an important crop, widely grown in tropical and subtropical regions of the world. It is highly susceptible to several biotic and abiotic stresses to which wild species are resistant. As a first step towards the introgression of these resistance genes into cultivated peanut, a linkage map based on microsatellite markers was constructed, using an F-2 population obtained from a cross between two diploid wild species with AA genome (A. duranensis and A. stenosperma). A total of 271 new microsatellite markers were developed in the present study from SSR-enriched genomic libraries, expressed sequence tags (ESTs), and by data-mining sequences available in GenBank. of these, 66 were polymorphic for cultivated peanut. The 271 new markers plus another 162 published for peanut were screened against both progenitors and 204 of these (47.1%) were polymorphic, with 170 codominant and 34 dominant markers. The 80 codominant markers segregating 1:2:1 (P < 0.05) were initially used to establish the linkage groups. Distorted and dominant markers were subsequently included in the map. The resulting linkage map consists of 11 linkage groups covering 1,230.89 cM of total map distance, with an average distance of 7.24 cM between markers. This is the first microsatellite-based map published for Arachis, and the first map based on sequences that are all currently publicly available. Because most markers used were derived from ESTs and genomic libraries made using methylation-sensitive restriction enzymes, about one-third of the mapped markers are genic. Linkage group ordering is being validated in other mapping populations, with the aim of constructing a transferable reference map for Arachis.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The data mining of Eucalyptus ESTs genome finds four clusters (EGCEST2257E11.g, EGBGRT3213F11.g, and EGCCFB1223H11.g) from highly conservative 14-3-3 protein family which modulates a wide variety of cellular processes. Multiple alignments were built from twenty four sequences of 14-3-3 proteins searched into the GenBank databases and into the four pools of Eucalyptus genome programs. The alignment has shown two regions highly conservative on the sequences corresponding to the motifs of protein phosphorylation and nine highly conservative regions on the sequence corresponding to the linkage regions of alpha helices structure based on three dimensional of dimer functional structure. The differences of amino acid into the structural and functional domains of 14-3-3 plant protein were identified and can explain the functional diversity of different isoforms. The phylogenic protein trees were built by the maximum parsimony and neighborjoining procedures of Clustal X alignments and PAUP software for phylogenic analysis.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Conselho Nacional de Desenvolvimento Científico e Tecnológico (CNPq)

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Genome-wide association studies have failed to establish common variant risk for the majority of common human diseases. The underlying reasons for this failure are explained by recent studies of resequencing and comparison of over 1200 human genomes and 10 000 exomes, together with the delineation of DNA methylation patterns (epigenome) and full characterization of coding and noncoding RNAs (transcriptome) being transcribed. These studies have provided the most comprehensive catalogues of functional elements and genetic variants that are now available for global integrative analysis and experimental validation in prospective cohort studies. With these datasets, researchers will have unparalleled opportunities for the alignment, mining, and testing of hypotheses for the roles of specific genetic variants, including copy number variations, single nucleotide polymorphisms, and indels as the cause of specific phenotypes and diseases. Through the use of next-generation sequencing technologies for genotyping and standardized ontological annotation to systematically analyze the effects of genomic variation on humans and model organism phenotypes, we will be able to find candidate genes and new clues for disease’s etiology and treatment. This article describes essential concepts in genetics and genomic technologies as well as the emerging computational framework to comprehensively search websites and platforms available for the analysis and interpretation of genomic data.