Biblioteca Digital

994 resultados para Gene annotation

Post-partum anoestrus in tropical beef cattle: A systems approach combining gene expression and genome-wide association results

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The length of the post-partum anoestrous interval affects reproductive efficiency in many tropical beef cattle herds. In this study, results from genome-wide association studies (Experiment 1: GWAS) and gene expression (Experiment 2: microarray) were combined in a systems approach to reveal genetic markers, genes and pathways underlying the physiology of post-partum anoestrus in tropically adapted cattle. The microarray study measured the expression of 13,964 genes in the hypothalamus of Brahman cows. A total of 366 genes were differentially expressed (DE) in the post-partum period, when acyclic cows were compared to cows that had resumed ovarian cycles. Associated markers (P < 0.05) from a high density GWAS pointed to 2829 genes that were associated with post-partum anoestrous interval (PPAI) in two populations of beef cattle: Brahman and Tropical composite. Together the experiments provided evidence for 63 genes that are likely to influence the resumption of ovulation post-partum in tropically adapted beef cattle. Functional annotation analysis revealed that some of the 63 genes have known roles in hormonal activity, energy balance and neuronal synapse plasticity. Polymorphisms within candidate genes identified by this systems approach could have biological significance in post-partum anoestrus and help select Zebu (Bos indicus) influenced cattle with genetic potential for shorter post-partum anoestrus. Crown Copyright (C) 2014 Published by Elsevier B.V. All rights reserved.

Genome-wide metabolic (re-) annotation of Kluyveromyces lactis

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Background: Even before having its genome sequence published in 2004, Kluyveromyces lactis had long been considered a model organism for studies in genetics and physiology. Research on Kluyveromyces lactis is quite advanced and this yeast species is one of the few with which it is possible to perform formal genetic analysis. Nevertheless, until now, no complete metabolic functional annotation has been performed to the proteins encoded in the Kluyveromyces lactis genome. Results: In this work, a new metabolic genome-wide functional re-annotation of the proteins encoded in the Kluyveromyces lactis genome was performed, resulting in the annotation of 1759 genes with metabolic functions, and the development of a methodology supported by merlin (software developed in-house). The new annotation includes novelties, such as the assignment of transporter superfamily numbers to genes identified as transporter proteins. Thus, the genes annotated with metabolic functions could be exclusively enzymatic (1410 genes), transporter proteins encoding genes (301 genes) or have both metabolic activities (48 genes). The new annotation produced by this work largely surpassed the Kluyveromyces lactis currently available annotations. A comparison with KEGG’s annotation revealed a match with 844 (~90%) of the genes annotated by KEGG, while adding 850 new gene annotations. Moreover, there are 32 genes with annotations different from KEGG. Conclusions: The methodology developed throughout this work can be used to re-annotate any yeast or, with a little tweak of the reference organism, the proteins encoded in any sequenced genome. The new annotation provided by this study offers basic knowledge which might be useful for the scientific community working on this model yeast, because new functions have been identified for the so-called metabolic genes. Furthermore, it served as the basis for the reconstruction of a compartmentalized, genome-scale metabolic model of Kluyveromyces lactis, which is currently being finished.

A clustering method for robust and reliable large scale functional and structural protein sequence annotation

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Bioinformatics, in the last few decades, has played a fundamental role to give sense to the huge amount of data produced. Obtained the complete sequence of a genome, the major problem of knowing as much as possible of its coding regions, is crucial. Protein sequence annotation is challenging and, due to the size of the problem, only computational approaches can provide a feasible solution. As it has been recently pointed out by the Critical Assessment of Function Annotations (CAFA), most accurate methods are those based on the transfer-by-homology approach and the most incisive contribution is given by cross-genome comparisons. In the present thesis it is described a non-hierarchical sequence clustering method for protein automatic large-scale annotation, called “The Bologna Annotation Resource Plus” (BAR+). The method is based on an all-against-all alignment of more than 13 millions protein sequences characterized by a very stringent metric. BAR+ can safely transfer functional features (Gene Ontology and Pfam terms) inside clusters by means of a statistical validation, even in the case of multi-domain proteins. Within BAR+ clusters it is also possible to transfer the three dimensional structure (when a template is available). This is possible by the way of cluster-specific HMM profiles that can be used to calculate reliable template-to-target alignments even in the case of distantly related proteins (sequence identity < 30%). Other BAR+ based applications have been developed during my doctorate including the prediction of Magnesium binding sites in human proteins, the ABC transporters superfamily classification and the functional prediction (GO terms) of the CAFA targets. Remarkably, in the CAFA assessment, BAR+ placed among the ten most accurate methods. At present, as a web server for the functional and structural protein sequence annotation, BAR+ is freely available at http://bar.biocomp.unibo.it/bar2.0.

Investigation of cell fate decision in Schizosaccharomyces pombe by ncRNA annotation and statistical analyses

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Studies in the fission yeast Schizosaccharomyces pombe (S. pombe) have done much to inform the view of heterochromatin and its control by the RNA interference (RNAi) machinery. Using cDNA synthesised from poly(A)-enriched RNA samples, numerous novel ncRNA loci were discovered, and the 50 and 30 ends of many other genes were refined in previous studies. Although some of these transcripts may encode novel proteins the function of the majority is yet to be determined. The authors have used strand-specific deep sequencing of RNA, irrespective of poly(A) status, to reveal a highly structured antisense programme that modulates gene expression to dictate cell fate decisions during sexual differentiation. They show that an extensive and elaborate array of ncRNA production accompanies sexual differentiation in the fission yeast S. pombe. Experimental manipulation suggests that these transcripts specifically regulate the function of the target genes.

Plant protein-coding gene families: emerging bioinformatics approaches

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Protein-coding gene families are sets of similar genes with a shared evolutionary origin and, generally, with similar biological functions. In plants, the size and role of gene families has been only partially addressed. However, suitable bioinformatics tools are being developed to cluster the enormous number of sequences currently available in databases. Specifically, comparative genomic databases promise to become powerful tools for gene family annotation in plant clades. In this review, I evaluate the data retrieved from various gene family databases, the ease with which they can be extracted and how useful the extracted information is.

RefSeq and LocusLink: NCBI gene-centered resources

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Thousands of genes have been painstakingly identified and characterized a few genes at a time. Many thousands more are being predicted by large scale cDNA and genomic sequencing projects, with levels of evidence ranging from supporting mRNA sequence and comparative genomics to computing ab initio models. This, coupled with the burgeoning scientific literature, makes it critical to have a comprehensive directory for genes and reference sequences for key genomes. The NCBI provides two resources, LocusLink and RefSeq, to meet these needs. LocusLink organizes information around genes to generate a central hub for accessing gene-specific information for fruit fly, human, mouse, rat and zebrafish. RefSeq provides reference sequence standards for genomes, transcripts and proteins; human, mouse and rat mRNA RefSeqs, and their corresponding proteins, are discussed here. Together, RefSeq and LocusLink provide a non-redundant view of genes and other loci to support research on genes and gene families, variation, gene expression and genome annotation. Additional information about LocusLink and RefSeq is available at http://www.ncbi.nlm.nih.gov/LocusLink/.

The TIGR Gene Indices: analysis of gene transcript sequences in highly sampled eukaryotic species

Relevância:

30.00% 30.00%

Publicador:

Resumo:

While genome sequencing projects are advancing rapidly, EST sequencing and analysis remains a primary research tool for the identification and categorization of gene sequences in a wide variety of species and an important resource for annotation of genomic sequence. The TIGR Gene Indices (http://www.tigr.org/tdb/tgi.shtml) are a collection of species-specific databases that use a highly refined protocol to analyze EST sequences in an attempt to identify the genes represented by that data and to provide additional information regarding those genes. Gene Indices are constructed by first clustering, then assembling EST and annotated gene sequences from GenBank for the targeted species. This process produces a set of unique, high-fidelity virtual transcripts, or Tentative Consensus (TC) sequences. The TC sequences can be used to provide putative genes with functional annotation, to link the transcripts to mapping and genomic sequence data, to provide links between orthologous and paralogous genes and as a resource for comparative sequence analysis.

Mendel-GFDb and Mendel-ESTS: databases of plant gene families and ESTs annotated with gene family numbers and gene family names

Relevância:

30.00% 30.00%

Publicador:

Resumo:

There is no control over the information provided with sequences when they are deposited in the sequence databases. Consequently mistakes can seed the incorrect annotation of other sequences. Grouping genes into families and applying controlled annotation overcomes the problems of incorrect annotation associated with individual sequences. Two databases (http://www.mendel.ac.uk) were created to apply controlled annotation to plant genes and plant ESTs: Mendel-GFDb is a database of plant protein (gene) families based on gapped-BLAST analysis of all sequences in the SWISS-PROT family of databases. Sequences are aligned (ClustalW) and identical and similar residues shaded. The families are visually curated to ensure that one or more criteria, for example overall relatedness and/or domain similarity relate all sequences within a family. Sequence families are assigned a ‘Gene Family Number’ and a unified description is developed which best describes the family and its members. If authority exists the gene family is assigned a ‘Gene Family Name’. This information is placed in Mendel-GFDb. Mendel-ESTS is primarily a database of plant ESTs, which have been compared to Mendel-GFDb, completely sequenced genomes and domain databases. This approach associated ESTs with individual sequences and the controlled annotation of gene families and protein domains; the information being placed in Mendel-ESTS. The controlled annotation applied to genes and ESTs provides a basis from which a plant transcription database can be developed.

Comparative organization of cattle chromosome 5 revealed by comparative mapping by annotation and sequence similarity and radiation hybrid mapping

Relevância:

30.00% 30.00%

Publicador:

Resumo:

A whole genome cattle-hamster radiation hybrid cell panel was used to construct a map of 54 markers located on bovine chromosome 5 (BTA5). Of the 54 markers, 34 are microsatellites selected from the cattle linkage map and 20 are genes. Among the 20 mapped genes, 10 are new assignments that were made by using the comparative mapping by annotation and sequence similarity strategy. A LOD-3 radiation hybrid framework map consisting of 21 markers was constructed. The relatively low retention frequency of markers on this chromosome (19%) prevented unambiguous ordering of the other 33 markers. The length of the map is 398.7 cR, corresponding to a ratio of ≈2.8 cR5,000/cM. Type I genes were binned for comparison of gene order among cattle, humans, and mice. Multiple internal rearrangements within conserved syntenic groups were apparent upon comparison of gene order on BTA5 and HSA12 and HSA22. A similarly high number of rearrangements were observed between BTA5 and MMU6, MMU10, and MMU15. The detailed comparative map of BTA5 should facilitate identification of genes affecting economically important traits that have been mapped to this chromosome and should contribute to our understanding of mammalian chromosome evolution.

Development and evaluation of an automated annotation pipeline and cDNA annotation system

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Manual curation has long been held to be the gold standard for functional annotation of DNA sequence. Our experience with the annotation of more than 20,000 full-length cDNA sequences revealed problems with this approach, including inaccurate and inconsistent assignment of gene names, as well as many good assignments that were difficult to reproduce using only computational methods. For the FANTOM2 annotation of more than 60,000 cDNA clones, we developed a number of methods and tools to circumvent some of these problems, including an automated annotation pipeline that provides high-quality preliminary annotation for each sequence by introducing an uninformative filter that eliminates uninformative annotations, controlled vocabularies to accurately reflect both the functional assignments and the evidence supporting them, and a highly refined, Web-based manual annotation tool that allows users to view a wide array of sequence analyses and to assign gene names and putative functions using a consistent nomenclature. The ultimate utility of our approach is reflected in the low rate of reassignment of automated assignments by manual curation. Based on these results, we propose a new standard for large-scale annotation, in which the initial automated annotations are manually investigated and then computational methods are iteratively modified and improved based on the results of manual curation.

A Mixture model with random-effects components for clustering correlated gene-expression profiles

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Motivation: The clustering of gene profiles across some experimental conditions of interest contributes significantly to the elucidation of unknown gene function, the validation of gene discoveries and the interpretation of biological processes. However, this clustering problem is not straightforward as the profiles of the genes are not all independently distributed and the expression levels may have been obtained from an experimental design involving replicated arrays. Ignoring the dependence between the gene profiles and the structure of the replicated data can result in important sources of variability in the experiments being overlooked in the analysis, with the consequent possibility of misleading inferences being made. We propose a random-effects model that provides a unified approach to the clustering of genes with correlated expression levels measured in a wide variety of experimental situations. Our model is an extension of the normal mixture model to account for the correlations between the gene profiles and to enable covariate information to be incorporated into the clustering process. Hence the model is applicable to longitudinal studies with or without replication, for example, time-course experiments by using time as a covariate, and to cross-sectional experiments by using categorical covariates to represent the different experimental classes. Results: We show that our random-effects model can be fitted by maximum likelihood via the EM algorithm for which the E(expectation) and M(maximization) steps can be implemented in closed form. Hence our model can be fitted deterministically without the need for time-consuming Monte Carlo approximations. The effectiveness of our model-based procedure for the clustering of correlated gene profiles is demonstrated on three real datasets, representing typical microarray experimental designs, covering time-course, repeated-measurement and cross-sectional data. In these examples, relevant clusters of the genes are obtained, which are supported by existing gene-function annotation. A synthetic dataset is considered too.

Definition and spatial annotation of the dynamic secretome during early kidney development

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The term secretome has been defined as a set of secreted proteins (Grimmond et al. [2003] Genome Res 13:1350-1359). The term secreted protein encompasses all proteins exported from the cell including growth factors, extracellular proteinases, morphogens, and extracellular matrix molecules. Defining the genes encoding secreted proteins that change in expression during organogenesis, the dynamic secretome, is likely to point to key drivers of morphogenesis. Such secreted proteins are involved in the reciprocal interactions between the ureteric bud (UB) and the metanephric mesenchyme (AM) that occur during organogenesis of the metanephros. Some key metanephric secreted proteins have been identified, but many remain to be determined. In this study, microarray expression profiling of E10.5, E11.5, and E13.5 kidney and consensus bioinformatic analysis were used to define a dynamic secretome of early metanephric development. In situ hybridisation was used to confirm microarray results and clarify spatial expression patterns for these genes. Forty-one secreted factors were dynamically expressed between the E10.5 and E13.5 timeframe profiled, and 25 of these factors had not previously been implicated in kidney development. A text-based anatomical ontology was used to spatially annotate the expression pattern of these genes in cultured metanephric explants.

Evolutionary conservation and murine embryonic expression of the gene encoding the SERTA domain-containing protein CDCA4 (HEPP)

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Cdca4 (Hepp) was originally identified as a gene expressed specifically in hematopoietic progenitor cells as opposed to hematopoietic stem cells. More recently, it has been shown to stimulate p53 activity and also lead to p53-independent growth inhibition when overexpressed. We independently isolated the murine Cdca4 gene in a genomic expression-based screen for genes involved in mammalian craniofacial development, and show that Cdca4 is expressed in a spatio-temporally restricted pattern during mouse embryogenesis. In addition to expression in the facial primordia including the pharyngeal arches, Cdca4 is expressed in the developing limb buds, brain, spinal cord, dorsal root ganglia, teeth, eye and hair follicles. Along with a small number of proteins from a range of species, the predicted CDCA4 protein contains a novel SERTA motif in addition to cyclin A-binding and PHD bromodomain-binding regions of homology. While the function of the SERTA domain is unknown, proteins containing this domain have previously been linked to cell cycle progression and chromatin remodelling. Using in silico database mining we have extended the number of evolutionarily conserved orthologues of known SERTA domain proteins and identified an uncharacterised member of the SERTA domain family, SERTAD4, with orthologues to date in human, mouse, rat, dog, cow, Tetraodon and chicken. Immunolocalisation of transiently and stably transfected epitope-tagged CDCA4 protein in mammalian cells suggests that it resides predominantly in the nucleus throughout all stages of the cell cycle. (c) 2006 Elsevier B.V. All rights reserved.

Bioinformatic analysis for exploring relationships between genes and gene products

Relevância:

30.00% 30.00%

Publicador:

Resumo:

To carry out their specific roles in the cell, genes and gene products often work together in groups, forming many relationships among themselves and with other molecules. Such relationships include physical protein-protein interaction relationships, regulatory relationships, metabolic relationships, genetic relationships, and much more. With advances in science and technology, some high throughput technologies have been developed to simultaneously detect tens of thousands of pairwise protein-protein interactions and protein-DNA interactions. However, the data generated by high throughput methods are prone to noise. Furthermore, the technology itself has its limitations, and cannot detect all kinds of relationships between genes and their products. Thus there is a pressing need to investigate all kinds of relationships and their roles in a living system using bioinformatic approaches, and is a central challenge in Computational Biology and Systems Biology. This dissertation focuses on exploring relationships between genes and gene products using bioinformatic approaches. Specifically, we consider problems related to regulatory relationships, protein-protein interactions, and semantic relationships between genes. A regulatory element is an important pattern or "signal", often located in the promoter of a gene, which is used in the process of turning a gene "on" or "off". Predicting regulatory elements is a key step in exploring the regulatory relationships between genes and gene products. In this dissertation, we consider the problem of improving the prediction of regulatory elements by using comparative genomics data. With regard to protein-protein interactions, we have developed bioinformatics techniques to estimate support for the data on these interactions. While protein-protein interactions and regulatory relationships can be detected by high throughput biological techniques, there is another type of relationship called semantic relationship that cannot be detected by a single technique, but can be inferred using multiple sources of biological data. The contributions of this thesis involved the development and application of a set of bioinformatic approaches that address the challenges mentioned above. These included (i) an EM-based algorithm that improves the prediction of regulatory elements using comparative genomics data, (ii) an approach for estimating the support of protein-protein interaction data, with application to functional annotation of genes, (iii) a novel method for inferring functional network of genes, and (iv) techniques for clustering genes using multi-source data.

Transcriptional dynamics of the developing sweet cherry (Prunus avium L.) fruit: sequencing, annotation and expression profiling of exocarp-associated genes

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The exocarp, or skin, of fleshy fruit is a specialized tissue that protects the fruit, attracts seed dispersing fruit eaters, and has large economical relevance for fruit quality. Development of the exocarp involves regulated activities of many genes. This research analyzed global gene expression in the exocarp of developing sweet cherry (Prunus avium L., 'Regina'), a fruit crop species with little public genomic resources. A catalog of transcript models (contigs) representing expressed genes was constructed from de novo assembled short complementary DNA (cDNA) sequences generated from developing fruit between flowering and maturity at 14 time points. Expression levels in each sample were estimated for 34 695 contigs from numbers of reads mapping to each contig. Contigs were annotated functionally based on BLAST, gene ontology and InterProScan analyses. Coregulated genes were detected using partitional clustering of expression patterns. The results are discussed with emphasis on genes putatively involved in cuticle deposition, cell wall metabolism and sugar transport. The high temporal resolution of the expression patterns presented here reveals finely tuned developmental specialization of individual members of gene families. Moreover, the de novo assembled sweet cherry fruit transcriptome with 7760 full-length protein coding sequences and over 20 000 other, annotated cDNA sequences together with their developmental expression patterns is expected to accelerate molecular research on this important tree fruit crop.

«
1
2
3
4
5
6
7
8
...
66
67
»