873 resultados para Motif Discovery
Resumo:
Understanding the machinery of gene regulation to control gene expression has been one of the main focuses of bioinformaticians for years. We use a multi-objective genetic algorithm to evolve a specialized version of side effect machines for degenerate motif discovery. We compare some suggested objectives for the motifs they find, test different multi-objective scoring schemes and probabilistic models for the background sequence models and report our results on a synthetic dataset and some biological benchmarking suites. We conclude with a comparison of our algorithm with some widely used motif discovery algorithms in the literature and suggest future directions for research in this area.
Resumo:
MHC class II (MHCII) genes are transactivated by the NOD-like receptor (NLR) family member CIITA, which is recruited to SXY enhancers of MHCII promoters via a DNA-binding "enhanceosome" complex. NLRC5, another NLR protein, was recently found to control transcription of MHC class I (MHCI) genes. However, detailed understanding of NLRC5's target gene specificity and mechanism of action remained lacking. We performed ChIP-sequencing experiments to gain comprehensive information on NLRC5-regulated genes. In addition to classical MHCI genes, we exclusively identified novel targets encoding non-classical MHCI molecules having important functions in immunity and tolerance. ChIP-sequencing performed with Rfx5(-/-) cells, which lack the pivotal enhanceosome factor RFX5, demonstrated its strict requirement for NLRC5 recruitment. Accordingly, Rfx5-knockout mice phenocopy Nlrc5 deficiency with respect to defective MHCI expression. Analysis of B cell lines lacking RFX5, RFXAP, or RFXANK further corroborated the importance of the enhanceosome for MHCI expression. Although recruited by common DNA-binding factors, CIITA and NLRC5 exhibit non-redundant functions, shown here using double-deficient Nlrc5(-/-)CIIta(-/-) mice. These paradoxical findings were resolved by using a "de novo" motif-discovery approach showing that the SXY consensus sequence occupied by NLRC5 in vivo diverges significantly from that occupied by CIITA. These sequence differences were sufficient to determine preferential occupation and transactivation by NLRC5 or CIITA, respectively, and the S box was found to be the essential feature conferring NLRC5 specificity. These results broaden our knowledge on the transcriptional activities of NLRC5 and CIITA, revealing their dependence on shared enhanceosome factors but their recruitment to distinct enhancer motifs in vivo. Furthermore, we demonstrated selectivity of NLRC5 for genes encoding MHCI or related proteins, rendering it an attractive target for therapeutic intervention. NLRC5 and CIITA thus emerge as paradigms for a novel class of transcriptional regulators dedicated for transactivating extremely few, phylogenetically related genes.
Resumo:
Les facteurs de transcription sont des protéines spécialisées qui jouent un rôle important dans différents processus biologiques tel que la différenciation, le cycle cellulaire et la tumorigenèse. Ils régulent la transcription des gènes en se fixant sur des séquences d’ADN spécifiques (éléments cis-régulateurs). L’identification de ces éléments est une étape cruciale dans la compréhension des réseaux de régulation des gènes. Avec l’avènement des technologies de séquençage à haut débit, l’identification de tout les éléments fonctionnels dans les génomes, incluant gènes et éléments cis-régulateurs a connu une avancée considérable. Alors qu’on est arrivé à estimer le nombre de gènes chez différentes espèces, l’information sur les éléments qui contrôlent et orchestrent la régulation de ces gènes est encore mal définie. Grace aux techniques de ChIP-chip et de ChIP-séquençage il est possible d’identifier toutes les régions du génome qui sont liées par un facteur de transcription d’intérêt. Plusieurs approches computationnelles ont été développées pour prédire les sites fixés par les facteurs de transcription. Ces approches sont classées en deux catégories principales: les algorithmes énumératifs et probabilistes. Toutefois, plusieurs études ont montré que ces approches génèrent des taux élevés de faux négatifs et de faux positifs ce qui rend difficile l’interprétation des résultats et par conséquent leur validation expérimentale. Dans cette thèse, nous avons ciblé deux objectifs. Le premier objectif a été de développer une nouvelle approche pour la découverte des sites de fixation des facteurs de transcription à l’ADN (SAMD-ChIP) adaptée aux données de ChIP-chip et de ChIP-séquençage. Notre approche implémente un algorithme hybride qui combine les deux stratégies énumérative et probabiliste, afin d’exploiter les performances de chacune d’entre elles. Notre approche a montré ses performances, comparée aux outils de découvertes de motifs existants sur des jeux de données simulées et des jeux de données de ChIP-chip et de ChIP-séquençage. SAMD-ChIP présente aussi l’avantage d’exploiter les propriétés de distributions des sites liés par les facteurs de transcription autour du centre des régions liées afin de limiter la prédiction aux motifs qui sont enrichis dans une fenêtre de longueur fixe autour du centre de ces régions. Les facteurs de transcription agissent rarement seuls. Ils forment souvent des complexes pour interagir avec l’ADN pour réguler leurs gènes cibles. Ces interactions impliquent des facteurs de transcription dont les sites de fixation à l’ADN sont localisés proches les uns des autres ou bien médier par des boucles de chromatine. Notre deuxième objectif a été d’exploiter la proximité spatiale des sites liés par les facteurs de transcription dans les régions de ChIP-chip et de ChIP-séquençage pour développer une approche pour la prédiction des motifs composites (motifs composés par deux sites et séparés par un espacement de taille fixe). Nous avons testé ce module pour prédire la co-localisation entre les deux demi-sites ERE qui forment le site ERE, lié par le récepteur des œstrogènes ERα. Ce module a été incorporé à notre outil de découverte de motifs SAMD-ChIP.
Resumo:
Background: Current methods to find significantly under- and over-represented gene ontology (GO) terms in a set of genes consider the genes as equally probable balls in a bag, as may be appropriate for transcripts in micro-array data. However, due to the varying length of genes and intergenic regions, that approach is inappropriate for deciding if any GO terms are correlated with a set of genomic positions. Results: We present an algorithm - GONOME - that can determine which GO terms are significantly associated with a set of genomic positions given a genome annotated with (at least) the starts and ends of genes. We show that certain GO terms may appear to be significantly associated with a set of randomly chosen positions in the human genome if gene lengths are not considered, and that these same terms have been reported as significantly over-represented in a number of recent papers. This apparent over-representation disappears when gene lengths are considered, as GONOME does. For example, we show that, when gene length is taken into account, the term development is not significantly enriched in genes associated with human CpG islands, in contradiction to a previous report. We further demonstrate the efficacy of GONOME by showing that occurrences of the proteosome-associated control element (PACE) upstream activating sequence in the S. cerevisiae genome associate significantly to appropriate GO terms. An extension of this approach yields a whole-genome motif discovery algorithm that allows identification of many other promoter sequences linked to different types of genes, including a large group of previously unknown motifs significantly associated with the terms 'translation' and 'translational elongation'. Conclusion: GONOME is an algorithm that correctly extracts over-represented GO terms from a set of genomic positions. By explicitly considering gene size, GONOME avoids a systematic bias toward GO terms linked to large genes. Inappropriate use of existing algorithms that do not take gene size into account has led to erroneous or suspect conclusions. Reciprocally GONOME may be used to identify new features in genomes that are significantly associated with particular categories of genes.
Resumo:
Cyclotides are a fascinating family of plant-derived peptides characterized by their head-to-tail cyclized backbone and knotted arrangement of three disulfide bonds. This conserved structural architecture, termed the CCK (cyclic cystine knot), is responsible for their exceptional resistance to thermal, chemical and enzymatic degradation. Cyclotides have a variety of biological activities, but their insecticidal activities suggest that their primary function is in plant defence. In the present study, we determined the cyclotide content of the sweet violet Viola odorata, a member of the Violaceae family. We identified 30 cyclotides from the aerial parts and roots of this plant, 13 of which are novel sequences. The new sequences provide information about the natural diversity of cyclotides and the role of particular residues in defining structure and function. As many of the biological activities of cyclotides appear to be associated with membrane interactions, we used haemolytic activity as a marker of bioactivity for a selection of the new cyclotides. The new cyclotides were tested for their ability to resist proteolysis by a range of enzymes and, in common with other cyclotides, were completely resistant to trypsin, pepsin and thermolysin. The results show that while biological activity varies with the sequence, the proteolytic stability of the framework does not, and appears to be an inherent feature of the cyclotide framework. The structure of one of the new cyclotides, cycloviolacin O14, was determined and shown to contain the CCK motif. This study confirms that cyclotides may be regarded as a natural combinatorial template that displays a variety of peptide epitopes most likely targeted to a range of plant pests and pathogens.
Resumo:
Protein α-helical coiled coil structures that elicit antibody responses, which block critical functions of medically important microorganisms, represent a means for vaccine development. By using bioinformatics algorithms, a total of 50 antigens with α-helical coiled coil motifs orthologous to Plasmodium falciparum were identified in the P. vivax genome. The peptides identified in silico were chemically synthesized; circular dichroism studies indicated partial or high α-helical content. Antigenicity was evaluated using human sera samples from malaria-endemic areas of Colombia and Papua New Guinea. Eight of these fragments were selected and used to assess immunogenicity in BALB/c mice. ELISA assays indicated strong reactivity of serum samples from individuals residing in malaria-endemic regions and sera of immunized mice, with the α-helical coiled coil structures. In addition, ex vivo production of IFN-γ by murine mononuclear cells confirmed the immunogenicity of these structures and the presence of T-cell epitopes in the peptide sequences. Moreover, sera of mice immunized with four of the eight antigens recognized native proteins on blood-stage P. vivax parasites, and antigenic cross-reactivity with three of the peptides was observed when reacted with both the P. falciparum orthologous fragments and whole parasites. Results here point to the α-helical coiled coil peptides as possible P. vivax malaria vaccine candidates as were observed for P. falciparum. Fragments selected here warrant further study in humans and non-human primate models to assess their protective efficacy as single components or assembled as hybrid linear epitopes.
Resumo:
Several macrocyclic peptides (similar to 30 amino acids), with diverse biological activities, have been isolated from the Rubiaceae and Violaceae plant families over recent years. We have significantly expanded the range of known macrocyclic peptides with the discovery of 16 novel peptides from extracts of Viola hederaceae, Viola odorata and Oldenlandia affinis. The Viola plants had not previously been examined for these peptides and thus represent novel species in which these unusual macrocyclic peptides are produced. Further, we have determined the three-dimensional struc ture of one of these novel peptides, cycloviolacin O1, using H-1 NMR spectroscopy. The structure consists of a distorted triple-stranded beta-sheet and a cystine-knot arrangement of the disulfide bonds. This structure is similar to kalata B1 and circulin A, the only two macrocyclic peptides for which a structure was available, suggesting that despite the sequence variation throughout the peptides they form a family in which the overall fold is conserved. We refer to these peptides as the cyclotide family and their embedded topology as the cyclic cystine knot (CCK) motif. The unique cyclic and knotted nature of these molecules makes them a fascinating example of topologically complex proteins. Examination of the sequences reveals they can be separated into two subfamilies, one of which tends to contain a larger number of positively charged residues and has a bracelet-like circularization of the backbone. The second subfamily contains a backbone twist due to a cis-Pro peptide bond and may conceptually be regarded as a molecular Moebius strip. Here we define the structural features of the two apparent subfamilies of the CCK peptides which may be significant for the likely defense related role of these peptides within plants. (C) 1999 Academic Press.
Resumo:
Circular disulfide-rich polypeptides were unknown a decade ago but over recent years a large family of such molecules has been discovered, which we now refer to as the cyclotides. They are typically about 30 amino acids in size, contain an N- to C-cyclised backbone and incorporate three disulfide bonds arranged in a cystine knot motif. In this motif, an embedded ring in the structure formed by two disulfide bonds and their connecting backbone segments is penetrated by the third disulfide bond. The combination of this knotted and strongly braced structure with a circular backbone renders the cyclotides impervious to enzymatic breakdown and makes them exceptionally stable. This article describes the discovery of the cyclotides in plants from the Rubiaceae and Violaceae families, their chemical synthesis, folding, structural characterisation, and biosynthetic origin. The cyclotides have a diverse range of biological applications, ranging from uterotonic action, to anti-HIV and neurotensin antagonism. Certain plants from which they are derived have a history of uses in native medicine, with activity being observed after oral ingestion of a tea made from the plants. This suggests the possibility that the cyclotides may be orally bioavailable. They therefore have a range of potential applications as a stable peptide framework.
Rapid identification of malaria vaccine candidates based on alpha-helical coiled coil protein motif.
Resumo:
To identify malaria antigens for vaccine development, we selected alpha-helical coiled coil domains of proteins predicted to be present in the parasite erythrocytic stage. The corresponding synthetic peptides are expected to mimic structurally "native" epitopes. Indeed the 95 chemically synthesized peptides were all specifically recognized by human immune sera, though at various prevalence. Peptide specific antibodies were obtained both by affinity-purification from malaria immune sera and by immunization of mice. These antibodies did not show significant cross reactions, i.e., they were specific for the original peptide, reacted with native parasite proteins in infected erythrocytes and several were active in inhibiting in vitro parasite growth. Circular dichroism studies indicated that the selected peptides assumed partial or high alpha-helical content. Thus, we demonstrate that the bioinformatics/chemical synthesis approach described here can lead to the rapid identification of molecules which target biologically active antibodies, thus identifying suitable vaccine candidates. This strategy can be, in principle, extended to vaccine discovery in a wide range of other pathogens.
Resumo:
A combined computational and experimental polymorph search was undertaken to establish the crystal forms of 7-fluoroisatin, a simple molecule with no reported crystal structures, to evaluate the value of crystal structure prediction studies as an aid to solid form discovery. Three polymorphs were found in a manual crystallisation screen, as well as two solvates. Form I ( P2(1)/c, Z0 1), found from the majority of solvent evaporation experiments, corresponded to the most stable form in the computational search of Z0 1 structures. Form III ( P21/ a, Z0 2) is probably a metastable form, which was only found concomitantly with form I, and has the same dimeric R2 2( 8) hydrogen bonding motif as form I and the majority of the computed low energy structures. However, the most thermodynamically stable polymorph, form II ( P1 , Z0 2), has an expanded four molecule R 4 4( 18) hydrogen bonding motif, which could not have been found within the routine computational study. The computed relative energies of the three forms are not in accord with experimental results. Thus, the experimental finding of three crystalline polymorphs of 7- fluoroisatin illustrates the many challenges for computational screening to be a tool for the experimental crystal engineer, in contrast to the results for an analogous investigation of 5- fluoroisatin.
Resumo:
The cyclotides are a family of small disulfide rich proteins that have a cyclic peptide backbone and a cystine knot formed by three conserved disulfide bonds. The combination of these two structural motifs contributes to the exceptional chemical, thermal and enzymatic stability of the cyclotides, which retain bioactivity after boiling. They were initially discovered based on native medicine or screening studies associated with some of their various activities, which include uterotonic action, anti-HIV activity, neurotensin antagonism, and cytotoxicity. They are present in plants from the Rubiaceae, Violaceae and Cucurbitaccae families and their natural function in plants appears to be in host defense: they have potent activity against certain insect pests and they also have antimicrobial activity. There are currently around 50 published sequences of cyclotides and their rate of discovery has been increasing over recent years. Ultimately the family may comprise thousands of members. This article describes the background to the discovery of the cyclotides, their structural characterization, chemical synthesis, genetic origin, biological activities and potential applications in the pharmaceutical and agricultural industries. Their unique topological features make them interesting from a protein folding perspective. Because of their highly stable peptide framework they might make useful templates in drug design programs, and their insecticidal activity opens the possibility of applications in crop protection.
Resumo:
This project identified a novel family of six 66-68 residue peptides from the venom of two Australian funnel-web spiders, Hadronyche sp. 20 and H. infensa: Orchid Beach (Hexathelidae: Atracinae), that appear to undergo N- and/or C-terminal post-translational modifications and conform to an ancestral protein fold. These peptides all show significant amino acid sequence homology to atracotoxin-Hvf17 (ACTX-Hvf17), a non-toxic peptide isolated from the venom of H. versuta, and a variety of AVIT family proteins including mamba intestinal toxin 1 (MIT1) and its mammalian and piscine orthologs prokineticin 1 (PK1) and prokineticin 2 PK2). These AVIT family proteins target prokineticin receptors involved in the sensitization of nociceptors and gastrointestinal smooth muscle activation. Given their sequence homology to MITI, we have named these spider venom peptides the MIT-like atracotoxin (ACTX) family. Using isolated rat stomach fundus or guinea-pia ileum organ bath preparations we have shown that the prototypical ACTX-Hvf17, at concentrations up to 1 mu M, did not stimulate smooth muscle contractility, nor did it inhibit contractions induced by human PK1 (hPK1). The peptide also lacked activity on other isolated smooth muscle preparations including rat aorta. Furthermore, a FLIPR Ca2+ flux assay using HEK293 cells expressing prokineticin receptors showed that ACTX-Hvf17 fails to activate or block hPK1 or hPK2 receptors. Therefore, while the MIT-like ACTX family appears to adopt the ancestral disulfide-directed beta-hairpin protein fold of MIT1, a motif believed to be shared by other AVIT family peptides, variations in the amino acid sequence and surface charge result in a loss of activity on prokineticin receptors. (c) 2005 Elsevier Inc. All rights reserved.
Resumo:
Cyclotides are mini-proteins of 28-37 amino acid residues that have the unusual feature of a head-to-tail cyclic backbone surrounding a cystine knot. This molecular architecture gives the cyclotides heightened resistance to thermal, chemical and enzymatic degradation and has prompted investigations into their use as scaffolds in peptide therapeutics. There are now more than 80 reported cyclotide sequences from plants in the families Rubiaceae, Violaceae and Cucurbitaceae, with a wide variety of biological activities observed. However, potentially limiting the development of cyclotide-based therapeutics is a lack of understanding of the mechanism by which these peptides are cyclized in vivo. Until now, no linear versions of cyclotides have been reported, limiting our understanding of the cyclization mechanism. This study reports the discovery of a naturally occurring linear cyclotide, violacin A, from the plant Viola odorata and discusses the implications for in vivo cyclization of peptides. The elucidation of the cDNA clone of violacin A revealed a point mutation that introduces a stop codon, which inhibits the translation of a key Asn residue that is thought to be required for cyclization. The three-dimensional solution structure of violacin A was determined and found to adopt the cystine knot fold of native cyclotides. Enzymatic stability assays on violacin A indicate that despite an increase in the flexibility of the structure relative to cyclic counterparts, the cystine knot preserves the overall stability of the molecule. (c) 2006 Elsevier Ltd. All rights reserved.
Resumo:
Cyclotides are peptides from plants of the Rubiaceae and Violaceae families that have the unusual characteristic of a macrocylic backbone. They are further characterized by their incorporation of a cystine knot in which two disulfides, along with the intervening backbone residues, form a ring through which a third disulfide is threaded. The cyclotides have been found in every Violaceae species screened to date but are apparently present in only a few Rubiaceae species. The selective distribution reported so far raises questions about the evolution of the cyclotides within the plant kingdom. In this study, we use a combined bioinformatics and expression analysis approach to elucidate the evolution and distribution of the cyclotides in the plant kingdom and report the discovery of related sequences widespread in the Poaceae family, including crop plants such as rice ( Oryza sativa), maize ( Zea mays), and wheat ( Triticum aestivum), which carry considerable economic and social importance. The presence of cyclotide-like sequences within these plants suggests that the cyclotides may be derived from an ancestral gene of great antiquity. Quantitative RT-PCR was used to show that two of the discovered cyclotide-like genes from rice and barley ( Hordeum vulgare) have tissue-specific expression patterns.
Resumo:
The cyclotides are a family of circular proteins with a range of biological activities and potential pharmaceutical and agricultural applications. The biosynthetic mechanism of cyclization is unknown and the discovery of novel sequences may assist in achieving this goal. In the present study, we have isolated a new cyclotide from Oldenlandia affinis, kalata B8, which appears to be a hybrid of the two major subfamilies (Mobius and bracelet) of currently known cyclotides. We have determined the three-dimensional structure of kalata B8 and observed broadening of resonances directly involved in the cystine knot motif, suggesting flexibility in this region despite it being the core structural element of the cyclotides. The cystine knot motif is widespread throughout Nature and inherently stable, making this apparent flexibility a surprising result. Further-more, there appears to be isomerization of the peptide backbone at an Asp-Gly sequence in the region involved in the cyclization process. Interestingly, such isomerization has been previously characterized in related cyclic knottins from Momordica cochinchinensis that have no sequence similarity to kalata B8 apart from the six conserved cysteine residues and may result from a common mechanism of cyclization. Kalata B8 also provides insight into the structure-activity relationships of cyclotides as it displays anti-HIV activity but lacks haemolytic activity. The 'uncoupling' of these two activities has not previously been observed for the cyclotides and may be related to the unusual hydrophilic nature of the peptide.