979 resultados para Protein evolution
Resumo:
BACKGROUND: Along the chromosome of the obligate intracellular bacteria Protochlamydia amoebophila UWE25, we recently described a genomic island Pam100G. It contains a tra unit likely involved in conjugative DNA transfer and lgrE, a 5.6-kb gene similar to five others of P. amoebophila: lgrA to lgrD, lgrF. We describe here the structure, regulation and evolution of these proteins termed LGRs since encoded by "Large G+C-Rich" genes. RESULTS: No homologs to the whole protein sequence of LGRs were found in other organisms. Phylogenetic analyses suggest that serial duplications producing the six LGRs occurred relatively recently and nucleotide usage analyses show that lgrB, lgrE and lgrF were relocated on the chromosome. The C-terminal part of LGRs is homologous to Leucine-Rich Repeats domains (LRRs). Defined by a cumulative alignment score, the 5 to 18 concatenated octacosapeptidic (28-meric) LRRs of LGRs present all a predicted alpha-helix conformation. Their closest homologs are the 28-residue RI-like LRRs of mammalian NODs and the 24-meres of some Ralstonia and Legionella proteins. Interestingly, lgrE, which is present on Pam100G like the tra operon, exhibits Pfam domains related to DNA metabolism. CONCLUSION: Comparison of the LRRs, enable us to propose a parsimonious evolutionary scenario of these domains driven by adjacent concatenations of LRRs. Our model established on bacterial LRRs can be challenged in eucaryotic proteins carrying less conserved LRRs, such as NOD proteins and Toll-like receptors.
Resumo:
Gene duplication was prevalent during hominoid evolution, yet little is known about the functional fate of new ape gene copies. We characterized the CDC14B cell cycle gene and the functional evolution of its hominoid-specific daughter gene, CDC14Bretro. We found that CDC14B encodes four different splice isoforms that show different subcellular localizations (nucleus or microtubule-associated) and functional properties. A microtubular CDC14B variant spawned CDC14Bretro through retroposition in the hominoid ancestor 18-25 million years ago (Mya). CDC14Bretro evolved brain-/testis-specific expression after the duplication event and experienced a short period of intense positive selection in the African ape ancestor 7-12 Mya. Using resurrected ancestral protein variants, we demonstrate that by virtue of amino acid substitutions in distinct protein regions during this time, the subcellular localization of CDC14Bretro progressively shifted from the association with microtubules (stabilizing them) to an association with the endoplasmic reticulum. CDC14Bretro evolution represents a paradigm example of rapid, selectively driven subcellular relocalization, thus revealing a novel mode for the emergence of new gene function
Resumo:
INTRODUCTION: Oxidative stress is involved in the development of secondary tissue damage and organ failure. Micronutrients contributing to the antioxidant (AOX) defense exhibit low plasma levels during critical illness. The aim of this study was to investigate the impact of early AOX micronutrients on clinical outcome in intensive care unit (ICU) patients with conditions characterized by oxidative stress. METHODS: We conducted a prospective, randomized, double-blind, placebo-controlled, single-center trial in patients admitted to a university hospital ICU with organ failure after complicated cardiac surgery, major trauma, or subarachnoid hemorrhage. Stratification by diagnosis was performed before randomization. The intervention was intravenous supplements for 5 days (selenium 270 microg, zinc 30 mg, vitamin C 1.1 g, and vitamin B1 100 mg) with a double-loading dose on days 1 and 2 or placebo. RESULTS: Two hundred patients were included (102 AOX and 98 placebo). While age and gender did not differ, brain injury was more severe in the AOX trauma group (P = 0.019). Organ function endpoints did not differ: incidence of acute kidney failure and sequential organ failure assessment score decrease were similar (-3.2 +/- 3.2 versus -4.2 +/- 2.3 over the course of 5 days). Plasma concentrations of selenium, zinc, and glutathione peroxidase, low on admission, increased significantly to within normal values in the AOX group. C-reactive protein decreased faster in the AOX group (P = 0.039). Infectious complications did not differ. Length of hospital stay did not differ (16.5 versus 20 days), being shorter only in surviving AOX trauma patients (-10 days; P = 0.045). CONCLUSION: The AOX intervention did not reduce early organ dysfunction but significantly reduced the inflammatory response in cardiac surgery and trauma patients, which may prove beneficial in conditions with an intense inflammation. TRIALS REGISTRATION: Clinical Trials.gov RCT Register: NCT00515736.
Resumo:
BACKGROUND: Complete mitochondrial genome sequences have become important tools for the study of genome architecture, phylogeny, and molecular evolution. Despite the rapid increase in available mitogenomes, the taxonomic sampling often poorly reflects phylogenetic diversity and is often also biased to represent deeper (family-level) evolutionary relationships. RESULTS: We present the first fully sequenced ant (Hymenoptera: Formicidae) mitochondrial genomes. We sampled four mitogenomes from three species of fire ants, genus Solenopsis, which represent various evolutionary depths. Overall, ant mitogenomes appear to be typical of hymenopteran mitogenomes, displaying a general A+T-bias. The Solenopsis mitogenomes are slightly more compact than other hymentoperan mitogenomes (~15.5 kb), retaining all protein coding genes, ribosomal, and transfer RNAs. We also present evidence of recombination between the mitogenomes of the two conspecific Solenopsis mitogenomes. Finally, we discuss potential ways to improve the estimation of phylogenies using complete mitochondrial genome sequences. CONCLUSIONS: The ant mitogenome presents an important addition to the continued efforts in studying hymenopteran mitogenome architecture, evolution, and phylogenetics. We provide further evidence that the sampling across many taxonomic levels (including conspecifics and congeners) is useful and important to gain detailed insights into mitogenome evolution. We also discuss ways that may help improve the use of mitogenomes in phylogenetic analyses by accounting for non-stationary and non-homogeneous evolution among branches.
Resumo:
The genomic era has revealed that the large repertoire of observed animal phenotypes is dependent on changes in the expression patterns of a finite number of genes, which are mediated by a plethora of transcription factors (TFs) with distinct specificities. The dimerization of TFs can also increase the complexity of a genetic regulatory network manifold, by combining a small number of monomers into dimers with distinct functions. Therefore, studying the evolution of these dimerizing TFs is vital for understanding how complexity increased during animal evolution. We focus on the second largest family of dimerizing TFs, the basic-region leucine zipper (bZIP), and infer when it expanded and how bZIP DNA-binding and dimerization functions evolved during the major phases of animal evolution. Specifically, we classify the metazoan bZIPs into 19 families and confirm the ancient nature of at least 13 of these families, predating the split of the cnidaria. We observe fixation of a core dimerization network in the last common ancestor of protostomes-deuterostomes. This was followed by an expansion of the number of proteins in the network, but no major dimerization changes in interaction partners, during the emergence of vertebrates. In conclusion, the bZIPs are an excellent model with which to understand how DNA binding and protein interactions of TFs evolved during animal evolution.
Resumo:
With the advancement of high-throughput sequencing and dramatic increase of available genetic data, statistical modeling has become an essential part in the field of molecular evolution. Statistical modeling results in many interesting discoveries in the field, from detection of highly conserved or diverse regions in a genome to phylogenetic inference of species evolutionary history Among different types of genome sequences, protein coding regions are particularly interesting due to their impact on proteins. The building blocks of proteins, i.e. amino acids, are coded by triples of nucleotides, known as codons. Accordingly, studying the evolution of codons leads to fundamental understanding of how proteins function and evolve. The current codon models can be classified into three principal groups: mechanistic codon models, empirical codon models and hybrid ones. The mechanistic models grasp particular attention due to clarity of their underlying biological assumptions and parameters. However, they suffer from simplified assumptions that are required to overcome the burden of computational complexity. The main assumptions applied to the current mechanistic codon models are (a) double and triple substitutions of nucleotides within codons are negligible, (b) there is no mutation variation among nucleotides of a single codon and (c) assuming HKY nucleotide model is sufficient to capture essence of transition- transversion rates at nucleotide level. In this thesis, I develop a framework of mechanistic codon models, named KCM-based model family framework, based on holding or relaxing the mentioned assumptions. Accordingly, eight different models are proposed from eight combinations of holding or relaxing the assumptions from the simplest one that holds all the assumptions to the most general one that relaxes all of them. The models derived from the proposed framework allow me to investigate the biological plausibility of the three simplified assumptions on real data sets as well as finding the best model that is aligned with the underlying characteristics of the data sets. -- Avec l'avancement de séquençage à haut débit et l'augmentation dramatique des données géné¬tiques disponibles, la modélisation statistique est devenue un élément essentiel dans le domaine dé l'évolution moléculaire. Les résultats de la modélisation statistique dans de nombreuses découvertes intéressantes dans le domaine de la détection, de régions hautement conservées ou diverses dans un génome de l'inférence phylogénétique des espèces histoire évolutive. Parmi les différents types de séquences du génome, les régions codantes de protéines sont particulièrement intéressants en raison de leur impact sur les protéines. Les blocs de construction des protéines, à savoir les acides aminés, sont codés par des triplets de nucléotides, appelés codons. Par conséquent, l'étude de l'évolution des codons mène à la compréhension fondamentale de la façon dont les protéines fonctionnent et évoluent. Les modèles de codons actuels peuvent être classés en trois groupes principaux : les modèles de codons mécanistes, les modèles de codons empiriques et les hybrides. Les modèles mécanistes saisir une attention particulière en raison de la clarté de leurs hypothèses et les paramètres biologiques sous-jacents. Cependant, ils souffrent d'hypothèses simplificatrices qui permettent de surmonter le fardeau de la complexité des calculs. Les principales hypothèses retenues pour les modèles actuels de codons mécanistes sont : a) substitutions doubles et triples de nucleotides dans les codons sont négligeables, b) il n'y a pas de variation de la mutation chez les nucléotides d'un codon unique, et c) en supposant modèle nucléotidique HKY est suffisant pour capturer l'essence de taux de transition transversion au niveau nucléotidique. Dans cette thèse, je poursuis deux objectifs principaux. Le premier objectif est de développer un cadre de modèles de codons mécanistes, nommé cadre KCM-based model family, sur la base de la détention ou de l'assouplissement des hypothèses mentionnées. En conséquence, huit modèles différents sont proposés à partir de huit combinaisons de la détention ou l'assouplissement des hypothèses de la plus simple qui détient toutes les hypothèses à la plus générale qui détend tous. Les modèles dérivés du cadre proposé nous permettent d'enquêter sur la plausibilité biologique des trois hypothèses simplificatrices sur des données réelles ainsi que de trouver le meilleur modèle qui est aligné avec les caractéristiques sous-jacentes des jeux de données. Nos expériences montrent que, dans aucun des jeux de données réelles, tenant les trois hypothèses mentionnées est réaliste. Cela signifie en utilisant des modèles simples qui détiennent ces hypothèses peuvent être trompeuses et les résultats de l'estimation inexacte des paramètres. Le deuxième objectif est de développer un modèle mécaniste de codon généralisée qui détend les trois hypothèses simplificatrices, tandis que d'informatique efficace, en utilisant une opération de matrice appelée produit de Kronecker. Nos expériences montrent que sur un jeux de données choisis au hasard, le modèle proposé de codon mécaniste généralisée surpasse autre modèle de codon par rapport à AICc métrique dans environ la moitié des ensembles de données. En outre, je montre à travers plusieurs expériences que le modèle général proposé est biologiquement plausible.
Resumo:
In addition to differences in protein-coding gene sequences, changes in expression resulting from mutations in regulatory sequences have long been hypothesized to be responsible for phenotypic differences between species. However, unlike comparison of genome sequences, few studies, generally restricted to pairwise comparisons of closely related mammalian species, have assessed between-species differences at the transcriptome level. They reported that gene expression evolves at different rates in various organs and in a pattern that is overall consistent with neutral models of evolution. In the first part of my thesis, I investigated the evolution of gene expression in therian mammals (i.e.7 placental and marsupials), based on microarray data from human, mouse and the gray short-tailed opossum (Monodelphis domestica). In addition to autosomal genes, a special focus was given to the evolution of X-linked genes. The therian X chromosome was recently shown to be younger than previously thought and to harbor a specific gene content (e.g., genes involved in brain or reproductive functions) that is thought to have been shaped by specific sex-related evolutionary forces. Sex chromosomes derive from ordinary autosomes and their differentiation led to the degeneration of the Y chromosome (in mammals) or W chromosome (in birds). Consequently, X- or Z-linked genes differ in gene dose between males and females such that the heterogametic sex has half the X/Z gene dose compared to the ancestral state. To cope with this dosage imbalance, mammals have been reported to have evolved mechanisms of dosage compensation.¦In the first project, I could first show that transcriptomes evolve at different rates in different organs. Out of the five tissues I investigated, the testis is the most rapidly evolving organ at the gene expression level while the brain has the most conserved transcriptome. Second, my analyses revealed that mammalian gene expression evolution is compatible with a neutral model, where the rates of change in gene expression levels is linked to the efficiency of purifying selection in a given lineage, which, in turn, is determined by the long-term effective population size in that lineage. Thus, the rate of DNA sequence evolution, which could be expected to determine the rate of regulatory sequence change, does not seem to be a major determinant of the rate of gene expression evolution. Thus, most gene expression changes seem to be (slightly) deleterious. Finally, X-linked genes seem to have experienced elevated rates of gene expression change during the early stage of X evolution. To further investigate the evolution of mammalian gene expression, we generated an extensive RNA-Seq gene expression dataset for nine mammalian species and a bird. The analyses of this dataset confirmed the patterns previously observed with microarrays and helped to significantly deepen our view on gene expression evolution.¦In a specific project based on these data, I sought to assess in detail patterns of evolution of dosage compensation in amniotes. My analyses revealed the absence of male to female dosage compensation in monotremes and its presence in marsupials and, in addition, confirmed patterns previously described for placental mammals and birds. I then assessed the global level of expression of X/Z chromosomes and contrasted this with its ancestral gene expression levels estimated from orthologous autosomal genes in species with non-homologous sex chromosomes. This analysis revealed a lack of up-regulation for placental mammals, the level of expression of X-linked genes being proportional to gene dose. Interestingly, the ancestral gene expression level was at least partially restored in marsupials as well as in the heterogametic sex of monotremes and birds. Finally, I investigated alternative mechanisms of dosage compensation and found that gene duplication did not seem to be a widespread mechanism to restore the ancestral gene dose. However, I could show that placental mammals have preferentially down-regulated autosomal genes interacting with X-linked genes which underwent gene expression decrease, and thus identified a novel alternative mechanism of dosage compensation.
Resumo:
Background: The RPS4 gene codifies for ribosomal protein S4, a very well-conserved protein present in all kingdoms. In primates, RPS4 is codified by two functional genes located on both sex chromosomes: the RPS4X and RPS4Y genes. In humans, RPS4Y is duplicated and the Y chromosome therefore carries a third functional paralog: RPS4Y2, which presents a testis-specific expression pattern. Results: DNA sequence analysis of the intronic and cDNA regions of RPS4Y genes from species covering the entire primate phylogeny showed that the duplication event leading to the second Y-linked copy occurred after the divergence of New World monkeys, about 35 million years ago. Maximum likelihood analyses of the synonymous and non-synonymous substitutions revealed that positive selection was acting on RPS4Y2 gene in the human lineage, which represents the first evidence of positive selection on a ribosomal protein gene. Putative positive amino acid replacements affected the three domains of the protein: one of these changes is located in the KOW protein domain and affects the unique invariable position of this motif, and might thus have a dramatic effect on the protein function.Conclusion: Here, we shed new light on the evolutionary history of RPS4Y gene family, especially on that of RPS4Y2. The results point that the RPS4Y1 gene might be maintained to compensate gene dosage between sexes, while RPS4Y2 might have acquired a new function, at least in the lineage leading to humans.
Resumo:
Chemoreception is a biological process essential for the survival of animals, as it allows the recognition of important volatile cues for the detection of food, egg-laying substrates, mates or predators, among other purposes. Furthermore, its role in pheromone detection may contribute to evolutionary processes such as reproductive isolation and speciation. This key role in several vital biological processes makes chemoreception a particularly interesting system for studying the role of natural selection in molecular adaptation. Two major gene families are involved in the perireceptor events of the chemosensory system: the odorant-binding protein (OBP) and chemosensory protein (CSP) families. Here, we have conducted an exhaustive comparative genomic analysis of these gene families in twenty Arthropoda species. We show that the evolution of the OBP and CSP gene families is highly dynamic, with a high number of gains and losses of genes, pseudogenes and independent origins of subfamilies. Taken together, our data clearly support the birth-and-death model for the evolution of these gene families with an overall high gene-turnover rate. Moreover, we show that the genome organization of the two families is significantly more clustered than expected by chance and, more important, that this pattern appears to be actively maintained across the Drosophila phylogeny. Finally, we suggest the homologous nature of the OBP and CSP gene families, dating back their MRCA (most recent common ancestor) to 380¿420 Mya, and we propose a scenario for the origin and diversification of these families.
Resumo:
BACKGROUND: The expansion of amino acid repeats is determined by a high mutation rate and can be increased or limited by selection. It has been suggested that recent expansions could be associated with the potential of adaptation to new environments. In this work, we quantify the strength of this association, as well as the contribution of potential confounding factors. RESULTS: Mammalian positively selected genes have accumulated more recent amino acid repeats than other mammalian genes. However, we found little support for an accelerated evolutionary rate as the main driver for the expansion of amino acid repeats. The most significant predictors of amino acid repeats are gene function and GC content. There is no correlation with expression level. CONCLUSIONS: Our analyses show that amino acid repeat expansions are causally independent from protein adaptive evolution in mammalian genomes. Relaxed purifying selection or positive selection do not associate with more or more recent amino acid repeats. Their occurrence is slightly favoured by the sequence context but mainly determined by the molecular function of the gene.
Resumo:
Ever since the pre-molecular era, the birth of new genes with novel functions has been considered to be a major contributor to adaptive evolutionary innovation. Here, I review the origin and evolution of new genes and their functions in eukaryotes, an area of research that has made rapid progress in the past decade thanks to the genomics revolution. Indeed, recent work has provided initial whole-genome views of the different types of new genes for a large number of different organisms. The array of mechanisms underlying the origin of new genes is compelling, extending way beyond the traditionally well-studied source of gene duplication. Thus, it was shown that novel genes also regularly arose from messenger RNAs of ancestral genes, protein-coding genes metamorphosed into new RNA genes, genomic parasites were co-opted as new genes, and that both protein and RNA genes were composed from scratch (i.e., from previously nonfunctional sequences). These mechanisms then also contributed to the formation of numerous novel chimeric gene structures. Detailed functional investigations uncovered different evolutionary pathways that led to the emergence of novel functions from these newly minted sequences and, with respect to animals, attributed a potentially important role to one specific tissue--the testis--in the process of gene birth. Remarkably, these studies also demonstrated that novel genes of the various types significantly impacted the evolution of cellular, physiological, morphological, behavioral, and reproductive phenotypic traits. Consequently, it is now firmly established that new genes have indeed been major contributors to the origin of adaptive evolutionary novelties.
Resumo:
In recent years, protein-ligand docking has become a powerful tool for drug development. Although several approaches suitable for high throughput screening are available, there is a need for methods able to identify binding modes with high accuracy. This accuracy is essential to reliably compute the binding free energy of the ligand. Such methods are needed when the binding mode of lead compounds is not determined experimentally but is needed for structure-based lead optimization. We present here a new docking software, called EADock, that aims at this goal. It uses an hybrid evolutionary algorithm with two fitness functions, in combination with a sophisticated management of the diversity. EADock is interfaced with the CHARMM package for energy calculations and coordinate handling. A validation was carried out on 37 crystallized protein-ligand complexes featuring 11 different proteins. The search space was defined as a sphere of 15 A around the center of mass of the ligand position in the crystal structure, and on the contrary to other benchmarks, our algorithm was fed with optimized ligand positions up to 10 A root mean square deviation (RMSD) from the crystal structure, excluding the latter. This validation illustrates the efficiency of our sampling strategy, as correct binding modes, defined by a RMSD to the crystal structure lower than 2 A, were identified and ranked first for 68% of the complexes. The success rate increases to 78% when considering the five best ranked clusters, and 92% when all clusters present in the last generation are taken into account. Most failures could be explained by the presence of crystal contacts in the experimental structure. Finally, the ability of EADock to accurately predict binding modes on a real application was illustrated by the successful docking of the RGD cyclic pentapeptide on the alphaVbeta3 integrin, starting far away from the binding pocket.
Resumo:
Abstract : Post-translational modifications such as proteolytic processing, phosphorylation, and glycosylation, add extra layers of complexity to proteomes and allow a finely tuned regulation of the activity of many proteins. The evolutionarily conserved cell-cycle and transcriptional regulator HCP-] is regulated by proteolytic maturation via which a stable heterodirneric complex of two cleaved subunits is formed from a single precursor protein. The human HCF-1 precursor is cleaved at six nearly identical 26 amino acid sequence repeats, called HCF-1pro repeats, which represent uncommon protease recognition sites dedicated to human HCF-1 proteolysis. This proteolytic maturation process is conserved in vertebrate HCF-1 homologues and is essential for the functions of the human protein in cell-cycle regulation; the mechanisms that execute and control HCF-1 proteolysis, however, remain poorly understood. In this dissertation I investigate the mechanisms of proteolytic maturation of HCF-1 proteins in different species. I show that the Drosophila homolog of human HCF-1, called dHCP, is proteolytically cleaved via a different mechanism than human HCF-1. dHCP is processed by the same protease, called Taspase], which cleaves one of the key developmental regulators in flies, the Trithorax protein. Maturation of HCP proteins via Taspase] cleavage is probably not particular to dHCP as many invertebrate HCP proteins, particularly insects and flatworms, possess Taspase] recognition sites. In contrast, the vertebrate HCF-1 proteins lack Taspase] recognition sites and the HCF-1pro repeats are not Taspase1 substrates, suggesting that multiple mechanisms for HCF-1 proteolytic maturation have appeared during evolution. I also show that the proteolytic activity responsible for the cleavage of the HCP- 1pro repeats is very difficult to characterize, being resistant to most protease inhibitors and very sensitive to biochemical fractionation. Moreover, the HCF-1pro repeats represent complex protease recognition sites and I demonstrate that, in addition to be the HCF-1 cleavage sites, these repeated sequences, also recruit the OG1cNAc transferase OGT. The OGT protein and the OG1cNAc modification of HCF-1 are both important for HCF-1pro repeat proteolysis. Interestingly, a human recombinant OGT purified from insect cells is able to induce cleavage of a HCF-1pro-repeat precursor in vitro, indicating that OGT either (i) induces HCF-1 autoproteolysis,(ii) is the HCF-1pro- repeat proteolytic activity itself, or (iii) physically associates with a proteolytic activity that is conserved in insect cells. In any case, OGT plays an important role in HCF-1 proteolytic maturation and perhaps a broader role in HCF-1 biological function. Résumé : Les modifications post-traductionelles pomme le clivage protéolytique, la phosphorylation, et la glycosylation, augmentent significativement la complexité des protéomes et permettent une régulation fine de l'activité de beaucoup de protéines. La protéine HCF-1, qui est un régulateur du cycle cellulaire et de la transcription, est elle- même régulée par clivage protéolytique. La protéine HCF-1 est en effet coupée en deux sous-unités qui s'associent l'une a l'autre pour former la protéine mature. Le précurseur de la protéine HCF-1 humaine est clivé à six sites correspondant à six séquences répétées nommées les HCF-1pro repeats, chacune composée de 26 acide aminés. Les HCF-1pro- repeats ne ressemblent ai aucune séquence de clivage protéolytique connue et sont présentes seulement dans les protéines HCF-1 chez les vertébrés. Bien que la maturation protéolytique d'HCF-1 soit essentielle pour les activités de cette protéine pendant le cycle cellulaire, les mécanismes qui la contrôlent restent inconnus. Au cours de mon travail de thèse, j'ai analysé les mécanismes de clivage protéolytique des protéines HCF dans différentes espèces. J'ai montré que la protéine de Drosophile homologue d'HCF-1 humaine nommée dHCF est clivée par une protéase nommée Taspase1. Ainsi, dHCF est clivé par la même protéase que celle qui induit la maturation protéolytique d'un des principaux facteurs du développement chez la mouche, la protéine Trithorax. La maturation de dHCF via le clivage par la Taspase1 n'est pas spécifique à la mouche, mais est probablement étendu à plusieurs protéines HCF chez les invertébrés, surtout dans les familles des insectes et des plathehninthes, car ces protéines HCF présentent des sites de reconnaissance pour la Taspasel. Par contre, les protéines HCF-1 chez les vertébrés n'ont pas de sites de reconnaissance pour la Taspasel et cela suggère que différents mécanismes de maturation des protéines HCF- ls ont apparu au cours de l'évolution. J'ai montré aussi que les HCF-1pro-repeats sont clivés par une activité protéolytique très difficile a identifier, car elle est résistante à la plupart des inhibiteurs de protéases, mais elle est très sensible au fractionnement biochimique. En plus, les HCF-1pro-repeats sont un site de protéolyse complexe qui ne sert pas seulement au clivage des protéines HCF- chez les vertébrés mais aussi à recruter l'enzyme responsable de la O- GlcNAcylation nommée OGT. La protéine OGT et la O-GlcNAcylatio d'HCF-1 sont toutes les deux importantes pour le clivage protéolytique des HCF1pro-repeats. Curieusement, la protéine OGT humaine produite dans des cellules d'insectes est capable de cliver les HCF-1pro repeats in vitro et cela suggère que OGT soit (i) induit le clivage autocatalytique cl'HCF-1, soit (ii) est elle-même l'activité protéolytique qui clive HCF4, soit (iii) est associée à une activité protéolytique conservée dans les cellules d'insectes qui a été co-purifiée avec OGT. En conclusion, OGT joue un rôle important dans la maturation protéolytique d'HCF-1 et peut-être aussi un rôle plus large dans les fonctions biologiques de la protéine HCF-1.
Resumo:
We present here a draft genome sequence of the red jungle fowl, Gallus gallus. Because the chicken is a modern descendant of the dinosaurs and the first non-mammalian amniote to have its genome sequenced, the draft sequence of its genome--composed of approximately one billion base pairs of sequence and an estimated 20,000-23,000 genes--provides a new perspective on vertebrate genome evolution, while also improving the annotation of mammalian genomes. For example, the evolutionary distance between chicken and human provides high specificity in detecting functional elements, both non-coding and coding. Notably, many conserved non-coding sequences are far from genes and cannot be assigned to defined functional classes. In coding regions the evolutionary dynamics of protein domains and orthologous groups illustrate processes that distinguish the lineages leading to birds and mammals. The distinctive properties of avian microchromosomes, together with the inferred patterns of conserved synteny, provide additional insights into vertebrate chromosome architecture.
Resumo:
Arising from either retrotransposition or genomic duplication of functional genes, pseudogenes are "genomic fossils" valuable for exploring the dynamics and evolution of genes and genomes. Pseudogene identification is an important problem in computational genomics, and is also critical for obtaining an accurate picture of a genome's structure and function. However, no consensus computational scheme for defining and detecting pseudogenes has been developed thus far. As part of the ENCyclopedia Of DNA Elements (ENCODE) project, we have compared several distinct pseudogene annotation strategies and found that different approaches and parameters often resulted in rather distinct sets of pseudogenes. We subsequently developed a consensus approach for annotating pseudogenes (derived from protein coding genes) in the ENCODE regions, resulting in 201 pseudogenes, two-thirds of which originated from retrotransposition. A survey of orthologs for these pseudogenes in 28 vertebrate genomes showed that a significant fraction ( approximately 80%) of the processed pseudogenes are primate-specific sequences, highlighting the increasing retrotransposition activity in primates. Analysis of sequence conservation and variation also demonstrated that most pseudogenes evolve neutrally, and processed pseudogenes appear to have lost their coding potential immediately or soon after their emergence. In order to explore the functional implication of pseudogene prevalence, we have extensively examined the transcriptional activity of the ENCODE pseudogenes. We performed systematic series of pseudogene-specific RACE analyses. These, together with complementary evidence derived from tiling microarrays and high throughput sequencing, demonstrated that at least a fifth of the 201 pseudogenes are transcribed in one or more cell lines or tissues.