924 resultados para Coding Sequences


Relevância:

60.00% 60.00%

Publicador:

Resumo:

A pimenteira-do-reino (Piper nigrum L.) constitui uma das espécies de pimenta mais amplamente utilizadas no mundo, pertencendo à família Piperaceae, a qual compreende cerca de 1400 espécies distribuídas principalmente no continente americano e sudeste da Ásia, onde esta cultura originou. A pimenteira-do-reino foi introduzida no Brasil no século XVII, e tornou-se uma cultura de importância econômica desde 1933. O Estado do Pará é o principal produto brasileiro de pimenta-do-reino, contudo sua produção vem sendo afetada pela doença fusariose causada pelo fungo Fusarium solani f. sp. piperis. Estudos prévios revelaram a identificação de sequencias de cDNA diferencialmente expressas durante a interação da pimenteira-do-reino com o F. solani f. sp. piperis. Entre elas, uma sequencia de cDNA parcial que codifica para uma proteína transportadora de lipídeos (LTP), a qual é conhecida por seu importante papel na defesa de plantas contra patógenos e insetos. Desta forma, o objetivo principal deste trabalho foi isolar e caracterizar as sequencias de cDNA e genômica de uma LTP de pimenteira-do-reino, denominada PnLTP. O cDNA completo da PnLTP isolado por meio de experimentos de RACE apresentou 621 bp com 32 pb and 235 bp nas regiões não traduzidas 5‘ e 3‘, respectivamente. Este cDNA contem uma ORF de 354 bp codificando uma proteína deduzida de 117 resíduos de aminoácidos que apresentou alta identidade com LTPs de outras espécies vegetais. Análises das sequencias revelou que a PnLTP contem um potencial peptídeo sinal na extremidade amino-terminal e oito resíduos de cisteína preditos por formar quatro pontes de dissulfeto, as quais poderiam contribuir para a estabilidade desta proteína. O alinhamento entre as sequencias de cDNA e genômica revelou a ausência de introns na região codificante do gene PnLTP, o que está de acordo ao encontrado em outros genes de LTPs de plantas. Por último, a PnLTP madura foi expressa em sistema bacteriano. Experimentos adicionais serão realizados com o objetivo de avaliar a habilidade da PnLTP recombinante em inibir o crescimento do F. solani f. sp. piperis.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES)

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Pós-graduação em Microbiologia Agropecuária - FCAV

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP)

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Vibrio campbellii PEL22A was isolated from open ocean water in the Abrolhos Bank. The genome of PEL22A consists of 6,788,038 bp (the GC content is 45%). The number of coding sequences (CDS) is 6,359, as determined according to the Rapid Annotation using Subsystem Technology (RAST) server. The number of ribosomal genes is 80, of which 68 are tRNAs and 12 are rRNAs. V. campbellii PEL22A contains genes related to virulence and fitness, including a complete proteorhodopsin cluster, complete type II and III secretion systems, incomplete type I, IV, and VI secretion systems, a hemolysin, and CTX Phi.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

This work describes the effects of the cell surface display of a synthetic phytochelatin in the highly metal tolerant bacterium Cupriavidus metallidurans CH34. The EC20sp synthetic phytochelatin gene was fused between the coding sequences of the signal peptide (SS) and of the autotransporter beta-domain of the Neisseria gonorrhoeae IgA protease precursor (IgA beta), which successfully targeted the hybrid protein toward the C. metallidurans outer membrane. The expression of the SS-EC20sp-IgA beta gene fusion was driven by a modified version of the Bacillus subtilis mrgA promoter showing high level basal gene expression that is further enhanced by metal presence in C. metallidurans. The recombinant strain showed increased ability to immobilize Pb2+, Zn2+, Cu2+, Cd2+, Mn2+, and Ni2+ ions from the external medium when compared to the control strain. To ensure plasmid stability and biological containment, the MOB region of the plasmid was replaced by the E. coli hok/sok coding sequence.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Abstract Background From shotgun libraries used for the genomic sequencing of the phytopathogenic bacterium Xanthomonas axonopodis pv. citri (XAC), clones that were representative of the largest possible number of coding sequences (CDSs) were selected to create a DNA microarray platform on glass slides (XACarray). The creation of the XACarray allowed for the establishment of a tool that is capable of providing data for the analysis of global genome expression in this organism. Findings The inserts from the selected clones were amplified by PCR with the universal oligonucleotide primers M13R and M13F. The obtained products were purified and fixed in duplicate on glass slides specific for use in DNA microarrays. The number of spots on the microarray totaled 6,144 and included 768 positive controls and 624 negative controls per slide. Validation of the platform was performed through hybridization of total DNA probes from XAC labeled with different fluorophores, Cy3 and Cy5. In this validation assay, 86% of all PCR products fixed on the glass slides were confirmed to present a hybridization signal greater than twice the standard deviation of the deviation of the global median signal-to-noise ration. Conclusions Our validation of the XACArray platform using DNA-DNA hybridization revealed that it can be used to evaluate the expression of 2,365 individual CDSs from all major functional categories, which corresponds to 52.7% of the annotated CDSs of the XAC genome. As a proof of concept, we used this platform in a previously work to verify the absence of genomic regions that could not be detected by sequencing in related strains of Xanthomonas.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Cardiac morphogenesis is a complex process governed by evolutionarily conserved transcription factors and signaling molecules. The Drosophila cardiac tube is linear, made of 52 pairs of cardiomyocytes (CMs), which express specific transcription factor genes that have human homologues implicated in Congenital Heart Diseases (CHDs) (NKX2-5, GATA4 and TBX5). The Drosophila cardiac tube is linear and composed of a rostral portion named aorta and a caudal one called heart, distinguished by morphological and functional differences controlled by Hox genes, key regulators of axial patterning. Overexpression and inactivation of the Hox gene abdominal-A (abd-A), which is expressed exclusively in the heart, revealed that abd-A controls heart identity. The aim of our work is to isolate the heart-specific cisregulatory sequences of abd-A direct target genes, the realizator genes granting heart identity. In each segment of the heart, four pairs of cardiomyocytes (CMs) express tinman (tin), homologous to NKX2-5, and acquire strong contractile and automatic rhythmic activities. By tyramide amplified FISH, we found that seven genes, encoding ion channels, pumps or transporters, are specifically expressed in the Tin-CMs of the heart. We initially used online available tools to identify their heart-specific cisregutatory modules by looking for Conserved Non-coding Sequences containing clusters of binding sites for various cardiac transcription factors, including Hox proteins. Based on these data we generated several reporter gene constructs and transgenic embryos, but none of them showed reporter gene expression in the heart. In order to identify additional abd-A target genes, we performed microarray experiments comparing the transcriptomes of aorta versus heart and identified 144 genes overexpressed in the heart. In order to find the heart-specific cis-regulatory regions of these target genes we developed a new bioinformatic approach where prediction is based on pattern matching and ordered statistics. We first retrieved Conserved Noncoding Sequences from the alignment between the D.melanogaster and D.pseudobscura genomes. We scored for combinations of conserved occurrences of ABD-A, ABD-B, TIN, PNR, dMEF2, MADS box, T-box and E-box sites and we ranked these results based on two independent strategies. On one hand we ranked the putative cis-regulatory sequences according to best scored ABD-A biding sites, on the other hand we scored according to conservation of binding sites. We integrated and ranked again the two lists obtained independently to produce a final rank. We generated nGFP reporter construct flies for in vivo validation. We identified three 1kblong heart-specific enhancers. By in vivo and in vitro experiments we are determining whether they are direct abd-A targets, demonstrating the role of a Hox gene in the realization of heart identity. The identified abd-A direct target genes may be targets also of the NKX2-5, GATA4 and/or TBX5 homologues tin, pannier and Doc genes, respectively. The identification of sequences coregulated by a Hox protein and the homologues of transcription factors causing CHDs, will provide a mean to test whether these factors function as Hox cofactors granting cardiac specificity to Hox proteins, increasing our knowledge on the molecular mechanisms underlying CHDs. Finally, it may be investigated whether these Hox targets are involved in CHDs.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Motivation An actual issue of great interest, both under a theoretical and an applicative perspective, is the analysis of biological sequences for disclosing the information that they encode. The development of new technologies for genome sequencing in the last years, opened new fundamental problems since huge amounts of biological data still deserve an interpretation. Indeed, the sequencing is only the first step of the genome annotation process that consists in the assignment of biological information to each sequence. Hence given the large amount of available data, in silico methods became useful and necessary in order to extract relevant information from sequences. The availability of data from Genome Projects gave rise to new strategies for tackling the basic problems of computational biology such as the determination of the tridimensional structures of proteins, their biological function and their reciprocal interactions. Results The aim of this work has been the implementation of predictive methods that allow the extraction of information on the properties of genomes and proteins starting from the nucleotide and aminoacidic sequences, by taking advantage of the information provided by the comparison of the genome sequences from different species. In the first part of the work a comprehensive large scale genome comparison of 599 organisms is described. 2,6 million of sequences coming from 551 prokaryotic and 48 eukaryotic genomes were aligned and clustered on the basis of their sequence identity. This procedure led to the identification of classes of proteins that are peculiar to the different groups of organisms. Moreover the adopted similarity threshold produced clusters that are homogeneous on the structural point of view and that can be used for structural annotation of uncharacterized sequences. The second part of the work focuses on the characterization of thermostable proteins and on the development of tools able to predict the thermostability of a protein starting from its sequence. By means of Principal Component Analysis the codon composition of a non redundant database comprising 116 prokaryotic genomes has been analyzed and it has been showed that a cross genomic approach can allow the extraction of common determinants of thermostability at the genome level, leading to an overall accuracy in discriminating thermophilic coding sequences equal to 95%. This result outperform those obtained in previous studies. Moreover, we investigated the effect of multiple mutations on protein thermostability. This issue is of great importance in the field of protein engineering, since thermostable proteins are generally more suitable than their mesostable counterparts in technological applications. A Support Vector Machine based method has been trained to predict if a set of mutations can enhance the thermostability of a given protein sequence. The developed predictor achieves 88% accuracy.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

The objective of this work is to characterize the genome of the chromosome 1 of A.thaliana, a small flowering plants used as a model organism in studies of biology and genetics, on the basis of a recent mathematical model of the genetic code. I analyze and compare different portions of the genome: genes, exons, coding sequences (CDS), introns, long introns, intergenes, untranslated regions (UTR) and regulatory sequences. In order to accomplish the task, I transformed nucleotide sequences into binary sequences based on the definition of the three different dichotomic classes. The descriptive analysis of binary strings indicate the presence of regularities in each portion of the genome considered. In particular, there are remarkable differences between coding sequences (CDS and exons) and non-coding sequences, suggesting that the frame is important only for coding sequences and that dichotomic classes can be useful to recognize them. Then, I assessed the existence of short-range dependence between binary sequences computed on the basis of the different dichotomic classes. I used three different measures of dependence: the well-known chi-squared test and two indices derived from the concept of entropy i.e. Mutual Information (MI) and Sρ, a normalized version of the “Bhattacharya Hellinger Matusita distance”. The results show that there is a significant short-range dependence structure only for the coding sequences whose existence is a clue of an underlying error detection and correction mechanism. No doubt, further studies are needed in order to assess how the information carried by dichotomic classes could discriminate between coding and noncoding sequence and, therefore, contribute to unveil the role of the mathematical structure in error detection and correction mechanisms. Still, I have shown the potential of the approach presented for understanding the management of genetic information.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Welche genetische Unterschiede machen uns verschieden von unseren nächsten Verwandten, den Schimpansen, und andererseits so ähnlich zu den Schimpansen? Was wir untersuchen und auch verstehen wollen, ist die komplexe Beziehung zwischen den multiplen genetischen und epigenetischen Unterschieden, deren Interaktion mit diversen Umwelt- und Kulturfaktoren in den beobachteten phänotypischen Unterschieden resultieren. Um aufzuklären, ob chromosomale Rearrangements zur Divergenz zwischen Mensch und Schimpanse beigetragen haben und welche selektiven Kräfte ihre Evolution geprägt haben, habe ich die kodierenden Sequenzen von 2 Mb umfassenden, die perizentrischen Inversionsbruchpunkte flankierenden Regionen auf den Chromosomen 1, 4, 5, 9, 12, 17 und 18 untersucht. Als Kontrolle dienten dabei 4 Mb umfassende kollineare Regionen auf den rearrangierten Chromosomen, welche mindestens 10 Mb von den Bruchpunktregionen entfernt lagen. Dabei konnte ich in den Bruchpunkten flankierenden Regionen im Vergleich zu den Kontrollregionen keine höhere Proteinevolutionsrate feststellen. Meine Ergebnisse unterstützen nicht die chromosomale Speziationshypothese für Mensch und Schimpanse, da der Anteil der positiv selektierten Gene (5,1% in den Bruchpunkten flankierenden Regionen und 7% in den Kontrollregionen) in beiden Regionen ähnlich war. Durch den Vergleich der Anzahl der positiv und negativ selektierten Gene per Chromosom konnte ich feststellen, dass Chromosom 9 die meisten und Chromosom 5 die wenigsten positiv selektierten Gene in den Bruchpunkt flankierenden Regionen und Kontrollregionen enthalten. Die Anzahl der negativ selektierten Gene (68) war dabei viel höher als die Anzahl der positiv selektierten Gene (17). Eine bioinformatische Analyse von publizierten Microarray-Expressionsdaten (Affymetrix Chip U95 und U133v2) ergab 31 Gene, die zwischen Mensch und Schimpanse differentiell exprimiert sind. Durch Untersuchung des dN/dS-Verhältnisses dieser 31 Gene konnte ich 7 Gene als negativ selektiert und nur 1 Gen als positiv selektiert identifizieren. Dieser Befund steht im Einklang mit dem Konzept, dass Genexpressionslevel unter stabilisierender Selektion evolvieren. Die meisten positiv selektierten Gene spielen überdies eine Rolle bei der Fortpflanzung. Viele dieser Speziesunterschiede resultieren eher aus Änderungen in der Genregulation als aus strukturellen Änderungen der Genprodukte. Man nimmt an, dass die meisten Unterschiede in der Genregulation sich auf transkriptioneller Ebene manifestieren. Im Rahmen dieser Arbeit wurden die Unterschiede in der DNA-Methylierung zwischen Mensch und Schimpanse untersucht. Dazu wurden die Methylierungsmuster der Promotor-CpG-Inseln von 12 Genen im Cortex von Menschen und Schimpansen mittels klassischer Bisulfit-Sequenzierung und Bisulfit-Pyrosequenzierung analysiert. Die Kandidatengene wurden wegen ihrer differentiellen Expressionsmuster zwischen Mensch und Schimpanse sowie wegen Ihrer Assoziation mit menschlichen Krankheiten oder dem genomischen Imprinting ausgewählt. Mit Ausnahme einiger individueller Positionen zeigte die Mehrzahl der analysierten Gene keine hohe intra- oder interspezifische Variation der DNA-Methylierung zwischen den beiden Spezies. Nur bei einem Gen, CCRK, waren deutliche intraspezifische und interspezifische Unterschiede im Grad der DNA-Methylierung festzustellen. Die differentiell methylierten CpG-Positionen lagen innerhalb eines repetitiven Alu-Sg1-Elements. Die Untersuchung des CCRK-Gens liefert eine umfassende Analyse der intra- und interspezifischen Variabilität der DNA-Methylierung einer Alu-Insertion in eine regulatorische Region. Die beobachteten Speziesunterschiede deuten darauf hin, dass die Methylierungsmuster des CCRK-Gens wahrscheinlich in Adaption an spezifische Anforderungen zur Feinabstimmung der CCRK-Regulation unter positiver Selektion evolvieren. Der Promotor des CCRK-Gens ist anfällig für epigenetische Modifikationen durch DNA-Methylierung, welche zu komplexen Transkriptionsmustern führen können. Durch ihre genomische Mobilität, ihren hohen CpG-Anteil und ihren Einfluss auf die Genexpression sind Alu-Insertionen exzellente Kandidaten für die Förderung von Veränderungen während der Entwicklungsregulation von Primatengenen. Der Vergleich der intra- und interspezifischen Methylierung von spezifischen Alu-Insertionen in anderen Genen und Geweben stellt eine erfolgversprechende Strategie dar.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Da nicht-synonyme tumorspezifische Punktmutationen nur in malignen Geweben vorkommen und das veränderte Proteinprodukt vom Immunsystem als „fremd“ erkannt werden kann, stellen diese einen bisher ungenutzten Pool von Zielstrukturen für die Immuntherapie dar. Menschliche Tumore können individuell bis zu tausenden nicht-synonymer Punktmutationen in ihrem Genom tragen, welche nicht der zentralen Immuntoleranz unterliegen. Ziel der vorliegenden Arbeit war die Hypothese zu untersuchen, dass das Immunsystem in der Lage sein sollte, mutierte Epitope auf Tumorzellen zu erkennen und zu klären, ob auf dieser Basis eine wirksame mRNA (RNA) basierte anti-tumorale Vakzinierung etabliert werden kann. Hierzu wurde von Ugur Sahin und Kollegen, das gesamte Genom des murinen B16-F10 Melanoms sequenziert und bioinformatisch analysiert. Im Rahmen der NGS Sequenzierung wurden mehr als 500 nicht-synonyme Punktmutationen identifiziert, von welchen 50 Mutationen selektiert und durch Sanger Sequenzierung validiert wurden. rnNach der Etablierung des immunologischen Testsysteme war eine Hauptfragestellung dieser Arbeit, die selektierten nicht-synonyme Punktmutationen in einem in vivo Ansatz systematisch auf Antigenität zu testen. Für diese Studien wurden mutierte Sequenzen in einer Länge von 27 Aminosäuren genutzt, in denen die mutierte Aminosäure zentral positioniert war. Durch die Länge der Peptide können prinzipiell alle möglichen MHC Klasse-I und -II Epitope abgedeckt werden, welche die Mutation enthalten. Eine Grundidee des Projektes Ansatzes ist es, einen auf in vitro transkribierter RNA basierten oligotopen Impfstoff zu entwickeln. Daher wurden die Impfungen naiver Mäuse sowohl mit langen Peptiden, als auch in einem unabhängigen Ansatz mit peptidkodierender RNA durchgeführt. Die Immunphänotypisierung der Impfstoff induzierten T-Zellen zeigte, dass insgesamt 16 der 50 (32%) mutierten Sequenzen eine T-Zellreaktivität induzierten. rnDie Verwendung der vorhergesagten Epitope in therapeutischen Vakzinierungsstudien bestätigten die Hypothese das mutierte Neo-Epitope potente Zielstrukturen einer anti-tumoralen Impftherapie darstellen können. So wurde in therapeutischen Tumorstudien gezeigt, dass auf Basis von RNA 9 von 12 bestätigten Epitopen einen anti-tumoralen Effekt zeigte.rnÜberaschenderweise wurde bei einem MHC Klasse-II restringierten mutiertem Epitop (Mut-30) sowohl in einem subkutanen, als auch in einem unabhängigen therapeutischen Lungenmetastasen Modell ein starker anti-tumoraler Effekt auf B16-F10 beobachtet, der dieses Epitop als neues immundominantes Epitop für das B16-F10 Melanom etabliert. Um den immunologischen Mechanismus hinter diesem Effekt näher zu untersuchen wurde in verschieden Experimenten die Rolle von CD4+, CD8+ sowie NK-Zellen zu verschieden Zeitpunkten der Tumorentwicklung untersucht. Die Analyse des Tumorgewebes ergab, eine signifikante erhöhte Frequenz von NK-Zellen in den mit Mut-30 RNA vakzinierten Tieren. Das NK Zellen in der frühen Phase der Therapie eine entscheidende Rolle spielen wurde anhand von Depletionsstudien bestätigt. Daran anschließend wurde gezeigt, dass im fortgeschrittenen Tumorstadium die NK Zellen keinen weiteren relevanten Beitrag zum anti-tumoralen Effekt der RNA Vakzinierung leisten, sondern die Vakzine induzierte adaptive Immunantwort. Durch die Isolierung von Lymphozyten aus dem Tumorgewebe und deren Einsatz als Effektorzellen im IFN-γ ELISPOT wurde nachgewiesen, dass Mut-30 spezifische T-Zellen das Tumorgewebe infiltrieren und dort u.a. IFN-γ sekretieren. Dass diese spezifische IFN-γ Ausschüttung für den beobachteten antitumoralen Effekt eine zentrale Rolle einnimmt wurde unter der Verwendung von IFN-γ -/- K.O. Mäusen bestätigt.rnDas Konzept der individuellen RNA basierten mutationsspezifischen Vakzine sieht vor, nicht nur mit einem mutations-spezifischen Epitop, sondern mit mehreren RNA-kodierten Mutationen Patienten zu impfen um der Entstehung von „escape“-Mutanten entgegenzuwirken. Da es nur Erfahrung mit der Herstellung und Verabreichung von Monotop-RNA gab, also RNA die für ein Epitop kodiert, war eine wichtige Fragestellungen, inwieweit Oligotope, welche die mutierten Sequenzen sequentiell durch Linker verbunden als Fusionsprotein kodieren, Immunantworten induzieren können. Hierzu wurden Pentatope mit variierender Position des einzelnen Epitopes hinsichtlich ihrer in vivo induzierten T-Zellreaktivitäten charakterisiert. Die Experimente zeigten, dass es möglich ist, unabhängig von der Position im Pentatop eine Immunantwort gegen ein Epitop zu induzieren. Des weiteren wurde beobachtet, dass die induzierten T-Zellfrequenzen nach Pentatop Vakzinierung im Vergleich zur Nutzung von Monotopen signifikant gesteigert werden kann.rnZusammenfassend wurde im Rahmen der vorliegenden Arbeit präklinisch erstmalig nachgewiesen, dass nicht-synonyme Mutationen eine numerisch relevante Quelle von Zielstrukturen für die anti-tumorale Immuntherapie darstellen. Überraschenderweise zeigte sich eine dominante Induktion MHC-II restringierter Immunantworten, welche partiell in der Lage waren massive Tumorabstoßungsreaktionen zu induzieren. Im Sinne einer Translation der gewonnenen Erkenntnisse wurde ein RNA basiertes Oligotop-Format etabliert, welches Eingang in die klinische Testung des Konzeptes fand.rn

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Protein is an essential component for life, and its synthesis is mediated by codons in any organisms on earth. While some codons encode the same amino acid, their usage is often highly biased. There are many factors that can cause the bias, but a potential effect of mononucleotide repeats, which are known to be highly mutable, on codon usage and codon pair preference is largely unknown. In this study we performed a genomic survey on the relationship between mononucleotide repeats and codon pair bias in 53 bacteria, 68 archaea, and 13 eukaryotes. By distinguishing the codon pair bias from the codon usage bias, four general patterns were revealed: strong avoidance of five or six mononucleotide repeats in codon pairs; lower observed/expected (o/e) ratio for codon pairs with C or G repeats (C/G pairs) than that with A or T repeats (A/T pairs); a negative correlation between genomic GC contents and the o/e ratios, particularly for C/G pairs; and avoidance of C/G pairs in highly conserved genes. These results support natural selection against long mononucleotide repeats, which could induce frameshift mutations in coding sequences. The fact that these patterns are found in all kingdoms of life suggests that this is a general phenomenon in living organisms. Thus, long mononucleotide repeats may play an important role in base composition and genetic stability of a gene and gene functions.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Different codons encoding the same amino acid are not used equally in protein-coding sequences. In bacteria, there is a bias towards codons with high translation rates. This bias is most pronounced in highly expressed proteins, but a recent study of synthetic GFP-coding sequences did not find a correlation between codon usage and GFP expression, suggesting that such correlation in natural sequences is not a simple property of translational mechanisms. Here, we investigate the effect of evolutionary forces on codon usage. The relation between codon bias and protein abundance is quantitatively analyzed based on the hypothesis that codon bias evolved to ensure the efficient usage of ribosomes, a precious commodity for fast growing cells. An explicit fitness landscape is formulated based on bacterial growth laws to relate protein abundance and ribosomal load. The model leads to a quantitative relation between codon bias and protein abundance, which accounts for a substantial part of the observed bias for E. coli. Moreover, by providing an evolutionary link, the ribosome load model resolves the apparent conflict between the observed relation of protein abundance and codon bias in natural sequences and the lack of such dependence in a synthetic gfp library. Finally, we show that the relation between codon usage and protein abundance can be used to predict protein abundance from genomic sequence data alone without adjustable parameters.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Different codons encoding the same amino acid are not used equally in protein-coding sequences. In bacteria, there is a bias towards codons with high translation rates. This bias is most pronounced in highly expressed proteins, but a recent study of synthetic GFP-coding sequences did not find a correlation between codon usage and GFP expression, suggesting that such correlation in natural sequences is not a simple property of translational mechanisms. Here, we investigate the effect of evolutionary forces on codon usage. The relation between codon bias and protein abundance is quantitatively analyzed based on the hypothesis that codon bias evolved to ensure the efficient usage of ribosomes, a precious commodity for fast growing cells. An explicit fitness landscape is formulated based on bacterial growth laws to relate protein abundance and ribosomal load. The model leads to a quantitative relation between codon bias and protein abundance, which accounts for a substantial part of the observed bias for E. coli. Moreover, by providing an evolutionary link, the ribosome load model resolves the apparent conflict between the observed relation of protein abundance and codon bias in natural sequences and the lack of such dependence in a synthetic gfp library. Finally, we show that the relation between codon usage and protein abundance can be used to predict protein abundance from genomic sequence data alone without adjustable parameters.