969 resultados para coding sequence
Resumo:
A large proportion of the variation in traits between individuals can be attributed to variation in the nucleotide sequence of the genome. The most commonly studied traits in human genetics are related to disease and disease susceptibility. Although scientists have identified genetic causes for over 4,000 monogenic diseases, the underlying mechanisms of many highly prevalent multifactorial inheritance disorders such as diabetes, obesity, and cardiovascular disease remain largely unknown. Identifying genetic mechanisms for complex traits has been challenging because most of the variants are located outside of protein-coding regions, and determining the effects of such non-coding variants remains difficult. In this dissertation, I evaluate the hypothesis that such non-coding variants contribute to human traits and diseases by altering the regulation of genes rather than the sequence of those genes. I will specifically focus on studies to determine the functional impacts of genetic variation associated with two related complex traits: gestational hyperglycemia and fetal adiposity. At the genomic locus associated with maternal hyperglycemia, we found that genetic variation in regulatory elements altered the expression of the HKDC1 gene. Furthermore, we demonstrated that HKDC1 phosphorylates glucose in vitro and in vivo, thus demonstrating that HKDC1 is a fifth human hexokinase gene. At the fetal-adiposity associated locus, we identified variants that likely alter VEPH1 expression in preadipocytes during differentiation. To make such studies of regulatory variation high-throughput and routine, we developed POP-STARR, a novel high throughput reporter assay that can empirically measure the effects of regulatory variants directly from patient DNA. By combining targeted genome capture technologies with STARR-seq, we assayed thousands of haplotypes from 760 individuals in a single experiment. We subsequently used POP-STARR to identify three key features of regulatory variants: that regulatory variants typically have weak effects on gene expression; that the effects of regulatory variants are often coordinated with respect to disease-risk, suggesting a general mechanism by which the weak effects can together have phenotypic impact; and that nucleotide transversions have larger impacts on enhancer activity than transitions. Together, the findings presented here demonstrate successful strategies for determining the regulatory mechanisms underlying genetic associations with human traits and diseases, and value of doing so for driving novel biological discovery.
Resumo:
Genome-wide association studies (GWAS) have identified several risk variants for late-onset Alzheimer's disease (LOAD)1, 2. These common variants have replicable but small effects on LOAD risk and generally do not have obvious functional effects. Low-frequency coding variants, not detected by GWAS, are predicted to include functional variants with larger effects on risk. To identify low-frequency coding variants with large effects on LOAD risk, we carried out whole-exome sequencing (WES) in 14 large LOAD families and follow-up analyses of the candidate variants in several large LOAD case–control data sets. A rare variant in PLD3 (phospholipase D3; Val232Met) segregated with disease status in two independent families and doubled risk for Alzheimer’s disease in seven independent case–control series with a total of more than 11,000 cases and controls of European descent. Gene-based burden analyses in 4,387 cases and controls of European descent and 302 African American cases and controls, with complete sequence data for PLD3, reveal that several variants in this gene increase risk for Alzheimer’s disease in both populations. PLD3 is highly expressed in brain regions that are vulnerable to Alzheimer’s disease pathology, including hippocampus and cortex, and is expressed at significantly lower levels in neurons from Alzheimer’s disease brains compared to control brains. Overexpression of PLD3 leads to a significant decrease in intracellular amyloid-β precursor protein (APP) and extracellular Aβ42 and Aβ40 (the 42- and 40-residue isoforms of the amyloid-β peptide), and knockdown of PLD3 leads to a significant increase in extracellular Aβ42 and Aβ40. Together, our genetic and functional data indicate that carriers of PLD3 coding variants have a twofold increased risk for LOAD and that PLD3 influences APP processing. This study provides an example of how densely affected families may help to identify rare variants with large effects on risk for disease or other complex traits.
Resumo:
Bacillus amyloliquefaciens H57 is a bacterium isolated from lucerne for its ability to prevent feed spoilage. Further interest developed when ruminants fed with H57-inoculated hay showed increased weight gain and nitrogen retention relative to controls, suggesting a probiotic effect. The near complete genome of H57 is ~3.96 Mb comprising 16 contigs. Within the genome there are 3,836 protein coding genes, an estimated sixteen rRNA genes and 69 tRNA genes. H57 has the potential to synthesise four different lipopeptides and four polyketide compounds, which are known antimicrobials. This antimicrobial capacity may facilitate the observed probiotic effect.
Resumo:
The under-reporting of cases of infectious diseases is a substantial impediment to the control and management of infectious diseases in both epidemic and endemic contexts. Information about infectious disease dynamics can be recovered from sequence data using time-varying coalescent approaches, and phylodynamic models have been developed in order to reconstruct demographic changes of the numbers of infected hosts through time. In this study I have demonstrated the general concordance between empirically observed epidemiological incidence data and viral demography inferred through analysis of foot-and-mouth disease virus VP1 coding sequences belonging to the CATHAY topotype over large temporal and spatial scales. However a more precise and robust relationship between the effective population size (
Resumo:
Sequence problems belong to the most challenging interdisciplinary topics of the actuality. They are ubiquitous in science and daily life and occur, for example, in form of DNA sequences encoding all information of an organism, as a text (natural or formal) or in form of a computer program. Therefore, sequence problems occur in many variations in computational biology (drug development), coding theory, data compression, quantitative and computational linguistics (e.g. machine translation). In recent years appeared some proposals to formulate sequence problems like the closest string problem (CSP) and the farthest string problem (FSP) as an Integer Linear Programming Problem (ILPP). In the present talk we present a general novel approach to reduce the size of the ILPP by grouping isomorphous columns of the string matrix together. The approach is of practical use, since the solution of sequence problems is very time consuming, in particular when the sequences are long.
Resumo:
Trans-splicing is a common phenomenon in nematodes and kinetoplastids, and it has also been reported in other organisms, including humans. Up to now, all in silico strategies to find evidence of trans-splicing in humans have required that the candidate sequences follow the consensus splicing site rules (spliceosome-mediated mechanism). However, this criterion is not supported by the best human experimental evidence, which, except in a single case, do not follow canonical splicing sites. Moreover, recent findings describe a novel alternative tRNA mediated trans-splicing mechanism, which prescinds the spliceosome machinery. In order to answer the question, ?Are there hybrid mRNAs in sequence databanks, whose characteristics resemble those of the best human experimental evidence??, we have developed a methodology that successfully identified 16 hybrid mRNAs which might be instances of interchromosomal trans-splicing. Each hybrid mRNA is formed by a trans-spliced region (TSR), which was successfully mapped either onto known genes or onto a human endogenous retrovirus (HERV-K) transcript which supports their transcription. The existence of these hybrid mRNAs indicates that trans-splicing may be more widespread than believed. Furthermore, non-canonical splice site patterns suggest that infrequent splicing sites may occur under special conditions, or that an alternative trans-splicing mechanism is involved. Finally, our candidates are supposedly from normal tissue, and a recent study has reported that trans-splicing may occur not only in malignant tissues, but in normal tissues as well. Our methodology can be applied to 5'-UTR, coding sequences and 3'-UTR in order to find new candidates for a posteriori experimental confirmation.
Resumo:
Cutaneous melanoma (CM) is a potentially lethal form of skin cancer and its most important histopathologic factor for staging is Breslow thickness (BT). Its correct determination is fundamental for pathologists. A deeper understanding of the molecular processes guiding CM pathogenesis could improve diagnosis, treatment and prognosis. MicroRNAs (miRNAs) play a key role in CM biology. The firs aim was to investigate miRNA expression in reference to BT assessment. We found that the combined miRNA expression of miR-21-5p and miR-146a-5p above or below 1.5 was significantly associated with overall survival and successfully identified all superficially spreading melanoma (SSM) patients with relapsing suggesting that the combined assessment of these miRNAs expression could aid in SSM staging. Secondly, we focus on multiple primary melanoma (MPM) patients, which develop multiple primary melanomas in their lifetime, and represent a model of high-risk CM occurrence. We explored the miRNome of single CM and MPM: CM and MPM present several dysregulated miRNAs, including key miRNAs involved in epithelial-mesenchymal transition. A different miRNA profile was observed between 1st and 2nd melanoma from the same patient. MiRNA target analysis revealed a more differentiated and less invasive status of MPMs compared to CMs. This characterization of the miRNA regulatory network of MPMs highlights molecular features differentiating this subtype from CM. Recently, NGS experiments revealed the existence of miRNA variants (isomiRs) with different length and sequence. We identified a shorter 3’isoform as tenfold over-represented compared to the canonical form of miR-125a-5p. Target analysis revealed that miRNA shortening could change the pattern of target gene regulation. Finally, we study miRNA and isomiR dysregulation in benign nevi (BN) and CM and in CM and melanoma metastasis. The reported non-random dysregulation of specific isomiRs contributes to the understanding of the complex melanoma pathogenesis and serves as the basis for further functional studies.
Resumo:
One of the great challenges of the scientific community on theories of genetic information, genetic communication and genetic coding is to determine a mathematical structure related to DNA sequences. In this paper we propose a model of an intra-cellular transmission system of genetic information similar to a model of a power and bandwidth efficient digital communication system in order to identify a mathematical structure in DNA sequences where such sequences are biologically relevant. The model of a transmission system of genetic information is concerned with the identification, reproduction and mathematical classification of the nucleotide sequence of single stranded DNA by the genetic encoder. Hence, a genetic encoder is devised where labelings and cyclic codes are established. The establishment of the algebraic structure of the corresponding codes alphabets, mappings, labelings, primitive polynomials (p(x)) and code generator polynomials (g(x)) are quite important in characterizing error-correcting codes subclasses of G-linear codes. These latter codes are useful for the identification, reproduction and mathematical classification of DNA sequences. The characterization of this model may contribute to the development of a methodology that can be applied in mutational analysis and polymorphisms, production of new drugs and genetic improvement, among other things, resulting in the reduction of time and laboratory costs.
Resumo:
Bacillus safensis is a microorganism recognized for its biotechnological and industrial potential due to its interesting enzymatic portfolio. Here, as a means of gathering information about the importance of this species in oil biodegradation, we report a draft genome sequence of a strain isolated from petroleum.
Resumo:
Avian pathogenic Escherichia coli (APEC) strains belong to a category that is associated with colibacillosis, a serious illness in the poultry industry worldwide. Additionally, some APEC groups have recently been described as potential zoonotic agents. In this work, we compared APEC strains with extraintestinal pathogenic E. coli (ExPEC) strains isolated from clinical cases of humans with extra-intestinal diseases such as urinary tract infections (UTI) and bacteremia. PCR results showed that genes usually found in the ColV plasmid (tsh, iucA, iss, and hlyF) were associated with APEC strains while fyuA, irp-2, fepC sitDchrom, fimH, crl, csgA, afa, iha, sat, hlyA, hra, cnf1, kpsMTII, clpVSakai and malX were associated with human ExPEC. Both categories shared nine serogroups (O2, O6, O7, O8, O11, O19, O25, O73 and O153) and seven sequence types (ST10, ST88, ST93, ST117, ST131, ST155, ST359, ST648 and ST1011). Interestingly, ST95, which is associated with the zoonotic potential of APEC and is spread in avian E. coli of North America and Europe, was not detected among 76 APEC strains. When the strains were clustered based on the presence of virulence genes, most ExPEC strains (71.7%) were contained in one cluster while most APEC strains (63.2%) segregated to another. In general, the strains showed distinct genetic and fingerprint patterns, but avian and human strains of ST359, or ST23 clonal complex (CC), presented more than 70% of similarity by PFGE. The results demonstrate that some zoonotic-related STs (ST117, ST131, ST10CC, ST23CC) are present in Brazil. Also, the presence of moderate fingerprint similarities between ST359 E. coli of avian and human origin indicates that strains of this ST are candidates for having zoonotic potential.
Resumo:
A Bacillus cereus strain, FT9, isolated from a hot spring in the midwest region of Brazil, had its entire genome sequenced.
Resumo:
A monomeric basic PLA2 (PhTX-II) of 14149.08 Da molecular weight was purified to homogeneity from Porthidium hyoprora venom. Amino acid sequence by in tandem mass spectrometry revealed that PhTX-II belongs to Asp49 PLA2 enzyme class and displays conserved domains as the catalytic network, Ca2+-binding loop and the hydrophobic channel of access to the catalytic site, reflected in the high catalytic activity displayed by the enzyme. Moreover, PhTX-II PLA2 showed an allosteric behavior and its enzymatic activity was dependent on Ca2+. Examination of PhTX-II PLA2 by CD spectroscopy indicated a high content of alpha-helical structures, similar to the known structure of secreted phospholipase IIA group suggesting a similar folding. PhTX-II PLA2 causes neuromuscular blockade in avian neuromuscular preparations with a significant direct action on skeletal muscle function, as well as, induced local edema and myotoxicity, in mice. The treatment of PhTX-II by BPB resulted in complete loss of their catalytic activity that was accompanied by loss of their edematogenic effect. On the other hand, enzymatic activity of PhTX-II contributes to this neuromuscular blockade and local myotoxicity is dependent not only on enzymatic activity. These results show that PhTX-II is a myotoxic Asp49 PLA2 that contributes with toxic actions caused by P. hyoprora venom.
Resumo:
OBJECTIVE: To determine the timing and sequence of eruption of primary teeth in children with complete bilateral cleft lip and palate. MATERIAL AND METHODS: This cross-sectional study was conducted at the Hospital for Rehabilitation of Craniofacial Anomalies of the University of São Paulo, Bauru, SP, Brazil, with a sample of 395 children (128 girls and 267 boys) aged 0 to 48 months, with complete bilateral cleft lip and palate. RESULTS: Children with complete bilateral clefts presented a higher mean age of eruption of all primary teeth for both arches and both genders, compared to children without clefts. This difference was statistically signifcant for all teeth, except for the maxillary first molar. Mean age of eruption of most teeth was lower for girls compared to boys. The greatest delay was found for the maxillary lateral incisor, which was the eighth tooth of children with clefts of both genders. Analyzing by gender, the maxillary lateral incisor was the eighth tooth to erupt in girls and the last in boys. CONCLUSION: The results suggest an interference of the cleft on the timing and sequence of eruption of primary teeth.
Resumo:
BACKGROUND: Rett syndrome (RS) is a severe neurodevelopmental X-linked dominant disorder caused by mutations in the MECP2 gene. PURPOSE: To search for point mutations on the MECP2 gene and to establish a correlation between the main point mutations found and the phenotype. METHOD: Clinical evaluation of 105 patients, following a standard protocol. Detection of point mutations on the MECP2 gene was performed on peripheral blood DNA by sequencing the coding region of the gene. RESULTS: Classical RS was seen in 68% of the patients. Pathogenic point mutations were found in 64.1% of all patients and in 70.42% of those with the classical phenotype. Four new sequence variations were found, and their nature suggests patogenicity. Genotype-phenotype correlations were performed. CONCLUSION: Detailed clinical descriptions and identification of the underlying genetic alterations of this Brazilian RS population add to our knowledge of genotype/phenotype correlations, guiding the implementation of mutation searching programs.
Resumo:
Non-coding RNAs (ncRNAs) were recently given much higher attention due to technical advances in sequencing which expanded the characterization of transcriptomes in different organisms. ncRNAs have different lengths (22 nt to >1, 000 nt) and mechanisms of action that essentially comprise a sophisticated gene expression regulation network. Recent publication of schistosome genomes and transcriptomes has increased the description and characterization of a large number of parasite genes. Here we review the number of predicted genes and the coverage of genomic bases in face of the public ESTs dataset available, including a critical appraisal of the evidence and characterization of ncRNAs in schistosomes. We show expression data for ncRNAs in Schistosoma mansoni. We analyze three different microarray experiment datasets: (1) adult worms' large-scale expression measurements; (2) differentially expressed S. mansoni genes regulated by a human cytokine (TNF-α) in a parasite culture; and (3) a stage-specific expression of ncRNAs. All these data point to ncRNAs involved in different biological processes and physiological responses that suggest functionality of these new players in the parasite's biology. Exploring this world is a challenge for the scientists under a new molecular perspective of host-parasite interactions and parasite development.