991 resultados para Ordered Gene Problems
Resumo:
Ordered gene problems are a very common classification of optimization problems. Because of their popularity countless algorithms have been developed in an attempt to find high quality solutions to the problems. It is also common to see many different types of problems reduced to ordered gene style problems as there are many popular heuristics and metaheuristics for them due to their popularity. Multiple ordered gene problems are studied, namely, the travelling salesman problem, bin packing problem, and graph colouring problem. In addition, two bioinformatics problems not traditionally seen as ordered gene problems are studied: DNA error correction and DNA fragment assembly. These problems are studied with multiple variations and combinations of heuristics and metaheuristics with two distinct types or representations. The majority of the algorithms are built around the Recentering- Restarting Genetic Algorithm. The algorithm variations were successful on all problems studied, and particularly for the two bioinformatics problems. For DNA Error Correction multiple cases were found with 100% of the codes being corrected. The algorithm variations were also able to beat all other state-of-the-art DNA Fragment Assemblers on 13 out of 16 benchmark problem instances.
Resumo:
In the past decade, the advent of efficient genome sequencing tools and high-throughput experimental biotechnology has lead to enormous progress in the life science. Among the most important innovations is the microarray tecnology. It allows to quantify the expression for thousands of genes simultaneously by measurin the hybridization from a tissue of interest to probes on a small glass or plastic slide. The characteristics of these data include a fair amount of random noise, a predictor dimension in the thousand, and a sample noise in the dozens. One of the most exciting areas to which microarray technology has been applied is the challenge of deciphering complex disease such as cancer. In these studies, samples are taken from two or more groups of individuals with heterogeneous phenotypes, pathologies, or clinical outcomes. these samples are hybridized to microarrays in an effort to find a small number of genes which are strongly correlated with the group of individuals. Eventhough today methods to analyse the data are welle developed and close to reach a standard organization (through the effort of preposed International project like Microarray Gene Expression Data -MGED- Society [1]) it is not unfrequant to stumble in a clinician's question that do not have a compelling statistical method that could permit to answer it.The contribution of this dissertation in deciphering disease regards the development of new approaches aiming at handle open problems posed by clinicians in handle specific experimental designs. In Chapter 1 starting from a biological necessary introduction, we revise the microarray tecnologies and all the important steps that involve an experiment from the production of the array, to the quality controls ending with preprocessing steps that will be used into the data analysis in the rest of the dissertation. While in Chapter 2 a critical review of standard analysis methods are provided stressing most of problems that In Chapter 3 is introduced a method to adress the issue of unbalanced design of miacroarray experiments. In microarray experiments, experimental design is a crucial starting-point for obtaining reasonable results. In a two-class problem, an equal or similar number of samples it should be collected between the two classes. However in some cases, e.g. rare pathologies, the approach to be taken is less evident. We propose to address this issue by applying a modified version of SAM [2]. MultiSAM consists in a reiterated application of a SAM analysis, comparing the less populated class (LPC) with 1,000 random samplings of the same size from the more populated class (MPC) A list of the differentially expressed genes is generated for each SAM application. After 1,000 reiterations, each single probe given a "score" ranging from 0 to 1,000 based on its recurrence in the 1,000 lists as differentially expressed. The performance of MultiSAM was compared to the performance of SAM and LIMMA [3] over two simulated data sets via beta and exponential distribution. The results of all three algorithms over low- noise data sets seems acceptable However, on a real unbalanced two-channel data set reagardin Chronic Lymphocitic Leukemia, LIMMA finds no significant probe, SAM finds 23 significantly changed probes but cannot separate the two classes, while MultiSAM finds 122 probes with score >300 and separates the data into two clusters by hierarchical clustering. We also report extra-assay validation in terms of differentially expressed genes Although standard algorithms perform well over low-noise simulated data sets, multi-SAM seems to be the only one able to reveal subtle differences in gene expression profiles on real unbalanced data. In Chapter 4 a method to adress similarities evaluation in a three-class prblem by means of Relevance Vector Machine [4] is described. In fact, looking at microarray data in a prognostic and diagnostic clinical framework, not only differences could have a crucial role. In some cases similarities can give useful and, sometimes even more, important information. The goal, given three classes, could be to establish, with a certain level of confidence, if the third one is similar to the first or the second one. In this work we show that Relevance Vector Machine (RVM) [2] could be a possible solutions to the limitation of standard supervised classification. In fact, RVM offers many advantages compared, for example, with his well-known precursor (Support Vector Machine - SVM [3]). Among these advantages, the estimate of posterior probability of class membership represents a key feature to address the similarity issue. This is a highly important, but often overlooked, option of any practical pattern recognition system. We focused on Tumor-Grade-three-class problem, so we have 67 samples of grade I (G1), 54 samples of grade 3 (G3) and 100 samples of grade 2 (G2). The goal is to find a model able to separate G1 from G3, then evaluate the third class G2 as test-set to obtain the probability for samples of G2 to be member of class G1 or class G3. The analysis showed that breast cancer samples of grade II have a molecular profile more similar to breast cancer samples of grade I. Looking at the literature this result have been guessed, but no measure of significance was gived before.
Resumo:
This paper aims to show, analyze and solve the problems related to the translation of the book with meaning-bond alphabetically ordered chapter titles La vita non è in ordine alfabetico, by the Italian writer Andrea Bajani. The procedure is inevitable for a possible translation of the book, and it is necessary to have a preliminary pattern to follow. After translating the whole book, not only is a revision fundamental, but a restructure and reorganization of the collection may be required, hence, what this thesis offers is a scheme to start from, together with an analysis of the possible problems that may arise, and a useful method to find a solution.
Resumo:
Due to the imprecise nature of biological experiments, biological data is often characterized by the presence of redundant and noisy data. This may be due to errors that occurred during data collection, such as contaminations in laboratorial samples. It is the case of gene expression data, where the equipments and tools currently used frequently produce noisy biological data. Machine Learning algorithms have been successfully used in gene expression data analysis. Although many Machine Learning algorithms can deal with noise, detecting and removing noisy instances from the training data set can help the induction of the target hypothesis. This paper evaluates the use of distance-based pre-processing techniques for noise detection in gene expression data classification problems. This evaluation analyzes the effectiveness of the techniques investigated in removing noisy data, measured by the accuracy obtained by different Machine Learning classifiers over the pre-processed data.
Resumo:
Background: The inference of gene regulatory networks (GRNs) from large-scale expression profiles is one of the most challenging problems of Systems Biology nowadays. Many techniques and models have been proposed for this task. However, it is not generally possible to recover the original topology with great accuracy, mainly due to the short time series data in face of the high complexity of the networks and the intrinsic noise of the expression measurements. In order to improve the accuracy of GRNs inference methods based on entropy (mutual information), a new criterion function is here proposed. Results: In this paper we introduce the use of generalized entropy proposed by Tsallis, for the inference of GRNs from time series expression profiles. The inference process is based on a feature selection approach and the conditional entropy is applied as criterion function. In order to assess the proposed methodology, the algorithm is applied to recover the network topology from temporal expressions generated by an artificial gene network (AGN) model as well as from the DREAM challenge. The adopted AGN is based on theoretical models of complex networks and its gene transference function is obtained from random drawing on the set of possible Boolean functions, thus creating its dynamics. On the other hand, DREAM time series data presents variation of network size and its topologies are based on real networks. The dynamics are generated by continuous differential equations with noise and perturbation. By adopting both data sources, it is possible to estimate the average quality of the inference with respect to different network topologies, transfer functions and network sizes. Conclusions: A remarkable improvement of accuracy was observed in the experimental results by reducing the number of false connections in the inferred topology by the non-Shannon entropy. The obtained best free parameter of the Tsallis entropy was on average in the range 2.5 <= q <= 3.5 (hence, subextensive entropy), which opens new perspectives for GRNs inference methods based on information theory and for investigation of the nonextensivity of such networks. The inference algorithm and criterion function proposed here were implemented and included in the DimReduction software, which is freely available at http://sourceforge.net/projects/dimreduction and http://code.google.com/p/dimreduction/.
Resumo:
in the Apis mellifera post-genomic era, RNAi protocols have been used in functional approaches. However, sample manipulation and invasive methods such as injection of double-stranded RNA (dsRNA) can compromise physiology and survival. To circumvent these problems, we developed a non-invasive method for honeybee gene knockdown, using a well-established vitellogenin RNAi system as a model. Second instar larvae received dsRNA for vitellogenin (dsVg-RNA) in their natural diet. For exogenous control, larvae received dsRNA for GFP (dsGFP-RNA). Untreated larvae formed another control group. Around 60% of the treated larvae naturally developed until adult emergence when 0.5 mu g of dsVg-RNA or dsGFP-RNA was offered while no larvae that received 3.0 mu g of dsRNA reached pupal stages. Diet dilution did not affect the removal rates. Viability depends not only on the delivered doses but also on the internal conditions of colonies. The weight of treated and untreated groups showed no statistical differences. This showed that RNAi ingestion did not elicit drastic collateral effects. Approximately 90% of vitellogenin transcripts from 7-day-old workers were silenced compared to controls. A large number of samples are handled in a relatively short time and smaller quantities of RNAi molecules are used compared to invasive methods. These advantages culminate in a versatile and a cost-effective approach. (c) 2008 Elsevier Ltd. All rights reserved.
Resumo:
Plant transformation is now a core research tool in plant biology and a practical tool for cultivar improvement. There are verified methods for stable introduction of novel genes into the nuclear genomes of over 120 diverse plant species. This review examines the criteria to verify plant transformation; the biological and practical requirements for transformation systems; the integration of tissue culture, gene transfer, selection, and transgene expression strategies to achieve transformation in recalcitrant species; and other constraints to plant transformation including regulatory environment, public perceptions, intellectual property, and economics. Because the costs of screening populations showing diverse genetic changes can far exceed the costs of transformation, it is important to distinguish absolute and useful transformation efficiencies. The major technical challenge facing plant transformation biology is the development of methods and constructs to produce a high proportion of plants showing predictable transgene expression without collateral genetic damage. This will require answers to a series of biological and technical questions, some of which are defined.
Resumo:
Background The continued increase in tuberculosis (TB) rates and the appearance of extremely resistant Mycobacterium tuberculosis strains (XDR-TB) worldwide are some of the great problems of public health. In this context, DNA immunotherapy has been proposed as an effective alternative that could circumvent the limitations of conventional drugs. Nonetheless, the molecular events underlying these therapeutic effects are poorly understood. Methods We characterized the transcriptional signature of lungs from mice infected with M. tuberculosis and treated with heat shock protein 65 as a genetic vaccine (DNAhsp65) combining microarray and real-time polymerase chain reaction analysis. The gene expression data were correlated with the histopathological analysis of lungs. Results The differential modulation of a high number of genes allowed us to distinguish DNAhsp65-treated from nontreated animals (saline and vector-injected mice). Functional analysis of this group of genes suggests that DNAhsp65 therapy could not only boost the T helper (Th)1 immune response, but also could inhibit Th2 cytokines and regulate the intensity of inflammation through fine tuning of gene expression of various genes, including those of interleukin-17, lymphotoxin A, tumour necrosis factor-cl, interleukin-6, transforming growth factor-beta, inducible nitric oxide synthase and Foxp3. In addition, a large number of genes and expressed sequence tags previously unrelated to DNA-therapy were identified. All these findings were well correlated with the histopathological lesions presented in the lungs. Conclusions The effects of DNA therapy are reflected in gene expression modulation; therefore, the genes identified as differentially expressed could be considered as transcriptional biomarkers of DNAhsp65 immunotherapy against TB. The data have important implications for achieving a better understanding of gene-based therapies. Copyright (C) 2008 John Wiley & Sons, Ltd.
Resumo:
In microarray studies, the application of clustering techniques is often used to derive meaningful insights into the data. In the past, hierarchical methods have been the primary clustering tool employed to perform this task. The hierarchical algorithms have been mainly applied heuristically to these cluster analysis problems. Further, a major limitation of these methods is their inability to determine the number of clusters. Thus there is a need for a model-based approach to these. clustering problems. To this end, McLachlan et al. [7] developed a mixture model-based algorithm (EMMIX-GENE) for the clustering of tissue samples. To further investigate the EMMIX-GENE procedure as a model-based -approach, we present a case study involving the application of EMMIX-GENE to the breast cancer data as studied recently in van 't Veer et al. [10]. Our analysis considers the problem of clustering the tissue samples on the basis of the genes which is a non-standard problem because the number of genes greatly exceed the number of tissue samples. We demonstrate how EMMIX-GENE can be useful in reducing the initial set of genes down to a more computationally manageable size. The results from this analysis also emphasise the difficulty associated with the task of separating two tissue groups on the basis of a particular subset of genes. These results also shed light on why supervised methods have such a high misallocation error rate for the breast cancer data.
Resumo:
A obesidade e a diabetes mellitus tipo 2 (DM2) são considerados dois grandes problemas de saúde pública. A má alimentação e a falta de atividade física encontram-se entre os principais desencadeadores de um crescente número de indivíduos obesos, diabéticos e com sensibilidade à insulina diminuída. Este aumento tem motivado a comunidade científica a investigar cada vez mais para o elevado contributo da herança genética associada aos fatores sociais e nutricionais. O gene dos recetores ativados por proliferadores do peroxissoma gama 2 (PPARγ2) desempenha um papel importante no metabolismo lipídico. Uma vez que o PPARγ2 é maioritariamente expresso no tecido adiposo, uma redução moderada da sua atividade tem influência na sensibilidade à insulina, diabetes, e outros parâmetros metabólicos. Vários estudos sugerem que tanto fatores genéticos como fatores ambientais (tais como a dieta), poderão estar envolvidos na formação de padrões associados ao polimorfismo Pro12Ala com a composição corporal em diferentes populações humanas. Os diversos estudos genéticos envolvendo o estudo do polimorfismo Pro12Ala do PPARγ2 na suscetibilidade de possuir risco de diabetes e obesidade em várias populações têm proposto conclusões diversas. Em alguns parece haver mais associações do que outros e, às vezes, não demonstram sequer associação. Desta forma, o presente trabalho teve como objectivo contribuir para a elucidação do impacto do polimorfismo Pro12Ala do PPARγ2 na resistência à insulina associada à DM2 e na obesidade, mediante estudo sistematizado da literatura existente até à data, através de meta análise. Do total de uma pesquisa de 63 publicações, foram incluídos 32 artigos no presente estudo, sendo que destes 25 foram incluídos na síntese qualitativa e 11 incluídos na sintese quantitativa. No presente trabalho pode-se concluir que existe evidência estatística que suporta a hipótese de que o polimorfismo Pro12Ala do PPARγ2 pode ser considerado um fator protetor para a DM2 [p <0,05 e OR (odds ratio) 0,702, com IC (intervalos de confiança) com valores que nunca incluem o 1]. No entanto, e mediante os mesmos pressupostos, o mesmo polimorfismo pode ser considerado um fator de risco ao desenvolvimento de obesidade, pela evidência estatística [p <0,05 e OR de 1,196, com IC com valores que nunca incluem o 1].
Resumo:
PURPOSE: To establish the Southern blotting technique using hybridization with a nonradioactive probe to detect large rearrangements of CYP21A2 in a Brazilian cohort with congenital adrenal hyperplasia due to 21-hydroxylase deficiency (CAH-21OH). METHOD: We studied 42 patients, 2 of them related, comprising 80 non-related alleles. DNA samples were obtained from peripheral blood, digested by restriction enzyme Taq I, submitted to Southern blotting and hybridized with biotin-labeled probes. RESULTS: This method was shown to be reliable with results similar to the radioactive-labeling method. We found CYP21A2 deletion (2.5%), large gene conversion (8.8%), CYP21AP deletion (3.8%), and CYP21A1P duplication (6.3%). These frequencies were similar to those found in our previous study in which a large number of cases were studied. Good hybridization patterns were achieved with a smaller amount of DNA (5 mug), and fragment signs were observed after 5 minutes to 1 hour of exposure. CONCLUSIONS: We established a non-radioactive (biotin) Southern blot/hybridization methodology for CYP21A2 large rearrangements with good results. Despite being more arduous, this technique is faster, requires a smaller amount of DNA, and most importantly, avoids problems with the use of radioactivity.
Resumo:
In longitudinal studies of disease, patients may experience several events through a follow-up period. In these studies, the sequentially ordered events are often of interest and lead to problems that have received much attention recently. Issues of interest include the estimation of bivariate survival, marginal distributions and the conditional distribution of gap times. In this work we consider the estimation of the survival function conditional to a previous event. Different nonparametric approaches will be considered for estimating these quantities, all based on the Kaplan-Meier estimator of the survival function. We explore the finite sample behavior of the estimators through simulations. The different methods proposed in this article are applied to a data set from a German Breast Cancer Study. The methods are used to obtain predictors for the conditional survival probabilities as well as to study the influence of recurrence in overall survival.
Resumo:
MOTIVATION: In silico modeling of gene regulatory networks has gained some momentum recently due to increased interest in analyzing the dynamics of biological systems. This has been further facilitated by the increasing availability of experimental data on gene-gene, protein-protein and gene-protein interactions. The two dynamical properties that are often experimentally testable are perturbations and stable steady states. Although a lot of work has been done on the identification of steady states, not much work has been reported on in silico modeling of cellular differentiation processes. RESULTS: In this manuscript, we provide algorithms based on reduced ordered binary decision diagrams (ROBDDs) for Boolean modeling of gene regulatory networks. Algorithms for synchronous and asynchronous transition models have been proposed and their corresponding computational properties have been analyzed. These algorithms allow users to compute cyclic attractors of large networks that are currently not feasible using existing software. Hereby we provide a framework to analyze the effect of multiple gene perturbation protocols, and their effect on cell differentiation processes. These algorithms were validated on the T-helper model showing the correct steady state identification and Th1-Th2 cellular differentiation process. AVAILABILITY: The software binaries for Windows and Linux platforms can be downloaded from http://si2.epfl.ch/~garg/genysis.html.
Resumo:
Background. Microglia and astrocytes respond to homeostatic disturbances with profound changes of gene expression. This response, known as glial activation or neuroinflammation, can be detrimental to the surrounding tissue. The transcription factor CCAAT/enhancer binding protein ß (C/EBPß) is an important regulator of gene expression in inflammation but little is known about its involvement in glial activation. To explore the functional role of C/EBPß in glial activation we have analyzed pro-inflammatory gene expression and neurotoxicity in murine wild type and C/EBPß-null glial cultures. Methods. Due to fertility and mortality problems associated with the C/EBPß-null genotype we developed a protocol to prepare mixed glial cultures from cerebral cortex of a single mouse embryo with high yield. Wild-type and C/EBPß-null glial cultures were compared in terms of total cell density by Hoechst-33258 staining; microglial content by CD11b immunocytochemistry; astroglial content by GFAP western blot; gene expression by quantitative real-time PCR, western blot, immunocytochemistry and Griess reaction; and microglial neurotoxicity by estimating MAP2 content in neuronal/microglial cocultures. C/EBPß DNA binding activity was evaluated by electrophoretic mobility shift assay and quantitative chromatin immunoprecipitation. Results. C/EBPß mRNA and protein levels, as well as DNA binding, were increased in glial cultures by treatment with lipopolysaccharide (LPS) or LPS + interferon ¿ (IFN¿). Quantitative chromatin immunoprecipitation showed binding of C/EBPß to pro-inflammatory gene promoters in glial activation in a stimulus- and gene-dependent manner. In agreement with these results, LPS and LPS+IFN¿ induced different transcriptional patterns between pro-inflammatory cytokines and NO synthase-2 genes. Furthermore, the expressions of IL-1ß and NO synthase-2, and consequent NO production, were reduced in the absence of C/EBPß. In addition, neurotoxicity elicited by LPS+IFN¿-treated microglia co-cultured with neurons was completely abolished by the absence of C/EBPß in microglia.
Resumo:
Integrating viral vectors hold great promise as gene transfer vectors for gene therapy purposes because they allow maintaining long-term expression of the therapeutic transgene throughout cell divisions. However, many issues related to integration of the provirus remain as a substantial risk for patients. The use of chromatin insulators has been proposed as a possible solution to problems raised by the integration of the vector.