929 resultados para GENE NETWORK INTERACTIONS
Resumo:
Assaying a large number of genetic markers from patients in clinical trials is now possible in order to tailor drugs with respect to efficacy. The statistical methodology for analysing such massive data sets is challenging. The most popular type of statistical analysis is to use a univariate test for each genetic marker, once all the data from a clinical study have been collected. This paper presents a sequential method for conducting an omnibus test for detecting gene-drug interactions across the genome, thus allowing informed decisions at the earliest opportunity and overcoming the multiple testing problems from conducting many univariate tests. We first propose an omnibus test for a fixed sample size. This test is based on combining F-statistics that test for an interaction between treatment and the individual single nucleotide polymorphism (SNP). As SNPs tend to be correlated, we use permutations to calculate a global p-value. We extend our omnibus test to the sequential case. In order to control the type I error rate, we propose a sequential method that uses permutations to obtain the stopping boundaries. The results of a simulation study show that the sequential permutation method is more powerful than alternative sequential methods that control the type I error rate, such as the inverse-normal method. The proposed method is flexible as we do not need to assume a mode of inheritance and can also adjust for confounding factors. An application to real clinical data illustrates that the method is computationally feasible for a large number of SNPs. Copyright (c) 2007 John Wiley & Sons, Ltd.
Resumo:
Human mesenchymal stem cells (MSC) are powerful sources for cell therapy in regenerative medicine. The long time cultivation can result in replicative senescence or can be related to the emergence of chromosomal alterations responsible for the acquisition of tumorigenesis features in vitro. In this study, for the first time, the expression profile of MSC with a paracentric chromosomal inversion (MSC/inv) was compared to normal karyotype (MSC/n) in early and late passages. Furthermore, we compared the transcriptome of each MSC in early passages with late passages. MSC used in this study were obtained from the umbilical vein of three donors, two MSC/n and one MSC/inv. After their cryopreservation, they have been expanded in vitro until reached senescence. Total RNA was extracted using the RNeasy mini kit (Qiagen) and marked with the GeneChip ® 3 IVT Express Kit (Affymetrix Inc.). Subsequently, the fragmented aRNA was hybridized on the microarranjo Affymetrix Human Genome U133 Plus 2.0 arrays (Affymetrix Inc.). The statistical analysis of differential gene expression was performed between groups MSC by the Partek Genomic Suite software, version 6.4 (Partek Inc.). Was considered statistically significant differences in expression to p-value Bonferroni correction ˂.01. Only signals with fold change ˃ 3.0 were included in the list of differentially expressed. Differences in gene expression data obtained from microarrays were confirmed by Real Time RT-PCR. For the interpretation of biological expression data were used: IPA (Ingenuity Systems) for analysis enrichment functions, the STRING 9.0 for construction of network interactions; Cytoscape 2.8 to the network visualization and analysis bottlenecks with the aid of the GraphPad Prism 5.0 software. BiNGO Cytoscape pluggin was used to access overrepresentation of Gene Ontology categories in Biological Networks. The comparison between senescent and young at each group of MSC has shown that there is a difference in the expression parttern, being higher in the senescent MSC/inv group. The results also showed difference in expression profiles between the MSC/inv versus MSC/n, being greater when they are senescent. New networks were identified for genes related to the response of two of MSC over cultivation time. Were also identified genes that can coordinate functional categories over represented at networks, such as CXCL12, SFRP1, xvi EGF, SPP1, MMP1 e THBS1. The biological interpretation of these data suggests that the population of MSC/inv has different constitutional characteristics, related to their potential for differentiation, proliferation and response to stimuli, responsible for a distinct process of replicative senescence in MSC/inv compared to MSC/n. The genes identified in this study are candidates for biomarkers of cellular senescence in MSC, but their functional relevance in this process should be evaluated in additional in vitro and/or in vivo assays
Resumo:
Autism is a neurodevelopmental disorder characterized by impaired social interaction and communication accompanied with repetitive behavioral patterns and unusual stereotyped interests. Autism is considered a highly heterogeneous disorder with diverse putative causes and associated factors giving rise to variable ranges of symptomatology. Incidence seems to be increasing with time, while the underlying pathophysiological mechanisms remain virtually uncharacterized (or unknown). By systematic review of the literature and a systems biology approach, our aims were to examine the multifactorial nature of autism with its broad range of severity, to ascertain the predominant biological processes, cellular components, and molecular functions integral to the disorder, and finally, to elucidate the most central contributions (genetic and/or environmental) in silico. With this goal, we developed an integrative network model for gene-environment interactions (GENVI model) where calcium (Ca2+) was shown to be its most relevant node. Moreover, considering the present data from our systems biology approach together with the results from the differential gene expression analysis of cerebellar samples from autistic patients, we believe that RAC1, in particular, and the RHO family of GTPases, in general, could play a critical role in the neuropathological events associated with autism. © 2013 Springer Science+Business Media New York.
Resumo:
Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP)
Resumo:
Background: In the analysis of effects by cell treatment such as drug dosing, identifying changes on gene network structures between normal and treated cells is a key task. A possible way for identifying the changes is to compare structures of networks estimated from data on normal and treated cells separately. However, this approach usually fails to estimate accurate gene networks due to the limited length of time series data and measurement noise. Thus, approaches that identify changes on regulations by using time series data on both conditions in an efficient manner are demanded. Methods: We propose a new statistical approach that is based on the state space representation of the vector autoregressive model and estimates gene networks on two different conditions in order to identify changes on regulations between the conditions. In the mathematical model of our approach, hidden binary variables are newly introduced to indicate the presence of regulations on each condition. The use of the hidden binary variables enables an efficient data usage; data on both conditions are used for commonly existing regulations, while for condition specific regulations corresponding data are only applied. Also, the similarity of networks on two conditions is automatically considered from the design of the potential function for the hidden binary variables. For the estimation of the hidden binary variables, we derive a new variational annealing method that searches the configuration of the binary variables maximizing the marginal likelihood. Results: For the performance evaluation, we use time series data from two topologically similar synthetic networks, and confirm that our proposed approach estimates commonly existing regulations as well as changes on regulations with higher coverage and precision than other existing approaches in almost all the experimental settings. For a real data application, our proposed approach is applied to time series data from normal Human lung cells and Human lung cells treated by stimulating EGF-receptors and dosing an anticancer drug termed Gefitinib. In the treated lung cells, a cancer cell condition is simulated by the stimulation of EGF-receptors, but the effect would be counteracted due to the selective inhibition of EGF-receptors by Gefitinib. However, gene expression profiles are actually different between the conditions, and the genes related to the identified changes are considered as possible off-targets of Gefitinib. Conclusions: From the synthetically generated time series data, our proposed approach can identify changes on regulations more accurately than existing methods. By applying the proposed approach to the time series data on normal and treated Human lung cells, candidates of off-target genes of Gefitinib are found. According to the published clinical information, one of the genes can be related to a factor of interstitial pneumonia, which is known as a side effect of Gefitinib.
Resumo:
Background The genetic mechanisms underlying interindividual blood pressure variation reflect the complex interplay of both genetic and environmental variables. The current standard statistical methods for detecting genes involved in the regulation mechanisms of complex traits are based on univariate analysis. Few studies have focused on the search for and understanding of quantitative trait loci responsible for gene × environmental interactions or multiple trait analysis. Composite interval mapping has been extended to multiple traits and may be an interesting approach to such a problem. Methods We used multiple-trait analysis for quantitative trait locus mapping of loci having different effects on systolic blood pressure with NaCl exposure. Animals studied were 188 rats, the progenies of an F2 rat intercross between the hypertensive and normotensive strain, genotyped in 179 polymorphic markers across the rat genome. To accommodate the correlational structure from measurements taken in the same animals, we applied univariate and multivariate strategies for analyzing the data. Results We detected a new quantitative train locus on a region close to marker R589 in chromosome 5 of the rat genome, not previously identified through serial analysis of individual traits. In addition, we were able to justify analytically the parametric restrictions in terms of regression coefficients responsible for the gain in precision with the adopted analytical approach. Conclusion Future work should focus on fine mapping and the identification of the causative variant responsible for this quantitative trait locus signal. The multivariable strategy might be valuable in the study of genetic determinants of interindividual variation of antihypertensive drug effectiveness.
Resumo:
Abstract Background To understand the molecular mechanisms underlying important biological processes, a detailed description of the gene products networks involved is required. In order to define and understand such molecular networks, some statistical methods are proposed in the literature to estimate gene regulatory networks from time-series microarray data. However, several problems still need to be overcome. Firstly, information flow need to be inferred, in addition to the correlation between genes. Secondly, we usually try to identify large networks from a large number of genes (parameters) originating from a smaller number of microarray experiments (samples). Due to this situation, which is rather frequent in Bioinformatics, it is difficult to perform statistical tests using methods that model large gene-gene networks. In addition, most of the models are based on dimension reduction using clustering techniques, therefore, the resulting network is not a gene-gene network but a module-module network. Here, we present the Sparse Vector Autoregressive model as a solution to these problems. Results We have applied the Sparse Vector Autoregressive model to estimate gene regulatory networks based on gene expression profiles obtained from time-series microarray experiments. Through extensive simulations, by applying the SVAR method to artificial regulatory networks, we show that SVAR can infer true positive edges even under conditions in which the number of samples is smaller than the number of genes. Moreover, it is possible to control for false positives, a significant advantage when compared to other methods described in the literature, which are based on ranks or score functions. By applying SVAR to actual HeLa cell cycle gene expression data, we were able to identify well known transcription factor targets. Conclusion The proposed SVAR method is able to model gene regulatory networks in frequent situations in which the number of samples is lower than the number of genes, making it possible to naturally infer partial Granger causalities without any a priori information. In addition, we present a statistical test to control the false discovery rate, which was not previously possible using other gene regulatory network models.
Resumo:
Despite current enthusiasm for investigation of gene-gene interactions and gene-environment interactions, the essential issue of how to define and detect gene-environment interactions remains unresolved. In this report, we define gene-environment interactions as a stochastic dependence in the context of the effects of the genetic and environmental risk factors on the cause of phenotypic variation among individuals. We use mutual information that is widely used in communication and complex system analysis to measure gene-environment interactions. We investigate how gene-environment interactions generate the large difference in the information measure of gene-environment interactions between the general population and a diseased population, which motives us to develop mutual information-based statistics for testing gene-environment interactions. We validated the null distribution and calculated the type 1 error rates for the mutual information-based statistics to test gene-environment interactions using extensive simulation studies. We found that the new test statistics were more powerful than the traditional logistic regression under several disease models. Finally, in order to further evaluate the performance of our new method, we applied the mutual information-based statistics to three real examples. Our results showed that P-values for the mutual information-based statistics were much smaller than that obtained by other approaches including logistic regression models.
Resumo:
A wealth of genetic associations for cardiovascular and metabolic phenotypes in humans has been accumulating over the last decade, in particular a large number of loci derived from recent genome wide association studies (GWAS). True complex disease-associated loci often exert modest effects, so their delineation currently requires integration of diverse phenotypic data from large studies to ensure robust meta-analyses. We have designed a gene-centric 50 K single nucleotide polymorphism (SNP) array to assess potentially relevant loci across a range of cardiovascular, metabolic and inflammatory syndromes. The array utilizes a "cosmopolitan" tagging approach to capture the genetic diversity across approximately 2,000 loci in populations represented in the HapMap and SeattleSNPs projects. The array content is informed by GWAS of vascular and inflammatory disease, expression quantitative trait loci implicated in atherosclerosis, pathway based approaches and comprehensive literature searching. The custom flexibility of the array platform facilitated interrogation of loci at differing stringencies, according to a gene prioritization strategy that allows saturation of high priority loci with a greater density of markers than the existing GWAS tools, particularly in African HapMap samples. We also demonstrate that the IBC array can be used to complement GWAS, increasing coverage in high priority CVD-related loci across all major HapMap populations. DNA from over 200,000 extensively phenotyped individuals will be genotyped with this array with a significant portion of the generated data being released into the academic domain facilitating in silico replication attempts, analyses of rare variants and cross-cohort meta-analyses in diverse populations. These datasets will also facilitate more robust secondary analyses, such as explorations with alternative genetic models, epistasis and gene-environment interactions.
Resumo:
Withdrawal reflexes of the mollusk Aplysia exhibit sensitization, a simple form of long-term memory (LTM). Sensitization is due, in part, to long-term facilitation (LTF) of sensorimotor neuron synapses. LTF is induced by the modulatory actions of serotonin (5-HT). Pettigrew et al. developed a computational model of the nonlinear intracellular signaling and gene network that underlies the induction of 5-HT-induced LTF. The model simulated empirical observations that repeated applications of 5-HT induce persistent activation of protein kinase A (PKA) and that this persistent activation requires a suprathreshold exposure of 5-HT. This study extends the analysis of the Pettigrew model by applying bifurcation analysis, singularity theory, and numerical simulation. Using singularity theory, classification diagrams of parameter space were constructed, identifying regions with qualitatively different steady-state behaviors. The graphical representation of these regions illustrates the robustness of these regions to changes in model parameters. Because persistent protein kinase A (PKA) activity correlates with Aplysia LTM, the analysis focuses on a positive feedback loop in the model that tends to maintain PKA activity. In this loop, PKA phosphorylates a transcription factor (TF-1), thereby increasing the expression of an ubiquitin hydrolase (Ap-Uch). Ap-Uch then acts to increase PKA activity, closing the loop. This positive feedback loop manifests multiple, coexisting steady states, or multiplicity, which provides a mechanism for a bistable switch in PKA activity. After the removal of 5-HT, the PKA activity either returns to its basal level (reversible switch) or remains at a high level (irreversible switch). Such an irreversible switch might be a mechanism that contributes to the persistence of LTM. The classification diagrams also identify parameters and processes that might be manipulated, perhaps pharmacologically, to enhance the induction of memory. Rational drug design, to affect complex processes such as memory formation, can benefit from this type of analysis.
Resumo:
This investigation examined the clonal dynamics of B-cell expression and evaluated the role of idiotype network interactions in shaping the expressed secondary B-cell repertoire. Three interrelated experimental approaches were applied. The first approach was designed to distinguish between regulatory influences controlled by the major histocompatibility complex (MHC) and regulatory influences controlled by non-MHC factors including the idiotype network. This approach consisted of studies on the clonal dynamics and heterogeneity of the expressed IgG antibody repertoire of BALB/c mice. The second approach involved the analysis of the clonal dynamics of antibody responses of outbred rabbits. This analysis was coupled with studies to detect the occurrence and activity of constituents of the idiotype network. In the third approach the transfer of rabbit lymphocytes from immunized donors to MHC matched naive recipients was used to examine the effects of recipient non-MHC immunoregulatory influences on the expression of donor memory B-cells. Although many memory B cells were unaffected by non-MHC influences, these data show that non-MHC immunoregulatory influences can affect the expression of B-cells in the secondary response of inbred mice and outbred rabbits. The results also indicate that most IgG antibody responses are heterogeneous and are characterized by a stable group of dominant clonotypes. Clonal dominance and B-cell memory were found to be established early in an immune response. The expression of B memory clones appeared to be favored over the expression of virgin B cells. The injection of anti-tetanus antibody induced the antigen independent production of anti-tetanus antibody, probably through idiotypic mechanisms. These results demonstrate that both antibody and antigen can affect the expressed B-ceIl repertoire. Thus, idiotypic interactions are capable of influencing the expression of B-cells and these findings support the existence and function of an idiotype network with strong immunoregulatory potential. ^
Resumo:
Epidemiological studies have led to the hypothesis that major risk factors for developing diseases such as hypertension, cardiovascular disease and adult-onset diabetes are established during development. This developmental programming hypothesis proposes that exposure to an adverse stimulus or insult at critical, sensitive periods of development can induce permanent alterations in normal physiological processes that lead to increased disease risk later in life. For cancer, inheritance of a tumor suppressor gene defect confers a high relative risk for disease development. However, these defects are rarely 100% penetrant. Traditionally, gene-environment interactions are thought to contribute to the penetrance of tumor suppressor gene defects by facilitating or inhibiting the acquisition of additional somatic mutations required for tumorigenesis. The studies presented herein identify developmental programming as a distinctive type of gene-environment interaction that can enhance the penetrance of a tumor suppressor gene defect in adult life. Using rats predisposed to uterine leiomyoma due to a germ-line defect in one allele of the tuberous sclerosis complex 2 (Tsc-2) tumor suppressor gene, these studies show that early-life exposure to the xenoestrogen, diethylstilbestrol (DES), during development of the uterus increased tumor incidence, multiplicity and size in genetically predisposed animals, but failed to induce tumors in wild-type rats. Uterine leiomyomas are ovarian-hormone dependent tumors that develop from the uterine myometrium. DES exposure was shown to developmentally program the myometrium, causing increased expression of estrogen-responsive genes prior to the onset of tumors. Loss of function of the normal Tsc-2 allele remained the rate-limiting event for tumorigenesis; however, tumors that developed in exposed animals displayed an enhanced proliferative response to ovarian steroid hormones relative to tumors that developed in unexposed animals. Furthermore, the studies presented herein identify developmental periods during which target tissues are maximally susceptible to developmental programming. These data suggest that exposure to environmental factors during critical periods of development can permanently alter normal physiological tissue responses and thus lead to increased disease risk in genetically susceptible individuals. ^
Resumo:
Genome-wide association studies (GWAS) have rapidly become a standard method for disease gene discovery. Many recent GWAS indicate that for most disorders, only a few common variants are implicated and the associated SNPs explain only a small fraction of the genetic risk. The current study incorporated gene network information into gene-based analysis of GWAS data for Crohn's disease (CD). The purpose was to develop statistical models to boost the power of identifying disease-associated genes and gene subnetworks by maximizing the use of existing biological knowledge from multiple sources. The results revealed that Markov random field (MRF) based mixture model incorporating direct neighborhood information from a single gene network is not efficient in identifying CD-related genes based on the GWAS data. The incorporation of solely direct neighborhood information might lead to the low efficiency of these models. Alternative MRF models looking beyond direct neighboring information are necessary to be developed in the future for the purpose of this study.^
Resumo:
My dissertation focuses on developing methods for gene-gene/environment interactions and imprinting effect detections for human complex diseases and quantitative traits. It includes three sections: (1) generalizing the Natural and Orthogonal interaction (NOIA) model for the coding technique originally developed for gene-gene (GxG) interaction and also to reduced models; (2) developing a novel statistical approach that allows for modeling gene-environment (GxE) interactions influencing disease risk, and (3) developing a statistical approach for modeling genetic variants displaying parent-of-origin effects (POEs), such as imprinting. In the past decade, genetic researchers have identified a large number of causal variants for human genetic diseases and traits by single-locus analysis, and interaction has now become a hot topic in the effort to search for the complex network between multiple genes or environmental exposures contributing to the outcome. Epistasis, also known as gene-gene interaction is the departure from additive genetic effects from several genes to a trait, which means that the same alleles of one gene could display different genetic effects under different genetic backgrounds. In this study, we propose to implement the NOIA model for association studies along with interaction for human complex traits and diseases. We compare the performance of the new statistical models we developed and the usual functional model by both simulation study and real data analysis. Both simulation and real data analysis revealed higher power of the NOIA GxG interaction model for detecting both main genetic effects and interaction effects. Through application on a melanoma dataset, we confirmed the previously identified significant regions for melanoma risk at 15q13.1, 16q24.3 and 9p21.3. We also identified potential interactions with these significant regions that contribute to melanoma risk. Based on the NOIA model, we developed a novel statistical approach that allows us to model effects from a genetic factor and binary environmental exposure that are jointly influencing disease risk. Both simulation and real data analyses revealed higher power of the NOIA model for detecting both main genetic effects and interaction effects for both quantitative and binary traits. We also found that estimates of the parameters from logistic regression for binary traits are no longer statistically uncorrelated under the alternative model when there is an association. Applying our novel approach to a lung cancer dataset, we confirmed four SNPs in 5p15 and 15q25 region to be significantly associated with lung cancer risk in Caucasians population: rs2736100, rs402710, rs16969968 and rs8034191. We also validated that rs16969968 and rs8034191 in 15q25 region are significantly interacting with smoking in Caucasian population. Our approach identified the potential interactions of SNP rs2256543 in 6p21 with smoking on contributing to lung cancer risk. Genetic imprinting is the most well-known cause for parent-of-origin effect (POE) whereby a gene is differentially expressed depending on the parental origin of the same alleles. Genetic imprinting affects several human disorders, including diabetes, breast cancer, alcoholism, and obesity. This phenomenon has been shown to be important for normal embryonic development in mammals. Traditional association approaches ignore this important genetic phenomenon. In this study, we propose a NOIA framework for a single locus association study that estimates both main allelic effects and POEs. We develop statistical (Stat-POE) and functional (Func-POE) models, and demonstrate conditions for orthogonality of the Stat-POE model. We conducted simulations for both quantitative and qualitative traits to evaluate the performance of the statistical and functional models with different levels of POEs. Our results showed that the newly proposed Stat-POE model, which ensures orthogonality of variance components if Hardy-Weinberg Equilibrium (HWE) or equal minor and major allele frequencies is satisfied, had greater power for detecting the main allelic additive effect than a Func-POE model, which codes according to allelic substitutions, for both quantitative and qualitative traits. The power for detecting the POE was the same for the Stat-POE and Func-POE models under HWE for quantitative traits.
Resumo:
The genomic era brought by recent advances in the next-generation sequencing technology makes the genome-wide scans of natural selection a reality. Currently, almost all the statistical tests and analytical methods for identifying genes under selection was performed on the individual gene basis. Although these methods have the power of identifying gene subject to strong selection, they have limited power in discovering genes targeted by moderate or weak selection forces, which are crucial for understanding the molecular mechanisms of complex phenotypes and diseases. Recent availability and rapid completeness of many gene network and protein-protein interaction databases accompanying the genomic era open the avenues of exploring the possibility of enhancing the power of discovering genes under natural selection. The aim of the thesis is to explore and develop normal mixture model based methods for leveraging gene network information to enhance the power of natural selection target gene discovery. The results show that the developed statistical method, which combines the posterior log odds of the standard normal mixture model and the Guilt-By-Association score of the gene network in a naïve Bayes framework, has the power to discover moderate/weak selection gene which bridges the genes under strong selection and it helps our understanding the biology under complex diseases and related natural selection phenotypes.^