35 resultados para Exploratory Multivariate Statistical Methods
Resumo:
The response of Arabidopsis to stress caused by mechanical wounding was chosen as a model to compare the performances of high resolution quadrupole-time-of-flight (Q-TOF) and single stage Orbitrap (Exactive Plus) mass spectrometers in untargeted metabolomics. Both instruments were coupled to ultra-high pressure liquid chromatography (UHPLC) systems set under identical conditions. The experiment was divided in two steps: the first analyses involved sixteen unwounded plants, half of which were spiked with pure standards that are not present in Arabidopsis. The second analyses compared the metabolomes of mechanically wounded plants to unwounded plants. Data from both systems were extracted using the same feature detection software and submitted to unsupervised and supervised multivariate analysis methods. Both mass spectrometers were compared in terms of number and identity of detected features, capacity to discriminate between samples, repeatability and sensitivity. Although analytical variability was lower for the UHPLC-Q-TOF, generally the results for the two detectors were quite similar, both of them proving to be highly efficient at detecting even subtle differences between plant groups. Overall, sensitivity was found to be comparable, although the Exactive Plus Orbitrap provided slightly lower detection limits for specific compounds. Finally, to evaluate the potential of the two mass spectrometers for the identification of unknown markers, mass and spectral accuracies were calculated on selected identified compounds. While both instruments showed excellent mass accuracy (<2.5ppm for all measured compounds), better spectral accuracy was recorded on the Q-TOF. Taken together, our results demonstrate that comparable performances can be obtained at acquisition frequencies compatible with UHPLC on Q-TOF and Exactive Plus MS, which may thus be equivalently used for plant metabolomics.
Resumo:
Analyzing the relationship between the baseline value and subsequent change of a continuous variable is a frequent matter of inquiry in cohort studies. These analyses are surprisingly complex, particularly if only two waves of data are available. It is unclear for non-biostatisticians where the complexity of this analysis lies and which statistical method is adequate.With the help of simulated longitudinal data of body mass index in children,we review statistical methods for the analysis of the association between the baseline value and subsequent change, assuming linear growth with time. Key issues in such analyses are mathematical coupling, measurement error, variability of change between individuals, and regression to the mean. Ideally, it is better to rely on multiple repeated measurements at different times and a linear random effects model is a standard approach if more than two waves of data are available. If only two waves of data are available, our simulations show that Blomqvist's method - which consists in adjusting for measurement error variance the estimated regression coefficient of observed change on baseline value - provides accurate estimates. The adequacy of the methods to assess the relationship between the baseline value and subsequent change depends on the number of data waves, the availability of information on measurement error, and the variability of change between individuals.
Resumo:
PURPOSE: To analyze final long-term survival and clinical outcomes from the randomized phase III study of sunitinib in gastrointestinal stromal tumor patients after imatinib failure; to assess correlative angiogenesis biomarkers with patient outcomes. EXPERIMENTAL DESIGN: Blinded sunitinib or placebo was given daily on a 4-week-on/2-week-off treatment schedule. Placebo-assigned patients could cross over to sunitinib at disease progression/study unblinding. Overall survival (OS) was analyzed using conventional statistical methods and the rank-preserving structural failure time (RPSFT) method to explore cross-over impact. Circulating levels of angiogenesis biomarkers were analyzed. RESULTS: In total, 243 patients were randomized to receive sunitinib and 118 to placebo, 103 of whom crossed over to open-label sunitinib. Conventional statistical analysis showed that OS converged in the sunitinib and placebo arms (median 72.7 vs. 64.9 weeks; HR, 0.876; P = 0.306) as expected, given the cross-over design. RPSFT analysis estimated median OS for placebo of 39.0 weeks (HR, 0.505, 95% CI, 0.262-1.134; P = 0.306). No new safety concerns emerged with extended sunitinib treatment. No consistent associations were found between the pharmacodynamics of angiogenesis-related plasma proteins during sunitinib treatment and clinical outcome. CONCLUSIONS: The cross-over design provided evidence of sunitinib clinical benefit based on prolonged time to tumor progression during the double-blind phase of this trial. As expected, following cross-over, there was no statistical difference in OS. RPSFT analysis modeled the absence of cross-over, estimating a substantial sunitinib OS benefit relative to placebo. Long-term sunitinib treatment was tolerated without new adverse events.
Resumo:
SUMMARY Heavy metal presence in the environment is a serious concern since some of them can be toxic to plants, animals and humans once accumulated along the food chain. Cadmium (Cd) is one of the most toxic heavy metal. It is naturally present in soils at various levels and its concentration can be increased by human activities. Several plants however have naturally developed strategies allowing them to grow on heavy metal enriched soils. One of them consists in the accumulation and sequestration of heavy metals in the above-ground biomass. Some plants present in addition an extreme strategy by which they accumulate a limited number of heavy metals in their shoots in amounts 100 times superior to those expected for a non-accumulating plant in the same conditions. Understanding the genetic basis of the hyperaccumulation trait - particularly for Cd - remains an important challenge which may lead to biotechnological applications in the soil phytoremediation. In this thesis, Thlaspi caerulescens J. & C. Presl (Brassicaceae) was used as a model plant to study the Cd hyperaccumulation trait, owing to its physiological and genetic characteristics. Twenty-four wild populations were sampled in different regions of Switzerland. They were characterized for environmental and soil parameters as well as intrinsic characteristics of plants (i.e. metal concentrations in shoots). They were as well genetically characterized by AFLPs, plastid DNA polymorphism and genes markers (CAPS and microsatellites) mainly developed in this thesis. Some of the investigated genes were putatively linked to the Cd hyperaccumulation trait. Since the study of the Cd hyperaccumulation in the field is important as it allows the identification of patterns of selection, the present work offered a methodology to define the Cd hyperaccumulation capacity of populations from different habitats permitting thus their comparison in the field. We showed that Cd, Zn, Fe and Cu accumulations were linked and that populations with higher Cd hyperaccumulation capacity had higher shoot and reproductive fitness. Using our genetic data, statistical methods (Beaumont & Nichols's procedure, partial Mantel tests) were applied to identify genomic signatures of natural selection related to the Cd hyperaccumulation capacity. A significant genetic difference between populations related to their Cd hyperaccumulation capacity was revealed based on somè specific markers (AFLP and candidate genes). Polymorphism at the gene encoding IRTl (Iron-transporter also participating to the transport of Zn) was suggested as explaining part of the variation in Cd hyperaccumulation capacity of populations supporting previous physiological investigations. RÉSUMÉ La présence de métaux lourds dans l'environnement est un phénomène préoccupant. En effet, certains métaux lourds - comme le cadmium (Cd) -sont toxiques pour les plantes, les animaux et enfin, accumulés le long de la chaîne alimentaire, pour les hommes. Le Cd est naturellement présent dans le sol et sa concentration peut être accrue par différentes activités humaines. Certaines plantes ont cependant développé des stratégies leur permettant de pousser sur des sols contaminés en métaux lourds. Parmi elles, certaines accumulent et séquestrent les métaux lourds dans leurs parties aériennes. D`autres présentent une stratégie encore plus extrême. Elles accumulent un nombre limité de métaux lourds en quantités 100 fois supérieures à celles attendues pour des espèces non-accumulatrices sous de mêmes conditions. La compréhension des bases génétiques de l'hyperaccumulation -particulièrement celle du Cd - représente un défi important avec des applications concrètes en biotechnologies, tout particulièrement dans le but appliqué de la phytoremediation des sols contaminés. Dans cette thèse, Thlaspi caerulescens J. & C. Presl (Brassicaceae) a été utilisé comme modèle pour l'étude de l'hyperaccumulation du Cd de par ses caractéristiques physiologiques et génétiques. Vingt-quatre populations naturelles ont été échantillonnées en Suisse et pour chacune d'elles les paramètres environnementaux, pédologique et les caractéristiques intrinsèques aux plantes (concentrations en métaux lourds) ont été déterminés. Les populations ont été caractérisées génétiquement par des AFLP, des marqueurs chloroplastiques et des marqueurs de gènes spécifiques, particulièrement ceux potentiellement liés à l'hyperaccumulation du Cd (CAPS et microsatellites). La plupart ont été développés au cours de cette thèse. L'étude de l'hyperaccumulation du Cd en conditions naturelles est importante car elle permet d'identifier la marque, éventuelle de sélection naturelle. Ce travail offre ainsi une méthodologie pour définir et comparer la capacité des populations à hyperaccumuler le Cd dans différents habitats. Nous avons montré que les accumulations du Cd, Zn, Fe et Cu sont liées et que les populations ayant une grande capacité d'hyperaccumuler le Cd ont également une meilleure fitness végétative et reproductive. Des méthodes statistiques (l'approche de Beaumont & Nichols, tests de Martel partiels) ont été utilisées sur les données génétiques pour identifier la signature génomique de la sélection naturelle liée à la capacité d'hyperaccumuler le Cd. Une différenciation génétique des populations liée à leur capacité d'hyperaccumuler le Cd a été mise en évidence sur certains marqueurs spécifiques. En accord avec les études physiologiques connues, le polymorphisme au gène codant IRT1 (un transporteur de Fe impliqué dans le transport du Zn) pourrait expliquer une partie de la variance de la capacité des populations à hyperaccumuler le Cd.
Resumo:
Understanding brain reserve in preclinical stages of neurodegenerative disorders allows determination of which brain regions contribute to normal functioning despite accelerated neuronal loss. Besides the recruitment of additional regions, a reorganisation and shift of relevance between normally engaged regions are a suggested key mechanism. Thus, network analysis methods seem critical for investigation of changes in directed causal interactions between such candidate brain regions. To identify core compensatory regions, fifteen preclinical patients carrying the genetic mutation leading to Huntington's disease and twelve controls underwent fMRI scanning. They accomplished an auditory paced finger sequence tapping task, which challenged cognitive as well as executive aspects of motor functioning by varying speed and complexity of movements. To investigate causal interactions among brain regions a single Dynamic Causal Model (DCM) was constructed and fitted to the data from each subject. The DCM parameters were analysed using statistical methods to assess group differences in connectivity, and the relationship between connectivity patterns and predicted years to clinical onset was assessed in gene carriers. In preclinical patients, we found indications for neural reserve mechanisms predominantly driven by bilateral dorsal premotor cortex, which increasingly activated superior parietal cortices the closer individuals were to estimated clinical onset. This compensatory mechanism was restricted to complex movements characterised by high cognitive demand. Additionally, we identified task-induced connectivity changes in both groups of subjects towards pre- and caudal supplementary motor areas, which were linked to either faster or more complex task conditions. Interestingly, coupling of dorsal premotor cortex and supplementary motor area was more negative in controls compared to gene mutation carriers. Furthermore, changes in the connectivity pattern of gene carriers allowed prediction of the years to estimated disease onset in individuals. Our study characterises the connectivity pattern of core cortical regions maintaining motor function in relation to varying task demand. We identified connections of bilateral dorsal premotor cortex as critical for compensation as well as task-dependent recruitment of pre- and caudal supplementary motor area. The latter finding nicely mirrors a previously published general linear model-based analysis of the same data. Such knowledge about disease specific inter-regional effective connectivity may help identify foci for interventions based on transcranial magnetic stimulation designed to stimulate functioning and also to predict their impact on other regions in motor-associated networks.
Resumo:
Port-a-Cath© (PAC) are totally implantable devices that offer an easy and long term access to venous circulation. They have been extensively used for intravenous therapy administration and are particularly well suited for chemotherapy in oncologic patients. Previous comparative studies have shown that these devices have the lowest catheter-related bloodstream infection rates among all intravascular access systems. However, bloodstream infection (BSI) still remains a major issue of port use and epidemiology data for PAC-associated BSI (PABSI) rates differ strongly depending on studies. Also, current literature about PABSI risk factors is scarce and sometimes controversial. Such heterogeneity may depend on type of studied population and local factors. Therefore, the aim of this study was to describe local epidemiology and risk factors for PABSI in adult patients in our tertiary- care university hospital. We conducted a retrospective cohort study in order to describe local epidemiology. We also performed a nested case-control study to identify local risk factors of PABSI. We analyzed medical files of adult patients who had a PAC implanted between January 1st, 2008 and December 31st, 2009 and looked for PABSI occurrence before May 1st, 2011 to define cases. Thirty nine PABSI occurred in this population with an attack rate of 5.8%. We estimated an incidence rate of 0.08/1000 PAC-days using the case-control study. PABSI causative agents were mainly Gram positive cocci (62%). We identified three predictive factors of PABSI by multivariate statistical analysis: neutropenia on outcome date (Odds Ratio [OR]: 4.05; 95% confidence interval [CI]:1.05- 15.66; p=0.042), diabetes (OR: 11.53; 95% CI: 1.07-124.70; p=0.044) and having another infection than PABSI on outcome date (OR: 6.35; 95% CI: 1.50-26.86; p=0.012). Patients suffering from acute or renal failure (OR: 4.26; 95% CI: 0.94-19.21; p=0.059) or wearing another invasive device (OR: 2.99; 95%CI:0.96-9.31; p=0.059) did not have a statistically increased risk for developing a PABSI according to classical threshold (p<0.05) but nevertheless remained close to significance. Our study demonstrated that local epidemiology and microbiology of PABSI in our institution was similar to previous reports. A larger prospective study is required to confirm our results or to test preventive measures.
Resumo:
Summary [résumé français voir ci-dessous] From the beginning of the 20th century the world population has been confronted with the human immune deficiency virus 1 (HIV-1). This virus has the particularity to mutate fast, and could thus evade and adapt to the human host. Our closest evolutionary related organisms, the non-human primates, are less susceptible to HIV-1. In a broader sense, primates are differentially susceptible to various retrovirus. Species specificity may be due to genetic differences among primates. In the present study we applied evolutionary and comparative genetic techniques to characterize the evolutionary pattern of host cellular determinants of HIV-1 pathogenesis. The study of the evolution of genes coding for proteins participating to the restriction or pathogenesis of HIV-1 may help understanding the genetic basis of modern human susceptibility to infection. To perform comparative genetics analysis, we constituted a collection of primate DNA and RNA to allow generation of de novo sequence of gene orthologs. More recently, release to the public domain of two new primate complete genomes (bornean orang-utan and common marmoset) in addition of the three previously available genomes (human, chimpanzee and Rhesus monkey) help scaling up the evolutionary and comparative genome analysis. Sequence analysis used phylogenetic and statistical methods for detecting molecular adaptation. We identified different selective pressures acting on host proteins involved in HIV-1 pathogenesis. Proteins with HIV-1 restriction properties in non-human primates were under strong positive selection, in particular in regions of interaction with viral proteins. These regions carried key residues for the antiviral activity. Proteins of the innate immunity presented an evolutionary pattern of conservation (purifying selection) but with signals of relaxed constrain if we compared them to the average profile of purifying selection of the primate genomes. Large scale analysis resulted in patterns of evolutionary pressures according to molecular function, biological process and cellular distribution. The data generated by various analyses served to guide the ancestral reconstruction of TRIM5a a potent antiviral host factor. The resurrected TRIM5a from the common ancestor of Old world monkeys was effective against HIV-1 and the recent resurrected hominoid variants were more effective against other retrovirus. Thus, as the result of trade-offs in the ability to restrict different retrovirus, human might have been exposed to HIV-1 at a time when TRIM5a lacked the appropriate specific restriction activity. The application of evolutionary and comparative genetic tools should be considered for the systematical assessment of host proteins relevant in viral pathogenesis, and to guide biological and functional studies. Résumé La population mondiale est confrontée depuis le début du vingtième siècle au virus de l'immunodéficience humaine 1 (VIH-1). Ce virus a un taux de mutation particulièrement élevé, il peut donc s'évader et s'adapter très efficacement à son hôte. Les organismes évolutivement le plus proches de l'homme les primates nonhumains sont moins susceptibles au VIH-1. De façon générale, les primates répondent différemment aux rétrovirus. Cette spécificité entre espèces doit résider dans les différences génétiques entre primates. Dans cette étude nous avons appliqué des techniques d'évolution et de génétique comparative pour caractériser le modèle évolutif des déterminants cellulaires impliqués dans la pathogenèse du VIH- 1. L'étude de l'évolution des gènes, codant pour des protéines impliquées dans la restriction ou la pathogenèse du VIH-1, aidera à la compréhension des bases génétiques ayant récemment rendu l'homme susceptible. Pour les analyses de génétique comparative, nous avons constitué une collection d'ADN et d'ARN de primates dans le but d'obtenir des nouvelles séquences de gènes orthologues. Récemment deux nouveaux génomes complets ont été publiés (l'orang-outan du Bornéo et Marmoset commun) en plus des trois génomes déjà disponibles (humain, chimpanzé, macaque rhésus). Ceci a permis d'améliorer considérablement l'étendue de l'analyse. Pour détecter l'adaptation moléculaire nous avons analysé les séquences à l'aide de méthodes phylogénétiques et statistiques. Nous avons identifié différentes pressions de sélection agissant sur les protéines impliquées dans la pathogenèse du VIH-1. Des protéines avec des propriétés de restriction du VIH-1 dans les primates non-humains présentent un taux particulièrement haut de remplacement d'acides aminés (sélection positive). En particulier dans les régions d'interaction avec les protéines virales. Ces régions incluent des acides aminés clé pour l'activité de restriction. Les protéines appartenant à l'immunité inné présentent un modèle d'évolution de conservation (sélection purifiante) mais avec des traces de "relaxation" comparé au profil général de sélection purifiante du génome des primates. Une analyse à grande échelle a permis de classifier les modèles de pression évolutive selon leur fonction moléculaire, processus biologique et distribution cellulaire. Les données générées par les différentes analyses ont permis la reconstruction ancestrale de TRIM5a, un puissant facteur antiretroviral. Le TRIM5a ressuscité, correspondant à l'ancêtre commun entre les grands singes et les groupe des catarrhiniens, est efficace contre le VIH-1 moderne. Les TRIM5a ressuscités plus récents, correspondant aux ancêtres des grands singes, sont plus efficaces contre d'autres rétrovirus. Ainsi, trouver un compromis dans la capacité de restreindre différents rétrovirus, l'homme aurait été exposé au VIH-1 à une période où TRIM5a manquait d'activité de restriction spécifique contre celui-ci. L'application de techniques d'évolution et de génétique comparative devraient être considérées pour l'évaluation systématique de protéines impliquées dans la pathogenèse virale, ainsi que pour guider des études biologiques et fonctionnelles
Adenocarcinoma of the pancreas: Comparative single centre analysis between ductal and mucinous type.
Resumo:
1. Background¦Adenocarcinomas of the pancreas are exocrine tumors, originate from ductal system, including two morphologically distinct entities: the ductal adenocarcinoma and mucinous adenocarcinoma. Ductal adenocarcinoma is by far the most frequent malignant tumor in the pancreas, representing at least about 90% of all pancreas cancers. It is associated with very poor prognosis, due to the fact that actually there are no any biological markers or diagnostic tools for identification of the disease at an early stage. Most of the time the disease is extensive with vascular and nerves involvement or with metastatic spread at the time of diagnosis (1). The median survival is less than 5% at 5 years, placing it, at the fifth leading cause of death by cancer in the world (2). The mucinous form of pancreatic adenocarcinoma is less frequent, and seems to have a better prognosis with about 57% survival at 5 years (1)(3)(4).¦Each morphologic type of pancreatic adenocarcinoma is associated with particular preneoplastic lesions. Two types of preneoplastic lesions are described: firstly, pancreatic intra-epithelial neoplasia (PanIN) which affects the small and peripheral pancreatic ducts, and the intraductal papillary-mucinous neoplasm (IPMN) interested the main pancreatic ducts and its principal branches. Both of preneoplastic lesions lead by different mechanisms to the pancreatic adenocarcinoma (1)(2)(3)(4)(5)(6)(7)(8)(9)(10).¦The purpose of our study consists in a retrospective analysis of various clinical and histo-morphological parameters in order to assess a difference in survival between these two morphological types of pancreatic adenocarcinomas.¦1.2 Material and methods¦We conducted a retrospective analysis including 35 patients, (20 men and 15 women), beneficed the surgical treatment for pancreas adenocarcinoma at the Surgical Department of University Hospital in Lausanne. The patients involved in our study have been treated between 2003 and 2008, permitting at least 5-years mean follow up. For each patient the following parameters were analysed: age, gender, type of operation, type of preneoplastic lesions, TNM stage, histological grade of the tumor, vascular invasion, lymphatic and perineural invasion, resection margins, and adjuvant treatment.¦The results from these observations were included in a univariate and multivariate statistical analysis and compared with overall survival, as well as specific survival for each morphologic subtype of adenocarcinoma.¦As a low number of mucinous adenocarcinomas (n=5) was insufficient to conduct a pertinent statistical analysis, we compared the data obtained from adenocarcinomas developed on PanIN with adenocarcinomas developed on IPMN including both, ductal or mucinous types.¦1.3 Result¦Our results show that adenocarcinomas developed on pre-existing IPMN including both morphologic types (ductal and mucinous form) are associated with a better survival and prognosis than adenocarciomas developed on PanIN.¦1.4 Conclusion¦This study reflects that the most relevant parameter in survival in pancreatic adenocarcinoma seems to be the type of preneoplastic lesion. The significant difference in survival was noted between adenocarcinomas developing on PanIN as compared to adenocarcinomas developed on IPMN precursor lesions. Ductal adenocarcinomas developped on IPMN present significantly longer survival than those developed on PanIN lesions (P value= 0,01). Therefore we can suggest that the histological type of preneoplastic lesion rather than the histological type of adenocarcinoma should be the determinant prognosis factor in survival of pancreatic adenocarcinoma.
Resumo:
Introduction: The aim of this study is to compare the walking activity of a cohort of individuals before and after total ankle arthroplasty (TAA). Methods: Nineteen consecutive patients (ten males and nine females) with mean age of 58.72, selected for TAA between January and June 2006, were prospectively reviewed with the use of a dedicated ambulatory activity-monitoring device to assess their natural ambulatory activity. Patients were tested in the community for two weeks duration, one month prior to and at least eighteen months after surgery. The ambulatory parameters were assessed through measurement of the number of steps at different cadence, and the time spent walking at different walking paces. Data were analyzed by using specific statistical methods. Results: This study revealed a significant improvement in the number of steps walked at normal cadence (b = 331.63, p = .00) and significantly reduced at low cadence (b = -402.52, p = .00) and medium cadence (b = -386.29, p = .00), before and after TAA. However, there are no significant different between two phases of assessment in term of time spent walking. Conclusion: These quantitative data allow a clear comparative assessment of walking ability following TAR and demonstrates that this intervention improves patient's walking pace.
Resumo:
Anthropogenic emissions of metals from sources such as smelters are an international problem, but there is limited published information on emissions from Australian smelters. The objective of this study was to investigate the regional distribution of heavy metals in soils in the vicinity of the industrial complex of Port Kembla, NSW, Australia, which comprises a copper smelter, steelworks and associated industries. Soil samples (n=25) were collected at the depths of 0-5 and 5-20 cm, air dried and sieved to < 2 mm. Aqua regia extractable amounts of As, Cr, Cu, Ph and Zn were analysed by inductively coupled plasma mass spectrometry (lCP-MS) and inductively coupled plasma atomic emission spectrometry (ICP-AES). Outliers were identified from background levels by statistical methods. Mean background levels at a depth of 0-5 cm were estimated at 3.2 mg/kg As, 12 mg/kg Cr, 49 mg/kg Cu, 20 mg/kg Ph and 42 mg/kg Zn. Outliers for elevated As and Cu values were mainly present within 4 km from the Port Kembla industrial complex, but high Ph at two sites and high Zn concentrations were found at six sites up to 23 km from Port Kembla. Chromium concentrations were not anomalous close to the industrial complex. There was no significant difference of metal concentrations at depths of 0-5 and 5-20 cm, except for Ph and Zn. Copper and As concentrations in the soils are probably related to the concentrations in the parent rock. From this investigation, the extent of the contamination emanating from the Port Kembla industrial complex is limited to 1-13 km, but most likely <4 km, depending on the element; the contamination at the greater distance may not originate from the industrial complex. (C) 2003 Elsevier B.V. All rights reserved.
Resumo:
It is estimated that around 230 people die each year due to radon (222Rn) exposure in Switzerland. 222Rn occurs mainly in closed environments like buildings and originates primarily from the subjacent ground. Therefore it depends strongly on geology and shows substantial regional variations. Correct identification of these regional variations would lead to substantial reduction of 222Rn exposure of the population based on appropriate construction of new and mitigation of already existing buildings. Prediction of indoor 222Rn concentrations (IRC) and identification of 222Rn prone areas is however difficult since IRC depend on a variety of different variables like building characteristics, meteorology, geology and anthropogenic factors. The present work aims at the development of predictive models and the understanding of IRC in Switzerland, taking into account a maximum of information in order to minimize the prediction uncertainty. The predictive maps will be used as a decision-support tool for 222Rn risk management. The construction of these models is based on different data-driven statistical methods, in combination with geographical information systems (GIS). In a first phase we performed univariate analysis of IRC for different variables, namely the detector type, building category, foundation, year of construction, the average outdoor temperature during measurement, altitude and lithology. All variables showed significant associations to IRC. Buildings constructed after 1900 showed significantly lower IRC compared to earlier constructions. We observed a further drop of IRC after 1970. In addition to that, we found an association of IRC with altitude. With regard to lithology, we observed the lowest IRC in sedimentary rocks (excluding carbonates) and sediments and the highest IRC in the Jura carbonates and igneous rock. The IRC data was systematically analyzed for potential bias due to spatially unbalanced sampling of measurements. In order to facilitate the modeling and the interpretation of the influence of geology on IRC, we developed an algorithm based on k-medoids clustering which permits to define coherent geological classes in terms of IRC. We performed a soil gas 222Rn concentration (SRC) measurement campaign in order to determine the predictive power of SRC with respect to IRC. We found that the use of SRC is limited for IRC prediction. The second part of the project was dedicated to predictive mapping of IRC using models which take into account the multidimensionality of the process of 222Rn entry into buildings. We used kernel regression and ensemble regression tree for this purpose. We could explain up to 33% of the variance of the log transformed IRC all over Switzerland. This is a good performance compared to former attempts of IRC modeling in Switzerland. As predictor variables we considered geographical coordinates, altitude, outdoor temperature, building type, foundation, year of construction and detector type. Ensemble regression trees like random forests allow to determine the role of each IRC predictor in a multidimensional setting. We found spatial information like geology, altitude and coordinates to have stronger influences on IRC than building related variables like foundation type, building type and year of construction. Based on kernel estimation we developed an approach to determine the local probability of IRC to exceed 300 Bq/m3. In addition to that we developed a confidence index in order to provide an estimate of uncertainty of the map. All methods allow an easy creation of tailor-made maps for different building characteristics. Our work is an essential step towards a 222Rn risk assessment which accounts at the same time for different architectural situations as well as geological and geographical conditions. For the communication of 222Rn hazard to the population we recommend to make use of the probability map based on kernel estimation. The communication of 222Rn hazard could for example be implemented via a web interface where the users specify the characteristics and coordinates of their home in order to obtain the probability to be above a given IRC with a corresponding index of confidence. Taking into account the health effects of 222Rn, our results have the potential to substantially improve the estimation of the effective dose from 222Rn delivered to the Swiss population.
Resumo:
Genetic variants influence the risk to develop certain diseases or give rise to differences in drug response. Recent progresses in cost-effective, high-throughput genome-wide techniques, such as microarrays measuring Single Nucleotide Polymorphisms (SNPs), have facilitated genotyping of large clinical and population cohorts. Combining the massive genotypic data with measurements of phenotypic traits allows for the determination of genetic differences that explain, at least in part, the phenotypic variations within a population. So far, models combining the most significant variants can only explain a small fraction of the variance, indicating the limitations of current models. In particular, researchers have only begun to address the possibility of interactions between genotypes and the environment. Elucidating the contributions of such interactions is a difficult task because of the large number of genetic as well as possible environmental factors.In this thesis, I worked on several projects within this context. My first and main project was the identification of possible SNP-environment interactions, where the phenotypes were serum lipid levels of patients from the Swiss HIV Cohort Study (SHCS) treated with antiretroviral therapy. Here the genotypes consisted of a limited set of SNPs in candidate genes relevant for lipid transport and metabolism. The environmental variables were the specific combinations of drugs given to each patient over the treatment period. My work explored bioinformatic and statistical approaches to relate patients' lipid responses to these SNPs, drugs and, importantly, their interactions. The goal of this project was to improve our understanding and to explore the possibility of predicting dyslipidemia, a well-known adverse drug reaction of antiretroviral therapy. Specifically, I quantified how much of the variance in lipid profiles could be explained by the host genetic variants, the administered drugs and SNP-drug interactions and assessed the predictive power of these features on lipid responses. Using cross-validation stratified by patients, we could not validate our hypothesis that models that select a subset of SNP-drug interactions in a principled way have better predictive power than the control models using "random" subsets. Nevertheless, all models tested containing SNP and/or drug terms, exhibited significant predictive power (as compared to a random predictor) and explained a sizable proportion of variance, in the patient stratified cross-validation context. Importantly, the model containing stepwise selected SNP terms showed higher capacity to predict triglyceride levels than a model containing randomly selected SNPs. Dyslipidemia is a complex trait for which many factors remain to be discovered, thus missing from the data, and possibly explaining the limitations of our analysis. In particular, the interactions of drugs with SNPs selected from the set of candidate genes likely have small effect sizes which we were unable to detect in a sample of the present size (<800 patients).In the second part of my thesis, I performed genome-wide association studies within the Cohorte Lausannoise (CoLaus). I have been involved in several international projects to identify SNPs that are associated with various traits, such as serum calcium, body mass index, two-hour glucose levels, as well as metabolic syndrome and its components. These phenotypes are all related to major human health issues, such as cardiovascular disease. I applied statistical methods to detect new variants associated with these phenotypes, contributing to the identification of new genetic loci that may lead to new insights into the genetic basis of these traits. This kind of research will lead to a better understanding of the mechanisms underlying these pathologies, a better evaluation of disease risk, the identification of new therapeutic leads and may ultimately lead to the realization of "personalized" medicine.
Resumo:
The R package EasyStrata facilitates the evaluation and visualization of stratified genome-wide association meta-analyses (GWAMAs) results. It provides (i) statistical methods to test and account for between-strata difference as a means to tackle gene-strata interaction effects and (ii) extended graphical features tailored for stratified GWAMA results. The software provides further features also suitable for general GWAMAs including functions to annotate, exclude or highlight specific loci in plots or to extract independent subsets of loci from genome-wide datasets. It is freely available and includes a user-friendly scripting interface that simplifies data handling and allows for combining statistical and graphical functions in a flexible fashion. AVAILABILITY: EasyStrata is available for free (under the GNU General Public License v3) from our Web site www.genepi-regensburg.de/easystrata and from the CRAN R package repository cran.r-project.org/web/packages/EasyStrata/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Resumo:
PURPOSE: To identify clinical risk factors for Dravet syndrome (DS) in a population of children with status epilepticus (SE). MATERIAL AND METHODS: Children aged between 1 month and 16 years with at least one episode of SE were referred from 6 pediatric neurology centers in Switzerland. SE was defined as a clinical seizure lasting for more than 30min without recovery of normal consciousness. The diagnosis of DS was considered likely in previously healthy patients with seizures of multiple types starting before 1 year and developmental delay on follow-up. The presence of a SCN1A mutation was considered confirmatory for the diagnosis. Data such as gender, age at SE, SE clinical presentation and recurrence, additional seizure types and epilepsy diagnosis were collected. SCN1A analyses were performed in all patients, initially with High Resolution Melting Curve Analysis (HRMCA) and then by direct sequencing on selected samples with an abnormal HRMCA. Clinical and genetic findings were compared between children with DS and those with another diagnosis, and statistical methods were applied for significance analysis. RESULTS: 71 children with SE were included. Ten children had DS, and 61 had another diagnosis. SCN1A mutations were found in 12 of the 71 patients (16.9%; ten with DS, and two with seizures in a Generalized Epilepsy with Febrile Seizures+(GEFS+) context). The median age at first SE was 8 months in patients with DS, and 41 months in those with another epilepsy syndrome (p<0.001). Nine of the 10 DS patients had their initial SE before 18 months. Among the 26 patients aged 18 months or less at initial SE, the risk of DS was significantly increased for patients with two or more episodes (56.3%), as compared with those who had only one episode (0.0%) (p=0.005). CONCLUSION: In a population of children with SE, patients most likely to have DS are those who present their initial SE episode before 18 months, and who present with recurrent SE episodes.
Resumo:
RESUME Les fibres textiles sont des produits de masse utilisés dans la fabrication de nombreux objets de notre quotidien. Le transfert de fibres lors d'une action délictueuse est dès lors extrêmement courant. Du fait de leur omniprésence dans notre environnement, il est capital que l'expert forensique évalue la valeur de l'indice fibres. L'interprétation de l'indice fibres passe par la connaissance d'un certain nombre de paramètres, comme la rareté des fibres, la probabilité de leur présence par hasard sur un certain support, ainsi que les mécanismes de transfert et de persistance des fibres. Les lacunes les plus importantes concernent les mécanismes de transfert des fibres. A ce jour, les nombreux auteurs qui se sont penchés sur le transfert de fibres ne sont pas parvenus à créer un modèle permettant de prédire le nombre de fibres que l'on s'attend à retrouver dans des circonstances de contact données, en fonction des différents paramètres caractérisant ce contact et les textiles mis en jeu. Le but principal de cette recherche est de démontrer que la création d'un modèle prédictif du nombre de fibres transférées lors d'un contact donné est possible. Dans le cadre de ce travail, le cas particulier du transfert de fibres d'un tricot en laine ou en acrylique d'un conducteur vers le dossier du siège de son véhicule a été étudié. Plusieurs caractéristiques des textiles mis en jeu lors de ces expériences ont été mesurées. Des outils statistiques (régression linéaire multiple) ont ensuite été utilisés sur ces données afin d'évaluer l'influence des caractéristiques des textiles donneurs sur le nombre de fibres transférées et d'élaborer un modèle permettant de prédire le nombre de fibres qui vont être transférées à l'aide des caractéristiques influençant significativement le transfert. Afin de faciliter la recherche et le comptage des fibres transférées lors des expériences de transfert, un appareil de recherche automatique des fibres (liber finder) a été utilisé dans le cadre de cette recherche. Les tests d'évaluation de l'efficacité de cet appareil pour la recherche de fibres montrent que la recherche automatique est globalement aussi efficace qu'une recherche visuelle pour les fibres fortement colorées. Par contre la recherche automatique perd de son efficacité pour les fibres très pâles ou très foncées. Une des caractéristiques des textiles donneurs à étudier est la longueur des fibres. Afin de pouvoir évaluer ce paramètre, une séquence d'algorithmes de traitement d'image a été implémentée. Cet outil permet la mesure de la longueur d'une fibre à partir de son image numérique à haute résolution (2'540 dpi). Les tests effectués montrent que les mesures ainsi obtenues présentent une erreur de l'ordre du dixième de millimètre, ce qui est largement suffisant pour son utilisation dans le cadre de cette recherche. Les résultats obtenus suite au traitement statistique des résultats des expériences de transfert ont permis d'aboutir à une modélisation du phénomène du transfert. Deux paramètres sont retenus dans le modèle: l'état de la surface du tissu donneur et la longueur des fibres composant le tissu donneur. L'état de la surface du tissu est un paramètre tenant compte de la quantité de fibres qui se sont détachées de la structure du tissu ou qui sont encore faiblement rattachées à celle-ci. En effet, ces fibres sont les premières à se transférer lors d'un contact, et plus la quantité de ces fibres par unité de surface est importante, plus le nombre de fibres transférées sera élevé. La longueur des fibres du tissu donneur est également un paramètre important : plus les fibres sont longues, mieux elles sont retenues dans la structure du tissu et moins elles se transféreront. SUMMARY Fibres are mass products used to produce numerous objects encountered everyday. The transfer of fibres during a criminal action is then very common. Because fibres are omnipresent in our environment, the forensic expert has to evaluate the value of the fibre evidence. To interpret fibre evidence, the expert has to know some parameters as frequency of fibres,' probability of finding extraneous fibres by chance on a given support, and transfer and persistence mechanisms. Fibre transfer is one of the most complex parameter. Many authors studied fibre transfer mechanisms but no model has been created to predict the number of fibres transferred expected in a given type of contact according to parameters as characteristics of the contact and characteristics of textiles. The main purpose of this research is to demonstrate that it is possible to create a model to predict the number of fibres transferred during a contact. In this work, the particular case of the transfer of fibres from a knitted textile in wool or in acrylic of a driver to the back of a carseat has been studied. Several characteristics of the textiles used for the experiments were measured. The data obtained were then treated with statistical tools (multiple linear regression) to evaluate the influence of the donor textile characteristics on the number of úbers transferred, and to create a model to predict this number of fibres transferred by an equation containing the characteristics having a significant influence on the transfer. To make easier the searching and the counting of fibres, an apparatus of automatic search. of fibers (fiber finder) was used. The tests realised to evaluate the efficiency of the fiber finder shows that the results obtained are generally as efficient as for visual search for well-coloured fibres. However, the efficiency of automatic search decreases for pales and dark fibres. One characteristic of the donor textile studied was the length of the fibres. To measure this parameter, a sequence of image processing algorithms was implemented. This tool allows to measure the length of a fibre from it high-resolution (2'540 dpi) numerical image. The tests done shows that the error of the measures obtained are about some tenths of millimetres. This precision is sufficient for this research. The statistical methods applied on the transfer experiment data allow to create a model of the transfer phenomenon. Two parameters are included in the model: the shedding capacity of the donor textile surface and the length of donor textile fibres. The shedding capacity of the donor textile surface is a parameter estimating the quantity of fibres that are not or slightly attached to the structure of the textile. These fibres are easily transferred during a contact, and the more this quantity of fibres is high, the more the number of fibres transferred during the contact is important. The length of fibres is also an important parameter: the more the fibres are long, the more they are attached in the structure of the textile and the less they are transferred during the contact.