959 resultados para Frequent pattern mining
Resumo:
La tesi da me svolta durante questi ultimi sei mesi è stata sviluppata presso i laboratori di ricerca di IMA S.p.a.. IMA (Industria Macchine Automatiche) è una azienda italiana che naque nel 1961 a Bologna ed oggi riveste il ruolo di leader mondiale nella produzione di macchine automatiche per il packaging di medicinali. Vorrei subito mettere in luce che in tale contesto applicativo l’utilizzo di algoritmi di data-mining risulta essere ostico a causa dei due ambienti in cui mi trovo. Il primo è quello delle macchine automatiche che operano con sistemi in tempo reale dato che non presentano a pieno le risorse di cui necessitano tali algoritmi. Il secondo è relativo alla produzione di farmaci in quanto vige una normativa internazionale molto restrittiva che impone il tracciamento di tutti gli eventi trascorsi durante l’impacchettamento ma che non permette la visione al mondo esterno di questi dati sensibili. Emerge immediatamente l’interesse nell’utilizzo di tali informazioni che potrebbero far affiorare degli eventi riconducibili a un problema della macchina o a un qualche tipo di errore al fine di migliorare l’efficacia e l’efficienza dei prodotti IMA. Lo sforzo maggiore per riuscire ad ideare una strategia applicativa è stata nella comprensione ed interpretazione dei messaggi relativi agli aspetti software. Essendo i dati molti, chiusi, e le macchine con scarse risorse per poter applicare a dovere gli algoritmi di data mining ho provveduto ad adottare diversi approcci in diversi contesti applicativi: • Sistema di identificazione automatica di errore al fine di aumentare di diminuire i tempi di correzione di essi. • Modifica di un algoritmo di letteratura per la caratterizzazione della macchina. La trattazione è così strutturata: • Capitolo 1: descrive la macchina automatica IMA Adapta della quale ci sono stati forniti i vari file di log. Essendo lei l’oggetto di analisi per questo lavoro verranno anche riportati quali sono i flussi di informazioni che essa genera. • Capitolo 2: verranno riportati degli screenshoot dei dati in mio possesso al fine di, tramite un’analisi esplorativa, interpretarli e produrre una formulazione di idee/proposte applicabili agli algoritmi di Machine Learning noti in letteratura. • Capitolo 3 (identificazione di errore): in questo capitolo vengono riportati i contesti applicativi da me progettati al fine di implementare una infrastruttura che possa soddisfare il requisito, titolo di questo capitolo. • Capitolo 4 (caratterizzazione della macchina): definirò l’algoritmo utilizzato, FP-Growth, e mostrerò le modifiche effettuate al fine di poterlo impiegare all’interno di macchine automatiche rispettando i limiti stringenti di: tempo di cpu, memoria, operazioni di I/O e soprattutto la non possibilità di aver a disposizione l’intero dataset ma solamente delle sottoporzioni. Inoltre verranno generati dei DataSet per il testing di dell’algoritmo FP-Growth modificato.
Resumo:
Detecting bugs as early as possible plays an important role in ensuring software quality before shipping. We argue that mining previous bug fixes can produce good knowledge about why bugs happen and how they are fixed. In this paper, we mine the change history of 717 open source projects to extract bug-fix patterns. We also manually inspect many of the bugs we found to get insights into the contexts and reasons behind those bugs. For instance, we found out that missing null checks and missing initializations are very recurrent and we believe that they can be automatically detected and fixed.
Resumo:
Mémoire numérisé par la Direction des bibliothèques de l'Université de Montréal.
Resumo:
Mémoire numérisé par la Direction des bibliothèques de l'Université de Montréal.
Resumo:
Although cartilaginous tumors have low microvascular density, vessels are important for the provision of nutrition so that the tumor can grow and generate metastasis. The aim of this study was to assess the value of the vascular pattern classification as a prognostic tool in chondrosarcomas (CSs) and its relation with vascular endothelial growth factor (VEGF) expression. This was a retrospective study of 21 enchondromas and 57 conventional CSs. Clinical data and outcome were retrieved from medical files. CSs histologic grades (on a scale of 1 to 3) were determined according to the World Health Organization classification. The vascular pattern (on a scale of A to C) was assessed through CD34, according to Kalinski. CD105 and VEGF were also evaluated. Poor outcome was significantly associated with vascular pattern groups B and C. Higher vascular pattern were 6.5 times more frequent in moderate-grade and high-grade CSs than in grade 1 CS. On multivariate analysis, a clear correlation was found between VEGF overexpression and B/C vascular patterns. Only 18 (benign and malignant) tumors stained for CD105. The results point to the use of the vascular pattern classification as a prognostic tool in CSs and to differentiate low-grade from moderate-grade/high-grade CSs. Vascular pattern might be also used to complement histologic grade, VEGF immunostaining, and microvascular density, for indicating a patient's prognosis. Low-grade CSs develop under low neoangiogenesis, which conforms to the slow growth rate of these tumors.
Resumo:
Human respiratory syncytial virus (HRSV) is the major cause of lower respiratory tract infections in children under 5 years of age and the elderly, causing annual disease outbreaks during the fall and winter. Multiple lineages of the HRSVA and HRSVB serotypes co-circulate within a single outbreak and display a strongly temporal pattern of genetic variation, with a replacement of dominant genotypes occurring during consecutive years. In the present study we utilized phylogenetic methods to detect and map sites subject to adaptive evolution in the G protein of HRSVA and HRSVB. A total of 29 and 23 amino acid sites were found to be putatively positively selected in HRSVA and HRSVB, respectively. Several of these sites defined genotypes and lineages within genotypes in both groups, and correlated well with epitopes previously described in group A. Remarkably, 18 of these positively selected tended to revert in time to a previous codon state, producing a ""flipflop'' phylogenetic pattern. Such frequent evolutionary reversals in HRSV are indicative of a combination of frequent positive selection, reflecting the changing immune status of the human population, and a limited repertoire of functionally viable amino acids at specific amino acid sites.
Resumo:
The principle of using induction rules based on spatial environmental data to model a soil map has previously been demonstrated Whilst the general pattern of classes of large spatial extent and those with close association with geology were delineated small classes and the detailed spatial pattern of the map were less well rendered Here we examine several strategies to improve the quality of the soil map models generated by rule induction Terrain attributes that are better suited to landscape description at a resolution of 250 m are introduced as predictors of soil type A map sampling strategy is developed Classification error is reduced by using boosting rather than cross validation to improve the model Further the benefit of incorporating the local spatial context for each environmental variable into the rule induction is examined The best model was achieved by sampling in proportion to the spatial extent of the mapped classes boosting the decision trees and using spatial contextual information extracted from the environmental variables.
Resumo:
One hundred forty-two women with polycystic ovary syndrome (PCOS) with an average body mass index (BMI) of 29.1 kg/m(2) and average age of 25.12 years were studied. By BMI, 30.2% were normal, 38.0% were overweight and 31.6% were obese. Thirty-one eumenorrheic women matched for BMI and age, with no evidence of hyperandrogenism, were recruited as controls. The incidence of dyslipidemia in the PCOS group was twice that of the Control group (76.1% versus 32.25%). The most frequent abnormalities were low high-density lipoprotein cholesterol (HDL-C; 57.6%) and high triglyceride (TG) (28.3%). HDL-C was significantly lower in all subgroups of women with PCOS when compared to the subgroups of normal women. No significant differences were seen in the total cholesterol (p = 0.307), low-density lipoprotein cholesterol (LDL-C; p = 0.283) and TGs (p = 0.113) levels among the subgroups. An independent effect on HDL-C was detected for glucose (p = 0.004) and fasting insulin (p = 0.01); on TG for age (p = 0.003) and homeostatic model assessment insulin resistance (p = 0.03) and on total cholesterol and LDL-C for age (p = 0.02 and p = 0.033, respectively). In conclusion, dyslipidemia is common in women with PCOS, mainly due to low HDL-C levels. BMI has a significant impact on this abnormality.
Resumo:
The molecular pathology of meningiomas and shwannomas involve the inactivation of the NF2 gene to generate grade I tumors. Genomic losses at 1p and 14q are observed in both neoplasms, although more frequently in meningiomas. The inactivation of unidentified genes located in these regions appears associated with tumor progression in meningiomas, but no clues to its molecular/clinical meaning are available in schwannomas. Recent microarray gene expression studies have demonstrated the existence of molecular subgroups in both entities. In the present study, we correlated the presence of genomic deletions at 1p, 14q, and 22q with the expression patterns of 96 tumor-related genes obtained by cDNA low-density microarrays in a series of 65 tumors including 42 meningiomas and 23 schwannomas. Two expression pattern groups were identified by cDNA mycroarray analysis when compared to the expression pattern in normal control RNA in both meningiomas and schwannomas, each one with patterns similar and different from the normal control. Meningioma and schwannoma subgroups differed in the expression of 38 and 16 genes, respectively. Using MLPA and microsatellites, we identified genomic losses at 1p, 14q, and 22q at nonrandom frequencies (12.5-69%) in meningiomas and schwannomas. Losses at 22q were almost equally frequent in both molecular expression subgroups in both neoplasms. However, deletions at 1p and 14q accumulated in meningiomas with a gene expression pattern different from the normal pattern, whereas the inverse situation occurred in schwannomas. Those anomalies characterized the schwannomas with expression pattern similar to the normal control. These findings suggest that deletions at 1p and 14q enhance the development of an abnormal tumor-related gene expression pattern in meningiomas, but this fact is not corroborated in schwannomas. (C) 2010 Elsevier Inc. All rights reserved.
Resumo:
Trabalho de Projeto apresentado como requisito parcial para obtenção do grau de Mestre em Estatística e Gestão de Informação
Resumo:
BACKGROUND Spain shows the highest bladder cancer incidence rates in men among European countries. The most important risk factors are tobacco smoking and occupational exposure to a range of different chemical substances, such as aromatic amines. METHODS This paper describes the municipal distribution of bladder cancer mortality and attempts to "adjust" this spatial pattern for the prevalence of smokers, using the autoregressive spatial model proposed by Besag, York and Molliè, with relative risk of lung cancer mortality as a surrogate. RESULTS It has been possible to compile and ascertain the posterior distribution of relative risk for bladder cancer adjusted for lung cancer mortality, on the basis of a single Bayesian spatial model covering all of Spain's 8077 towns. Maps were plotted depicting smoothed relative risk (RR) estimates, and the distribution of the posterior probability of RR>1 by sex. Towns that registered the highest relative risks for both sexes were mostly located in the Provinces of Cadiz, Seville, Huelva, Barcelona and Almería. The highest-risk area in Barcelona Province corresponded to very specific municipal areas in the Bages district, e.g., Suría, Sallent, Balsareny, Manresa and Cardona. CONCLUSION Mining/industrial pollution and the risk entailed in certain occupational exposures could in part be dictating the pattern of municipal bladder cancer mortality in Spain. Population exposure to arsenic is a matter that calls for attention. It would be of great interest if the relationship between the chemical quality of drinking water and the frequency of bladder cancer could be studied.
Resumo:
Data mining can be defined as the extraction of previously unknown and potentially useful information from large datasets. The main principle is to devise computer programs that run through databases and automatically seek deterministic patterns. It is applied in different fields of application, e.g., remote sensing, biometry, speech recognition, but has seldom been applied to forensic case data. The intrinsic difficulty related to the use of such data lies in its heterogeneity, which comes from the many different sources of information. The aim of this study is to highlight potential uses of pattern recognition that would provide relevant results from a criminal intelligence point of view. The role of data mining within a global crime analysis methodology is to detect all types of structures in a dataset. Once filtered and interpreted, those structures can point to previously unseen criminal activities. The interpretation of patterns for intelligence purposes is the final stage of the process. It allows the researcher to validate the whole methodology and to refine each step if necessary. An application to cutting agents found in illicit drug seizures was performed. A combinatorial approach was done, using the presence and the absence of products. Methods coming from the graph theory field were used to extract patterns in data constituted by links between products and place and date of seizure. A data mining process completed using graphing techniques is called ``graph mining''. Patterns were detected that had to be interpreted and compared with preliminary knowledge to establish their relevancy. The illicit drug profiling process is actually an intelligence process that uses preliminary illicit drug classes to classify new samples. Methods proposed in this study could be used \textit{a priori} to compare structures from preliminary and post-detection patterns. This new knowledge of a repeated structure may provide valuable complementary information to profiling and become a source of intelligence.
Resumo:
Expression data contribute significantly to the biological value of the sequenced human genome, providing extensive information about gene structure and the pattern of gene expression. ESTs, together with SAGE libraries and microarray experiment information, provide a broad and rich view of the transcriptome. However, it is difficult to perform large-scale expression mining of the data generated by these diverse experimental approaches. Not only is the data stored in disparate locations, but there is frequent ambiguity in the meaning of terms used to describe the source of the material used in the experiment. Untangling semantic differences between the data provided by different resources is therefore largely reliant on the domain knowledge of a human expert. We present here eVOC, a system which associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. We have curated and annotated 7016 cDNA libraries represented in dbEST, as well as 104 SAGE libraries,with expression information,and provide this as an integrated, public resource that allows the linking of transcripts and libraries with expression terms. Both the vocabularies and the vocabulary-annotated libraries can be retrieved from http://www.sanbi.ac.za/evoc/. Several groups are involved in developing this resource with the aim of unifying transcript expression information.
Resumo:
Background: To compare the characteristics and prognostic features of ischemic stroke in patients with diabetes and without diabetes, and to determine the independent predictors of in-hospital mortality in people with diabetes and ischemic stroke.Methods: Diabetes was diagnosed in 393 (21.3%) of 1,840 consecutive patients with cerebral infarction included in a prospective stroke registry over a 12-year period. Demographic characteristics, cardiovascular risk factors, clinical events, stroke subtypes, neuroimaging data, and outcome in ischemic stroke patients with and without diabetes were compared. Predictors of in-hospital mortality in diabetic patients with ischemic stroke were assessed by multivariate analysis. Results: People with diabetes compared to people without diabetes presented more frequently atherothrombotic stroke (41.2% vs 27%) and lacunar infarction (35.1% vs 23.9%) (P < 0.01). The in-hospital mortality in ischemic stroke patients with diabetes was 12.5% and 14.6% in those without (P = NS). Ischemic heart disease, hyperlipidemia, subacute onset, 85 years old or more, atherothrombotic and lacunar infarcts, and thalamic topography were independently associated with ischemic stroke in patients with diabetes, whereas predictors of in-hospital mortality included the patient's age, decreased consciousness, chronic nephropathy, congestive heart failure and atrial fibrillation. Conclusion: Ischemic stroke in people with diabetes showed a different clinical pattern from those without diabetes, with atherothrombotic stroke and lacunar infarcts being more frequent. Clinical factors indicative of the severity of ischemic stroke available at onset have a predominant influence upon in-hospital mortality and may help clinicians to assess prognosis more accurately.
Resumo:
Background: The Valais's cancer registry (RVsT) of the Observatoire valaisan de le santé (OVS) and the department of oncology of Valais's Hospital conducted a study on the epidemiology and pattern of care of colorectal cancer in Valais. Colorectal cancer is the third cause of death by cancer in Switzerland with about 1600 deaths per year. It is the third most frequent cancer for males and the second most frequent for females in Valais. The number of new colorectal cancer cases (average per year) increased between 1989 and 2009 for males as well as for females in Valais. The number of colorectal cancer death cases (average per year) slightly increased between 1989 and 2009 for males as well as for females in Valais. Age-standardized rates of incidence were stable for males and females in Valais and in Switzerland between 1989 and 2009, while age-standardized rates of mortality decreased for males and females in Valais and Switzerland. Results: 774 cases were recorded (59% males). Median age at diagnosis was 70 years old. Most of cancers were invasive (79%) and the main localization was the colon (71%). The most frequent mode of detection was a consultation for non emergency symptoms (75%), but almost 10% of patients consulted in emergency. 82% of patients were treated within 30 days from diagnosis. 90% of the patients were treated by surgery alone or with combined treatment. The first treatment was surgery, including endoscopic resection in 86% of the cases. The treatment was different according to the localization and the stage of the cancer. Survival rate was 95% at 30 days and 79% at one year. The survival was dependent on the stage and the age at diagnosis. Cox model shows an association between mortality and age (better survival for young people) and between mortality and stage (better survival for the lower stages). Methods: RVsT collects information on all cancer cases since 1989 for people registered in the communes of Valais. RVsT has an authorization to collect non anonymized data. All new incident cancers are coded according to the International Classification of Diseases for Oncology (ICD-O-3) and the stages are coded according to the TNM classification. We studied all cases of in situ and invasive colorectal cancers diagnosed between 2006 and 2009 and registered routinely at the RVsT. We checked for data completeness and if necessary sent questionnaires to avoid missing data. A distance of 15 cm has been chosen to delimitate the colon (sigmoid) and the rectal cancers. We made an active follow-up for vital status to have a valid survival analysis. We analyzed the characteristics of the tumors according to age, sex, localization and stage with stata 9 software. Kaplan-Meier curves were generated and Cox model were fitted to analyze survival. Conclusion: The characteristics of patients and tumors and the one year survival were similar to those observed in Switzerland and some European countries. Patterns of care were close to those recommended in guidelines. Routine data recorded in a cancer registry can be used, not only to provide general statistics, but also to help clinicians assess local practices.