954 resultados para Clusters analysis
Resumo:
A methodology based on data mining techniques to support the analysis of zonal prices in real transmission networks is proposed in this paper. The mentioned methodology uses clustering algorithms to group the buses in typical classes that include a set of buses with similar LMP values. Two different clustering algorithms have been used to determine the LMP clusters: the two-step and K-means algorithms. In order to evaluate the quality of the partition as well as the best performance algorithm adequacy measurements indices are used. The paper includes a case study using a Locational Marginal Prices (LMP) data base from the California ISO (CAISO) in order to identify zonal prices.
Resumo:
Here, we report the molecular analysis of two independent 5S rRNA clusters found in the intergenic region of two ubiquitin genomic clones isolated from Tetrahymena pyriformis. Each cluster contains two 120-bp-long coding regions organized in tandem with 142/145-bp-long spacers.
Resumo:
Dissertação de Mestrado em Gestão de Empresas/MBA.
Resumo:
International Scientific Forum, ISF 2013, ISF 2013, 12-14 December 2013, Tirana.
Resumo:
Cluster analysis for categorical data has been an active area of research. A well-known problem in this area is the determination of the number of clusters, which is unknown and must be inferred from the data. In order to estimate the number of clusters, one often resorts to information criteria, such as BIC (Bayesian information criterion), MML (minimum message length, proposed by Wallace and Boulton, 1968), and ICL (integrated classification likelihood). In this work, we adopt the approach developed by Figueiredo and Jain (2002) for clustering continuous data. They use an MML criterion to select the number of clusters and a variant of the EM algorithm to estimate the model parameters. This EM variant seamlessly integrates model estimation and selection in a single algorithm. For clustering categorical data, we assume a finite mixture of multinomial distributions and implement a new EM algorithm, following a previous version (Silvestre et al., 2008). Results obtained with synthetic datasets are encouraging. The main advantage of the proposed approach, when compared to the above referred criteria, is the speed of execution, which is especially relevant when dealing with large data sets.
Resumo:
OBJECTIVE: To identify clusters of the major occurrences of leprosy and their associated socioeconomic and demographic factors. METHODS: Cases of leprosy that occurred between 1998 and 2007 in São José do Rio Preto (southeastern Brazil) were geocodified and the incidence rates were calculated by census tract. A socioeconomic classification score was obtained using principal component analysis of socioeconomic variables. Thematic maps to visualize the spatial distribution of the incidence of leprosy with respect to socioeconomic levels and demographic density were constructed using geostatistics. RESULTS: While the incidence rate for the entire city was 10.4 cases per 100,000 inhabitants annually between 1998 and 2007, the incidence rates of individual census tracts were heterogeneous, with values that ranged from 0 to 26.9 cases per 100,000 inhabitants per year. Areas with a high leprosy incidence were associated with lower socioeconomic levels. There were identified clusters of leprosy cases, however there was no association between disease incidence and demographic density. There was a disparity between the places where the majority of ill people lived and the location of healthcare services. CONCLUSIONS: The spatial analysis techniques utilized identified the poorer neighborhoods of the city as the areas with the highest risk for the disease. These data show that health departments must prioritize politico-administrative policies to minimize the effects of social inequality and improve the standards of living, hygiene, and education of the population in order to reduce the incidence of leprosy.
Resumo:
This paper analyzes the Portuguese short-run business cycles over the last 150 years and presents the multidimensional scaling (MDS) for visualizing the results. The analytical and numerical assessment of this long-run perspective reveals periods with close connections between the macroeconomic variables related to government accounts equilibrium, balance of payments equilibrium, and economic growth. The MDS method is adopted for a quantitative statistical analysis. In this way, similarity clusters of several historical periods emerge in the MDS maps, namely, in identifying similarities and dissimilarities that identify periods of prosperity and crises, growth, and stagnation. Such features are major aspects of collective national achievement, to which can be associated the impact of international problems such as the World Wars, the Great Depression, or the current global financial crisis, as well as national events in the context of broad political blueprints for the Portuguese society in the rising globalization process.
Resumo:
Global warming and the associated climate changes are being the subject of intensive research due to their major impact on social, economic and health aspects of the human life. Surface temperature time-series characterise Earth as a slow dynamics spatiotemporal system, evidencing long memory behaviour, typical of fractional order systems. Such phenomena are difficult to model and analyse, demanding for alternative approaches. This paper studies the complex correlations between global temperature time-series using the Multidimensional scaling (MDS) approach. MDS provides a graphical representation of the pattern of climatic similarities between regions around the globe. The similarities are quantified through two mathematical indices that correlate the monthly average temperatures observed in meteorological stations, over a given period of time. Furthermore, time dynamics is analysed by performing the MDS analysis over slices sampling the time series. MDS generates maps describing the stations’ locus in the perspective that, if they are perceived to be similar to each other, then they are placed on the map forming clusters. We show that MDS provides an intuitive and useful visual representation of the complex relationships that are present among temperature time-series, which are not perceived on traditional geographic maps. Moreover, MDS avoids sensitivity to the irregular distribution density of the meteorological stations.
Resumo:
Biosignals analysis has become widespread, upstaging their typical use in clinical settings. Electrocardiography (ECG) plays a central role in patient monitoring as a diagnosis tool in today's medicine and as an emerging biometric trait. In this paper we adopt a consensus clustering approach for the unsupervised analysis of an ECG-based biometric records. This type of analysis highlights natural groups within the population under investigation, which can be correlated with ground truth information in order to gain more insights about the data. Preliminary results are promising, for meaningful clusters are extracted from the population under analysis. © 2014 EURASIP.
Resumo:
We propose a graphical method to visualize possible time-varying correlations between fifteen stock market values. The method is useful for observing stable or emerging clusters of stock markets with similar behaviour. The graphs, originated from applying multidimensional scaling techniques (MDS), may also guide the construction of multivariate econometric models.
Resumo:
ABSTRACT OBJECTIVE To describe the spatial distribution of avoidable hospitalizations due to tuberculosis in the municipality of Ribeirao Preto, SP, Brazil, and to identify spatial and space-time clusters for the risk of occurrence of these events. METHODS This is a descriptive, ecological study that considered the hospitalizations records of the Hospital Information System of residents of Ribeirao Preto, SP, Southeastern Brazil, from 2006 to 2012. Only the cases with recorded addresses were considered for the spatial analyses, and they were also geocoded. We resorted to Kernel density estimation to identify the densest areas, local empirical Bayes rate as the method for smoothing the incidence rates of hospital admissions, and scan statistic for identifying clusters of risk. Softwares ArcGis 10.2, TerraView 4.2.2, and SaTScanTM were used in the analysis. RESULTS We identified 169 hospitalizations due to tuberculosis. Most were of men (n = 134; 79.2%), averagely aged 48 years (SD = 16.2). The predominant clinical form was the pulmonary one, which was confirmed through a microscopic examination of expectorated sputum (n = 66; 39.0%). We geocoded 159 cases (94.0%). We observed a non-random spatial distribution of avoidable hospitalizations due to tuberculosis concentrated in the northern and western regions of the municipality. Through the scan statistic, three spatial clusters for risk of hospitalizations due to tuberculosis were identified, one of them in the northern region of the municipality (relative risk [RR] = 3.4; 95%CI 2.7–4,4); the second in the central region, where there is a prison unit (RR = 28.6; 95%CI 22.4–36.6); and the last one in the southern region, and area of protection for hospitalizations (RR = 0.2; 95%CI 0.2–0.3). We did not identify any space-time clusters. CONCLUSIONS The investigation showed priority areas for the control and surveillance of tuberculosis, as well as the profile of the affected population, which shows important aspects to be considered in terms of management and organization of health care services targeting effectiveness in primary health care.
Resumo:
Forest fires dynamics is often characterized by the absence of a characteristic length-scale, long range correlations in space and time, and long memory, which are features also associated with fractional order systems. In this paper a public domain forest fires catalogue, containing information of events for Portugal, covering the period from 1980 up to 2012, is tackled. The events are modelled as time series of Dirac impulses with amplitude proportional to the burnt area. The time series are viewed as the system output and are interpreted as a manifestation of the system dynamics. In the first phase we use the pseudo phase plane (PPP) technique to describe forest fires dynamics. In the second phase we use multidimensional scaling (MDS) visualization tools. The PPP allows the representation of forest fires dynamics in two-dimensional space, by taking time series representative of the phenomena. The MDS approach generates maps where objects that are perceived to be similar to each other are placed on the map forming clusters. The results are analysed in order to extract relationships among the data and to better understand forest fires behaviour.
Resumo:
This paper analyses forest fires in the perspective of dynamical systems. Forest fires exhibit complex correlations in size, space and time, revealing features often present in complex systems, such as the absence of a characteristic length-scale, or the emergence of long range correlations and persistent memory. This study addresses a public domain forest fires catalogue, containing information of events for Portugal, during the period from 1980 up to 2012. The data is analysed in an annual basis, modelling the occurrences as sequences of Dirac impulses with amplitude proportional to the burnt area. First, we consider mutual information to correlate annual patterns. We use visualization trees, generated by hierarchical clustering algorithms, in order to compare and to extract relationships among the data. Second, we adopt the Multidimensional Scaling (MDS) visualization tool. MDS generates maps where each object corresponds to a point. Objects that are perceived to be similar to each other are placed on the map forming clusters. The results are analysed in order to extract relationships among the data and to identify forest fire patterns.
Resumo:
RAPD markers have been used for the analysis of genetic differentiation of Aedes aegypti, because they allow the study of genetic relationships among populations. The aim of this study was to identify populations in different geographic regions of the São Paulo State in order to understand the infestation pattern of A. aegypti. The dendrogram constructed with the combined data set of the RAPD patterns showed that the mosquitoes were segregated into two major clusters. Mosquitoes from the Western region of the São Paulo State constituted one cluster and the other was composed of mosquitoes from a laboratory strain and from a coastal city, where the largest Latin American port is located. These data are in agreement with the report on the infestation in the São Paulo State. The genetic proximity was greater between mosquitoes whose geographic origin was closer. However, mosquitoes from the coastal city were genetically closer to laboratory-reared mosquitoes than to field-collected mosquitoes from the São Paulo State. The origin of the infestation in this place remains unclear, but certainly it is related to mosquitoes of origins different from those that infested the West and North region of the State in the 80's.
Resumo:
The aim of this research was to evaluate the protein polymorphism degree among seventy-five C. albicans strains from healthy children oral cavities of five socioeconomic categories from eight schools (private and public) in Piracicaba city, São Paulo State, in order to identify C. albicans subspecies and their similarities in infantile population groups and to establish their possible dissemination route. Cell cultures were grown in YEPD medium, collected by centrifugation, and washed with cold saline solution. The whole-cell proteins were extracted by cell disruption, using glass beads and submitted to SDS-PAGE technique. After electrophoresis, the protein bands were stained with Coomassie-blue and analyzed by statistics package NTSYS-pc version 1.70 software. Similarity matrix and dendrogram were generated by using the Dice similarity coefficient and UPGMA algorithm, respectively, which made it possible to evaluate the similarity or intra-specific polymorphism degrees, based on whole-cell protein fingerprinting of C. albicans oral isolates. A total of 13 major phenons (clusters) were analyzed, according to their homogeneous (socioeconomic category and/or same school) and heterogeneous (distinct socioeconomic categories and/or schools) characteristics. Regarding to the social epidemiological aspect, the cluster composition showed higher similarities (0.788 < S D < 1.0) among C. albicans strains isolated from healthy children independent of their socioeconomic bases (high, medium, or low). Isolates of high similarity were not found in oral cavities from healthy children of social stratum A and D, B and D, or C and E. This may be explained by an absence of a dissemination route among these children. Geographically, some healthy children among identical and different schools (private and public) also are carriers of similar strains but such similarity was not found among other isolates from children from certain schools. These data may reflect a restricted dissemination route of these microorganisms in some groups of healthy scholars, which may be dependent of either socioeconomic categories or geographic site of each child. In contrast to the higher similarity, the lower similarity or higher polymorphism degree (0.499 < S D < 0.788) of protein profiles was shown in 23 (30.6%) C. albicans oral isolates. Considering the social epidemiological aspect, 42.1%, 41.7%, 26.6%, 23.5%, and 16.7% were isolates from children concerning to socioeconomic categories A, D, C, B, and E, respectively, and geographically, 63.6%, 50%, 33.3%, 33.3%, 30%, 25%, and 14.3% were isolates from children from schools LAE (Liceu Colégio Albert Einstein), MA (E.E.P.S.G. "Prof. Elias de Melo Ayres"), CS (E.E.P.G. "Prof. Carlos Sodero"), AV (Alphaville), HF (E.E.P.S.G. "Honorato Faustino), FMC (E.E.P.G. "Prof. Francisco Mariano da Costa"), and MEP (E.E.P.S.G. "Prof. Manasses Ephraim Pereira), respectively. Such results suggest a higher protein polymorphism degree among some strains isolated from healthy children independent of their socioeconomic strata or geographic sites. Complementary studies, involving healthy students and their families, teachers, servants, hygiene and nutritional habits must be done in order to establish the sources of such colonization patterns in population groups of healthy children. The whole-cell protein profile obtained by SDS-PAGE associated with computer-assisted numerical analysis may provide additional criteria for the taxonomic and epidemiological studies of C. albicans.