966 resultados para Inverse frequent set mining


Relevância:

30.00% 30.00%

Publicador:

Resumo:

In this paper a support vector machine (SVM) approach for characterizing the feasible parameter set (FPS) in non-linear set-membership estimation problems is presented. It iteratively solves a regression problem from which an approximation of the boundary of the FPS can be determined. To guarantee convergence to the boundary the procedure includes a no-derivative line search and for an appropriate coverage of points on the FPS boundary it is suggested to start with a sequential box pavement procedure. The SVM approach is illustrated on a simple sine and exponential model with two parameters and an agro-forestry simulation model.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

We consider the Dirichlet boundary-value problem for the Helmholtz equation, Au + x2u = 0, with Imx > 0. in an hrbitrary bounded or unbounded open set C c W. Assuming continuity of the solution up to the boundary and a bound on growth a infinity, that lu(x)l < Cexp (Slxl), for some C > 0 and S~< Imx, we prove that the homogeneous problem has only the trivial salution. With this resnlt we prove uniqueness results for direct and inverse problems of scattering by a bounded or infinite obstacle.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

In the context of environmental valuation of natural disasters, an important component of the evaluation procedure lies in determining the periodicity of events. This paper explores alternative methodologies for determining such periodicity, illustrating the advantages and the disadvantages of the separate methods and their comparative predictions. The procedures employ Bayesian inference and explore recent advances in computational aspects of mixtures methodology. The procedures are applied to the classic data set of Maguire et al (Biometrika, 1952) which was subsequently updated by Jarrett (Biometrika, 1979) and which comprise the seminal investigations examining the periodicity of mining disasters within the United Kingdom, 1851-1962.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The modeling and analysis of lifetime data is an important aspect of statistical work in a wide variety of scientific and technological fields. Good (1953) introduced a probability distribution which is commonly used in the analysis of lifetime data. For the first time, based on this distribution, we propose the so-called exponentiated generalized inverse Gaussian distribution, which extends the exponentiated standard gamma distribution (Nadarajah and Kotz, 2006). Various structural properties of the new distribution are derived, including expansions for its moments, moment generating function, moments of the order statistics, and so forth. We discuss maximum likelihood estimation of the model parameters. The usefulness of the new model is illustrated by means of a real data set. (c) 2010 Elsevier B.V. All rights reserved.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Tendo como motivação o desenvolvimento de uma representação gráfica de redes com grande número de vértices, útil para aplicações de filtro colaborativo, este trabalho propõe a utilização de superfícies de coesão sobre uma base temática multidimensionalmente escalonada. Para isso, utiliza uma combinação de escalonamento multidimensional clássico e análise de procrustes, em algoritmo iterativo que encaminha soluções parciais, depois combinadas numa solução global. Aplicado a um exemplo de transações de empréstimo de livros pela Biblioteca Karl A. Boedecker, o algoritmo proposto produz saídas interpretáveis e coerentes tematicamente, e apresenta um stress menor que a solução por escalonamento clássico.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The main goal of our research was to search for SSRs in the Eucalyptus EST FORESTs database (using a software for mining SSR-motifs). With this objective, we created a database for cataloging Eucalyptus EST-derived SSRs, and developed a bioinformatics tool, named Satellyptus, for finding and analyzing microsatellites in the Eucalyptus EST database. The search for microsatellites in the FORESTs database containing 71,115 Eucalyptus EST sequences (52.09 Mb) revealed 20,530 SSRs in 15,621 ESTs. The SSR abundance detected on the Eucalyptus ESTs database (29% or one microsatellite every four sequences) is considered very high for plants. Amongst the categories of SSR motifs, the dimeric (37%) and trimeric ones (33%) predominated. The AG/CT motif was the most frequent (35.15%) followed by the trimeric CCG/CGG (12.81%). From a random sample of 1,217 sequences, 343 microsatellites in 265 SSR-containing sequences were identified. Approximately 48% of these ESTs containing microsatellites were homologous to proteins with known biological function. Most of the microsatellites detected in Eucalyptus ESTs were positioned at either the 5 or 3 end. Our next priority involves the design of flanking primers for codominant SSR loci, which could lead to the development of a set of microsatellite-based markers suitable for marker-assisted Eucalyptus breeding programs.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Fundação de Amparo à Pesquisa do Estado de São Paulo (FAPESP)

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The code STATFLUX, implementing a new and simple statistical procedure for the calculation of transfer coefficients in radionuclide transport to animals and plants, is proposed. The method is based on the general multiple-compartment model, which uses a system of linear equations involving geometrical volume considerations. Flow parameters were estimated by employing two different least-squares procedures: Derivative and Gauss-Marquardt methods, with the available experimental data of radionuclide concentrations as the input functions of time. The solution of the inverse problem, which relates a given set of flow parameter with the time evolution of concentration functions, is achieved via a Monte Carlo Simulation procedure.Program summaryTitle of program: STATFLUXCatalogue identifier: ADYS_v1_0Program summary URL: http://cpc.cs.qub.ac.uk/summaries/ADYS_v1_0Program obtainable from: CPC Program Library, Queen's University of Belfast, N. IrelandLicensing provisions: noneComputer for which the program is designed and others on which it has been tested: Micro-computer with Intel Pentium III, 3.0 GHzInstallation: Laboratory of Linear Accelerator, Department of Experimental Physics, University of São Paulo, BrazilOperating system: Windows 2000 and Windows XPProgramming language used: Fortran-77 as implemented in Microsoft Fortran 4.0. NOTE: Microsoft Fortran includes non-standard features which are used in this program. Standard Fortran compilers such as, g77, f77, ifort and NAG95, are not able to compile the code and therefore it has not been possible for the CPC Program Library to test the program.Memory, required to execute with typical data: 8 Mbytes of RAM memory and 100 MB of Hard disk memoryNo. of bits in a word: 16No. of lines in distributed program, including test data, etc.: 6912No. of bytes in distributed Program, including test data, etc.: 229 541Distribution format: tar.gzNature of the physical problem: the investigation of transport mechanisms for radioactive substances, through environmental pathways, is very important for radiological protection of populations. One such pathway, associated with the food chain, is the grass-animal-man sequence. The distribution of trace elements in humans and laboratory animals has been intensively studied over the past 60 years [R.C. Pendlenton, C.W. Mays, R.D. Lloyd, A.L. Brooks, Differential accumulation of iodine-131 from local fallout in people and milk, Health Phys. 9 (1963) 1253-1262]. In addition, investigations on the incidence of cancer in humans, and a possible causal relationship to radioactive fallout, have been undertaken [E.S. Weiss, M.L. Rallison, W.T. London, W.T. Carlyle Thompson, Thyroid nodularity in southwestern Utah school children exposed to fallout radiation, Amer. J. Public Health 61 (1971) 241-249; M.L. Rallison, B.M. Dobyns, F.R. Keating, J.E. Rall, F.H. Tyler, Thyroid diseases in children, Amer. J. Med. 56 (1974) 457-463; J.L. Lyon, M.R. Klauber, J.W. Gardner, K.S. Udall, Childhood leukemia associated with fallout from nuclear testing, N. Engl. J. Med. 300 (1979) 397-402]. From the pathways of entry of radionuclides in the human (or animal) body, ingestion is the most important because it is closely related to life-long alimentary (or dietary) habits. Those radionuclides which are able to enter the living cells by either metabolic or other processes give rise to localized doses which can be very high. The evaluation of these internally localized doses is of paramount importance for the assessment of radiobiological risks and radiological protection. The time behavior of trace concentration in organs is the principal input for prediction of internal doses after acute or chronic exposure. The General Multiple-Compartment Model (GMCM) is the powerful and more accepted method for biokinetical studies, which allows the calculation of concentration of trace elements in organs as a function of time, when the flow parameters of the model are known. However, few biokinetics data exist in the literature, and the determination of flow and transfer parameters by statistical fitting for each system is an open problem.Restriction on the complexity of the problem: This version of the code works with the constant volume approximation, which is valid for many situations where the biological half-live of a trace is lower than the volume rise time. Another restriction is related to the central flux model. The model considered in the code assumes that exist one central compartment (e.g., blood), that connect the flow with all compartments, and the flow between other compartments is not included.Typical running time: Depends on the choice for calculations. Using the Derivative Method the time is very short (a few minutes) for any number of compartments considered. When the Gauss-Marquardt iterative method is used the calculation time can be approximately 5-6 hours when similar to 15 compartments are considered. (C) 2006 Elsevier B.V. All rights reserved.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The multi-relational Data Mining approach has emerged as alternative to the analysis of structured data, such as relational databases. Unlike traditional algorithms, the multi-relational proposals allow mining directly multiple tables, avoiding the costly join operations. In this paper, is presented a comparative study involving the traditional Patricia Mine algorithm and its corresponding multi-relational proposed, MR-Radix in order to evaluate the performance of two approaches for mining association rules are used for relational databases. This study presents two original contributions: the proposition of an algorithm multi-relational MR-Radix, which is efficient for use in relational databases, both in terms of execution time and in relation to memory usage and the presentation of the empirical approach multirelational advantage in performance over several tables, which avoids the costly join operations from multiple tables. © 2011 IEEE.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

O depósito mineral de Sapucaia, situado no município de Bonito, região nordeste do Estado do Pará, é parte de um conjunto de ocorrências de fosfatos de alumínio lateríticos localizados predominantemente ao longo da zona costeira dos estados do Pará e Maranhão. Estes depósitos foram alvos de estudo desde o início do século passado, quando as primeiras descrições de “bauxitas fosforosas” foram mencionadas na região NW do Maranhão. Nas últimas décadas, com o crescimento acentuado da demanda por produtos fertilizantes pelo mercado agrícola mundial, diversos projetos de exploração mineral foram iniciados ou tiveram seus recursos ampliados no território brasileiro, dentre estes destaca-se a viabilização econômica de depósitos de fosfatos aluminosos, como o de Sapucaia, que vem a ser o primeiro projeto econômico mineral de produção e comercialização de termofosfatos do Brasil. Este trabalho teve como principal objetivo caracterizar a geologia, a constituição mineralógica e a geoquímica do perfil laterítico alumino-fosfático do morro Sapucaia. A macrorregião abrange terrenos dominados em sua maioria por rochas pré-cambrianas a paleozóicas, localmente definidas pela Formação Pirabas, Formação Barreiras, Latossolos e sedimentos recentes. A morfologia do depósito é caracterizada por um discreto morrote alongado que apresenta suaves e contínuos declives em suas bordas, e que tornam raras as exposições naturais dos horizontes do perfil laterítico. Desta forma, a metodologia aplicada para a caracterização do depósito tomou como base o programa de pesquisa geológica executada pela Fosfatar Mineração, até então detentora dos respectivos direitos minerais, onde foram disponibilizadas duas trincheiras e amostras de 8 testemunhos de sondagem. A amostragem limitou-se à extensão litológica do perfil laterítico, com a seleção de 44 amostras em intervalos médios de 1m, e que foram submetidas a uma rota de preparação e análise em laboratório. Em consonância com as demais ocorrências da região do Gurupi, os fosfatos de Sapucaia constituem um horizonte individualizado, de geometria predominantemente tabular, denominado simplesmente de horizonte de fosfatos de alumínio ou crosta aluminofosfática, que varia texturalmente de maciça a cavernosa, porosa a microporosa, que para o topo grada para uma crosta ferroalumino fosfática, tipo pele-de-onça, compacta a cavernosa, composta por nódulos de hematita e/ou goethita cimentados por fosfatos de alumínio, com características similares aos do horizonte de fosfatos subjacente. A crosta aluminofosfática, para a base do perfil, grada para um espesso horizonte argiloso caulinítico com níveis arenosos, que repousa sobre sedimentos heterolíticos intemperizados de granulação fina, aspecto argiloso, por vezes sericítico, intercalados por horizontes arenosos, e que não possuem correlação aparente com as demais rochas aflorantes da geologia na região. Aproximadamente 40% da superfície do morro é encoberta por colúvio composto por fragmentos mineralizados da crosta e por sedimentos arenosos da Formação Barreiras. Na crosta, os fosfatos de alumínio estão representados predominantemente pelo subgrupo da crandallita: i) série crandallita-goyazita (média de 57,3%); ii) woodhouseíta-svanbergita (média de 15,8%); e pela iii) wardita-millisita (média de 5,1%). Associados aos fosfatos encontram-se hematita, goethita, quartzo, caulinita, muscovita e anatásio, com volumes que variam segundo o horizonte laterítico correspondente. Como os minerais pesados em nível acessório a raro estão zircão, estaurolita, turmalina, anatásio, andalusita e silimanita. O horizonte de fosfatos, bem como a crosta ferroalumínio-fosfática, mostra-se claramente rica em P2O5, além de Fe2O3, CaO, Na2O, SrO, SO3, Th, Ta e em terras-raras leves como La e Ce em relação ao horizonte saprolítico. Os teores de SiO2 são consideravelmente elevados, porém muito inferiores aqueles identificados no horizonte argiloso sotoposto. No perfil como um todo, observa-se uma correlação inversa entre SiO2 e Al2O3; entre Al2O3 e Fe2O3, e positiva entre SiO2 e Fe2O3, que ratificam a natureza laterítica do perfil. Diferente do que é esperado para lateritos bauxíticos, os teores de P2O5, CaO, Na2O, SrO e SO3 são fortemente elevados, concentrações consideradas típicas de depósitos de fosfatos de alumínio ricos em crandallita-goyazita e woodhouseítasvanbergita. A sucessão dos horizontes, sua composição mineralógica, e os padrões geoquímicos permitem correlacionar o presente depósito com os demais fosfatos de alumínio da região, mais especificamente Jandiá (Pará) e Trauíra (Maranhão), bem como outros situados além do território brasileiro, indicando portanto, que os fosfatos de alumínio de Sapucaia são produtos da gênese de um perfil laterítico maturo e completo, cuja rocha fonte pode estar relacionada a rochas mineralizadas em fósforo, tais como as observadas na Formação Pimenteiras, parcialmente aflorante na borda da Bacia do Parnaíba. Possivelmente, o atual corpo de minério integrou a paleocosta do mar de Pirabas, uma vez que furos de sondagem às proximidades do corpo deixaram claro a relação de contato lateral entre estas unidades.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Background: Once multi-relational approach has emerged as an alternative for analyzing structured data such as relational databases, since they allow applying data mining in multiple tables directly, thus avoiding expensive joining operations and semantic losses, this work proposes an algorithm with multi-relational approach. Methods: Aiming to compare traditional approach performance and multi-relational for mining association rules, this paper discusses an empirical study between PatriciaMine - an traditional algorithm - and its corresponding multi-relational proposed, MR-Radix. Results: This work showed advantages of the multi-relational approach in performance over several tables, which avoids the high cost for joining operations from multiple tables and semantic losses. The performance provided by the algorithm MR-Radix shows faster than PatriciaMine, despite handling complex multi-relational patterns. The utilized memory indicates a more conservative growth curve for MR-Radix than PatriciaMine, which shows the increase in demand of frequent items in MR-Radix does not result in a significant growth of utilized memory like in PatriciaMine. Conclusion: The comparative study between PatriciaMine and MR-Radix confirmed efficacy of the multi-relational approach in data mining process both in terms of execution time and in relation to memory usage. Besides that, the multi-relational proposed algorithm, unlike other algorithms of this approach, is efficient for use in large relational databases.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The increase in new electronic devices had generated a considerable increase in obtaining spatial data information; hence these data are becoming more and more widely used. As well as for conventional data, spatial data need to be analyzed so interesting information can be retrieved from them. Therefore, data clustering techniques can be used to extract clusters of a set of spatial data. However, current approaches do not consider the implicit semantics that exist between a region and an object’s attributes. This paper presents an approach that enhances spatial data mining process, so they can use the semantic that exists within a region. A framework was developed, OntoSDM, which enables spatial data mining algorithms to communicate with ontologies in order to enhance the algorithm’s result. The experiments demonstrated a semantically improved result, generating more interesting clusters, therefore reducing manual analysis work of an expert.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Purpose: We sought to determine the mechanisms of downregulation of the airway transcription factor Foxa2 in lung cancer and the expression status of Foxa2 in non-small-cell lung cancer (NSCLC). Methods: A series of 25 lung cancer cell lines were evaluated for Foxa2 protein expression, FOXA2 mRNA levels, FOXA2 mutations, FOXA2 copy number changes and for evidence of FOXA2 promoter hypermethylation. In addition, 32 NSCLCs were sequenced for FOXA2 mutations and 173 primary NSCLC tumors evaluated for Foxa2 expression using an immunohistochemical assay. Results: Out of the 25 cell lines, 13 (52%) had undetectable FOXA2 mRNA. The expression of FOXA2 mRNA and Foxa2 protein were congruent in 19/22 cells (p = 0.001). FOXA2 mutations were not identified in primary NSCLCs and were infrequent in cell lines. Focal or broad chromosomal deletions involving FOXA2 were not present. The promoter region of FOXA2 had evidence of hypermethylation, with an inverse correlation between FOXA2 mRNA expression and presence of CpG dinucleotide methylation (p < 0.0001). In primary NSCLC tumor specimens, there was a high frequency of either absence (42/173, 24.2%) or no/low expression (96/173,55.4%) of Foxa2. In 130 patients with stage I NSCLC there was a trend towards decreased survival in tumors with no/low expression of Foxa2 (HR of 1.6, 95%CI 0.9-3.1; p = 0.122). Conclusions: Loss of expression of Foxa2 is frequent in lung cancer cell lines and NSCLCs. The main mechanism of downregulation of Foxa2 is epigenetic silencing through promoter hypermethylation. Further elucidation of the involvement of Foxa2 and other airway transcription factors in the pathogenesis of lung cancer may identify novel therapeutic targets. (C) 2012 Elsevier Ireland Ltd. All rights reserved.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Abstract Background Once multi-relational approach has emerged as an alternative for analyzing structured data such as relational databases, since they allow applying data mining in multiple tables directly, thus avoiding expensive joining operations and semantic losses, this work proposes an algorithm with multi-relational approach. Methods Aiming to compare traditional approach performance and multi-relational for mining association rules, this paper discusses an empirical study between PatriciaMine - an traditional algorithm - and its corresponding multi-relational proposed, MR-Radix. Results This work showed advantages of the multi-relational approach in performance over several tables, which avoids the high cost for joining operations from multiple tables and semantic losses. The performance provided by the algorithm MR-Radix shows faster than PatriciaMine, despite handling complex multi-relational patterns. The utilized memory indicates a more conservative growth curve for MR-Radix than PatriciaMine, which shows the increase in demand of frequent items in MR-Radix does not result in a significant growth of utilized memory like in PatriciaMine. Conclusion The comparative study between PatriciaMine and MR-Radix confirmed efficacy of the multi-relational approach in data mining process both in terms of execution time and in relation to memory usage. Besides that, the multi-relational proposed algorithm, unlike other algorithms of this approach, is efficient for use in large relational databases.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Given a large image set, in which very few images have labels, how to guess labels for the remaining majority? How to spot images that need brand new labels different from the predefined ones? How to summarize these data to route the user’s attention to what really matters? Here we answer all these questions. Specifically, we propose QuMinS, a fast, scalable solution to two problems: (i) Low-labor labeling (LLL) – given an image set, very few images have labels, find the most appropriate labels for the rest; and (ii) Mining and attention routing – in the same setting, find clusters, the top-'N IND.O' outlier images, and the 'N IND.R' images that best represent the data. Experiments on satellite images spanning up to 2.25 GB show that, contrasting to the state-of-the-art labeling techniques, QuMinS scales linearly on the data size, being up to 40 times faster than top competitors (GCap), still achieving better or equal accuracy, it spots images that potentially require unpredicted labels, and it works even with tiny initial label sets, i.e., nearly five examples. We also report a case study of our method’s practical usage to show that QuMinS is a viable tool for automatic coffee crop detection from remote sensing images.