959 resultados para frequent pattern mining


Relevância:

30.00% 30.00%

Publicador:

Resumo:

BACKGROUND: The annotation of protein post-translational modifications (PTMs) is an important task of UniProtKB curators and, with continuing improvements in experimental methodology, an ever greater number of articles are being published on this topic. To help curators cope with this growing body of information we have developed a system which extracts information from the scientific literature for the most frequently annotated PTMs in UniProtKB. RESULTS: The procedure uses a pattern-matching and rule-based approach to extract sentences with information on the type and site of modification. A ranked list of protein candidates for the modification is also provided. For PTM extraction, precision varies from 57% to 94%, and recall from 75% to 95%, according to the type of modification. The procedure was used to track new publications on PTMs and to recover potential supporting evidence for phosphorylation sites annotated based on the results of large scale proteomics experiments. CONCLUSIONS: The information retrieval and extraction method we have developed in this study forms the basis of a simple tool for the manual curation of protein post-translational modifications in UniProtKB/Swiss-Prot. Our work demonstrates that even simple text-mining tools can be effectively adapted for database curation tasks, providing that a thorough understanding of the working process and requirements are first obtained. This system can be accessed at http://eagl.unige.ch/PTM/.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Complex cortical malformations associated with mutations in tubulin genes are commonly referred to as "Tubulinopathies". To further characterize the mutation frequency and phenotypes associated with tubulin mutations, we studied a cohort of 60 foetal cases. Twenty-six tubulin mutations were identified, of which TUBA1A mutations were the most prevalent (19 cases), followed by TUBB2B (6 cases) and TUBB3 (one case). Three subtypes clearly emerged. The most frequent (n = 13) was microlissencephaly with corpus callosum agenesis, severely hypoplastic brainstem and cerebellum. The cortical plate was either absent (6/13), with a 2-3 layered pattern (5/13) or less frequently thickened (2/13), often associated with neuroglial overmigration (4/13). All cases had voluminous germinal zones and ganglionic eminences. The second subtype was lissencephaly (n = 7), either classical (4/7) or associated with cerebellar hypoplasia (3/7) with corpus callosum agenesis (6/7). All foetuses with lissencephaly and cerebellar hypoplasia carried distinct TUBA1A mutations, while those with classical lissencephaly harbored recurrent mutations in TUBA1A (3 cases) or TUBB2B (1 case). The third group was polymicrogyria-like cortical dysplasia (n = 6), consisting of asymmetric multifocal or generalized polymicrogyria with inconstant corpus callosum agenesis (4/6) and hypoplastic brainstem and cerebellum (3/6). Polymicrogyria was either unlayered or 4-layered with neuronal heterotopias (5/6) and occasional focal neuroglial overmigration (2/6). Three had TUBA1A mutations and 3 TUBB2B mutations. Foetal TUBA1A tubulinopathies most often consist in microlissencephaly or classical lissencephaly with corpus callosum agenesis, but polymicrogyria may also occur. Conversely, TUBB2B mutations are responsible for either polymicrogyria (4/6) or microlissencephaly (2/6).

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The Mississippi Valley-type (MVT) Pb-Zn ore district at Mezica is hosted by Middle to Upper Triassic platform carbonate rocks in the Northern Karavanke/Drau Range geotectonic units of the Eastern Alps, northeastern Slovenia. The mineralization at Mezica covers an area of 64 km(2) with more than 350 orebodies and numerous galena and sphalerite occurrences, which formed epigenetically, both conformable and discordant to bedding. While knowledge on the style of mineralization has grown considerably, the origin of discordant mineralization is still debated. Sulfur stable isotope analyses of 149 sulfide samples from the different types of orebodies provide new insights on the genesis of these mineralizations and their relationship. Over the whole mining district, sphalerite and galena have delta(34)S values in the range of -24.7 to -1.5% VCDT (-13.5 +/- 5.0%) and -24.7 to -1.4% (-10.7 +/- 5.9%), respectively. These values are in the range of the main MVT deposits of the Drau Range. All sulfide delta(34)S values are negative within a broad range, with delta(34)S(pyrite) < delta(34)S(sphalerite) < delta(34)S(galena) for both conformable and discordant orebodies, indicating isotopically heterogeneous H(2)S in the ore-forming fluids and precipitation of the sulfides at thermodynamic disequilibrium. This clearly supports that the main sulfide sulfur originates from bacterially mediated reduction (BSR) of Middle to Upper Triassic seawater sulfate or evaporite sulfate. Thermochemical sulfate reduction (TSR) by organic compounds contributed a minor amount of (34)S-enriched H(2)S to the ore fluid. The variations of delta(34)S values of galena and coarse-grained sphalerite at orefield scale are generally larger than the differences observed in single hand specimens. The progressively more negative delta(34)S values with time along the different sphalerite generations are consistent with mixing of different H(2)S sources, with a decreasing contribution of H(2)S from regional TSR, and an increase from a local H(2)S reservoir produced by BSR (i.e., sedimentary biogenic pyrite, organo-sulfur compounds). Galena in discordant ore (-11.9 to -1.7%; -7.0 +/- 2.7%, n=12) tends to be depleted in (34)S compared with conformable ore (-24.7 to -2.8%, -11.7 +/- 6.2%, n=39). A similar trend is observed from fine-crystalline sphalerite I to coarse open-space filling sphalerite II. Some variation of the sulfide delta(34)S values is attributed to the inherent variability of bacterial sulfate reduction, including metabolic recycling in a locally partially closed system and contribution of H(2)S from hydrolysis of biogenic pyrite and thermal cracking of organo-sulfur compounds. The results suggest that the conformable orebodies originated by mixing of hydrothermal saline metal-rich fluid with H(2)S-rich pore waters during late burial diagenesis, while the discordant orebodies formed by mobilization of the earlier conformable mineralization.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Recent advances in machine learning methods enable increasingly the automatic construction of various types of computer assisted methods that have been difficult or laborious to program by human experts. The tasks for which this kind of tools are needed arise in many areas, here especially in the fields of bioinformatics and natural language processing. The machine learning methods may not work satisfactorily if they are not appropriately tailored to the task in question. However, their learning performance can often be improved by taking advantage of deeper insight of the application domain or the learning problem at hand. This thesis considers developing kernel-based learning algorithms incorporating this kind of prior knowledge of the task in question in an advantageous way. Moreover, computationally efficient algorithms for training the learning machines for specific tasks are presented. In the context of kernel-based learning methods, the incorporation of prior knowledge is often done by designing appropriate kernel functions. Another well-known way is to develop cost functions that fit to the task under consideration. For disambiguation tasks in natural language, we develop kernel functions that take account of the positional information and the mutual similarities of words. It is shown that the use of this information significantly improves the disambiguation performance of the learning machine. Further, we design a new cost function that is better suitable for the task of information retrieval and for more general ranking problems than the cost functions designed for regression and classification. We also consider other applications of the kernel-based learning algorithms such as text categorization, and pattern recognition in differential display. We develop computationally efficient algorithms for training the considered learning machines with the proposed kernel functions. We also design a fast cross-validation algorithm for regularized least-squares type of learning algorithm. Further, an efficient version of the regularized least-squares algorithm that can be used together with the new cost function for preference learning and ranking tasks is proposed. In summary, we demonstrate that the incorporation of prior knowledge is possible and beneficial, and novel advanced kernels and cost functions can be used in algorithms efficiently.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Pancreatic adenocarcinoma is associated with a very poor prognosis, characterized with a 5-year survival rate of only 5%. Surgery is the only curative treatment for selected patients. Nevertheless, recurrence is very frequent. Identifying prognostic factors is thus warranted. Like numerous other tumors, adenocarcinomas are preceded by preneoplastic lesions. The role and the impact of these lesions remain unclear. This study aimed to assess the impact of the preneoplastic lesion pattern and histo-morphological features, on survival after pancreatic resection. Thirty-five patients who underwent pancreatic resection for pancreatic adenocarcinoma were identified from a prospective database of a single center, between 2003 and 2008. We considered demographics, tumor characteristics and type of treatment. The major outcome was survival. Analyzes were separated into two groups, according to the preneoplastic lesions: Pancreatic intraepithelial neoplasia (PanIN)-related carcinomas and intracanalar papillary mucinous neoplasia (IPMN)-related carcinomas. The former were more frequent, accounting for 63% (22/35). Moreover, they displayed more aggressive features, with a higher tumor stage (p = 0.01) and higher rate of positive lymph nodes (p = 0.019). Lymphatic (p = 0.009) and perinervous (p = 0.019) invasions were also more frequent. Survival was negatively influenced by PanIN preneoplastic lesions (p = 0.015), T3-4 tumor stage (p = 0.038), positive lymph nodes (p = 0.044), lymphatic (p = 0.019) and vascular (p = 0.029) invasions. Pancreatic adenocarcinoma displays different behavior according to its preneoplastic lesion. Indeed, PanIN-related adenocarcinoma showed more aggressive features and lower survival rate. Preneoplastic lesions may represent predictive factors for survival. Their role and predictive value should be investigated more thoroughly.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The virulence pattern of the isolates of Pyricularia grisea from commercial fields of the upland rice (Oryza sativa) cultivars 'Primavera' and 'BRS Bonança' was analyzed. A hundred and seventy monoconidial isolates of the pathogen virulent to 'Primavera' and 139 to 'BRS Bonança' collected from eight fields, during two years (2001-2003) were tested, under greenhouse conditions, on six newly released rice cultivars. Differences in virulence pattern were observed in pathogenic populations of 'Primavera' and 'BRS Bonança'. Isolates with virulence to improved cultivars were common in samples from farmers' fields in the absence of aloinfection. The virulence frequency of P. grisea isolates collected from 'Primavera'' to cultivars 'BRS Vencedora', 'BRS Colosso', 'BRS Liderança', 'BRS Soberana', 'BRS Curinga' and 'BRS Talento', was high in descending order. On the other hand, in the fungus population of 'BRS BRS Bonança' virulence frequency was high in 'BRS Talento', followed by 'BRS Curinga', 'BRS Vencedora', 'BRS Liderança', 'BRS Colosso' and 'BRS Soberana'. While virulence to 'BRS Talento' was rare among isolates from 'Primavera', it was most frequent in isolates of 'BRS Bonança'. The six improved rice cultivars permitted to differentiating agriculturally important virulences in the pathogen population which can be utilized in selecting breeding lines for specific resistance, in rice blast improvement program.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The aim of this study was to group temporal profiles of 10-day composites NDVI product by similarity, which was obtained by the SPOT Vegetation sensor, for municipalities with high soybean production in the state of Paraná, Brazil, in the 2005/2006 cropping season. Data mining is a valuable tool that allows extracting knowledge from a database, identifying valid, new, potentially useful and understandable patterns. Therefore, it was used the methods for clusters generation by means of the algorithms K-Means, MAXVER and DBSCAN, implemented in the WEKA software package. Clusters were created based on the average temporal profiles of NDVI of the 277 municipalities with high soybean production in the state and the best results were found with the K-Means algorithm, grouping the municipalities into six clusters, considering the period from the beginning of October until the end of March, which is equivalent to the crop vegetative cycle. Half of the generated clusters presented spectro-temporal pattern, a characteristic of soybeans and were mostly under the soybean belt in the state of Paraná, which shows good results that were obtained with the proposed methodology as for identification of homogeneous areas. These results will be useful for the creation of regional soybean "masks" to estimate the planted area for this crop.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

This study aimed to identify differences in swine vocalization pattern according to animal gender and different stress conditions. A total of 150 barrow males and 150 females (Dalland® genetic strain), aged 100 days, were used in the experiment. Pigs were exposed to different stressful situations: thirst (no access to water), hunger (no access to food), and thermal stress (THI exceeding 74). For the control treatment, animals were kept under a comfort situation (animals with full access to food and water, with environmental THI lower than 70). Acoustic signals were recorded every 30 minutes, totaling six samples for each stress situation. Afterwards, the audios were analyzed by Praat® 5.1.19 software, generating a sound spectrum. For determination of stress conditions, data were processed by WEKA® 3.5 software, using the decision tree algorithm C4.5, known as J48 in the software environment, considering cross-validation with samples of 10% (10-fold cross-validation). According to the Decision Tree, the acoustic most important attribute for the classification of stress conditions was sound Intensity (root node). It was not possible to identify, using the tested attributes, the animal gender by vocal register. A decision tree was generated for recognition of situations of swine hunger, thirst, and heat stress from records of sound intensity, Pitch frequency, and Formant 1.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

This thesis Entitled Environmental impact of Sand Mining :A case Study in the river catchments of vembanad lake southwest india.The entire study is addressed in nine chapters. Chapter l deals with the general introduction about rivers, problems of river sand mining, objectives, location of the study area and scope of the study. A detailed review on river classification, classic concepts in riverine studies, geological work of rivers and channel processes, importance of river ecosystems and its need for management are dealt in Chapter 2. Chapter 3 gives a comprehensive account of the study area - its location, administrative divisions, physiography, soil, geology, land use and living and non-living resources. The various methods adopted in the study are dealt in Chapter 4. Chapter 5 contains river characteristics like drainage, environmental and geologic setting, channel characteristics, river discharge and water quality of the study area. Chapter 6 gives an account of river sand mining (instream and floodplain mining) from the study area. The various environmental problems of river sand mining on the land adjoining the river banks, river channel, water, biotic and social / human environments of the area and data interpretation are presented in Chapter 7. Chapter 8 deals with the Environmental Impact Assessment (EIA) and Environmental Management Plan (EMP) of sand mining from the river catchments of Vembanad lake.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Data mining is one of the hottest research areas nowadays as it has got wide variety of applications in common man’s life to make the world a better place to live. It is all about finding interesting hidden patterns in a huge history data base. As an example, from a sales data base, one can find an interesting pattern like “people who buy magazines tend to buy news papers also” using data mining. Now in the sales point of view the advantage is that one can place these things together in the shop to increase sales. In this research work, data mining is effectively applied to a domain called placement chance prediction, since taking wise career decision is so crucial for anybody for sure. In India technical manpower analysis is carried out by an organization named National Technical Manpower Information System (NTMIS), established in 1983-84 by India's Ministry of Education & Culture. The NTMIS comprises of a lead centre in the IAMR, New Delhi, and 21 nodal centres located at different parts of the country. The Kerala State Nodal Centre is located at Cochin University of Science and Technology. In Nodal Centre, they collect placement information by sending postal questionnaire to passed out students on a regular basis. From this raw data available in the nodal centre, a history data base was prepared. Each record in this data base includes entrance rank ranges, reservation, Sector, Sex, and a particular engineering. From each such combination of attributes from the history data base of student records, corresponding placement chances is computed and stored in the history data base. From this data, various popular data mining models are built and tested. These models can be used to predict the most suitable branch for a particular new student with one of the above combination of criteria. Also a detailed performance comparison of the various data mining models is done.This research work proposes to use a combination of data mining models namely a hybrid stacking ensemble for better predictions. A strategy to predict the overall absorption rate for various branches as well as the time it takes for all the students of a particular branch to get placed etc are also proposed. Finally, this research work puts forward a new data mining algorithm namely C 4.5 * stat for numeric data sets which has been proved to have competent accuracy over standard benchmarking data sets called UCI data sets. It also proposes an optimization strategy called parameter tuning to improve the standard C 4.5 algorithm. As a summary this research work passes through all four dimensions for a typical data mining research work, namely application to a domain, development of classifier models, optimization and ensemble methods.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Knowledge discovery support environments include beside classical data analysis tools also data mining tools. For supporting both kinds of tools, a unified knowledge representation is needed. We show that concept lattices which are used as knowledge representation in Conceptual Information Systems can also be used for structuring the results of mining association rules. Vice versa, we use ideas of association rules for reducing the complexity of the visualization of Conceptual Information Systems.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

We present a new algorithm called TITANIC for computing concept lattices. It is based on data mining techniques for computing frequent itemsets. The algorithm is experimentally evaluated and compared with B. Ganter's Next-Closure algorithm.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Formal Concept Analysis is an unsupervised learning technique for conceptual clustering. We introduce the notion of iceberg concept lattices and show their use in Knowledge Discovery in Databases (KDD). Iceberg lattices are designed for analyzing very large databases. In particular they serve as a condensed representation of frequent patterns as known from association rule mining. In order to show the interplay between Formal Concept Analysis and association rule mining, we discuss the algorithm TITANIC. We show that iceberg concept lattices are a starting point for computing condensed sets of association rules without loss of information, and are a visualization method for the resulting rules.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

This paper provides an extended analysis of livelihood diversification in rural Tanzania, with special emphasis on artisanal and small-scale mining (ASM). Over the past decade, this sector of industry, which is labour-intensive and comprises an array of rudimentary and semi-mechanized operations, has become an indispensable economic activity throughout Sub-Saharan Africa, providing employment to a host of redundant public sector workers, retrenched large-scale mine labourers and poor farmers. In many of the region’s rural areas, it is overtaking subsistence agriculture as the primary industry. Such a pattern appears to be unfolding within the Morogoro and Mbeya regions of southern Tanzania, where findings from recent research suggest that a growing number of smallholder farmers are turning to ASM for employment and financial support. It is imperative that national rural development programmes take this trend into account and provide support to these people.