999 resultados para Ontology population
Resumo:
Software de lectura y población de ontología con información de DBpedia y Wikipedia.
Resumo:
Le dictionnaire LVF (Les Verbes Français) de J. Dubois et F. Dubois-Charlier représente une des ressources lexicales les plus importantes dans la langue française qui est caractérisée par une description sémantique et syntaxique très pertinente. Le LVF a été mis disponible sous un format XML pour rendre l’accès aux informations plus commode pour les applications informatiques telles que les applications de traitement automatique de la langue française. Avec l’émergence du web sémantique et la diffusion rapide de ses technologies et standards tels que XML, RDF/RDFS et OWL, il serait intéressant de représenter LVF en un langage plus formalisé afin de mieux l’exploiter par les applications du traitement automatique de la langue ou du web sémantique. Nous en présentons dans ce mémoire une version ontologique OWL en détaillant le processus de transformation de la version XML à OWL et nous en démontrons son utilisation dans le domaine du traitement automatique de la langue avec une application d’annotation sémantique développée dans GATE.
Resumo:
OntoTag - A Linguistic and Ontological Annotation Model Suitable for the Semantic Web
1. INTRODUCTION. LINGUISTIC TOOLS AND ANNOTATIONS: THEIR LIGHTS AND SHADOWS
Computational Linguistics is already a consolidated research area. It builds upon the results of other two major ones, namely Linguistics and Computer Science and Engineering, and it aims at developing computational models of human language (or natural language, as it is termed in this area). Possibly, its most well-known applications are the different tools developed so far for processing human language, such as machine translation systems and speech recognizers or dictation programs.
These tools for processing human language are commonly referred to as linguistic tools. Apart from the examples mentioned above, there are also other types of linguistic tools that perhaps are not so well-known, but on which most of the other applications of Computational Linguistics are built. These other types of linguistic tools comprise POS taggers, natural language parsers and semantic taggers, amongst others. All of them can be termed linguistic annotation tools.
Linguistic annotation tools are important assets. In fact, POS and semantic taggers (and, to a lesser extent, also natural language parsers) have become critical resources for the computer applications that process natural language. Hence, any computer application that has to analyse a text automatically and ‘intelligently’ will include at least a module for POS tagging. The more an application needs to ‘understand’ the meaning of the text it processes, the more linguistic tools and/or modules it will incorporate and integrate.
However, linguistic annotation tools have still some limitations, which can be summarised as follows:
1. Normally, they perform annotations only at a certain linguistic level (that is, Morphology, Syntax, Semantics, etc.).
2. They usually introduce a certain rate of errors and ambiguities when tagging. This error rate ranges from 10 percent up to 50 percent of the units annotated for unrestricted, general texts.
3. Their annotations are most frequently formulated in terms of an annotation schema designed and implemented ad hoc.
A priori, it seems that the interoperation and the integration of several linguistic tools into an appropriate software architecture could most likely solve the limitations stated in (1). Besides, integrating several linguistic annotation tools and making them interoperate could also minimise the limitation stated in (2). Nevertheless, in the latter case, all these tools should produce annotations for a common level, which would have to be combined in order to correct their corresponding errors and inaccuracies. Yet, the limitation stated in (3) prevents both types of integration and interoperation from being easily achieved.
In addition, most high-level annotation tools rely on other lower-level annotation tools and their outputs to generate their own ones. For example, sense-tagging tools (operating at the semantic level) often use POS taggers (operating at a lower level, i.e., the morphosyntactic) to identify the grammatical category of the word or lexical unit they are annotating. Accordingly, if a faulty or inaccurate low-level annotation tool is to be used by other higher-level one in its process, the errors and inaccuracies of the former should be minimised in advance. Otherwise, these errors and inaccuracies would be transferred to (and even magnified in) the annotations of the high-level annotation tool.
Therefore, it would be quite useful to find a way to
(i) correct or, at least, reduce the errors and the inaccuracies of lower-level linguistic tools;
(ii) unify the annotation schemas of different linguistic annotation tools or, more generally speaking, make these tools (as well as their annotations) interoperate.
Clearly, solving (i) and (ii) should ease the automatic annotation of web pages by means of linguistic tools, and their transformation into Semantic Web pages (Berners-Lee, Hendler and Lassila, 2001). Yet, as stated above, (ii) is a type of interoperability problem. There again, ontologies (Gruber, 1993; Borst, 1997) have been successfully applied thus far to solve several interoperability problems. Hence, ontologies should help solve also the problems and limitations of linguistic annotation tools aforementioned.
Thus, to summarise, the main aim of the present work was to combine somehow these separated approaches, mechanisms and tools for annotation from Linguistics and Ontological Engineering (and the Semantic Web) in a sort of hybrid (linguistic and ontological) annotation model, suitable for both areas. This hybrid (semantic) annotation model should (a) benefit from the advances, models, techniques, mechanisms and tools of these two areas; (b) minimise (and even solve, when possible) some of the problems found in each of them; and (c) be suitable for the Semantic Web. The concrete goals that helped attain this aim are presented in the following section.
2. GOALS OF THE PRESENT WORK
As mentioned above, the main goal of this work was to specify a hybrid (that is, linguistically-motivated and ontology-based) model of annotation suitable for the Semantic Web (i.e. it had to produce a semantic annotation of web page contents). This entailed that the tags included in the annotations of the model had to (1) represent linguistic concepts (or linguistic categories, as they are termed in ISO/DCR (2008)), in order for this model to be linguistically-motivated; (2) be ontological terms (i.e., use an ontological vocabulary), in order for the model to be ontology-based; and (3) be structured (linked) as a collection of ontology-based
Resumo:
Automated ontology population using information extraction algorithms can produce inconsistent knowledge bases. Confidence values assigned by the extraction algorithms may serve as evidence in helping to repair inconsistencies. The Dempster-Shafer theory of evidence is a formalism, which allows appropriate interpretation of extractors’ confidence values. This chapter presents an algorithm for translating the subontologies containing conflicts into belief propagation networks and repairing conflicts based on the Dempster-Shafer plausibility.
Resumo:
Genetic determinants of blood pressure are poorly defined. We undertook a large-scale, gene-centric analysis to identify loci and pathways associated with ambulatory systolic and diastolic blood pressure. We measured 24-hour ambulatory blood pressure in 2020 individuals from 520 white European nuclear families (the Genetic Regulation of Arterial Pressure of Humans in the Community Study) and genotyped their DNA using the Illumina HumanCVD BeadChip array, which contains ≈50 000 single nucleotide polymorphisms in >2000 cardiovascular candidate loci. We found a strong association between rs13306560 polymorphism in the promoter region of MTHFR and CLCN6 and mean 24-hour diastolic blood pressure; each minor allele copy of rs13306560 was associated with 2.6 mm Hg lower mean 24-hour diastolic blood pressure (P=1.2×10(-8)). rs13306560 was also associated with clinic diastolic blood pressure in a combined analysis of 8129 subjects from the Genetic Regulation of Arterial Pressure of Humans in the Community Study, the CoLaus Study, and the Silesian Cardiovascular Study (P=5.4×10(-6)). Additional analysis of associations between variants in gene ontology-defined pathways and mean 24-hour blood pressure in the Genetic Regulation of Arterial Pressure of Humans in the Community Study showed that cell survival control signaling cascades could play a role in blood pressure regulation. There was also a significant overrepresentation of rare variants (minor allele frequency: <0.05) among polymorphisms showing at least nominal association with mean 24-hour blood pressure indicating that a considerable proportion of its heritability may be explained by uncommon alleles. Through a large-scale gene-centric analysis of ambulatory blood pressure, we identified an association of a novel variant at the MTHFR/CLNC6 locus with diastolic blood pressure and provided new insights into the genetic architecture of blood pressure.
Resumo:
Background: The main goal of the present study was to analyse the genetic architecture of mRNA expression in muscle, a tissue with an outmost economic importance for pig breeders. Previous studies have used F2 crosses to detect porcine expression QTL (eQTL), so they contributed with data that mostly represents the between-breed component of eQTL variation. Herewith, we have analysed eQTL segregation in an outbred Duroc population using two groups of animals with divergent fatness profiles. This approach is particularly suitable to analyse the within-breed component of eQTL variation, with a special emphasis on loci involved in lipid metabolism. Methodology/Principal Findings: GeneChip Porcine Genome arrays (Affymetrix) were used to determine the mRNA expression levels of gluteus medius samples from 105 Duroc barrows. A whole-genome eQTL scan was carried out with a panel of 116 microsatellites. Results allowed us to detect 613 genome-wide significant eQTL unevenly distributed across the pig genome. A clear predominance of trans- over cis-eQTL, was observed. Moreover, 11 trans-regulatory hotspots affecting the expression levels of four to 16 genes were identified. A Gene Ontology study showed that regulatory polymorphisms affected the expression of muscle development and lipid metabolism genes. A number of positional concordances between eQTL and lipid trait QTL were also found, whereas limited evidence of a linear relationship between muscle fat deposition and mRNA levels of eQTL regulated genes was obtained. Conclusions/Significance: Our data provide substantial evidence that there is a remarkable amount of within-breed genetic variation affecting muscle mRNA expression. Most of this variation acts in trans and influences biological processes related with muscle development, lipid deposition and energy balance. The identification of the underlying causal mutations and the ascertainment of their effects on phenotypes would allow gaining a fundamental perspective about how complex traits are built at the molecular level.
Resumo:
Introduction: the statistical record used in the Field Academic Programs (PAC for it’s initials in Spanish) of Rehabilitation denotes generalities in the data conceptualization, which complicates the reliable guidance in making decisions and provides a low support for research in rehabilitation and disability. In response, the Research Group in Rehabilitation and Social Integration of Persons with Disabilities has worked on the creation of a registry to characterize the population seen by Rehabilitation PAC. This registry includes the use of the International Classification of Functioning, Disability and Health (ICF) of the WHO. Methodology: the proposed methodology includes two phases: the first one is a descriptive study and the second one involves performing methodology Methontology, which integrates the identification and development of ontology knowledge. This article contextualizes the progress made in the second phase. Results: the development of the registry in 2008, as an information system, included documentary review and the analysis of possible use scenarios to help guide the design and development of the SIDUR system. The system uses the ICF given that it is a terminology standardization that allows the reduction of ambiguity and that makes easier the transformation of health facts into data translatable to information systems. The record raises three categories and a total of 129 variables Conclusions: SIDUR facilitates accessibility to accurate and updated information, useful for decision making and research.
Resumo:
Background Levels of differentiation among populations depend both on demographic and selective factors: genetic drift and local adaptation increase population differentiation, which is eroded by gene flow and balancing selection. We describe here the genomic distribution and the properties of genomic regions with unusually high and low levels of population differentiation in humans to assess the influence of selective and neutral processes on human genetic structure. Methods Individual SNPs of the Human Genome Diversity Panel (HGDP) showing significantly high or low levels of population differentiation were detected under a hierarchical-island model (HIM). A Hidden Markov Model allowed us to detect genomic regions or islands of high or low population differentiation. Results Under the HIM, only 1.5% of all SNPs are significant at the 1% level, but their genomic spatial distribution is significantly non-random. We find evidence that local adaptation shaped high-differentiation islands, as they are enriched for non-synonymous SNPs and overlap with previously identified candidate regions for positive selection. Moreover there is a negative relationship between the size of islands and recombination rate, which is stronger for islands overlapping with genes. Gene ontology analysis supports the role of diet as a major selective pressure in those highly differentiated islands. Low-differentiation islands are also enriched for non-synonymous SNPs, and contain an overly high proportion of genes belonging to the 'Oncogenesis' biological process. Conclusions Even though selection seems to be acting in shaping islands of high population differentiation, neutral demographic processes might have promoted the appearance of some genomic islands since i) as much as 20% of islands are in non-genic regions ii) these non-genic islands are on average two times shorter than genic islands, suggesting a more rapid erosion by recombination, and iii) most loci are strongly differentiated between Africans and non-Africans, a result consistent with known human demographic history.
Resumo:
We show a new method for term extraction from a domain relevant corpus using natural language processing for the purposes of semi-automatic ontology learning. Literature shows that topical words occur in bursts. We find that the ranking of extracted terms is insensitive to the choice of population model, but calculating frequencies relative to the burst size rather than the document length in words yields significantly different results.
Resumo:
The aim of this study was to assess the quality of diet among the elderly and associations with socio-demographic variables, health-related behaviors, and diseases. A population-based cross-sectional study was conducted in a representative sample of 1,509 elderly participants in a health survey in Campinas, São Paulo State, Brazil. Food quality was assessed using the Revised Diet Quality Index (DQI-R). Mean index scores were estimated and a multiple regression model was employed for the adjusted analyses. The highest diet quality scores were associated with age 80 years or older, Evangelical religion, diabetes mellitus, and physical activity, while the lowest scores were associated with home environments shared with three or more people, smoking, and consumption of soft drinks and alcoholic beverages. The findings emphasize a general need for diet quality improvements in the elderly, specifically in subgroups with unhealthy behaviors, who should be targeted with comprehensive strategies.
Resumo:
The aim of the present study was to identify factors associated with the occurrence of falls among elderly adults in a population-based study (ISACamp 2008). A population-based cross-sectional study was carried out with two-stage cluster sampling. The sample was composed of 1,520 elderly adults living in the urban area of the city of Campinas, São Paulo, Brazil. The occurrence of falls was analyzed based on reports of the main accident occurred in the previous 12 months. Data on socioeconomic/demographic factors and adverse health conditions were tested for possible associations with the outcome. Prevalence ratios (PR) were estimated and adjusted for gender and age using the Poisson multiple regression analysis. Falls were more frequent, after adjustment for gender and age, among female elderly participants (PR = 2.39; 95% confidence interval (95% CI) 1.47 - 3.87), elderly adults (80 years old and older) (PR = 2.50; 95% CI 1.61 - 3.88), widowed (PR = 1.74; 95% CI 1.04 - 2.89) and among elderly adults who had rheumatism/arthritis/arthrosis (PR = 1.58; 95% CI 1.00 - 2.48), osteoporosis (PR = 1.71; 95% CI 1.18 - 2.49), asthma/bronchitis/emphysema (PR = 1,73; 95% CI 1.09 - 2.74), headache (PR = 1.59; 95% CI 1.07 - 2.38), mental common disorder (PR = 1.72; 95% CI 1.12 - 2.64), dizziness (PR = 2.82; 95% CI 1.98 - 4.02), insomnia (PR = 1.75; 95% CI 1.16 - 2.65), use of multiple medications (five or more) (PR = 2.50; 95% CI 1.12 - 5.56) and use of cane/walker (PR = 2.16; 95% CI 1.19 - 3,93). The present study shows segments of the elderly population who are more prone to falls through the identification of factors associated with this outcome. The findings can contribute to the planning of public health policies and programs addressed to the prevention of falls.
Resumo:
In oral and oropharyngeal squamous cell carcinoma (OCSCC and OPSCC) exist an association between clinical and histopathological parameters with cell proliferation, basal lamina, connective tissue degradation and surrounding stroma markers. We evaluated these associations in Chilean patients. A convenience sample of 37 cases of OCSCC (n=16) and OPSCC (n=21) was analyzed clinically (TNM, clinical stage) and histologically (WHO grade of differentiation, pattern of tumor invasion). We assessed the expression of p53, Ki67, HOXA1, HOXB7, type IV collagen (ColIV) and carcinoma-associated fibroblast (α-SMA-positive cells). Additionally we conducted a univariate/bivariate analysis to assess the relationship of these variables with survival rates. Males were mostly affected (56.2% OCSCC, 76.2% OPSCC). Patients were mainly diagnosed at III/IV clinical stages (68.8% OCSCC, 90.5% OPSCC) with a predominantly infiltrative pattern invasion (62.9% OCSCC, 57.1% OPSCC). Significant association between regional lymph nodes (N) and clinical stage with OCSCC-HOXB7 expression (Chi-Square test P < 0.05) was observed. In OPSCC a statistically significant association exists between p53, Ki67 with gender (Chi-Square test P < 0.05). In OCSCC and OPSCC was statistically significant association between ki67 with HOXA1, HOXB7, and between these last two antigens (Pearson's Correlation test P < 0.05). Furthermore OPSCC-p53 showed significant correlation when it was compared with α-SMA (Kendall's Tau-c test P < 0.05). Only OCSCC-pattern invasion and OPSCC-primary tumor (T) pattern resulted associated with survival at the end of the follow up period (Chi-Square Likelihood Ratio, P < 0.05). Clinical, histological and immunohistochemical features are similar to seen in other countries. Cancer proliferation markers were associated strongly from each other. Our sample highlights prognostic value of T and pattern of invasion, but the conclusions may be limited and should be considered with caution (small sample). Many cases were diagnosed in the advanced stages of the disease, which suggests that the diagnosis of OCSCC and OPSCC is made late.
Resumo:
This study sought to identify factors involved in access to the services of a basic health unit. It is a cross-sectional, population-based study involving 101 randomly-selected families residing in the area covered by the health unit. An adult resident of each household was interviewed. The response variable was whether or not the resident frequented the health unit if he/she or anyone in the family required assistance to resolve a health issue. The independent variables investigated were service provision aspects, demographic and socio-economic characteristics, individual habits, morbidities and use of the health unit. In addition to descriptive and univariate analysis, logistic regression was applied in the multivariate analysis. The results show that access to the basic health unit is associated with the treatment received previously (OR = 3,224) with accessibility (OR = 0,146) and micro-area of residence (OR = 10,918). These findings suggest that access is related to the impressions created by the care received at the health unit and is based on experiences with the service, but can also be strongly modulated by individual aspects and factors related to the territory.
Resumo:
To evaluate the prevalence and associated risk factors for urinary incontinence, as well as its association with multimorbidity among Brazilian women aged 50 or over. This was a secondary analysis of a cross-sectional population-based study including 622 women 50 years or older, conducted in the city of Campinas-SP-Brazil. The dependent variable was Urinary Incontinence (UI), defined as any complaint of urine loss. The independent variables were sociodemographic data, health-related habits, self-perception of health and functional capacity evaluation. Statistical analysis was carried out using the Chi-square test and Poisson regression. The mean age of the women was 64. UI was prevalent in 52.3% of these women: Mixed UI (26.6%), Urge UI (13.2%) and Stress UI (12.4%). Factors associated with a higher prevalence of UI were hypertension (OR 1.21, CI 1:01-1:47, P = 0.004), osteoarthritis (OR 1.24, CI 1:03-1:50, P = 0.022), physical activity ≥3 days/week (OR 1.21, CI 1:01-1:44, P = 0.039), BMI ≥ 25 at the time of the interview (OR 1.25, CI 1:04-1:49, P = 0.018), negative self-perception of health (OR 1.23, CI 1:06-1:44 P = 0.007) and limitations in daily living activities (PR 1:56 CI 1:16-2:10, P = 0.004). The prevalence of UI was high. Mixed incontinence was the most frequent type of UI. Many associated factors can be prevented or improved. Thus, health policies targeted at these combined factors could reduce their prevalence rate and possibly decrease the prevalence of UI. Neurourol. Urodynam. © 2014 Wiley Periodicals, Inc.
Resumo:
To determine the prevalence of the Papanicolaou exam among women aged 20 to 59 years in the city of Campinas (state of São Paulo, Brazil) and to analyze associations between this test and affiliation to private health insurance plans as well as socioeconomic/demographic variables and health-related behavior. To do so, a population-based, cross-sectional study was carried out. Statistical analyses took the study design into account. Despite the significant socioeconomic differences between women with and without private health plans, no differences between these groups were found regarding having been submitted to the Papanicolaou test. In fact no differences were found as to socioeconomic and health variables analyzed. Among all variables analyzed, only marital status was significantly associated with having undergone the test. The Brazilian public health system accounted for 55.7% of the exams. The present findings indicate social equity in the city of Campinas regarding the preventive exam for cervical cancer in the age group studied.