988 resultados para Annotated corpora


Relevância:

100.00% 100.00%

Publicador:

Resumo:

This paper presents the automatic extension to other languages of TERSEO, a knowledge-based system for the recognition and normalization of temporal expressions originally developed for Spanish. TERSEO was first extended to English through the automatic translation of the temporal expressions. Then, an improved porting process was applied to Italian, where the automatic translation of the temporal expressions from English and from Spanish was combined with the extraction of new expressions from an Italian annotated corpus. Experimental results demonstrate how, while still adhering to the rule-based paradigm, the development of automatic rule translation procedures allowed us to minimize the effort required for porting to new languages. Relying on such procedures, and without any manual effort or previous knowledge of the target language, TERSEO recognizes and normalizes temporal expressions in Italian with good results (72% precision and 83% recall for recognition).

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Ce travail porte sur la construction d’un corpus étalon pour l’évaluation automatisée des extracteurs de termes. Ces programmes informatiques, conçus pour extraire automatiquement les termes contenus dans un corpus, sont utilisés dans différentes applications, telles que la terminographie, la traduction, la recherche d’information, l’indexation, etc. Ainsi, leur évaluation doit être faite en fonction d’une application précise. Une façon d’évaluer les extracteurs consiste à annoter toutes les occurrences des termes dans un corpus, ce qui nécessite un protocole de repérage et de découpage des unités terminologiques. À notre connaissance, il n’existe pas de corpus annoté bien documenté pour l’évaluation des extracteurs. Ce travail vise à construire un tel corpus et à décrire les problèmes qui doivent être abordés pour y parvenir. Le corpus étalon que nous proposons est un corpus entièrement annoté, construit en fonction d’une application précise, à savoir la compilation d’un dictionnaire spécialisé de la mécanique automobile. Ce corpus rend compte de la variété des réalisations des termes en contexte. Les termes sont sélectionnés en fonction de critères précis liés à l’application, ainsi qu’à certaines propriétés formelles, linguistiques et conceptuelles des termes et des variantes terminologiques. Pour évaluer un extracteur au moyen de ce corpus, il suffit d’extraire toutes les unités terminologiques du corpus et de comparer, au moyen de métriques, cette liste à la sortie de l’extracteur. On peut aussi créer une liste de référence sur mesure en extrayant des sous-ensembles de termes en fonction de différents critères. Ce travail permet une évaluation automatique des extracteurs qui tient compte du rôle de l’application. Cette évaluation étant reproductible, elle peut servir non seulement à mesurer la qualité d’un extracteur, mais à comparer différents extracteurs et à améliorer les techniques d’extraction.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Sign language animations can lead to better accessibility of information and services for people who are deaf and have low literacy skills in spoken/written languages. Due to the distinct word-order, syntax, and lexicon of the sign language from the spoken/written language, many deaf people find it difficult to comprehend the text on a computer screen or captions on a television. Animated characters performing sign language in a comprehensible way could make this information accessible. Facial expressions and other non-manual components play an important role in the naturalness and understandability of these animations. Their coordination to the manual signs is crucial for the interpretation of the signed message. Software to advance the support of facial expressions in generation of sign language animation could make this technology more acceptable for deaf people. In this survey, we discuss the challenges in facial expression synthesis and we compare and critique the state of the art projects on generating facial expressions in sign language animations. Beginning with an overview of facial expressions linguistics, sign language animation technologies, and some background on animating facial expressions, a discussion of the search strategy and criteria used to select the five projects that are the primary focus of this survey follows. This survey continues on to introduce the work from the five projects under consideration. Their contributions are compared in terms of support for specific sign language, categories of facial expressions investigated, focus range in the animation generation, use of annotated corpora, input data or hypothesis for their approach, and other factors. Strengths and drawbacks of individual projects are identified in the perspectives above. This survey concludes with our current research focus in this area and future prospects.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

The exponential growth of the subjective information in the framework of the Web 2.0 has led to the need to create Natural Language Processing tools able to analyse and process such data for multiple practical applications. They require training on specifically annotated corpora, whose level of detail must be fine enough to capture the phenomena involved. This paper presents EmotiBlog – a fine-grained annotation scheme for subjectivity. We show the manner in which it is built and demonstrate the benefits it brings to the systems using it for training, through the experiments we carried out on opinion mining and emotion detection. We employ corpora of different textual genres –a set of annotated reported speech extracted from news articles, the set of news titles annotated with polarity and emotion from the SemEval 2007 (Task 14) and ISEAR, a corpus of real-life self-expressed emotion. We also show how the model built from the EmotiBlog annotations can be enhanced with external resources. The results demonstrate that EmotiBlog, through its structure and annotation paradigm, offers high quality training data for systems dealing both with opinion mining, as well as emotion detection.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

In this paper we present an automatic system for the extraction of syntactic semantic patterns applied to the development of multilingual processing tools. In order to achieve optimum methods for the automatic treatment of more than one language, we propose the use of syntactic semantic patterns. These patterns are formed by a verbal head and the main arguments, and they are aligned among languages. In this paper we present an automatic system for the extraction and alignment of syntactic semantic patterns from two manually annotated corpora, and evaluate the main linguistic problems that we must deal with in the alignment process.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

* The following text has been originally published in the Proceedings of the Language Recourses and Evaluation Conference held in Lisbon, Portugal, 2004, under the title of "Towards Intelligent Written Cultural Heritage Processing - Lexical processing". I present here a revised contribution of the aforementioned paper and I add here the latest efforts done in the Center for Computational Linguistic in Prague in the field under discussion.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The construction and use of multimedia corpora has been advocated for a while in the literature as one of the expected future application fields of Corpus Linguistics. This research project represents a pioneering experience aimed at applying a data-driven methodology to the study of the field of AVT, similarly to what has been done in the last few decades in the macro-field of Translation Studies. This research was based on the experience of Forlixt 1, the Forlì Corpus of Screen Translation, developed at the University of Bologna’s Department of Interdisciplinary Studies in Translation, Languages and Culture. As a matter of fact, in order to quantify strategies of linguistic transfer of an AV product, we need to take into consideration not only the linguistic aspect of such a product but all the meaning-making resources deployed in the filmic text. Provided that one major benefit of Forlixt 1 is the combination of audiovisual and textual data, this corpus allows the user to access primary data for scientific investigation, and thus no longer rely on pre-processed material such as traditional annotated transcriptions. Based on this rationale, the first chapter of the thesis sets out to illustrate the state of the art of research in the disciplinary fields involved. The primary objective was to underline the main repercussions on multimedia texts resulting from the interaction of a double support, audio and video, and, accordingly, on procedures, means, and methods adopted in their translation. By drawing on previous research in semiotics and film studies, the relevant codes at work in visual and acoustic channels were outlined. Subsequently, we concentrated on the analysis of the verbal component and on the peculiar characteristics of filmic orality as opposed to spontaneous dialogic production. In the second part, an overview of the main AVT modalities was presented (dubbing, voice-over, interlinguistic and intra-linguistic subtitling, audio-description, etc.) in order to define the different technologies, processes and professional qualifications that this umbrella term presently includes. The second chapter focuses diachronically on various theories’ contribution to the application of Corpus Linguistics’ methods and tools to the field of Translation Studies (i.e. Descriptive Translation Studies, Polysystem Theory). In particular, we discussed how the use of corpora can favourably help reduce the gap existing between qualitative and quantitative approaches. Subsequently, we reviewed the tools traditionally employed by Corpus Linguistics in regard to the construction of traditional “written language” corpora, to assess whether and how they can be adapted to meet the needs of multimedia corpora. In particular, we reviewed existing speech and spoken corpora, as well as multimedia corpora specifically designed to investigate Translation. The third chapter reviews Forlixt 1's main developing steps, from a technical (IT design principles, data query functions) and methodological point of view, by laying down extensive scientific foundations for the annotation methods adopted, which presently encompass categories of pragmatic, sociolinguistic, linguacultural and semiotic nature. Finally, we described the main query tools (free search, guided search, advanced search and combined search) and the main intended uses of the database in a pedagogical perspective. The fourth chapter lists specific compilation criteria retained, as well as statistics of the two sub-corpora, by presenting data broken down by language pair (French-Italian and German-Italian) and genre (cinema’s comedies, television’s soapoperas and crime series). Next, we concentrated on the discussion of the results obtained from the analysis of summary tables reporting the frequency of categories applied to the French-Italian sub-corpus. The detailed observation of the distribution of categories identified in the original and dubbed corpus allowed us to empirically confirm some of the theories put forward in the literature and notably concerning the nature of the filmic text, the dubbing process and Italian dubbed language’s features. This was possible by looking into some of the most problematic aspects, like the rendering of socio-linguistic variation. The corpus equally allowed us to consider so far neglected aspects, such as pragmatic, prosodic, kinetic, facial, and semiotic elements, and their combination. At the end of this first exploration, some specific observations concerning possible macrotranslation trends were made for each type of sub-genre considered (cinematic and TV genre). On the grounds of this first quantitative investigation, the fifth chapter intended to further examine data, by applying ad hoc models of analysis. Given the virtually infinite number of combinations of categories adopted, and of the latter with searchable textual units, three possible qualitative and quantitative methods were designed, each of which was to concentrate on a particular translation dimension of the filmic text. The first one was the cultural dimension, which specifically focused on the rendering of selected cultural references and on the investigation of recurrent translation choices and strategies justified on the basis of the occurrence of specific clusters of categories. The second analysis was conducted on the linguistic dimension by exploring the occurrence of phrasal verbs in the Italian dubbed corpus and by ascertaining the influence on the adoption of related translation strategies of possible semiotic traits, such as gestures and facial expressions. Finally, the main aim of the third study was to verify whether, under which circumstances, and through which modality, graphic and iconic elements were translated into Italian from an original corpus of both German and French films. After having reviewed the main translation techniques at work, an exhaustive account of possible causes for their non-translation was equally provided. By way of conclusion, the discussion of results obtained from the distribution of annotation categories on the French-Italian corpus, as well as the application of specific models of analysis allowed us to underline possible advantages and drawbacks related to the adoption of a corpus-based approach to AVT studies. Even though possible updating and improvement were proposed in order to help solve some of the problems identified, it is argued that the added value of Forlixt 1 lies ultimately in having created a valuable instrument, allowing to carry out empirically-sound contrastive studies that may be usefully replicated on different language pairs and several types of multimedia texts. Furthermore, multimedia corpora can also play a crucial role in L2 and translation teaching, two disciplines in which their use still lacks systematic investigation.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

We present here a checklist of recent marine bryozoans recorded in the literature from Brazil. The total number of species recorded is 346. The most diverse group is the order Cheilostomata with 271 species, followed by the order Ctenostomata, with 42 species, and the order Cyclostomata, with 33 species. Included in the checklist are records by state and citations for species with synonyms utilized in Brazilian works.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

Introduction. Erectile dysfunction (ED), as well as cardiovascular diseases (CVDs), is associated with endothelial dysfunction and increased levels of proinflammatory cytokines, such as tumor necrosis factor-alpha (TNF-alpha). Aim. We hypothesized that increased TNF-alpha levels impair cavernosal function. Methods. In vitro organ bath studies were used to measure cavernosal reactivity in mice infused with vehicle or TNF-alpha-(220 ng/kg/min) for 14 days. Gene expression of nitric oxide synthase isoforms was evaluated by real-time polymerase chain reaction. Results. Cavernosal strips from the TNF-alpha-infused mice displayed decreased nonadrenergic-noncholinergic (NANC)-induced relaxation (59.4 +/- 6.2 vs. control: 76.2 +/- 4.7; 16 Hz) compared with the control animals. These responses were associated with decreased gene expression of eNOS and nNOS (P < 0.05). Sympathetic-mediated, as well as phenylephrine (PE)-induced, contractile responses (PE-induced contraction; 1.32 +/- 0.06 vs. control: 0.9 +/- 0.09, mN) were increased in cavernosal strips from TNF-alpha-infused mice. Additionally, infusion of TNF-alpha increased cavernosal responses to endothelin-1 and endothelin receptor A subtype (ET(A)) receptor expression (P < 0.05) and slightly decreased tumor necrosis factor-alpha receptor 1 (TNFRI) expression (P=0.063). Conclusion. Corpora cavernosa from TNF-alpha-infused mice display increased contractile responses and decreased NANC nerve-mediated relaxation associated with decreased eNOS and nNOS gene expression. There changes may trigger ED and indicate that TNF-alpha plays a detrimental role in erectile function. Blockade of TNF-alpha actions may represent an alternative therapeutic approach for ED, especially in pathologic conditions associated with increased levels of this cytokine. Carneiro FS, Zemse S, Giachini FRC, Carneiro ZN, Lima W, Clinton Webb R, and Tostes RC. TNF-alpha infusion impairs corpora cavernosa reactivity. J Sex Med 2009;6(suppl 3):311-319.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

Erectile dysfunction is considered an early clinical manifestation of vascular disease and an independent risk factor for cardiovascular events associated with endothelial dysfunction and increased levels of pro-inflammatory cytokines. Tumor necrosis factor-alpha (TNF-alpha), a pro-inflammatory cytokine, suppresses endothelial nitric oxide synthase (eNOS) expression. Considering that nitric oxide (NO) is of critical importance in penile erection, we hypothesized that blockade of TNF-alpha actions would increase cavernosal smooth muscle relaxation. In vitro organ bath studies were used to measure cavernosal reactivity in wild type and TNF-alpha knockout (TNF-alpha KO) mice and NOS expression was evaluated by western blot. In addition, spontaneous erections (in vivo) were evaluated by videomonitoring the animals (30 minutes). Collagen and elastin expression were evaluated by Masson trichrome and Verhoff-van Gieson stain reaction, respectively. Corpora cavernosa from TNF-alpha KO mice exhibited increased NO-dependent relaxation, which was associated with increased eNOS and neuronal NOS (nNOS) cavernosal expression. Cavernosal strips from TNF-alpha KO mice displayed increased endothelium-dependent (97.4 +/- 5.3 vs. Control: 76.3 +/- 6.3, %) and nonadrenergic-noncholinergic (93.3 +/- 3.0 vs. Control: 67.5 +/- 16.0; 16 Hz) relaxation compared to control animals. These responses were associated with increased protein expression of eNOS and nNOS (P < 0.05). Sympathetic-mediated (0.69 +/- 0.16 vs. Control: 1.22 +/- 0.22; 16 Hz) as well as phenylephrine-induced contractile responses (1.6 +/- 0.1 vs. Control: 2.5 +/- 0.1, mN) were attenuated in cavernosal strips from TNF-alpha KO mice. Additionally, corpora cavernosa from TNF-alpha KO mice displayed increased collagen and elastin expression. In vivo experiments demonstrated that TNF-alpha KO mice display increased number of spontaneous erections. Corpora cavernosa from TNF-alpha KO mice display alterations that favor penile tumescence, indicating that TNF-alpha plays a detrimental role in erectile function. A key role for TNF-alpha in mediating endothelial dysfunction in ED is markedly relevant since we now have access to anti-TNF-alpha therapies. Carneiro FS, Sturgis LC, Giachini FRC, Carneiro ZN, Lima VV, Wynne BM, Martin SS, Brands MW, Tostes RC, and Webb RC. TNF-alpha knockout mice have increased corpora cavernosa relaxation. J Sex Med 2009;6:115-125.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

Polissema: Revista de Letras do ISCAP 2002/N.º 2 Linguagens

Relevância:

20.00% 20.00%

Publicador:

Resumo:

Este artigo apresenta uma pesquisa sobre a representação do discurso ficcional embasado na gramática sistêmico - funcional proposta por Halliday e na Lingüística de Corpus, utilizando-se o software WordSmith Tools. A análise focaliza a metafunção ideacional, realizada pelo sistema de transitividade, focalizando os processos mentais e a relação lógico - semântica da projeção. O objetivo da pesquisa foi observar como os pensamentos das personagens de um corpus ficcional são representados através dos verbos de elocução THINK e PENSAR, buscando descrever padrões textuais nos três romances que compõem o corpus.