Biblioteca Digital

2 resultados para Web as a Corpus

em Bulgarian Digital Mathematics Library at IMI-BAS

Automatic Identification of False Friends in Parallel Corpora: Statistical and Semantic Approach

Relevância:

60.00% 60.00%

Publicador:

Resumo:

False friends are pairs of words in two languages that are perceived as similar but have different meanings. We present an improved algorithm for acquiring false friends from sentence-level aligned parallel corpus based on statistical observations of words occurrences and co-occurrences in the parallel sentences. The results are compared with an entirely semantic measure for cross-lingual similarity between words based on using the Web as a corpus through analyzing the words’ local contexts extracted from the text snippets returned by searching in Google. The statistical and semantic measures are further combined into an improved algorithm for identification of false friends that achieves almost twice better results than previously known algorithms. The evaluation is performed for identifying cognates between Bulgarian and Russian but the proposed methods could be adopted for other language pairs for which parallel corpora and bilingual glossaries are available.

Veja mais

Web-application for Presentation of Bulgarian Language Heritage: Bilingual Digital Corpora and Dictionaries

Relevância:

30.00% 30.00%

Publicador:

Resumo:

The paper describes three software packages - the main components of a software system for processing and web-presentation of Bulgarian language resources – parallel corpora and bilingual dictionaries. The author briefly presents current versions of the core components “Dictionary” and “Corpus” as well as the recently developed component “Connection” that links both “Dictionary” and “Corpus”. The components main functionalities are described as well. Some examples of the usage of the system’s web-applications are included.

Veja mais

2 resultados para Web as a Corpus

em Bulgarian Digital Mathematics Library at IMI-BAS

Filtro por publicador

Automatic Identification of False Friends in Parallel Corpora: Statistical and Semantic Approach

Web-application for Presentation of Bulgarian Language Heritage: Bilingual Digital Corpora and Dictionaries