Towards improving WEBSOM with multi-word expressions


Autoria(s): Alves, Stefan Eduard Raposo
Contribuinte(s)

Marques, Nuno Cavalheiro

Silva, Joaquim

Data(s)

23/07/2013

23/07/2013

2013

Resumo

Dissertação para obtenção do Grau de Mestre em Engenharia Informática

Large quantities of free-text documents are usually rich in information and covers several topics. However, since their dimension is very large, searching and filtering data is an exhaustive task. A large text collection covers a set of topics where each topic is affiliated to a group of documents. This thesis presents a method for building a document map about the core contents covered in the collection. WEBSOM is an approach that combines document encoding methods and Self-Organising Maps (SOM) to generate a document map. However, this methodology has a weakness in the document encoding method because it uses single words to characterise documents. Single words tend to be ambiguous and semantically vague, so some documents can be incorrectly related. This thesis proposes a new document encoding method to improve the WEBSOM approach by using multi word expressions (MWEs) to describe documents. Previous research and ongoing experiments encourage us to use MWEs to characterise documents because these are semantically more accurate than single words and more descriptive.

Identificador

http://hdl.handle.net/10362/10169

Idioma(s)

eng

Publicador

Faculdade de Ciências e Tecnologia

Direitos

openAccess

Palavras-Chave #Self-Organising Maps (SOM) #Text mining #WEBSOM #Relevant expressions
Tipo

masterThesis