A methodology to compare dimensionality reduction algorithms in terms of loss of quality


Autoria(s): Gracia Berná, Antonio; González Tortosa, Santiago; Robles Forcada, Víctor; Menasalvas Ruiz, Ernestina
Data(s)

01/06/2014

31/12/1969

Resumo

Dimensionality Reduction (DR) is attracting more attention these days as a result of the increasing need to handle huge amounts of data effectively. DR methods allow the number of initial features to be reduced considerably until a set of them is found that allows the original properties of the data to be kept. However, their use entails an inherent loss of quality that is likely to affect the understanding of the data, in terms of data analysis. This loss of quality could be determinant when selecting a DR method, because of the nature of each method. In this paper, we propose a methodology that allows different DR methods to be analyzed and compared as regards the loss of quality produced by them. This methodology makes use of the concept of preservation of geometry (quality assessment criteria) to assess the loss of quality. Experiments have been carried out by using the most well-known DR algorithms and quality assessment criteria, based on the literature. These experiments have been applied on 12 real-world datasets. Results obtained so far show that it is possible to establish a method to select the most appropriate DR method, in terms of minimum loss of quality. Experiments have also highlighted some interesting relationships between the quality assessment criteria. Finally, the methodology allows the appropriate choice of dimensionality for reducing data to be established, whilst giving rise to a minimum loss of quality.

Formato

application/pdf

Identificador

http://oa.upm.es/25892/

Idioma(s)

eng

Relação

http://oa.upm.es/25892/1/INVE_MEM_2014_161554.pdf

http://www.sciencedirect.com/science/article/pii/S0020025514001741

info:eu-repo/semantics/altIdentifier/doi/10.1016/j.ins.2014.02.068

Direitos

http://creativecommons.org/licenses/by-nc-nd/3.0/es/

info:eu-repo/semantics/openAccess

Fonte

Information Sciences, ISSN 0020-0255, 2014-06, Vol. 270

Palavras-Chave #Informática
Tipo

info:eu-repo/semantics/article

Artículo

PeerReviewed