Automatic ICD-10 classification of cancers from free-text death certificates


Autoria(s): Koopman, Bevan; Zuccon, Guido; Nguyen, Anthony; Bergheim, Anton; Grayson, Narelle
Data(s)

01/11/2015

Resumo

Objective Death certificates provide an invaluable source for cancer mortality statistics; however, this value can only be realised if accurate, quantitative data can be extracted from certificates – an aim hampered by both the volume and variable nature of certificates written in natural language. This paper proposes an automatic classification system for identifying cancer related causes of death from death certificates. Methods Detailed features, including terms, n-grams and SNOMED CT concepts were extracted from a collection of 447,336 death certificates. These features were used to train Support Vector Machine classifiers (one classifier for each cancer type). The classifiers were deployed in a cascaded architecture: the first level identified the presence of cancer (i.e., binary cancer/nocancer) and the second level identified the type of cancer (according to the ICD-10 classification system). A held-out test set was used to evaluate the effectiveness of the classifiers according to precision, recall and F-measure. In addition, detailed feature analysis was performed to reveal the characteristics of a successful cancer classification model. Results The system was highly effective at identifying cancer as the underlying cause of death (F-measure 0.94). The system was also effective at determining the type of cancer for common cancers (F-measure 0.7). Rare cancers, for which there was little training data, were difficult to classify accurately (F-measure 0.12). Factors influencing performance were the amount of training data and certain ambiguous cancers (e.g., those in the stomach region). The feature analysis revealed a combination of features were important for cancer type classification, with SNOMED CT concept and oncology specific morphology features proving the most valuable. Conclusion The system proposed in this study provides automatic identification and characterisation of cancers from large collections of free-text death certificates. This allows organisations such as Cancer Registries to monitor and report on cancer mortality in a timely and accurate manner. In addition, the methods and findings are generally applicable beyond cancer classification and to other sources of medical text besides death certificates.

Formato

application/pdf

Identificador

http://eprints.qut.edu.au/91420/

Publicador

Elsevier Ireland Ltd.

Relação

http://eprints.qut.edu.au/91420/1/jmi2014_cancer_cascade_classification.pdf

DOI:10.1016/j.ijmedinf.2015.08.004

Koopman, Bevan, Zuccon, Guido, Nguyen, Anthony, Bergheim, Anton, & Grayson, Narelle (2015) Automatic ICD-10 classification of cancers from free-text death certificates. International Journal of Medical Informatics, 84(11), pp. 956-965.

Direitos

Copyright 2015 Elsevier Ireland Ltd.

This manuscript version is made available under the CC-BY-NC-ND 4.0 license http://creativecommons.org/licenses/by-nc-nd/4.0/

Fonte

Faculty of Science and Technology; School of Information Systems; Science & Engineering Faculty

Palavras-Chave #Cancer classification; Death certificates; Machine learning; Natural language processing
Tipo

Journal Article