Multilingual RECIST classification of radiology reports using supervised learning.

Mottin, L.; Goldman, J.P.; Jäggli, C.; Achermann, R.; Gobeill, J.; Knafou, J.; Ehrsam, J.; Wicky, A.; Gérard, C.L.; Schwenk, T.; Charrier, M.; Tsantoulis, P.; Lovis, C.; Leichtle, A.; Kiessling, M.K.; Michielin, O.; Pradervand, S.; Foufi, V.; Ruch, P.

doi:10.3389/fdgth.2023.1195017

Multilingual RECIST classification of radiology reports using supervised learning.

Details

Download: 37388252_BIB_E6AA8FD72F0E.pdf (6090.13 [Ko])
State: Public
Version: Final published version
License: CC BY 4.0

Serval ID

serval:BIB_E6AA8FD72F0E

Type

Article: article from journal or magazin.

Collection

Publications

Institution

UNIL/CHUV

Title

Multilingual RECIST classification of radiology reports using supervised learning.

Journal

Frontiers in digital health

Author(s)

Mottin L., Goldman J.P., Jäggli C., Achermann R., Gobeill J., Knafou J., Ehrsam J., Wicky A., Gérard C.L., Schwenk T., Charrier M., Tsantoulis P., Lovis C., Leichtle A., Kiessling M.K., Michielin O., Pradervand S., Foufi V., Ruch P.

ISSN

2673-253X (Electronic)

ISSN-L

2673-253X

Publication state

Published

Issued date

2023

Peer-reviewed

Oui

Volume

Pages

1195017

Language

english

Notes

Publication types: Journal Article
Publication Status: epublish

Abstract

The objective of this study is the exploration of Artificial Intelligence and Natural Language Processing techniques to support the automatic assignment of the four Response Evaluation Criteria in Solid Tumors (RECIST) scales based on radiology reports. We also aim at evaluating how languages and institutional specificities of Swiss teaching hospitals are likely to affect the quality of the classification in French and German languages.
In our approach, 7 machine learning methods were evaluated to establish a strong baseline. Then, robust models were built, fine-tuned according to the language (French and German), and compared with the expert annotation.
The best strategies yield average F1-scores of 90% and 86% respectively for the 2-classes (Progressive/Non-progressive) and the 4-classes (Progressive Disease, Stable Disease, Partial Response, Complete Response) RECIST classification tasks.
These results are competitive with the manual labeling as measured by Matthew's correlation coefficient and Cohen's Kappa (79% and 76%). On this basis, we confirm the capacity of specific models to generalize on new unseen data and we assess the impact of using Pre-trained Language Models (PLMs) on the accuracy of the classifiers.

Keywords

RECIST, language models, narrative text classification, radiology reports, supervised machine learning

URN

urn:nbn:ch:serval-BIB_E6AA8FD72F0E4

OAI-PMH

oai:serval.unil.ch:BIB_E6AA8FD72F0E

DOI

10.3389/fdgth.2023.1195017

Pubmed

37388252

Web of science

001033312400001