Natural Language Processing in Breast Imaging Reports: A Scoping Review with Implications for Low-Resource Clinical Languages

relationships.isProjectOf

relationships.isJournalIssueOf

Abstract

Breast imaging reports are commonly recorded as unstructured free-text documents, which limits their secondary use for large-scale clinical analysis, structured information extraction, and clinical decision support. These challenges are particularly important in morphologically rich and low-resource clinical languages, where linguistic variability, inconsistent terminology, and limited annotated corpora may reduce the direct applicability of existing natural language processing (NLP) approaches. This study presents a scoping review of NLP research on breast imaging reports, conducted and reported in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses Extension for Scoping Reviews (PRISMA-ScR). Attention is given to BI-RADS (Breast Imaging Reporting and Data System)-related analytical tasks, methodological trends, dataset characteristics, evaluation practices, and implications for low-resource clinical NLP settings, including Turkish. A comprehensive literature search was conducted across Web of Science, Scopus, PubMed, IEEE Xplore, and Google Scholar in February 2026. To reduce the risk of missing recent studies using transformer- and large language model (LLM)-related terminology, a targeted supplementary search was also conducted. Following screening and eligibility assessment, 39 studies were included in the final synthesis. The findings show that the literature is concentrated mainly on BI-RADS classification/annotation and information extraction tasks. Task-wise, BI-RADS classification/annotation was the most frequent category, followed by information extraction. Methodologically, the reviewed literature shows a shift from rule-based and traditional machine learning approaches toward transformer- and LLM-based methods. LLM-based studies were frequently represented, particularly among recent studies and those identified through the targeted supplementary search; therefore, their observed prominence should be interpreted cautiously. Despite these advances, the literature remains linguistically imbalanced and methodologically heterogeneous. English was the most frequently represented report or dataset language, whereas Turkish breast imaging NLP studies remained limited. Major challenges identified across the reviewed studies include dataset heterogeneity, inconsistent annotation practices, variable evaluation metrics, limited external validation, and incomplete reporting of reproducibility-related details. This scoping review provides a structured synthesis of methodological trends, task categories, dataset characteristics, evaluation practices, and reproducibility-related limitations in breast imaging NLP. Overall, the findings highlight the need for better documented datasets, standardised evaluation practices, transparent reporting, clinically grounded validation, and stronger research efforts for low-resource clinical language settings.

Description

Institutional Author Profiles

Keywords

BI-RADS, Low-Resource Clinical Languages, Natural Language Processing, Clinical NLP, Breast Imaging Reports, Computer Science, Breast Imaging, Artificial Intelligence, English Language

Fields of Science

Citation

WoS Q

Scopus Q

Volume

16

Issue

12

Start Page

5847

End Page

5847
PlumX Metrics
Captures

Mendeley Readers : 1

Google Scholar Logo
Google Scholar™
OpenAlex Logo
OpenAlex FWCI
4.18