Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Automated clinical coding is a critical component of healthcare information systems, yet it remains challenging due to the complexity of medical terminology and the hierarchical structure of classification systems such as ICD-10. Recent advances in Large Language Models (LLMs) have enabled the generation of ranked lists of candidate codes; however, individual models often exhibit variability and limited reliability when used in isolation. In this study, we investigate whether classical rank aggregation methods from Multi-Criteria Decision Analysis (MCDA) can improve decision quality by combining the ranked outputs of multiple LLMs for ICD-10 code assignment. We consider a decision framework in which each LLM is treated as an independent decision agent providing a ranked set of candidate codes, and evaluate three aggregation methods—Plurality Voting, Borda Count, and Reciprocal Rank Fusion (RRF)—on a dataset of 1,117 obstetric clinical notes in Brazilian Portuguese. Performance is assessed at both the three-character and full-code levels, focusing on top-k recommendations representative of practical decision-support scenarios, and robustness is evaluated using bootstrap resampling with 95\% confidence intervals. The results show that rank-based aggregation methods consistently improve decision performance compared to individual models, with methods that incorporate ranking information, particularly Borda Count and RRF, achieving higher Micro-F1 scores and demonstrating more stable behavior across evaluation settings. In contrast, simple voting schemes show inferior performance, highlighting the importance of leveraging ranking structure, while weighted aggregation strategies do not yield statistically significant improvements over unweighted methods. These findings demonstrate that rank aggregation provides a principled and effective approach for combining multiple imperfect decision agents in clinical coding tasks, and suggest that simple and robust aggregation strategies can improve decision consistency and reliability in real-world healthcare applications.</p>

Show More

Keywords

decision aggregation methods clinical models

Related Articles

PORE

About

Connect