Evidence map›Paper›PMID 42662633›Full record

SynthesisFrontiers in digital health2026

Artificial intelligence for automated ICD-10 coding: a systematic review of multi-label text classification in clinical narratives.

Kamonrat Tangudomkit, Sawrawit Chairat, Sitthichok Chaichulee

Abstract readSystematic Review
In one paragraph

Synthesis in Frontiers in digital health, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

3 authors.

Kamonrat TangudomkitDepartment of Biomedical Sciences and Biomedical Engineering, Faculty of Medicine, Prince of Songkla University, Songkhla, Thailand.
Sawrawit ChairatDepartment of Biomedical Sciences and Biomedical Engineering, Faculty of Medicine, Prince of Songkla University, Songkhla, Thailand.
Sitthichok ChaichuleeDepartment of Biomedical Sciences and Biomedical Engineering, Faculty of Medicine, Prince of Songkla University, Songkhla, Thailand.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: ICD-10 coding is an essential process in healthcare systems that supports clinical management, reimbursement, and health data analytics. However, the complexity of its hierarchical structure and the large number of available codes make manual coding limited in terms of time, cost, and consistency. Despite growing research in this area, evidence remains fragmented, particularly regarding real-world implementation readiness. Objective: To review and synthesize existing knowledge on algorithms, datasets, evaluation methods, and real-world implementation readiness of automatic ICD-10 coding systems. Methods: Eligible studies were original research articles, preprints, or conference papers published in English between January 1, 2020 and December 31, 2025, and retrieved from seven academic databases: Scopus, PubMed, Web of Science, IEEE Xplore, ACM Digital Library, arXiv, and Google Scholar. Studies were included if they investigated automatic ICD-10 coding from clinical text using machine learning, deep learning, transformer-based, or large language model (LLM) approaches. Methodological quality was assessed using a research-question-driven appraisal framework. This systematic review followed PRISMA 2020 guidance and was preregistered in the Open Science Framework (OSF) at https://osf.io/cegqk. Results: A total of 257 records were identified, of which 24 studies met the inclusion criteria and contributed 296 experimental evaluations overall. Study quality was high in 7 studies, moderate in 8, and limited by technical or methodological concerns in 9. Hybrid deep learning (Hybrid DL) was most often used as the main automated coding approach, while machine learning (ML) and rule-based approaches were mainly used as baselines. F1-macro was consistently lower than F1-micro among studies reporting both metrics. Hybrid DL showed the most stable performance under all-code or full-code evaluation, while AI model performance varied by the documents-per-label (D/L) ratio. Discussion: The evidence indicates continued technical progress, particularly through Hybrid DL and transformer-based approaches, while LLM-based methods remain emerging and less consistently effective for structured multi-label coding. The observed D/L-performance relationship suggested that AI model selection should consider dataset structure and label support, in addition to algorithmic complexity. Conclusion: AI-based automatic ICD-10 coding is a promising approach for clinical coding support. Future research should prioritize rare-label imbalance, reproducibility, explainability, and validation across diverse clinical settings. Systematic Review Registration: https://osf.io/cegqk.

Indexed as

automatic ICD-10 codingclass imbalanceclinical natural language processingdeep learningexplainabilitylarge language modelsmulti-label text classificationsystematic review

Identifiers

PMID42662633
PMCPMC13520067

What OpenQuestion holds

Textmetadata
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.