Evidence map›Paper›PMID 41595962›Full record

ArticleBioengineering (Basel, Switzerland)2025

Accurate Clinical Entity Recognition and Code Mapping of Anatomopathological Reports Using BioClinicalBERT Enhanced by Retrieval-Augmented Generation: A Hybrid Deep Learning Approach.

Hamida Abdaoui, Chamseddine Barki, Ismail Dergaa, Karima Tlili, Halil İbrahim Ceylan, Nicola Luigi Bragazzi, Andrea de Giorgio, Ridha Ben Salah, Hanene Boussi Rahmouni

Abstract read
In one paragraph

Article in Bioengineering (Basel, Switzerland), 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.

0numbers the graph read from it
0cells of the map it votes in
5citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

5 citing papers in PubMed.

  1. Article
  2. Article
  3. Article
  4. Review
  5. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Hamida AbdaouiLaboratory of Biophysics and Medical Technologies, Higher Institute of Medical Technologies of Tunis (ISTMT), University of Tunis El Manar, Tunis 1006, Tunisia.
Chamseddine BarkiLaboratory of Biophysics and Medical Technologies, Higher Institute of Medical Technologies of Tunis (ISTMT), University of Tunis El Manar, Tunis 1006, Tunisia.ORCID 0000-0002-6595-3211
Ismail DergaaHigher Institute of Sport and Physical Education of Ksar Said, University of Manouba, Manouba 2010, Tunisia.ORCID 0000-0001-8091-1856
Karima TliliDepartment of Pathology, Faculty of Medicine of Tunis, Military Hospital of Tunis, Tunis 1008, Tunisia.
Halil İbrahim CeylanPhysical Education of Sports Teaching Department, Faculty of Sports Sciences, Atatürk University, 25240 Erzurum, Türkiye.ORCID 0009-0005-2214-4667
Nicola Luigi BragazziLaboratory for Industrial and Applied Mathematics (LIAM), Department of Mathematics and Statistics, York University, Toronto, ON M3J 1P3, Canada.ORCID 0000-0001-8409-868X
Andrea de GiorgioArtificial Engineering, 80121 Naples, Italy.ORCID 0000-0001-6064-5634
Ridha Ben SalahLaboratory of Biophysics and Medical Technologies, Higher Institute of Medical Technologies of Tunis (ISTMT), University of Tunis El Manar, Tunis 1006, Tunisia.
Hanene Boussi RahmouniLaboratory of Biophysics and Medical Technologies, Higher Institute of Medical Technologies of Tunis (ISTMT), University of Tunis El Manar, Tunis 1006, Tunisia.

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

backgroundAnatomopathological reports are largely unstructured, which limits automated data extraction, interoperability, and large-scale research. Manual extraction and standardization are costly and difficult to scale.

objectiveWe developed and evaluated an automated pipeline for entity extraction and multi-ontology normalization of anatomopathological reports.

methodsA corpus of 560 reports from the Military Hospital of Tunis, Tunisia, was manually annotated for three entity types: sample type, test performed, and finding. The entity extraction utilized BioBERT v1.1, while the normalization combined BioClinicalBERT multi-label classification with retrieval-augmented generation, incorporating both dense and BM25 sparse retrieval over SNOMED CT, LOINC, and ICD-11. The performance was measured using precision, recall, F1-score, and statistical tests.

resultsBioBERT achieved high extraction performance (F1: 0.97 for the sample type, 0.98 for the test performed, and 0.93 for the finding; overall 0.963, 95% CI: 0.933-0.982), with low absolute errors. For terminology mapping, the combination of BioClinicalBERT and dense retrieval outperformed the standalone and BM25-based approaches (macro-F1: 0.6159 for SNOMED CT, 0.9294 for LOINC, and 0.7201 for ICD-11). Cohen's Kappa ranged from 0.7829 to 0.9773, indicating substantial to near-perfect agreement.

conclusionsThe pipeline provides robust automated extraction and multi-ontology coding of anatomopathological entities, supporting transformer-based named entity recognition with retrieval-augmented generation. However, given the limitations of this study, multi-institutional validation is needed before clinical deployment.

Indexed as

anatomopathological reportBioClinicalBERTcode mappingdeep learningICD-11LOINCnamed entity recognitionnatural language processingSNOMED CTtransformer architecture

Identifiers

PMID41595962
PMCPMC12838374

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.