Evidence map›Paper›PMID 34600486›Full record

SynthesisBMC medical imaging2021

The reporting quality of natural language processing studies: systematic review of studies of radiology reports.

Emma M Davidson, Michael T C Poon, Arlene Casey, Andreas Grivas, Daniel Duma, Hang Dong, Víctor Suárez-Paniagua, Claire Grover, Richard Tobin, Heather Whalley and 3 more

Open access · goldAbstract readSystematic Review
In one paragraph

Synthesis in BMC medical imaging, 2021. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 20 papers, 2 of them syntheses that pooled it.

0numbers the graph read from it
0cells of the map it votes in
20citing papers in PubMed, 2 pooled it
1.6field-weighted citation impact, top 14% of its field
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

20 citing papers in PubMed, 2 syntheses or guidelines pooled it, 37 citations in OpenAlex.

  1. Pooled it
  2. Pooled it
  3. Article
  4. Article
  5. Article
  6. Article
  7. Article
  8. Review
  9. Article
  10. Review
  11. Article
  12. Review
  13. Article
  14. Article
  15. Article
  16. Article
  17. Review
  18. Article
  19. Article
  20. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

13 authors at 3 institutions in 1 country.

Emma M Davidson *Centre for Clinical Brain Sciences, University of Edinburgh, Chancellor's Building, Little France, Edinburgh, EH16 4TJ, Scotland, UK. emma.davidson@ed.ac.uk.
Michael T C Poon *Centre for Medical Informatics, Usher Institute, University of Edinburgh, Edinburgh, Scotland, UK.
Arlene CaseySchool of Literatures, Languages and Cultures (LLC), University of Edinburgh, Edinburgh, Scotland, UK.
Andreas GrivasSchool of Literatures, Languages and Cultures (LLC), University of Edinburgh, Edinburgh, Scotland, UK.
Daniel DumaSchool of Literatures, Languages and Cultures (LLC), University of Edinburgh, Edinburgh, Scotland, UK.
Hang DongCentre for Medical Informatics, Usher Institute, University of Edinburgh, Edinburgh, Scotland, UK.
Víctor Suárez-PaniaguaCentre for Medical Informatics, Usher Institute, University of Edinburgh, Edinburgh, Scotland, UK.
Claire GroverInstitute for Language, Cognition and Computation, School of Informatics, University of Edinburgh, Edinburgh, Scotland, UK.
Richard TobinInstitute for Language, Cognition and Computation, School of Informatics, University of Edinburgh, Edinburgh, Scotland, UK.
Heather WhalleyCentre for Clinical Brain Sciences, University of Edinburgh, Chancellor's Building, Little France, Edinburgh, EH16 4TJ, Scotland, UK.
Honghan WuHealth Data Research UK, London, UK.
Beatrice Alex *School of Literatures, Languages and Cultures (LLC), University of Edinburgh, Edinburgh, Scotland, UK.
William Whiteley *Centre for Clinical Brain Sciences, University of Edinburgh, Chancellor's Building, Little France, Edinburgh, EH16 4TJ, Scotland, UK.
University of Edinburgh · GBHealth Data Research UK · GBEdinburgh Cancer Research · GB

Funding

Cancer Research UK 27589Chief Scientist Office SCAF/17/01Medical Research Council MC_PC_18029Medical Research Council MR/S004149/1Medical Research Council MR/S004149/2
6 · The paper itself

Abstract

backgroundAutomated language analysis of radiology reports using natural language processing (NLP) can provide valuable information on patients' health and disease. With its rapid development, NLP studies should have transparent methodology to allow comparison of approaches and reproducibility. This systematic review aims to summarise the characteristics and reporting quality of studies applying NLP to radiology reports.

methodsWe searched Google Scholar for studies published in English that applied NLP to radiology reports of any imaging modality between January 2015 and October 2019. At least two reviewers independently performed screening and completed data extraction. We specified 15 criteria relating to data source, datasets, ground truth, outcomes, and reproducibility for quality assessment. The primary NLP performance measures were precision, recall and F1 score.

resultsOf the 4,836 records retrieved, we included 164 studies that used NLP on radiology reports. The commonest clinical applications of NLP were disease information or classification (28%) and diagnostic surveillance (27.4%). Most studies used English radiology reports (86%). Reports from mixed imaging modalities were used in 28% of the studies. Oncology (24%) was the most frequent disease area. Most studies had dataset size > 200 (85.4%) but the proportion of studies that described their annotated, training, validation, and test set were 67.1%, 63.4%, 45.7%, and 67.7% respectively. About half of the studies reported precision (48.8%) and recall (53.7%). Few studies reported external validation performed (10.8%), data availability (8.5%) and code availability (9.1%). There was no pattern of performance associated with the overall reporting quality.

conclusionsThere is a range of potential clinical applications for NLP of radiology reports in health services and research. However, we found suboptimal reporting quality that precludes comparison, reproducibility, and replication. Our results support the need for development of reporting standards specific to clinical NLP studies.

Indexed as

Natural Language ProcessingRadiographyDatasets as TopicHumansRadiologyReproducibility of ResultsResearch ReportNatural language processingRadiology reportsSystematic review

Identifiers

PMID34600486
PMCPMC8487512
OpenAlexW3203041483

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.