Evidence map›Paper›PMID 40100270›Full record

ArticleJournal of medical Internet research2025

How to Design, Create, and Evaluate an Instruction-Tuning Dataset for Large Language Model Training in Health Care: Tutorial From a Clinical Perspective.

Wojciech Nazar, Grzegorz Nazar, Aleksandra Kamińska, Ludmila Danilowicz-Szymanowicz

Abstract read
In one paragraph

Article in Journal of medical Internet research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 7 papers.

0numbers the graph read from it
0cells of the map it votes in
7citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

7 citing papers in PubMed.

  1. Article
  2. Review
  3. Article
  4. Article
  5. Review
  6. Review
  7. Assessing Patient Education Materials for Colorectal Cancer Generated by Four Large Language Models: Readability, Quality, and Transparency Challenges.Journal of cancer education : the official journal of the American Association for Cancer Education · 2025
    Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

4 authors.

Wojciech NazarDepartment of Allergology, Faculty of Medicine, Gdańsk Medical University, Gdansk, Poland.ORCID https://orcid.org/0000-0002-8448-0800
Grzegorz NazarFaculty of Medicine, Gdańsk Medical University, Gdansk, Poland.ORCID https://orcid.org/0009-0006-2418-1282
Aleksandra KamińskaFaculty of Medicine, Gdańsk Medical University, Gdansk, Poland.ORCID https://orcid.org/0009-0002-8268-3804
Ludmila Danilowicz-SzymanowiczDepartment of Cardiology and Electrotherapy, Faculty of Medicine, Gdańsk Medical University, Gdansk, Poland.ORCID https://orcid.org/0000-0002-2269-1880

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

High-quality data are critical in health care, forming the cornerstone for accurate diagnoses, effective treatment plans, and reliable conclusions. Similarly, high-quality datasets underpin the development and performance of large language models (LLMs). Among these, instruction-tuning datasets (ITDs) used for instruction fine-tuning have been pivotal in enhancing LLM performance and generalization capabilities across diverse tasks. This tutorial provides a comprehensive guide to designing, creating, and evaluating ITDs for health care applications. Written from a clinical perspective, it aims to make the concepts accessible to a broad audience, especially medical practitioners. Key topics include identifying useful data sources, defining the characteristics of well-designed datasets, and crafting high-quality instruction-input-output examples. We explore practical approaches to dataset construction, examining the advantages and limitations of 3 primary methods: fully manual preparation by expert annotators, fully synthetic generation using artificial intelligence (AI), and an innovative hybrid approach in which experts draft the initial dataset and AI generates additional data. Moreover, we discuss strategies for metadata selection and human evaluation to ensure the quality and effectiveness of ITDs. By integrating these elements, this tutorial provides a structured framework for establishing ITDs. It bridges technical and clinical domains, supporting the continued interdisciplinary advancement of AI in medicine. Additionally, we address the limitations of current practices and propose future directions, emphasizing the need for a global, unified framework for ITDs. We also argue that artificial general intelligence (AGI), if realized, will not replace empirical research in medicine. AGI will depend on human-curated datasets to process and apply medical knowledge. At the same time, ITDs will likely remain the most effective method of supplying this knowledge to AGI, positioning them as a critical tool in AI-driven health care.

Indexed as

Delivery of Health CareArtificial IntelligenceHumansLarge Language Modelsevaluation frameworkgenerative artificial intelligencehealth careinstruction-tuning datasetslarge language modelstutorials

Identifiers

PMID40100270
PMCPMC11962319

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.