Evidence map›Paper›PMID 41032884›Full record

ArticleJMIR formative research2025

Application of Large Language Models in Data Analysis and Medical Education for Assisted Reproductive Technology: Comparative Study.

Noriyuki Okuyama, Mika Ishii, Yuriko Fukuoka, Hiromitsu Hattori, Yuta Kasahara, Tai Toshihiro, Koki Yoshinaga, Tomoko Hashimoto, Koichi Kyono

Abstract readComparative Study
In one paragraph

Article in JMIR formative research, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Noriyuki Okuyama *Kyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0000-0001-7347-5312
Mika IshiiKyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0009-0001-9009-915X
Yuriko FukuokaKyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0009-0000-0777-0559
Hiromitsu HattoriKyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0009-0001-0522-7760
Yuta KasaharaKyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0000-0002-6369-1086
Tai ToshihiroKyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0000-0002-4069-3724
Koki YoshinagaKyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0009-0006-0839-0981
Tomoko HashimotoKyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0000-0001-6178-7204
Koichi Kyono *Kyono ART Clinic Takanawa, Takanawa Court 5F, 3-13-1 Takanawa, Minato-ku, Tokyo, 108-0074, Japan, 81 364084708, 81 364084702.ORCID 0000-0001-5211-7778

Funding

No grant is acknowledged in the PubMed record.

6 · The paper itself

Abstract

Background: Recent studies have demonstrated that large language models exhibit exceptional performance in medical examinations. However, there is a lack of reports assessing their capabilities in specific domains or their application in practical data analysis using code interpreters. Furthermore, comparative analyses across different large language models have not been extensively conducted. Objective: The purpose of this study was to evaluate whether advanced artificial intelligence (AI) models can analyze data from template-based input and demonstrate basic knowledge of reproductive medicine. Four AI models (GPT-4, GPT-4o, Claude 3.5 Sonnet, and Gemini Pro 1.5) were evaluated for their data analytical capabilities through numerical calculations and graph rendering. Their knowledge of infertility treatment was assessed using 10 examination questions developed by experts. Methods: First, we uploaded data to the AI models and furnished instruction templates using the chat interface. The study investigated whether the AI models could perform pregnancy rate analysis and graph rendering, based on blastocyst grades according to Gardner criteria. Second, we assessed model diagnostic capabilities based on specialized knowledge. This evaluation used 10 questions derived from the Japanese Fertility Specialist Examination and the Embryologist Certification Exam, along with chromosome imaging. These materials were curated under the supervision of certified embryologists and fertility specialists. All procedures were repeated 10 times per AI model. Results: GPT-4o achieved grade A output (defined as achieving the objective with a single output attempt) in 9 out of 10 trials, outperforming GPT-4, which achieved grade A in 7 out of 10. The average processing times for data analysis were 26.8 (SD 3.7) seconds for GPT-4o and 36.7 (SD 3) seconds for GPT-4, whereas Claude failed in all 10 attempts. Gemini achieved an average processing time of 23 (SD 3) seconds and received grade A in 6 out of 10 trials, though occasional manual corrections were needed. Embryologists required an average of 358.3 (SD 9.7) seconds for the same tasks. In the knowledge-based assessment, GPT-4o, Claude, and Gemini achieved perfect scores (9/9) on multiple-choice questions, while GPT-4 showed a 60% (6/10) success rate on 1 question. None of the AI models could reliably diagnose chromosomal abnormalities from karyotype images, with the highest image diagnostic accuracy being 70% (7/10) for Claude and Gemini. Conclusions: This rapid processing demonstrates the potential for these AI models to significantly expedite data-intensive tasks in clinical settings. This performance underscores their potential utility as educational tools or decision support systems in reproductive medicine. However, none of the models were able to accurately interpret and diagnose using medical images.

Indexed as

Artificial IntelligenceData AnalysisEducation, MedicalLanguageReproductive Techniques, AssistedFemaleHumansJapanLarge Language ModelsPregnancyartificial intelligencedata analysiseducationinfertilitylarge language model

Identifiers

PMID41032884
PMCPMC12488165

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.