Evidence map›Paper›PMID 42375709›Full record

ArticleAlgorithms2025

Multimodal LLM vs. Human-Measured Features for AI Predictions of Autism in Home Videos.

Parnian Azizian, Mohammadmahdi Honarmand, Aditi Jaiswal, Aaron Kline, Kaitlyn Dunlap, Peter Washington, Dennis P Wall

Abstract read
In one paragraph

Article in Algorithms, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 1 paper.

0numbers the graph read from it
0cells of the map it votes in
1citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

1 citing paper in PubMed.

  1. Review
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Parnian AzizianDepartment of Mechanical Engineering, Stanford University, Stanford, CA 94305, USA.ORCID 0009-0006-6975-4953
Mohammadmahdi HonarmandDepartment of Mechanical Engineering, Stanford University, Stanford, CA 94305, USA.ORCID 0000-0002-5778-6054
Aditi JaiswalDepartment of Information and Computer Sciences, University of Hawaii at Manoa, Honolulu, HI 96822, USA.ORCID 0000-0003-1367-818X
Aaron KlineDepartment of Biomedical Data Science, Stanford University School of Medicine, Stanford, CA 94305, USA.ORCID 0000-0002-0077-5485
Kaitlyn DunlapDepartment of Biomedical Data Science, Stanford University School of Medicine, Stanford, CA 94305, USA.ORCID 0000-0003-4423-5269
Peter WashingtonDepartment of Medicine (Clinical Informatics and Digital Transformation), University of California San Francisco, San Francisco, CA 94143, USA.ORCID 0000-0003-3276-4411
Dennis P WallDepartment of Biomedical Data Science, Stanford University School of Medicine, Stanford, CA 94305, USA.ORCID 0000-0002-7889-9146

Funding

A Mobile Game for Domain Adaptation and Deep Learning in Autism HealthcareR01LM013364 · NLM · STANFORD UNIVERSITY · PI WALL, DENNIS PAUL · 2021 to 2025
$3.2M
Crowd-Powered Machine Learning to Diagnose ASD and ADHD in Adolescents from Digital Social InteractionsDP2EB035858 · NIBIB · UNIVERSITY OF HAWAII AT MANOA · PI Peter Washington · 2023 to 2026
$2.3M
An active learning framework for adaptive autism healthcareR01LM014342 · NLM · STANFORD UNIVERSITY · PI Dennis Paul Wall · 2023 to 2026
$1.9M
NIBIB NIH HHS DP2 EB035858NLM NIH HHS R01 LM013364NLM NIH HHS R01 LM014342
6 · The paper itself

Abstract

Autism diagnosis remains a critical healthcare challenge, with current assessments contributing to average diagnostic ages of 5 and extending to 8 in underserved populations. With the FDA approval of CanvasDx in 2021, the paradigm of human-in-the-loop AI diagnostics entered the pediatric market as the first medical device for clinically precise autism diagnosis at scale, while fully automated deep learning approaches have remained underdeveloped. However, the importance of early autism detection, ideally before 3 years of age, underscores the value of developing even more automated AI approaches, due to their potentials for scale, reach, and privacy. We present the first systematic evaluation of multimodal LLMs as direct replacements for human annotation in AI-based autism detection. Evaluating seven Gemini model variants (1.5-2.5 series) on 50 YouTube videos shows clear generational progression: version 1.5 models achieve 72-80% accuracy, version 2.0 models reach 80%, and version 2.5 models attain 85-90%, with the best model (2.5 Pro) achieving 89.6% classification accuracy using validated autism detection AI models (LR5)-comparable to the 88% clinical baseline and approaching crowdworker performance of 92-98%. The 24% improvement across two generations suggests the gap is closing. LLMs demonstrate high within-model consistency versus moderate human agreement, with distinct assessment strategies: LLMs focus on language/behavioral markers, crowdworkers prioritize social-emotional engagement, clinicians balance both. While LLMs have yet to match the highest-performing subset of human annotators in their ability to extract behavioral features that are useful for human-in-the-loop AI diagnosis, their rapid improvement and advantages in consistency, scalability, cost, and privacy position them as potentially viable alternatives for aiding diagnostic processes in the future.

Indexed as

artificial intelligenceautism spectrum disorderhuman-in-the-loop AImachine learningmultimodal large language modelsvideo-based screening

Identifiers

PMID42375709
PMCPMC13313062

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.