Evidence map›Paper›PMID 41275284›Full record

ArticleInternational journal of health geographics2025

H3-MOSAIC: multimodal generative AI for semantic place detection from high-frequency GPS on H3 grids in mental health geomatics.

Lingbo Liu, Rachel Franklin, Jiaee Cheong, Tianyue Cong, Jin Soo Byun, Allie Yubin Oh, John Torous

Abstract read
In one paragraph

Article in International journal of health geographics, 2025. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.

0numbers the graph read from it
0cells of the map it votes in
0citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

0 citing papers in PubMed.

No citing paper in PubMed yet.

4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

7 authors.

Lingbo LiuCenter for Geographic Analysis, Harvard University, Cambridge, MA, 02138, USA. lingboliu@fas.harvard.edu.ORCID 0000-0002-9876-8506
Rachel FranklinCenter for Geographic Analysis, Harvard University, Cambridge, MA, 02138, USA. rachel_franklin@cga.harvard.edu.ORCID 0000-0002-2614-4665
Jiaee CheongBeth Israel Deaconess Medical Center, Boston, MA, 02115, USA.
Tianyue CongBeth Israel Deaconess Medical Center, Boston, MA, 02115, USA.
Jin Soo ByunBeth Israel Deaconess Medical Center, Boston, MA, 02115, USA.
Allie Yubin OhBeth Israel Deaconess Medical Center, Boston, MA, 02115, USA.
John TorousBeth Israel Deaconess Medical Center, Boston, MA, 02115, USA. jtorous@bidmc.harvard.edu.ORCID 0000-0002-5362-7937

Funding

National Science Foundation 181143
6 · The paper itself

Abstract

backgroundMental-health geomatics require reliable ways to convert high-frequency GPS trajectories into meaningful place types that support indicators such as homestay, location entropy, and spatial extent of daily activities. Raw coordinates are typically noisy and carry little semantic information. We introduce H3-MOSAIC(H3-based Multimodal OSM-and-Satellite AI for Classification), a multimodal generative framework that fuses OpenStreetMap (OSM) building text and satellite imagery on H3 grids to infer place semantics from high-frequency GPS.

methodsRaw GPS was smoothed by minute-level speed filtering, then assigned to Level 10 H3 hexagons. Cells were retained if the mean speed was ≤ 1.2 m/s and the cumulative duration was ≥ 15 min, contiguous cells were merged, and home was defined as the cell with the longest dwell from 23:45 to 06:00. We compared text-only OSM classification with image-based and fused approaches across open-source models (DeepSeek, CLIP, LLaVA, Qwen-VL) and proprietary models (GPT-4o-mini, Gemini-2.5-flash-lite). Performance was assessed by accuracy, Cohen's kappa, precision, recall, F-measure, and confusion matrices. Day level associations between H3 semantic exposures and stress were examined by a random forest model and explainable methods.

resultsMultimodal methods outperformed single-modality baselines. In the 11-class task, accuracies were: CLIP 0.179, LLaVA 0.269, Qwen-VL 0.565, GPT-4o-mini 0.779, and Gemini-2.5-flash-lite 0.790. In the 5-class consolidation, accuracies rose to 0.702 (Qwen-VL), 0.849 (GPT-4o-mini), and 0.858 (Gemini-2.5-flash-lite). Text-only OSM baselines were lower (≈ 0.60-0.68). Across 3,845 hexagons with OSM text, closed-source models agreed on 79% of labels; disagreements concentrated in mixed-use, office, and green classes. Error modes reflected area-dominant versus keyword-triggered reasoning, hybrid-parcel ambiguity, tag sparsity, and symbolic artifacts. Stabilized semantics support more robust computation of homestay, entropy, and activity space and are suitable for privacy-aware, cross-city reuse. In a day-level case study, minutes at Home related to lower stress; Green showed a U-shaped pattern.

conclusionsH3-MOSAIC provides a scalable, auditable pipeline for semantic place detection from high-frequency GPS. Multimodal fusion markedly improves accuracy and consistency. Proprietary models are most robust on hard classes and open-source models are practical for coarse taxonomies. H3 day level exposures show stress patterns consistent with established mental health pathways, supporting face validity. The framework enables downstream exposure analyses with reduced misclassification and improved interpretability.

Indexed as

Artificial IntelligenceGeographic Information SystemsMental HealthSatellite ImagerySemanticsHumansH3 hexagonal gridsHigh-frequency GPSLarge language models (LLMs)Multimodal generative AISemantic place detectionVision-language models (VLMs)

Identifiers

PMID41275284
PMCPMC12640565

What OpenQuestion holds

Textmetadata
LicenceCC BY-NC-ND
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.