Evidence map›Paper›PMID 38909025›Full record

ArticleScientific reports2024

Synthetic data generation for a longitudinal cohort study - evaluation, method extension and reproduction of published data analysis results.

Lisa Kühnel, Julian Schneider, Ines Perrar, Tim Adams, Sobhan Moazemi, Fabian Prasser, Ute Nöthlings, Holger Fröhlich, Juliane Fluck

Abstract read
In one paragraph

Article in Scientific reports, 2024. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 4 papers.

0numbers the graph read from it
0cells of the map it votes in
4citing papers in PubMed
–field-weighted citation impact
1 · What the graph read from it

What it found

Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.

The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.

2 · The registry

The trial behind it

Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.

Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.

3 · Its place in the literature

Who cites it

4 citing papers in PubMed.

  1. Article
  2. Review
  3. Article
  4. Article
4 · The record

Corrections and comments

PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.

5 · Who and what money

Authors and funding

9 authors.

Lisa KühnelKnowledge Management, ZB MED - Information Centre for Life Sciences, 50931, Cologne, Germany. kuehnel@zbmed.de.
Julian SchneiderKnowledge Management, ZB MED - Information Centre for Life Sciences, 50931, Cologne, Germany.
Ines PerrarInstitute of Nutritional and Food Sciences - Nutritional Epidemiology, University of Bonn, 53115, Bonn, Germany.
Tim AdamsDepartment of Bioinformatics, Fraunhofer Institute for Algorithms and Scientific Computing SCAI, 53757, Sankt Augustin, Germany.
Sobhan MoazemiDepartment of Bioinformatics, Fraunhofer Institute for Algorithms and Scientific Computing SCAI, 53757, Sankt Augustin, Germany.
Fabian PrasserMedical Informatics Group, Berlin Institute of Health at Charité - Universitätsmedizin Berlin, 10117, Berlin, Germany.
Ute NöthlingsInstitute of Nutritional and Food Sciences - Nutritional Epidemiology, University of Bonn, 53115, Bonn, Germany.
Holger FröhlichDepartment of Bioinformatics, Fraunhofer Institute for Algorithms and Scientific Computing SCAI, 53757, Sankt Augustin, Germany.
Juliane FluckKnowledge Management, ZB MED - Information Centre for Life Sciences, 50931, Cologne, Germany.

Funding

Alzheimer's Disease Neuroimaging Initiative - SupplementU01AG024904 · NIA · NORTHERN CALIFORNIA INSTITUTE RES &EDUC · PI WEINER, MICHAEL W · 2004 to 2015
$121.0M
UC Davis Alzheimer's Disease Core CenterP30AG010129 · NIA · UNIVERSITY OF CALIFORNIA DAVIS · PI JOHNSON, DAVID K · 1991 to 2020
$28.3M
"MR Morphometrics and Cognitive Decline Rate in Large-Scale Aging Studies"K01AG030514 · NIA · UNIVERSITY OF CALIFORNIA AT DAVIS · PI CARMICHAEL, OWEN T. · 2008 to 2012
$495k
Bundesministerium für Ernährung und Landwirtschaft 2816HS024Deutsche Forschungsgemeinschaft 442326535NIA NIH HHS K01 AG030514NIA NIH HHS P30 AG010129NIA NIH HHS U01 AG024904
6 · The paper itself

Abstract

Access to individual-level health data is essential for gaining new insights and advancing science. In particular, modern methods based on artificial intelligence rely on the availability of and access to large datasets. In the health sector, access to individual-level data is often challenging due to privacy concerns. A promising alternative is the generation of fully synthetic data, i.e., data generated through a randomised process that have similar statistical properties as the original data, but do not have a one-to-one correspondence with the original individual-level records. In this study, we use a state-of-the-art synthetic data generation method and perform in-depth quality analyses of the generated data for a specific use case in the field of nutrition. We demonstrate the need for careful analyses of synthetic data that go beyond descriptive statistics and provide valuable insights into how to realise the full potential of synthetic datasets. By extending the methods, but also by thoroughly analysing the effects of sampling from a trained model, we are able to largely reproduce significant real-world analysis results in the chosen use case.

Indexed as

Data AnalysisArtificial IntelligenceHumansLongitudinal Studies

Identifiers

PMID38909025
PMCPMC11193715

What OpenQuestion holds

Textmetadata
LicenceCC BY
Read underepoch 390

Registered trials

None linked

Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.