ArticleDigital discovery2026
Censoring chemical data to mitigate dual use risk.
Article in Digital discovery, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Cited by 5 papers.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
5 citing papers in PubMed.
- Probing the limitations of multimodal language models for chemistry and materials research.Nature computational science · 2025Article
- A framework for evaluating the chemical knowledge and reasoning abilities of large language models against the expertise of chemists.Nature chemistry · 2025Article
- Augmenting large language models with chemistry tools.Nature machine intelligence · 2024Article
- From intuition to AI: evolution of small molecule representations in drug discovery.Briefings in bioinformatics · 2023Article
- 14 examples of how LLMs can transform materials science and chemistry: a reflection on a large language model hackathon.Digital discovery · 2023Review
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
3 authors.
Funding
Abstract
Machine learning models have dual use potential, potentially serving both beneficial and malicious purposes. The development of open-source models in chemistry has specifically surfaced dual use concerns around toxicological data and chemical warfare agents. We discuss a chain risk framework identifying three misuse pathways and corresponding mitigation strategies: inference-level, model-level, and data-level. At the data level, we introduce a noising method to increase prediction error in specific desired regions (sensitive regions). Our results show that selective noise induces variance and attenuation bias, whereas simply omitting sensitive data fails to prevent extrapolation. These findings hold for both molecular feature multilayer perceptrons and graph neural networks. Thus, noising molecular structures represents a step toward enabling safer sharing of potential dual use molecular data.
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.