ArticleNicotine & tobacco research : official journal of the Society for Research on Nicotine and Tobacco2026
Training a Smoking Status Probabilistic Model Using Cotinine Levels in a Large Claims Database.
Article in Nicotine & tobacco research : official journal of the Society for Research on Nicotine and Tobacco, 2026. The graph could read no effect estimate from its abstract, so it casts no vote on the map. Not yet cited in PubMed.
What it found
Each row is one number read from the abstract, on the scale the paper reported it, with its interval. Left of the dashed line favours the treatment, right favours the comparator. Under each row is the sentence it came from. New to these charts? A ten-minute tutorial.
The abstract states no effect estimate the extractor could read, or names no intervention and outcome on the map, so this paper lights no cell and moves no belief. It is still indexed, cited and linked below.
The trial behind it
Trials whose registry record cites this paper, or whose number appears in the abstract. A trial that started after this paper was published is citing it as background, not reporting it.
Neither the registry nor the abstract names a trial number. If this is a trial report, that itself is worth knowing.
Who cites it
0 citing papers in PubMed.
No citing paper in PubMed yet.
Corrections and comments
PubMed lists nothing against this paper. Absence here is not a guarantee, only a check that was made.
Authors and funding
5 authors.
Funding
Abstract
introductionSmoking status is an important confounder for many epidemiologic studies, yet it is not well documented in common sources of real-world data, including administrative claims. Probabilistic models can be used to create a proxy for smoking status, yet most published models have been trained using self-reported data. The objective of this study was to train a smoking status probabilistic model using cotinine values available in a large claims database.
methodsBeneficiaries were included if they had at least one cotinine measurement and were categorized as a "current smoker" if their serum or plasma cotinine value was ≥5 ng/mL or urine cotinine value was ≥30 ng/mL. Predictors were collected across one year prior to the cotinine assessment date. The model was fit using logistic regression with stepwise forward selection. Model performance was assessed using discrimination and calibration.
resultsThe final model yielded an area under the receiver operating characteristic curve of 0.77 (95%CI:0.75-0.78) and was well calibrated across most prediction deciles. The strongest predictors included diagnosis codes for smoking and drug abuse, and number of medications. The model was found to be highly specific, yet not sensitive at probability cutoffs ≥0.2.
conclusionsA smoking status model was developed and internally validated for application in claims data, using available cotinine values to define smoking status and found to have acceptable discrimination and calibration. The model is based on 26 predictors, fewer than other similar published smoking status models. External validation of the model should be a next step toward utilizing the model for epidemiological research. IMPLICATIONS: This study tests the utility of cotinine values to validate a smoking status probabilistic model, which has not been done in the literature to date. The results were robust to various cotinine levels used to define smoking status, per current guidance. The final model uses only 26 factors to predict smoking status, simplifying the application of the model in other claims databases.
Indexed as
Identifiers
What OpenQuestion holds
Registered trials
Read under generation 80e0d062 · epoch 390. Bibliography from PubMed, PubMed Central and OpenAlex; grants from NIH RePORTER; trial links from ClinicalTrials.gov; estimates, votes and beliefs from the OpenQuestion graph.