Yayın: Predicting smoking status from short voice recordings under small-sample constraints: A calibrated leave-one-speaker-out study
| dc.contributor.author | Aydogan, Yigit | |
| dc.contributor.author | Duygun, Oguzhan | |
| dc.contributor.author | Canturk, Ismail | |
| dc.date.accessioned | 2026-06-27T15:31:57Z | |
| dc.date.issued | 2026 | |
| dc.description.abstract | The feasibility of inferring smoking status from short voice recordings was examined under small-sample, speaker-independent constraints, emphasizing calibrated probabilities and decision utility for screening. Sustained/a/ phonations recorded on smartphones (44.1 kHz, mono) were analyzed from 64 unique speakers (30 smokers, 34 non-smokers; prevalence 0.469; one recording per speaker). Two representation families were compared: (i) a physiology-informed handcrafted prosody-spectral set (208 variables) summarizing perturbation, harmonicity/noise structure, spectral-energy distribution, and formants; and (ii) pretrained embeddings (YAMNet, wav2vec 2.0, WavLM) pooled to utterance vectors and classified with partial least squares plus logistic regression. Models were evaluated with strict leave-one-speaker-out validation, nested hyperparameter selection, fold-safe preprocessing, and within-fold Platt scaling. The handcrafted elastic-net logistic model (PS_ENet) achieved the strongest discrimination (AUC = 0.885), with accuracy 0.844, F1 0.828, average precision 0.894, and Brier score 0.193. Embedding baselines underperformed (AUC: YAM_PLS 0.475; W2V2_PLS 0.561; WAVLM_PLS 0.525). A probability-averaging ensemble favored sensitivity (recall 0.833) with AUC 0.797. Demographics alone were partially predictive (age+gender LOSO AUC = 0.708), but PS_ENet retained incremental discrimination under restricted permutation within age & times; gender strata and demographic matching. Speaker-level bootstrapping yielded PS_ENet AUC 0.886 (95% CI 0.790-0.962), and full-pipeline permutation testing supported discrimination beyond chance (p approximate to 0.005). Decision-curve analysis showed positive net benefit for thresholds 0.05-0.30 (exceeding treat-all for thresholds >= 0.15). Prospective planning suggested N approximate to 44 speakers for 80% power to detect AUC = 0.70 under the observed prevalence. | en |
| dc.description.uri | https://doi.org/10.1016/j.bspc.2026.109915 | |
| dc.identifier.doi | 10.1016/j.bspc.2026.109915 | |
| dc.identifier.eissn | 1746-8108 | |
| dc.identifier.issn | 1746-8094 | |
| dc.identifier.uri | https://hdl.handle.net/20.500.14981/71615 | |
| dc.identifier.volume | 119 | |
| dc.identifier.wos | 001699361300001 | |
| dc.language.iso | eng | |
| dc.publisher | ELSEVIER SCI LTD | |
| dc.relation.ispartof | BIOMEDICAL SIGNAL PROCESSING AND CONTROL | |
| dc.subject | Smoking detection | |
| dc.subject | Voice biomarkers | |
| dc.subject | Small-sample learning | |
| dc.subject | Calibration | |
| dc.subject | Decision-curve analysis | |
| dc.subject | LOSO evaluation | |
| dc.subject | Engineering | |
| dc.title | Predicting smoking status from short voice recordings under small-sample constraints: A calibrated leave-one-speaker-out study | |
| dc.type | Article | |
| dspace.entity.type | Publication | |
| local.import.source | WOS |