Abstract
Large-scale population surveys provide valuable information for studying child well-being, yet their structure often limits the direct application of machine-learning methods. The National Survey of Children’s Health (NSCH) is one of the most comprehensive datasets for monitoring children’s health and development in the United States, but the raw survey files contain logical skip patterns, categorical variables, and complex survey-design elements that require substantial preprocessing before predictive analysis can be performed. This study presents a curated machine-learning-ready benchmark dataset derived from the 2023 NSCH together with a fully reproducible computational pipeline for studying school-age child flourishing. The workflow constructs a binary flourishing outcome from four survey items related to curiosity, task persistence, emotional self-regulation, and interest in doing well in school. After restricting the sample to children aged 6–17 years and retaining only records with valid responses in all four outcome items, the final analytical dataset contained 32,934 observations. Feature selection based on mutual information computed on the training partition, combined with cross-validated subset-size selection, yielded a final benchmark subset of 150 predictors. Baseline experiments using logistic regression and random forest showed stable and reasonably strong predictive performance, with held-out ROC-AUC values around 0.84–0.85 and closely aligned cross-validation results. An exploratory comparison between weighted and unweighted learning further showed that survey weighting did not improve discriminative performance in this benchmark setting, although the magnitude of the effect was modest and model-dependent. By releasing both the curated benchmark dataset and the reproducible pipeline, this study provides a reusable resource for machine-learning research on child well-being and survey-based computational benchmarking.
| Original language | English |
|---|---|
| Article number | 103 |
| Journal | Data |
| Volume | 11 |
| Issue number | 5 |
| DOIs | |
| Publication status | Published - May 2026 |
Bibliographical note
Publisher Copyright:© 2026 by the authors.
Fingerprint
Dive into the research topics of 'NSCH-Flourishing-ML: A Curated Dataset and Reproducible Pipeline for Machine Learning Analysis of Child Flourishing'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver