Fordham · CISC 5800, Machine Learning · Spring 2026
Type 2 Diabetes Prediction
A comparison of four ML classifiers on 253K CDC BRFSS survey records, tuned for screening: minority-class recall rose from 76% to 86% at ROC-AUC ≈ 0.82.

- clean records from a 441K-row, 330-column survey
- 253Kclean records from a 441K-row, 330-column survey
- minority-class recall at a screening-appropriate threshold
- 76% → 86%minority-class recall at a screening-appropriate threshold
- ROC-AUC across all four model families
- ≈ 0.82ROC-AUC across all four model families
The question
Can routinely collected survey answers about health, lifestyle, and demographics flag people likely to have type 2 diabetes? For a screening tool, missing a true case costs far more than a false alarm, so the goal is high recall on the diabetic class, not raw accuracy. With roughly 84% of respondents non-diabetic, accuracy is a misleading target anyway.
Data and preprocessing
The data is the CDC Behavioral Risk Factor Surveillance System (BRFSS), an annual telephone survey of US adults: 441K records and 330 columns. I built a reproducible preprocessing pipeline that:
- Narrows the 330 columns to 21 literature-grounded features, including blood pressure, cholesterol, BMI, general health, physical activity, and age
- Cleans survey-specific missing and "refused" codes, recodes binary features, and rescales BMI
- Produces 253K clean records, split 80/20 with stratification, plus a class-balanced training set for the SVMs
The pipeline runs as standalone scripts (preprocessing.py → modeling.py) with dependencies locked through uv, so the full analysis can be reproduced from the raw download.
Models
I trained and compared four model families, tuning each with 5-fold cross-validation:
- Logistic regression, as an interpretable baseline with L1 regularization paths
- Linear SVM
- Kernel SVM (RBF)
- PyTorch MLP, iterated three times: Adam, then AdamW, then AdamW with a tuned positive-class weight

Tuning for screening
All four families converge to ROC-AUC ≈ 0.82. The data seems to cap how well any of them can rank risk. What differs is where each model operates on that curve.
For the MLP, I tuned the precision/recall trade-off directly using a pos_weight on the loss and early stopping on F2 score, which weights recall twice as heavily as precision. This raised minority-class recall from 76% to 86% at an operating point that makes sense for screening, where a flagged patient gets a follow-up test rather than a diagnosis.

Takeaways
- When model families converge on AUC, choosing the operating point is the real modeling decision, and it should come from the clinical use case.
- A well-regularized logistic regression is competitive with a neural network on tabular survey data, and it is far easier to explain to a clinician.