JC
All projects

Fordham · CISC 5800, Machine Learning · Spring 2026

Type 2 Diabetes Prediction

A comparison of four ML classifiers on 253K CDC BRFSS survey records, tuned for screening: minority-class recall rose from 76% to 86% at ROC-AUC ≈ 0.82.

PyTorchScikit-LearnClassificationImbalanced DataHealth
ROC curves for logistic regression, linear SVM, kernel SVM, and a PyTorch MLP, all near AUC 0.82.
clean records from a 441K-row, 330-column survey
253Kclean records from a 441K-row, 330-column survey
minority-class recall at a screening-appropriate threshold
76% → 86%minority-class recall at a screening-appropriate threshold
ROC-AUC across all four model families
≈ 0.82ROC-AUC across all four model families

The question

Can routinely collected survey answers about health, lifestyle, and demographics flag people likely to have type 2 diabetes? For a screening tool, missing a true case costs far more than a false alarm, so the goal is high recall on the diabetic class, not raw accuracy. With roughly 84% of respondents non-diabetic, accuracy is a misleading target anyway.

Data and preprocessing

The data is the CDC Behavioral Risk Factor Surveillance System (BRFSS), an annual telephone survey of US adults: 441K records and 330 columns. I built a reproducible preprocessing pipeline that:

  • Narrows the 330 columns to 21 literature-grounded features, including blood pressure, cholesterol, BMI, general health, physical activity, and age
  • Cleans survey-specific missing and "refused" codes, recodes binary features, and rescales BMI
  • Produces 253K clean records, split 80/20 with stratification, plus a class-balanced training set for the SVMs

The pipeline runs as standalone scripts (preprocessing.py → modeling.py) with dependencies locked through uv, so the full analysis can be reproduced from the raw download.

Models

I trained and compared four model families, tuning each with 5-fold cross-validation:

  1. Logistic regression, as an interpretable baseline with L1 regularization paths
  2. Linear SVM
  3. Kernel SVM (RBF)
  4. PyTorch MLP, iterated three times: Adam, then AdamW, then AdamW with a tuned positive-class weight

ROC curves for all four models

Tuning for screening

All four families converge to ROC-AUC ≈ 0.82. The data seems to cap how well any of them can rank risk. What differs is where each model operates on that curve.

For the MLP, I tuned the precision/recall trade-off directly using a pos_weight on the loss and early stopping on F2 score, which weights recall twice as heavily as precision. This raised minority-class recall from 76% to 86% at an operating point that makes sense for screening, where a flagged patient gets a follow-up test rather than a diagnosis.

Confusion matrix for the final MLP

Takeaways

  • When model families converge on AUC, choosing the operating point is the real modeling decision, and it should come from the clinical use case.
  • A well-regularized logistic regression is competitive with a neural network on tabular survey data, and it is far easier to explain to a clinician.
Next projectFinBERT Sentiment vs. Price-Based ML →