Healthcare data analysis · public dataset

Reading the overlap

I wanted to look at a diagnostic dataset the way I would look at an operational one: what separates the groups, where do they overlap, and what do we give up when we move a cutoff?

Wisconsin Diagnostic Breast Cancer569 samplesInteractive exploratory analysis
The question: If one measurement is used to flag a case for follow-up, how does the cutoff change the balance between catching malignant samples and flagging benign ones? This is an exercise in data analysis, not a clinical decision rule.
569Samples in this dataset
212Labeled malignant
357Labeled benign

Start with the distribution

Pick a measurement. Each bar shows how many samples fall in the same range; the two colors show whether the original dataset labels them malignant or benign. The important part is where both colors show up.

MalignantBenign

What I notice

Malignant median
Benign median

The values come from individual cell nuclei measured in fine-needle aspirate images. “Worst” is a dataset feature name, not a judgment about a person. The chart uses equal-width bins across the full observed range.

Now move the cutoff

Here I use worst radius alone. A sample at or above the selected value is flagged. Move the slider and watch what happens to the actual dataset counts. There is no single cutoff that gets everything right.

Sensitivity · malignant caught
Specificity · benign cleared
SensitivitySpecificity

At this cutoff

Malignant flagged
Malignant missed
Benign flagged
Benign cleared

Sensitivity = malignant flagged ÷ 212. Specificity = benign cleared ÷ 357. Counts are calculated on the same 569 samples used to explore the cutoff, so these are descriptive, in-sample results.

What I would take forward

1. The groups differ

Worst radius has a median of 20.59 for malignant samples and 13.35 for benign ones here. That is a useful signal to investigate, but the distributions still overlap.

2. The tradeoff matters

At a cutoff of 15.0, this simple rule flags 205 of 212 malignant samples and 75 of 357 benign samples. A lower cutoff catches more malignant samples and asks for more follow-up.

3. This is a starting point

I would want validation on new, representative data and input from clinicians before discussing real use. These are historical samples, not a screening population or a picture of current cancer rates.

How I did it

  1. Loaded the public Wisconsin Diagnostic Breast Cancer dataset through scikit-learn. It contains 569 labeled samples and 30 numerical features derived from images of fine-needle aspirates.
  2. Kept five measurements for the comparison above. Calculated the medians separately for malignant and benign labels, then placed all samples into the same equal-width histogram bins.
  3. For the threshold example, flagged each row where worst radius is at least the selected value. Counted true positives, false negatives, false positives, and true negatives against the original labels.
  4. Recalculated sensitivity and specificity from those counts. I did not train or validate a predictive model in this project.

Limits: The dataset is small and selected. Its class mix (212 malignant, 357 benign) should not be treated as population prevalence. There are no patient dates or demographics here for trend or equity analysis. Measurements and labels alone cannot establish clinical safety or causation.

Original dataset and documentation (UCI) · scikit-learn dataset details · Download the five-column analysis extract

The extract includes a row ID and diagnosis label, plus the five measurements used on this page. Charts and counts run in your browser from that same extract embedded below; no external data service is needed.