Abstract
<title>Abstract</title> <p>Fitzpatrick-17k is widely used to develop and audit dermatology artificial intelligence. Groh et al. released two crowd-consensus annotation sources (Scale AI and Centaur Labs) and reported that 85% of consensus labels were within one Fitzpatrick unit of each other. We reanalyzed the annotations using chance-corrected statistics. Among 15,230 images with two valid labels, within-one-unit agreement was 91.0%, whereas exact agreement was 47.9%. Cohen's kappa was 0.351 unweighted, 0.610 linear-weighted, and 0.786 quadratic-weighted. An exploratory analysis treating labels within one unit as agreement yielded a chance-corrected kappa of 0.808. Row-wise exact agreement, with Scale AI labels defining strata, was 87.6% for type I and 34.4%, 41.5%, 36.2%, 44.8%, and 55.8% for types II through VI. Among images either source assigned type V or VI, exact agreement was 40.0%. Centaur labels were lighter than Scale AI labels for 6,059 images and darker for 1,878. Type VI contained 61 malignant images under Scale AI labels and 45 under Centaur labels, yielding illustrative 95% Wilson intervals of 74.3%-92.0% and 71.2%-92.3%, respectively, around an assumed sensitivity of 0.85. For combined types I-II, the corresponding intervals were 82.9%-86.9% and 83.1%-86.8%. Fitzpatrick-stratified results on this benchmark are therefore sensitive to the annotation source, agreement criterion, and stratum size. Studies should identify the annotation source, define agreement and weighting, and provide per-stratum counts with confidence intervals. The preregistered protocol and primary analysis code are public.</p>