Abstract
<title>Abstract</title> <p>Skin diseases impose a major global disease burden, yet dermatologists remain scarce, especially in low- and middle-income countries. Foundation models could extend specialist-level care, but leading systems depend on private cohorts that most institutions cannot access, reproduce, afford or audit. Here we present UniDerm, a dermatology foundation model trained entirely on publicly available image–text data through supervision-denoising contrastive learning. Public skin images each carry several heterogeneous texts, and standard contrastive learning averages them into one representation, letting each text add noise to the others. UniDerm instead uses natural-language instructions to disentangle these targets and recover the fine-grained supervision that averaging discards. Unlike private-cohort systems, UniDerm is released as a fully open system: its model weights, code and the data recipe needed to reconstruct its training corpus are all public. Across 11 benchmarks spanning five continents, eight countries and the full Fitzpatrick range, UniDerm matches or exceeds private-cohort models: it reaches their full-label diagnostic accuracy using roughly a third of the labels, and detects malignancy at an AUROC of 0.946, above a 302-reader expert consensus. Its accuracy holds across skin tones, including the darker-skin lesions on which such systems most often fail. Because the approach is not specific to skin, it charts a route to specialist-level medical AI from public data that under-resourced communities can build, audit and own rather than purchase.</p>