Abstract
<title>Abstract</title> <p>Background: Longitudinal wound photographs can yield segmentation-derived healing signals, but their interpretation depends on segmentation transfer, visit timing, image framing, and correct handling of repeated observations. We audited these dependencies rather than developing a clinical prediction tool. Methods: A legacy DeepLabV3-MobileNetV3 segmentation checkpoint was internally audited on 200 labeled FUSeg development-validation images. It was then treated as a fixed, target-domain-unvalidated measurement operator in a local WoundsDB subset comprising 46 cases, 77 visits, and 79 scenes. Threshold endpoints at 20%, 30%, and 50% occupancy reduction were subjected to single-visit deletion and schedule thinning in seven cases with at least three visits. In a Bayuan series comprising 10 patients, 62 patient-days, and 105 images, a compact convolutional neural network (CNN) was evaluated by leave-one-patient-out temporal ranking over five seeds and compared with segmentation-derived occupancy. Label-permutation and untrained-network controls and deterministic 0.8x and 1.2x whole-image zoom were evaluated. Patient-level estimates used equal weighting and 20,000 patient bootstrap samples. Results: FUSeg mean Dice was 0.715 (image-bootstrap 95% confidence interval 0.683 to 0.745), with per-image values ranging from 0 to 1. The local WoundsDB copy lacked the published expert masks, so target-domain segmentation error was not estimable. Deleting one visit changed event status in 1 of 15 scenarios at each threshold, but changed event status or crossing interval in 7 of 15 to 8 of 15 scenarios and affected 6 of 7 cases. Bayuan patient-equal concordance was 0.690 (95% confidence interval 0.580 to 0.804) for the CNN and 0.705 (0.503 to 0.865) for occupancy; the paired difference was -0.015 (-0.219 to 0.227). Permuted-label and untrained controls yielded 0.473 and 0.518. At 1.2x zoom, CNN concordance changed by -0.021 (-0.059 to 0.006), whereas occupancy concordance changed by 0.152 (0.031 to 0.329). Conclusions: The available data support a failure-mode audit, not calibrated prognosis. Sparse visits frequently altered interval localization, the CNN showed no demonstrated incremental value over segmentation-derived occupancy, and framing changed the apparent occupancy signal. Target-domain reference masks, standardized image scale, and larger independently evaluated patient cohorts are required for clinical performance claims.</p>