Abstract
<title>Abstract</title> <p>Rule-based supervision can reduce unsafe clinical LLM outputs, but residual safety depends on deployment protocol and base model. We evaluated a five-gate sentinel across five models (Qwen 7B, Llama 8B, Meditron 7B, Gemma 4, Sonnet 4.5) on two clinical safety benchmarks under handoff- on and handoff-off protocols. Under handoff-on, all models achieved deploy violation rates at most 6%, though strict violation rates (counting blocked and escalated outputs) ranged from 1.5% to 49.5%, showing wide variation in how many outputs the sentinel blocks or escalates per model. Under handoff-off, residual risk stratified by base model (0.0–50%), driven by failure-mode profiles and repair conversion rate. A multisite extension (Llama 70B, N =2935 encounter-derived scenarios, three sites) confirmed 0.0% deploy violation rate under handoff-on (strict VR 2.0%) and 2.4% under handoff-off. Moreover, a separate test on HealthBench Hard (N =1000, an independent benchmark with different scenarios and grading) also reached 0.0% deploy VR under C2, and fine-tuning increased the allow rate sixfold when combined with supervision. An exploratory clinician-adjudication pilot (n=50) nonetheless distinguished automated safety-rule compliance from clinical adequacy: none of 20 intercepted outputs were rated harmful but all were flagged for missing critical information, while 5 of 30 allowed outputs were rated as potentially harmful, all involving clinical reasoning errors beyond the sentinel’s scope. Hence, deployment policy should reflect handoff availability, base-model failure mode, and the gap between automated compliance and clinician-judged adequacy.</p>