Abstract
<title>Abstract</title> <p>The practical study of digital abuse is severely restricted by a "Forensic Data Gap": privacy regulations rightfully protect users but prevent researchers from accessing real-world logs of coercive control. To overcome this, we developed a privacy-preserving framework that generates synthetic, clinically accurate abusive conversations. By anchoring prompts to DSM-5 criteria and embedding temporal tokens to simulate manipulative pacing, we created a highly realistic dataset. Expert clinical assessment confirmed its fidelity ($\kappa = 0.84$) and statistical tests revealed low discriminability between synthetic and real human dialogues. We trained a DistilBERT sequence model exclusively on this synthetic data to flag pattern-level risk. It successfully detected coercive control (F1-score of 0.75, AUC of 0.82), significantly outperforming standard lexical baselines. Crucially, the model successfully generalised to real-world benchmarks, achieving an F1-score of 0.72 on curated UK legal evidence and 0.65 on Reddit data, while remaining robust against generator-specific biases. This study demonstrates that clinically constrained synthetic data can safely bypass privacy bottlenecks, providing digital forensic investigators with a reliable, human-in-the-loop triage tool to detect psychological manipulation.</p>