Abstract
<jats:p>This pilot study evaluates the feasibility of using ChatGPT-4 for automated preliminary clinical reporting in myocardial perfusion scintigraphy (MPS), a key non-invasive imaging modality for assessing myocardial ischemia and infarction. A comparative analysis was conducted using 30 consecutive de-identified MPS cases spanning a broad spectrum of clinical scenarios, where structured clinical data were used to generate AI-based reports, which were then compared with reports prepared by experienced nuclear medicine physicians. Reports were evaluated using four criteria: clinical accuracy, report structure, terminological appropriateness, and overall comprehensibility, scored on a 5-point Likert scale. Results demonstrated that ChatGPT-4 performed strongly in report structure, terminological appropriateness, and overall comprehensibility, consistently producing well-organized and coherent reports. However, it showed lower performance in clinical accuracy, particularly in complex cases requiring advanced interpretation, where outputs were occasionally superficial or lacked specificity. Statistical analysis indicated a significant difference in clinical accuracy compared to physician reports (p = 0.002, exploratory analysis, n=30). Inter-observer agreement between evaluating physicians was substantial (Cohen&#039;s κ = 0.78). In conclusion, ChatGPT-4 shows promise as a supportive tool for preliminary reporting and medical education, but its limitations in higher-order clinical reasoning necessitate careful human oversight in clinical practice. These findings should be considered preliminary and require validation in adequately powered studies before any clinical implementation.</jats:p>