Abstract
<title>Abstract</title> <p>Objective To compare the information quality, source transparency, educational value, and readability of responses generated by four widely used artificial intelligence chatbots to patient questions about EPI and determine whether the content met the sixth-grade reading level recommended for patient education materials. Methods A web-based cross-sectional comparative design was used. Terms of global interest over the previous 5 years were identified through Medical Subject Headings and Google Trends. After deduplication and assessment of applicability, 20 core questions were developed. Each question was entered into ChatGPT 5.6, Copilot, Gemini 3.6 Flash, and Perplexity using a standardized procedure. Response quality was assessed with DISCERN, EQIP, the JAMA benchmarks, and the Global Quality Score (GQS). Readability was assessed with the Automated Readability Index (ARI), Gunning Fog Index (GFI), Flesch-Kincaid Grade Level (FKGL), Coleman-Liau Index, Simple Measure of Gobbledygook (SMOG), and Flesch Reading Ease Score (FRES). Quality scores are presented as medians and interquartile ranges. Models were compared with the Kruskal-Wallis test, and ε² was reported as the effect size. Readability results were compared with sixth-grade thresholds. Results A total of 80 responses were obtained. No overall differences among the four models were statistically significant for DISCERN, EQIP, JAMA, or GQS scores, with P values of 0.187, 0.529, 0.392, and 0.467, respectively. Median DISCERN scores ranged from 35.00 to 38.50, indicating poor treatment information quality or scores near the boundary between poor and fair quality. Median EQIP scores ranged from 52.50 to 55.00, indicating good quality with minor problems. The median GQS was 4.00 for all models, indicating favorable organization and usefulness. The median JAMA score was 0 for all models, as source attribution, authorship information, disclosures, and update dates were generally absent. Median values for all grade-level readability measures were substantially higher than 6, and median FRES values ranged from 13.04 to 30.76, all below 80. Perplexity and Copilot were easier to read on most measures, whereas Gemini had the greatest linguistic complexity. None of the four models met the recommended reading level. Conclusions The four artificial intelligence chatbots generated well-structured and broadly usable information about EPI. Treatment information quality and source transparency remained limited, and linguistic complexity substantially exceeded the recommended level for general patient education materials. These responses are currently best used as supplementary information after professional verification. Developers should provide authoritative citations, indicate the currency of guidelines, generate plain-language content, and adapt responses to patients' health literacy.</p>