Evaluating ChatGPT as a Patient Education Tool: Insights on Quality, Readability, and Reliability for Trigger Finger FAQs
Medical records-international medical journal (Online), no.1, pp.1-6, 2026 (TRDizin)
- Publication Type: Article / Article
- Volume: Issue: 1
- Publication Date: 2026
- Doi Number: 10.37990/medr.1730013
- Journal Name: Medical records-international medical journal (Online)
- Journal Indexes: TR DİZİN (ULAKBİM)
- Page Numbers: pp.1-6
- Open Archive Collection: AVESIS Open Access Collection
- Dokuz Eylül University Affiliated: Yes
Abstract
Aim: Trigger finger (TF), or stenosing tenosynovitis, causes pain, snapping, and finger locking. It greatly affects patients’ quality of life, prompting frequent inquiries to healthcare providers. ChatGPT, an AI language model, has gained popularity as a tool for patient education. This study evaluated the quality and readability of ChatGPT’s responses to common TF FAQs. Material and Methods: A set of FAQs regarding TF was developed based on reputable sources such as WebMD, Mayo Clinic, and NHS Trusts. Two experienced surgeons reviewed and refined the questions before submitting them to ChatGPT-4 for response generation. The quality of the responses was evaluated using the Global Quality Score (GQS) and DISCERN scale, while readability was assessed using the Flesch Reading Ease Score (FRES) and Flesch–Kincaid Grade Level (FKGL). Inter-rater reliability was determined using Cohen’s Kappa. Results: Overall quality was generally acceptable. The mean GQS was 3.8, with mean scores ranging from 3.0 to 4.5 across questions. The mean DISCERN score was 39.83 ± 8.43 (range 30.0–53.0), indicating overall fair quality, although some responses reached the good range. Inter-rater agreement was high (Cohen’s Kappa= 0.91). Readability was frequently suboptimal for patient education. The mean FRES was 39.98 (range 6.93–78.03), and only two responses met the recommended threshold of ≥60. The mean FKGL was 12.08 (range 6.43–18.30), and all responses except one exceeded the commonly recommended patient education level of grade 8. Conclusion: ChatGPT-4 produced generally acceptable trigger finger answers by GQS, but overall reliability (DISCERN) and read- ability (FRES/FKGL) were often inadequate for patient education. Clinician oversight and plain-language optimization are necessary before such outputs are used.