Thai University RankingsRESEARCH RADAR
← Back to research database
มีศักยภาพระดับโลก

Quality and readability of AI Chatbot responses to frequently asked questions from patients undergoing progressive collapsing foot deformity surgery: a comparative study of ChatGPT, Perplexity, and Gemini

IMPACT SIGNAL83/100
01

Information from the abstract

Background Patients increasingly turn to AI chatbots for medical information, including before complex orthopaedic procedures such as progressive collapsing foot deformity (PCFD) surgery. Whether these tools deliver content of sufficient quality and accessibility for preoperative patient education remains unclear, particularly across competing platforms. This study addressed three questions: (1) Do ChatGPT, Perplexity AI and Google Gemini differ in the accuracy, comprehensiveness and clarity of their responses to PCFD-related patient questions? (2) Do these platforms produce content meeting recommended readability thresholds for patient education? (3) Does the level of agreement among blinded foot and ankle surgeons rating the quality of chatbot responses vary depending on the platform used? Hypothesis The three AI chatbot platforms produce responses of comparable accuracy but differ significantly in readability, with none reaching the recommended readability thresholds for patient education materials. Patients and methods Cross-sectional comparative study. Twenty frequently asked questions regarding PCFD, covering disease understanding, conservative management, surgical planning and postoperative recovery, were submitted verbatim to ChatGPT (GPT-4o mini), Perplexity AI and Google Gemini (free versions, March 25, 2026). The 60 resulting responses were rated by three blinded foot and ankle surgeons on three 5-point Likert scales (accuracy, comprehensiveness, clarity). Readability was assessed using the Flesch Reading Ease (FRE) and Flesch-Kincaid Grade Level (FKGL). Inter-rater agreement used Kendall’s W; differences between platforms were analysed using Kruskal-Wallis tests with Bonferroni-corrected pairwise comparisons. Results All platforms produced responses rated accurate to very accurate. Perplexity achieved significantly higher accuracy than ChatGPT (p = 0.0003) and higher accuracy and clarity than Gemini (p = 0.0065 and p = 0.0075). No platform reached the recommended FRE ≥ 60 or FKGL ≤ 6 thresholds: median FKGL ranged from 13.1 (ChatGPT) to 21.4 (Perplexity), with ChatGPT producing the most readable and Perplexity the least readable content (p < 0.001). Inter-rater agreement was fair to substantial across platforms, lowest for Perplexity. Discussion AI chatbots produce generally accurate baseline information on PCFD surgery, with Perplexity showing significantly higher expert-rated accuracy and clarity than the other platforms—contrary to our hypothesis of comparable accuracy—while readability remains uniformly inadequate for all platforms, as hypothesized. These tools may serve as a supplementary source of information, but their inadequate readability suggests they are not yet suited to replace tailored, surgeon-led patient education. Level of evidence III; cross-sectional comparative study.

02

Why this record is monitored

This record has an Impact Signal of 83/100 based on recency, source, collaboration, and bibliographic signals. It prioritizes monitoring and is not a judgment of research quality.

Related topics: Artificial Intelligence in Healthcare and Education · Mobile Health and mHealth Applications · AI in Service Interactions

03

Thai researcher and institutional participation

Krit Rachayont · Thammasat University

04

Data limitations

This page is a bibliographic record based on abstract-level information, not a full analysis or quality assessment. Verify the DOI and original article before citation.