ORAL HEALTH EVIDENCE-BASED PRACTICE PROGRAM
View the CAT printer-friendly / share this CAT
spacer
Title ChatGPT is Less Accurate than Orthodontists and Faculty in Providing Orthodontic Advice
Clinical Question In patients seeking orthodontic advice, is Chat GPT more accurate than orthodontists and orthodontic residents in giving orthodontic advice?
Clinical Bottom Line ChatGPT demonstrated a relatively high level of accuracy and completeness when responding to orthodontic questions, with a median accuracy score of 4.9/6 and a completeness score of 2.4/3. While the AI showed strong potential, including the ability to address complex clinical cases, its answers were not consistently 100% accurate or complete. The validity of the evidence is moderate, as expert evaluations confirmed reliable performance in many scenarios, but also highlighted important limitations that prevent ChatGPT from replacing the expertise and judgment of orthodontists.
Best Evidence (you may view more info by clicking on the PubMed ID link)
PubMed ID Author / Year Patient Group Study type
(level of evidence)
#1) 38337430Hatia/202421 clinical open-ended questions generated by 10 specialized orthodontists from 10 Italian postgraduate orthodontics schoolsCross-Sectional
Key resultsChatGPT’s responses to orthodontic questions had a median accuracy of 4.9/6 and completeness of 2.4–2.5/3 across both open-ended and clinical case questions. Reviewers found that about 40%–46% of answers were entirely accurate and 50%–54% were entirely complete, showing the AI could handle even complex clinical scenarios. While the results indicate a moderately high level of reliability, the answers were not consistently perfect. The study concludes that ChatGPT shows promise as a supportive tool, but it is not yet advanced enough to replace the expertise of orthodontists.
#2) 39075513Dursun/202420 frequently-asked patient questions about clear aligners found via Google by three experienced orthodontists Cross Sectional
Key resultsThis study compared ChatGPT-3.5, ChatGPT-4, Gemini, and Copilot in answering frequently asked patient questions about clear aligners found via Google. Although ChatGPT-4 had the highest mean accuracy (Likert) score, the difference among chatbot models was not statistically significant (p > 0.05). However, Copilot demonstrated significantly higher reliability and quality (modified DISCERN and GQS; p < 0.05), while Gemini produced significantly more readable responses (FRES; p < 0.05). In terms of readability, Gemini had the highest Flesch Reading Ease Score (54.12 ± 10.27), indicating easier comprehension compared to ChatGPT-4 (43.88 ± 10.13), Copilot (41.72 ± 10.74), and ChatGPT-3.5 (38.39 ± 11.56), whose responses were considered difficult to read. There was no statistically significant difference between ChatGPT-4, Copilot and ChatGPT-3.5 (p > 0.05).
#3) 38195060Abu Arqub/2024111 questions from predefined domains and subdomains generated by three orthodontists Cross Sectional
Key resultsThe authors posed 111 clear aligner–related questions to ChatGPT and had five orthodontists rate their accuracy using a four-point scale. The mean accuracy score was 2.6 ± 1.1. They classified the AI answers as 58% “objectively true,” 18% “selected facts,” 9% “minimal facts,” and 15% “false.” The study noted several false claims in the AI’s answers (for example, suggesting aligners could reduce need for orthognathic surgery, improve airway function, etc.) and criticized the lack of citations and omissions of relevant information.
Evidence Search (“ChatGPT” OR “artificial intelligence” OR “large language model”) AND (“orthodontics” OR “dentistry” OR “dental advice”) AND (“accuracy” OR “information quality” OR “patient education”)
Comments on
The Evidence
The studies available are recent, small in sample size, and primarily cross-sectional. Validity is moderate: most studies blinded faculty reviewers to the source of answers, but the assessment criteria varied. Importance is limited by the lack of randomized or prospective studies. Nonetheless, all reviewed evidence consistently shows ChatGPT provides moderately accurate orthodontic information but falls short in clinical specificity, risk assessment, and treatment planning.
Applicability The evidence is applicable to patients seeking orthodontic information online. ChatGPT may be a helpful adjunct for general knowledge but cannot replace orthodontists or faculty when patient-specific advice is needed. The results generalize best to English-speaking patients with internet access. Clinical decision-making, cost estimation, and treatment planning remain inappropriate uses for ChatGPT. Finally, the results will not directly extend to low-literacy groups, non-English speakers, or people without web access.
Specialty/Discipline (Orthodontics)
Keywords ChatGPT; artificial intelligence; large language model; orthodontics; dentistry; clear aligners; patient education; information accuracy; information quality; readability; reliability; cross-sectional comparative study
ID# 3591
Date of submission: 10/15/2025spacer
E-mail reesel@livemail.uthscsa.edu
Author Logan Reese
Co-author(s)
Co-author(s) e-mail
Faculty mentor/Co-author Shaza Abass
Faculty mentor/Co-author e-mail abass@uthscsa.edu
Basic Science Rationale
(Mechanisms that may account for and/or explain the clinical question, i.e. is the answer to the clinical question consistent with basic biological, physical and/or behavioral science principles, laws and research?)
post a rationale
None available
spacer
Comments and Evidence-Based Updates on the CAT
(FOR PRACTICING DENTISTS', FACULTY, RESIDENTS and/or STUDENTS COMMENTS ON PUBLISHED CATs)
post a comment
None available
spacer

Return to Found CATs list