| Title |
ChatGPT is Less Accurate than Orthodontists and Faculty in Providing Orthodontic Advice |
| Clinical Question |
In patients seeking orthodontic advice, is Chat GPT more accurate than orthodontists and orthodontic residents in giving orthodontic advice? |
| Clinical Bottom Line |
ChatGPT demonstrated a relatively high level of accuracy and completeness when responding to orthodontic questions, with a median accuracy score of 4.9/6 and a completeness score of 2.4/3. While the AI showed strong potential, including the ability to address complex clinical cases, its answers were not consistently 100% accurate or complete. The validity of the evidence is moderate, as expert evaluations confirmed reliable performance in many scenarios, but also highlighted important limitations that prevent ChatGPT from replacing the expertise and judgment of orthodontists. |
| Best Evidence |
|
| PubMed ID |
Author / Year |
Patient Group |
Study type
(level of evidence) |
| 38337430 | Hatia/2024 | 21 clinical open-ended questions generated by 10 specialized orthodontists from 10 Italian postgraduate orthodontics schools | Cross-Sectional | | Key results | ChatGPT’s responses to orthodontic questions had a median accuracy of 4.9/6 and completeness of 2.4–2.5/3 across both open-ended and clinical case questions. Reviewers found that about 40%–46% of answers were entirely accurate and 50%–54% were entirely complete, showing the AI could handle even complex clinical scenarios. While the results indicate a moderately high level of reliability, the answers were not consistently perfect. The study concludes that ChatGPT shows promise as a supportive tool, but it is not yet advanced enough to replace the expertise of orthodontists. | | 39075513 | Dursun/2024 | 20 frequently-asked patient questions about clear aligners found via Google by three experienced orthodontists | Cross Sectional | | Key results | This study compared ChatGPT-3.5, ChatGPT-4, Gemini, and Copilot in answering frequently asked patient questions about clear aligners found via Google. Although ChatGPT-4 had the highest mean accuracy (Likert) score, the difference among chatbot models was not statistically significant (p > 0.05). However, Copilot demonstrated significantly higher reliability and quality (modified DISCERN and GQS; p < 0.05), while Gemini produced significantly more readable responses (FRES; p < 0.05). In terms of readability, Gemini had the highest Flesch Reading Ease Score (54.12 ± 10.27), indicating easier comprehension compared to ChatGPT-4 (43.88 ± 10.13), Copilot (41.72 ± 10.74), and ChatGPT-3.5 (38.39 ± 11.56), whose responses were considered difficult to read. There was no statistically significant difference between ChatGPT-4, Copilot and ChatGPT-3.5 (p > 0.05). | | 38195060 | Abu Arqub/2024 | 111 questions from predefined domains and subdomains generated by three orthodontists | Cross Sectional | | Key results | The authors posed 111 clear aligner–related questions to ChatGPT and had five orthodontists rate their accuracy using a four-point scale. The mean accuracy score was 2.6 ± 1.1. They classified the AI answers as 58% “objectively true,” 18% “selected facts,” 9% “minimal facts,” and 15% “false.” The study noted several false claims in the AI’s answers (for example, suggesting aligners could reduce need for orthognathic surgery, improve airway function, etc.) and criticized the lack of citations and omissions of relevant information. | |
| Evidence Search |
(“ChatGPT” OR “artificial intelligence” OR “large language model”) AND (“orthodontics” OR “dentistry” OR “dental advice”) AND (“accuracy” OR “information quality” OR “patient education”) |
Comments on
The Evidence |
The studies available are recent, small in sample size, and primarily cross-sectional. Validity is moderate: most studies blinded faculty reviewers to the source of answers, but the assessment criteria varied. Importance is limited by the lack of randomized or prospective studies. Nonetheless, all reviewed evidence consistently shows ChatGPT provides moderately accurate orthodontic information but falls short in clinical specificity, risk assessment, and treatment planning. |
| Applicability |
The evidence is applicable to patients seeking orthodontic information online. ChatGPT may be a helpful adjunct for general knowledge but cannot replace orthodontists or faculty when patient-specific advice is needed. The results generalize best to English-speaking patients with internet access. Clinical decision-making, cost estimation, and treatment planning remain inappropriate uses for ChatGPT. Finally, the results will not directly extend to low-literacy groups, non-English speakers, or people without web access. |
| Specialty |
(Orthodontics) |
| Keywords |
ChatGPT; artificial intelligence; large language model; orthodontics; dentistry; clear aligners; patient education; information accuracy; information quality; readability; reliability; cross-sectional comparative study
|
| ID# |
3591 |
| Date of submission |
10/15/2025 |
| E-mail |
reesel@livemail.uthscsa.edu |
| Author |
Logan Reese |
| Co-author(s) |
|
| Co-author(s) e-mail |
|
| Faculty mentor |
Shaza Abass |
| Faculty mentor e-mail |
abass@uthscsa.edu |
| |
|
Basic Science Rationale
(Mechanisms that may account for and/or explain the clinical question, i.e. is the answer to the clinical question consistent with basic biological, physical and/or behavioral science principles, laws and research?) |
| None available | |
 |
Comments and Evidence-Based Updates on the CAT
(FOR PRACTICING DENTISTS', FACULTY, RESIDENTS and/or STUDENTS COMMENTS ON PUBLISHED CATs) |
| None available | |