The Archives of Bone and Joint Surgery

The Archives of Bone and Joint Surgery

Artificial Intelligence and Frequently Asked Questions Related to Basal Joint Arthroplasty: A Comparative Study of Multiple Chatbots

Document Type : RESEARCH PAPER

Authors
1 Rothman Orthopaedics
2 Rothman Orthopaedic Institute
3 Sidney Kimmel Medical College
10.22038/abjs.2026.85929.3925
Abstract
Title of Paper: Artificial Intelligence and Frequently Asked Questions Related to Basal Joint Arthroplasty: A Comparative Study of Multiple Chatbots



Objectives

This study sought to compare the readability and accuracy of answers generated from three natural language processing chatbots in response to commonly asked questions regarding basal joint arthroplasty (BJA).



Methods

Using a previously described method, ten frequently asked (FAQ) style questions were posed to ChatGPT 3.5 and 4.0 and Gemini. Readability was calculated based on Flesch Reading Ease Scores (FRES) and Flesch-Kincaid Grade Levels (FKGL). Four raters assigned a response accuracy score (RAS) between 1 and 4 to each response and ranked the chatbots. The latter of which was tested for levels of agreement. A qualitative analysis was performed to compare these chatbot generated responses. The ‘best response’ for each question was determined based on the lowest RAS in addition to their ranking.



Results

In terms of readability, there was a trend towards greater ease of comprehension for Gemini responses based on FRES and FKGL scores. The combined average RAS for all three chatbots was 2.21 (±0.36). The question that received the best responses (average RAS = 1.33) was related to the costs associated with BJA. In contrast, the question that posed the greatest challenge for the chatbots (average RAS = 3.17) was related to alternative procedures to BJA. Agreement levels were poor among the raters, ranging from no true agreement to weak agreement. The chatbot that was most frequently ranked 1st was ChatGPT 4.0 (n=6), in contrast to Gemini that received the most instances of a 3rd ranking (n=6).



Conclusion

There was an overall satisfactory level of response accuracy to FAQ style questions regarding BJA among the three chatbots. Moving forward, physicians must be prepared to assist patients in sorting out potential inaccuracies that are sometimes generated by these artificial intelligence platforms.



Keywords: artificial intelligence, arthritis, arthroplasty, basal joint, chatbot



Level of Evidence: V
Keywords
Subjects


Articles in Press, Accepted Manuscript
Available Online from 06 October 2026

  • Receive Date 27 February 2025
  • Revise Date 13 September 2026
  • Accept Date 14 September 2026