Are AI chatbots ready for chikungunya public education? Evidence on validity, reliability, and readability.

Publication date: May 11, 2026

Chikungunya continues to expand geographically, driving demand for trustworthy, easy-to-read public guidance. Conversational AI systems are increasingly used for health information, yet their medical validity, reliability, and readability remain uneven. We evaluated four widely used chatbots (ChatGPT, Claude, DeepSeek, Gemini) on two task sets: (1) validity on a 50-item single-answer MCQ dataset about Chikungunya; and (2) reliability and readability on 13 core public-education questions derived from Google Trends “topics” and clinician. Reliability was scored with DISCERN, EQIP, GQS, and JAMA benchmarks by clinician raters; readability used ARI, CL, FKGL, FRES, GFI, and SMOG. Across three independent runs, Deepseek-V3. 2 achieved the highest MCQ accuracy (86. 7%); other models ranged near 72-78% accuracy. In paired brand-line comparisons, newer iterations outperformed predecessors: ChatGPT-5 vs. ChatGPT-4o showed a 5. 3-percentage-point gain, and Deepseek-V3. 2 vs. Deepseek-R1 showed an 8. 7-point gain in accuracy. Reliability (DISCERN, EQIP, GQS, and JAMA) differed significantly across models (p 

Open Access PDF

Concepts Keywords
Chatbots Artificial intelligence
Gemini Chatbots
Google Chikungunya
Outperformed Information quality
Readability

Semantics

Type Source Name
pathway REACTOME Reproduction
disease MESH included

Original Article

(Visited 12 times, 1 visits today)

Leave a Comment

Your email address will not be published. Required fields are marked *