Publication date: May 11, 2026
Chikungunya continues to expand geographically, driving demand for trustworthy, easy-to-read public guidance. Conversational AI systems are increasingly used for health information, yet their medical validity, reliability, and readability remain uneven. We evaluated four widely used chatbots (ChatGPT, Claude, DeepSeek, Gemini) on two task sets: (1) validity on a 50-item single-answer MCQ dataset about Chikungunya; and (2) reliability and readability on 13 core public-education questions derived from Google Trends “topics” and clinician. Reliability was scored with DISCERN, EQIP, GQS, and JAMA benchmarks by clinician raters; readability used ARI, CL, FKGL, FRES, GFI, and SMOG. Across three independent runs, Deepseek-V3. 2 achieved the highest MCQ accuracy (86. 7%); other models ranged near 72-78% accuracy. In paired brand-line comparisons, newer iterations outperformed predecessors: ChatGPT-5 vs. ChatGPT-4o showed a 5. 3-percentage-point gain, and Deepseek-V3. 2 vs. Deepseek-R1 showed an 8. 7-point gain in accuracy. Reliability (DISCERN, EQIP, GQS, and JAMA) differed significantly across models (p
Open Access PDF
| Concepts | Keywords |
|---|---|
| Chatbots | Artificial intelligence |
| Gemini | Chatbots |
| Chikungunya | |
| Outperformed | Information quality |
| Readability |
Semantics
| Type | Source | Name |
|---|---|---|
| pathway | REACTOME | Reproduction |
| disease | MESH | included |