ⓘ Educational information only — not medical advice. Nothing on this page diagnoses or treats any condition. In an emergency, call your local emergency number.
Guide · Patient tools

AI Symptom Checkers: Accuracy, Risks, and When Not to Use Them

Symptom checkers have improved since the landmark 2015 audit — but "improved" is not the same as reliable. Here is what the studies actually found — and the situations where using one is genuinely dangerous.

What the landmark study found

The most-cited evaluation is the 2015 BMJ audit by Semigran and colleagues, which tested 23 free online symptom checkers against 45 standardized patient vignettes[1]. Two numbers get mixed up constantly, so read carefully:

34%
Correct diagnosis listed first
57%
Appropriate triage advice overall
33–78%
Triage accuracy range across tools

The "34%" figure is diagnosis-listed-first accuracy — not triage. Triage advice was appropriate 57% of the time overall, and performance varied by urgency: 80% for emergencies, 55% for non-urgent cases, and 33% for self-care situations[1]. In other words, checkers were reasonably good at saying "go to the hospital now" but poor at the calm end — and they tended to be over-cautious, sending people to care they didn't need. A 2022 systematic review in npj Digital Medicine reached the same qualitative conclusion: diagnostic accuracy is generally low, triage accuracy moderate and highly variable between tools[2]. Newer AI-based checkers are improving, but no independent evaluation we could verify shows them approaching clinician-level reliability.

When a symptom checker is useful — and when it is dangerous

A useful pattern: use a checker to prepare for a conversation with a clinician — organizing symptoms, duration and history — not to decide whether you need one. The dangerous pattern is the reverse: using a checker to talk yourself out of seeking care. In the BMJ audit, checkers failed to triage roughly 1 in 5 emergencies appropriately[1]. No consumer checker should be used to decide whether to skip emergency care.

When not to trust a symptom checker

Severe or rapidly worsening symptoms; chest pain or difficulty breathing; symptoms in a child; anything following a head injury; or any situation where you would otherwise seek urgent care. In these cases, contact emergency services or a clinician directly — do not ask an app first.

The contrast worth knowing: NHS 111 online in England is a clinician-governed national triage pathway for people aged 5 and over — it feeds into real services (999, A&E, GPs, pharmacists) with nurse callbacks prioritized by severity[3]. That is a fundamentally different category from an unregulated chatbot.

AI mental health apps: the evidence, honestly

The research base is small and early. The foundational Woebot study (2017) randomized 70 students to a CBT-based chatbot or an ebook for two weeks and found self-reported reductions in depression and anxiety symptoms in the chatbot group[4] — promising, but small, short, self-reported, and no substitute for a therapy trial. Wysa has a 2024 peer-reviewed RCT among people with chronic diseases, and received an FDA Breakthrough Device designation in May 2022 — a designation is a fast-track flag, not an approval[5]. The apps themselves are honest about their limits: Wysa states it "cannot diagnose, treat, monitor or cure" and is not for crisis use.

There is one regulated exception: Rejoyn, cleared by the FDA in April 2024 as the first prescription digital therapeutic — a six-week program used alongside clinician-managed outpatient care for adults with major depressive disorder, prescribed by professionals[6]. Even Rejoyn is an adjunct, not a replacement. And on privacy: Mozilla's 2022 review found most mental-health apps failed basic privacy and security standards, with data-sharing practices described as "terrible"[7] — read the privacy policy before typing anything personal.

If you are in crisis

An app is not the right tool. Contact a crisis helpline, a trusted person, or your local emergency services immediately. Crisis care is human care.

ChatGPT and medical questions

Chatbots can sound astonishingly medical. ChatGPT performed at or near the passing threshold on USMLE-style exam questions in a 2023 study[8], and in another study evaluators preferred its answers to physicians' 79% of the time, rating them more empathetic[9]. Read that again: those studies measured exam answers and perceived warmth, not clinical safety. Meanwhile, a 2023 study found that 55% of GPT-3.5's generated citations — and 18% of GPT-4's — were fabricated, with substantive errors in many of the real ones[10]. A tool that invents references should never be your source for health decisions. The CDC puts it plainly: treat generative AI as "a drafting and synthesis aid, not an authority," because it "can confidently produce incorrect or fabricated content"[11]. And legally, chatbots are not medical devices: regulators base that status on the manufacturer's intended purpose, which for ChatGPT is not medicine[12].

Frequently asked questions

How accurate are AI symptom checkers?+
Modest and variable. The landmark audit found correct diagnosis listed first 34% of the time and appropriate triage 57% overall, with tools ranging from 33% to 78% on triage[1]. A 2022 systematic review concluded accuracy is generally low for diagnosis and moderate for triage[2]. Newer AI tools are improving, but treat all of them as conversation prep, not decision-makers.
Should I use ChatGPT as a symptom checker?+
No, not for decisions about your health. ChatGPT is not a medical device[12], fabricates references at documented rates[10], and the CDC says generative AI should be treated as a drafting aid, not an authority[11]. For triage, use official pathways like NHS 111 (England) or contact your clinician.
Can a symptom checker diagnose me?+
No consumer checker is licensed to diagnose. They can list possibilities — correctly, in only about a third of cases in the landmark audit[1] — but a diagnosis is a clinical act that requires a qualified professional.
Do AI mental health apps actually work?+
The evidence is early and limited: Woebot's foundational study was 70 students over two weeks[4], and Wysa's 2024 RCT exists but is one study[5]. The apps state they cannot diagnose or treat and are not for crises. They can support well-being routines; they are not therapy.
Are mental health apps safe for my data?+
Often not, per independent audits: Mozilla's 2022 review found most mental-health apps failed basic privacy and security standards[7]. Check what an app collects, shares and sells before disclosing sensitive information.

Sources

All claims verified September 2026. Study designs are stated in the text.

  1. Semigran HL, et al. "Evaluation of symptom checkers for self diagnosis and triage: audit study." BMJ 2015;351:h3480: bmj.com (abstract verified via bmjchina.com.cn)
  2. Wallace W, et al. "The diagnostic and triage accuracy of digital and online symptom checker tools: a systematic review." npj Digital Medicine 2022: pubmed.ncbi.nlm.nih.gov
  3. NHS — "When to use NHS 111 online or call 111": nhs.uk
  4. Fitzpatrick KK, Darcy A, Vierhile M. Woebot RCT, JMIR Mental Health 2017;4(2):e19: mental.jmir.org
  5. Wysa chronic-disease RCT, JMIR Formative Research 2024;8:e50025: formative.jmir.org; FDA Breakthrough designation (May 2022): healio.com
  6. Otsuka — Rejoyn FDA clearance press release, Apr 2, 2024: otsuka.co.jp
  7. The Verge — Mozilla mental-health app privacy report, May 2, 2022: theverge.com
  8. Kung TH, et al. "Performance of ChatGPT on USMLE." PLOS Digital Health 2023;2(2):e0000198: plos.org
  9. Ayers JW, et al. JAMA Internal Medicine 2023;183(6):589–596 (perceived quality/empathy, not safety): jamanetwork.com
  10. Walters WH, Wilder EI. "Fabrication and errors in the bibliographic citations generated by ChatGPT." Scientific Reports 2023: nature.com
  11. CDC — "Considerations for Generative AI in Public Health," Mar 13, 2026: cdc.gov
  12. MHRA — "Software and AI as a medical device," updated Feb 3, 2025: gov.uk