Top 10 AI Tools for Clinicians at the Point of Care (2026)

September 01, 2026

Reviewed

10 AI Tools for Clinicians at the Point of Care (2026)

More than twenty AI tools now answer a clinician's natural-language question at the point of care, returning a cited clinical answer in seconds. The peer-reviewed evidence on what those answers are worth is less flattering than the launch posts: unaugmented models fabricate between 18% and 69% of the references they cite (PMID 37679503; PMID 40206627), and in a randomised trial of 50 physicians, GPT-4 access did not improve diagnostic reasoning at all (76% vs 74%, p = 0.60; PMID 39466245). This is the map of the category, what the literature actually supports, and the one job none of these tools claims to do.

What an AI point-of-care tool for clinicians actually does

Strip the positioning and the category has one shape: a clinician asks a question in ordinary language and gets a short, cited answer at or near the bedside, in seconds rather than in a search session. Fourteen functionalities recur across the products surveyed:

  1. Natural-language question to synthesised answer in seconds, at the point of care.
  2. A corpus wider than journals: clinical practice guidelines, drug formularies, interaction databases, care pathways. Vera Health states 60M+ papers, guidelines and real-world care pathways; Elna lists guidelines, open-access research, systematic reviews, formularies and interaction databases.
  3. Statement-level citation linking back to the original source.
  4. Evidence grading attached to the answer. Only two tools in the surveyed set publish a formal grading system: Vera Health uses its own, DynaMed applies GRADE.
  5. Conversational follow-up.
  6. Glanceable output: tables, care pathways, structured summaries.
  7. Clinical calculators. Vera Health states 900+.
  8. Mobile-first delivery, including native iOS and Android apps.
  9. Credential verification as an access gate. OpenEvidence requires a US NPI number; Vera Health requires professional verification.
  10. Compliance posture: HIPAA, GDPR, SOC 2, NIST AI RMF, PHI handling.
  11. EHR integration, which is where seat-based enterprise licensing comes from.
  12. Clinical governance and society endorsement. Vera Health operates a named international Clinical Council and an ACEP partnership.
  13. Published benchmark scores, almost always self-reported.
  14. Free access for the verified clinician, funded by advertising, an existing network business, or enterprise seats.

Notice what is missing from that list: source integrity screening. Not one product in the category publishes a policy for excluding retracted papers or hijacked journals from the evidence it synthesises.

That gap defines a second job, and it is not the one these products are built for. A point-of-care engine answers the clinical question. Something else has to answer whether the literature behind that recommendation still stands: whether the trial was retracted, whether the cited passage says what the answer claims, whether the effect replicated. BioSkepsis does that second job, on the research side rather than at the bedside, and it is deliberately absent from the tables below because it is not a point-of-care tool. This post comes back to the division of labour once the category is laid out.

The AI-native clinical answer engines competing for physicians in 2026

These are the direct equivalents: products built from scratch to answer clinical questions, rather than reference libraries with a chat box added.

The 10 leading AI-native clinician answer engines, September 2026. Figures are vendor-reported unless stated.
Tool Market Access model What it is built around
OpenEvidence US only since April 2026 Free to verified clinicians, advertising-funded The category's scale player: 757,000+ verified physicians claimed, 20M+ monthly queries in January 2026, NEJM and JAMA content partnerships, PHI upload supported. Raised roughly $700M across four rounds in twelve months to a stated $12B valuation in January 2026.
Vera Health Global Free to licensed clinicians and students, verification required Retrieval across a stated 60M+ papers, guidelines and care pathways, plus 900+ clinical calculators. Publishes its own evidence grading, an international Clinical Council, an ACEP partnership, and HIPAA, GDPR, SOC 2 and NIST AI RMF posture.
Doximity Ask US Free to Doximity account holders Built on Pathway Medical technology, acquired for $63M in cash and equity in 2025 and renamed from DoxGPT in May 2026. Doximity reaches over 80% of US physicians and funds the free tier from an already profitable network business.
Elna Not stated Not stated Mobile-optimised web app, native apps stated as forthcoming; guidelines, open-access research, systematic reviews, formularies, interaction databases, visual pathways. No corpus size, user count, benchmark or compliance certification published.
Dr.Oracle US, app-store distribution Paid consumer subscription Cited answers from guidelines, research, FDA labels and case reports, with general and research modes. Positions explicitly as physician-owned and not pharma-funded.
Heidi Evidence Heidi markets Freemium, paid Evidence Plus tier Scribe-first company extending into evidence; launched February 2026 with close to two million queries reported since. Source Control lets an organisation add its own governance documents.
AMBOSS AI Mode (LiSA) DACH-strong, global Subscription Natural-language querying over a curated, physician-authored knowledge base with inline citations and stated limitations. A curated-content trust model rather than open retrieval.
iatroX UK-first, plus US, CA, AU exam markets Individual all-access tier Guideline-first conversational support with NICE and international guidance, structured reasoning, medicines, calculators, Q-banks, CPD and native apps.
Medwise AI NHS Enterprise only Integrates local Trust policies and formularies alongside national guidance. The only product in the set with an HRA-listed prospective pilot comparing AI search against manual intranet guideline search.
Glass Health US Free tier upward Encounter-native: listens to the consultation and returns real-time insights, a three-tier differential, assessment and plan, and cited clinical Q&A.

Also in the category, at smaller scale. Praxis Medicine (UK-first, backed by Balderton and Creandum, sourcing NICE, NICE CKS, NHS Digital and Europe PMC), Umbil (UK: NICE, CKS, SIGN and BNF, plus referral letters, SBAR and discharge summaries), ClariMed (Germany-specific guideline search across AWMF, NVL and S3, hospital licensing), DR.INFO (referenced answers and drug information across Europe, with an explicitly bounded library-tool posture), EvidenceHunt (Amsterdam, PICO and study-detail extraction, organisational sources alongside external evidence, 25,000 active users claimed) and Tali Medical Search (Canada, in-flow search bundled with a documentation product). Several are the right answer in their own jurisdiction, which is the point of the next section.

The April 2026 European discontinuity

OpenEvidence withdrew from the European Union and the United Kingdom in April 2026, citing regulatory uncertainty over the treatment of AI systems in those jurisdictions, including the EU Artificial Intelligence Act. A European clinician cannot use the category's largest product, and the European vendors, iatroX, Medwise AI, Praxis Medicine, ClariMed and Umbil, have been positioning against that gap ever since. Where you practise now determines which evidence engine you are allowed to consult.

Reference incumbents with an AI answer layer: UpToDate, DynaMed and ClinicalKey AI

The established clinical reference platforms are a separate structural category. They are editorially authored, paid, and mostly institutional, and they added generative answers later than the AI-native products did.

Incumbent clinical reference platforms with a generative answer layer
Platform Owner Access Distinguishing feature
UpToDate Expert AI Wolters Kluwer Paid, no general free tier Launched September 2025; answers drawn only from UpToDate's expert-authored content, with sources and rationale shown, EHR integration, and CME inside the Expert AI workflow as of March 2026.
DynaMed / DynaMedex EBSCO Paid Applies GRADE. One of only two products in the whole surveyed set with a formal, externally recognised evidence grading system.
ClinicalKey AI Elsevier Paid, institutional Reference incumbent with an AI answer layer over Elsevier's clinical content.
AMBOSS library AMBOSS Subscription Reference-first knowledge base the clinician searches manually, alongside the LiSA AI mode.

The trade is legible. Incumbents restrict the answer to content they commissioned and can defend, which caps hallucination but also caps recall; a question the editorial team has not written about returns nothing useful. AI-native products retrieve over a much larger surface and inherit everything wrong with that surface, including the retracted papers still sitting in it. Neither posture is free.

What the peer-reviewed evidence says about LLM diagnostic accuracy in clinical practice

Benchmark scores and clinical performance are not the same measurement, and the gap between them is the single most important finding in this literature.

On static examinations, models do well. Scaled models including Flan-PaLM and Med-PaLM reached 67.6% and above on MedQA and MultiMedQA (PMID 37438534). On curated cases, they are competitive with specialists: in rheumatology case assessments, ChatGPT-4 gave the correct top diagnosis in 35% of cases against 39% for rheumatologists (p = 0.30), and had the correct diagnosis in its top three in 60% against 55% (p = 0.38) (PMID 37742280). In emergency department triage, GPT-4 scored 1.76 of 2 points against 1.59 for resident physicians (p = 0.01) (PMID 38976865).

Put the same models into a workflow that resembles clinical work and performance collapses. Across 2,400 real patient encounters from MIMIC-IV covering appendicitis, cholecystitis, diverticulitis and pancreatitis, leading open-access models achieved mean diagnostic accuracies of 45.5% to 54.9% when required to request examinations, labs and imaging themselves (PMID 38965432). Given all the information upfront, they reached 58.8% to 67.8%, against 87.5% to 92.5% for the clinicians assessed on the same cases (p < 0.001) (PMID 38965432). On 1,000 unselected emergency department visits, GPT-4-turbo underperformed residents on admission recommendations (0.43 to 0.58 vs 0.83) and radiology requests (0.74 vs 0.79), while beating them on antibiotic prescribing status (0.82 to 0.83 vs 0.78) (PMID 39379357).

The failure mode nobody puts in a launch post: numbers

Given explicit reference ranges, the same models classified abnormally high laboratory results correctly only 24.1% to 50.1% of the time, and abnormally low results 26.5% to 70.2% of the time (PMID 38965432). In the same study they failed to consistently recommend emergent colectomy for perforated diverticulitis or surgical drainage for infected pancreatic necrosis, and omitted antibiotic coverage for appendicitis and colonoscopy surveillance after diverticulitis. A tool that reads a creatinine or a lactate wrongly is not a documentation inconvenience.

Prompt sensitivity is a safety property, not a usability quirk

Asking for the "main diagnosis" rather than the "final diagnosis" moved accuracy by between +8.7% and -10.6% on identical MIMIC-IV cases. Reordering identical physical, laboratory and imaging findings shifted diagnoses by up to 18.0%, and models sometimes performed worse with the complete clinical picture than with a single test (PMID 38965432). Two clinicians asking the same clinical question in slightly different words can receive materially different answers, and neither will see any signal that this happened.

Guideline concordance shows the same pattern. Across 104 prompts on breast, prostate and lung cancer, 34.3% of GPT-3.5 outputs recommending treatment contained at least one recommendation non-concordant with NCCN guidelines, and 12.5% recommended entirely hallucinated treatments, including localised therapy for advanced disease (PMID 37615976). Against AAOS osteoarthritis guidance, web GPT-4 reached 62.9% overall concordance, ranging from 30% to 77.5% depending on recommendation strength and prompting (PMID 38378899). In ten precision-oncology cases put to a molecular tumour board, unaugmented models scored F1 between 0.04 and 0.19 against expert recommendations, and 27 of 85 clinical trial identifiers ChatGPT produced did not exist (PMID 37976064).

Does clinical AI change physician decisions? The randomised trial evidence

Two randomised trials give the clearest read available, and they disagree in an informative way.

In the first, 50 attending and resident physicians worked six complex diagnostic vignettes. Access to GPT-4 alongside conventional resources produced no significant improvement in diagnostic reasoning: median 76% against 74%, adjusted difference 2 percentage points (95% CI -4 to 8, p = 0.60), with no significant change in case time (519 s vs 565 s, p = 0.20). GPT-4 working alone scored 92%, sixteen points above the control group (p = 0.03) (PMID 39466245). The model was better than the physicians it was given to, and handing it over changed nothing.

In the second, 92 physicians produced 400 case-responses across five complex management vignettes. Here GPT-4 access did improve management reasoning: 43.0% against 35.7%, mean difference 6.5 percentage points (95% CI 2.7 to 10.2, p < 0.001). It also cost time, an adjusted 119.3 s more per case (95% CI 17.4 to 221.2, p = 0.022), and did not change the rate of severe clinical harm (7.6% vs 7.5%). The authors' reading is that consulting the model functioned as a structured pause that prompted deeper reflection on patient and contextual factors (PMID 39910272).

Note what that implies for a point-of-care product. The measured benefit came from slowing down and reasoning, not from receiving an answer quickly, which sits awkwardly with a category whose central promise is seconds. It does not mean fast answers are worthless. It means the evidence for "AI made the clinician better" is currently strongest in exactly the mode the category is optimising away from, and that seconds-to-answer should not be reported as a clinical outcome.

Both trials also expose the oversight problem. Verifying a confident, fluent, well-formatted answer is real cognitive work, and clinicians are not reliably compensated with time to do it. The field's own recommendation is prospective trials with standardised reporting, such as the QUEST evaluation principles, before semi-autonomous models are embedded in point-of-care workflows (PMID 39333376). Medwise AI's HRA-listed prospective pilot is, as far as this survey found, the only such study running against a live product in this category.

Which biomedical AI tool fits which clinician or life-science researcher

OpenEvidence or Doximity AskUS clinicians answering questions at the bedside

Free at the point of use, NPI-gated, and integrated with the journal content and the professional network US physicians already use. OpenEvidence supports PHI upload; Doximity Ask reaches physicians inside an existing workflow. Neither is available to you in the EU or UK.

Vera Health, iatroX, Medwise AI, Praxis or ClariMedClinicians in Europe and the UK after the April 2026 withdrawal

Vera Health is global, free to verified clinicians and publishes the fullest compliance posture in the category. iatroX, Praxis and Umbil are NICE and CKS-first for UK practice. Medwise AI is the option when local Trust policy and formulary have to sit alongside national guidance, and it is running the category's only registered prospective comparison. ClariMed covers AWMF, NVL and S3 for German practice.

UpToDate Expert AI or DynaMedClinicians who want a graded, editorially defended answer

DynaMed applies GRADE; UpToDate answers only from content its own experts wrote and carries CME in the workflow. You pay for it, and recall is bounded by what the editorial team has covered. In exchange, the provenance question has a short answer.

BioSkepsisClinician-researchers and life-science researchers appraising the evidence itself

Different job, different tool. BioSkepsis reads full text across 40 million or more curated papers, analyses up to 1,000 retrieved papers per run, blocks retracted and hijacked-journal sources before weighing evidence, and scores every brief on a Trust Index across provenance, grounding, evidence strength, reproducibility, durability and reagent validity. It answers in research-session time, not seconds. Use it when the question is whether the literature behind a recommendation actually holds, not what the dose is.

Where BioSkepsis fits: appraising the biomedical literature behind a clinical recommendation

Every product mapped above answers the clinical question. None of them answers the question behind it: does the literature this recommendation rests on still hold. That is a different job, on a different clock, and it is the one BioSkepsis was built for.

Three failure modes make that second job necessary rather than optional. A cited paper can be retracted and still be cited for years afterwards. A citation can resolve to a real paper that does not support the sentence attached to it. And a finding can rest on a single group that nobody has replicated. A point-of-care engine optimised for seconds is not positioned to catch any of the three, and none of the ten publishes a policy that says it tries.

What BioSkepsis does instead: retrieval first, over 40 million or more curated biomedical papers, 1931 to present, updated weekly, read as full text including methods, controls and supplementary data rather than as abstracts. Retrieval is biology-native, keyed on Gene Ontology, MeSH terms and gene identifiers rather than raw text similarity. A single research session analyses up to 1,000 retrieved papers and shortlists up to 70 Sources. Every claim links back to the exact passage it came from, and where the evidence is insufficient the answer says so rather than filling the gap.

A general-purpose model asked the same question inverts that order. It composes the answer first and assembles a plausible reference list afterwards, which is why 55% of GPT-3.5 and 18% of GPT-4 citations in one controlled evaluation were fabricated outright, with a further 43% and 24% of the real ones carrying substantive bibliographic errors (PMID 37679503). Retrieving before writing is not a refinement of that approach. It is a different architecture.

Two mechanisms sit on top of retrieval. Integrity screening runs on every session: retracted papers and hijacked journals are blocked before evidence is weighed, not flagged after the answer is written, and corrections on cited papers are surfaced. The Trust Index scores each brief from 0 to 100 across six weighted facets, provenance, grounding, evidence strength, reproducibility, durability and reagent validity, alongside a citation grammar that marks each claim Direct, Derived or Indirect. BioSkepsis does not claim zero hallucinations; no retrieval-grounded system can. It claims that you can see where every sentence came from and how far to trust it.

Checking a point-of-care answer, in practice

A clinical engine returns a recommendation citing three trials. In BioSkepsis the appraisal runs in three steps: ask the same clinical question as a research question and let retrieval pull the full-text literature; read the Trust Index to see whether the answer rests on replicated work or on one group; open the citation grammar to check whether each supporting claim is Direct, Derived or Indirect. If one of the three trials was retracted, it never entered the synthesis. If the effect rests on a single laboratory, the reproducibility facet says so before you build on it.

The output is shaped for that job rather than for the bedside: a research brief and notebook, mechanistic link tables, a landscape narrative, and a connections graph over the citation network. Findings export as PDF, DOCX, Markdown or JSON, references in APA, Chicago, Harvard, Vancouver, BibTeX, RIS, JSON or CSV, with direct Zotero sync and LibKey resolution to institutional full text. Threads can be shared privately or published, and cloned by whoever receives them. For institutions, the Organization tier adds dedicated or private cloud, SSO, GDPR and HIPAA, and data isolation.

Point-of-care category norm compared with BioSkepsis
Dimension Point-of-care category BioSkepsis
User Practising clinician at the bedside Biomedical and life-science researcher
Corpus Literature plus guidelines, drug labels, formularies, interaction databases, care pathways 40M+ curated papers, 1931 to present, weekly updates, journal literature and preprints, full text rather than abstracts
Retrieval unit Ranked passages assembled into a short answer Full-text analysis of up to 1,000 retrieved papers, up to 70 shortlisted sources
Answer form Short clinical answer, tables, care pathways Research brief and notebook, mechanistic link tables, landscape narrative, connections graph
Latency Seconds Research-session length
Grading Vera Health's own grading; DynaMed applies GRADE; most publish none Trust Index 0 to 100 across six weighted facets, plus Direct / Derived / Indirect citation grammar
Integrity screening Not published by any product surveyed Standing screening on every run: retracted and hijacked-journal sources blocked, corrections flagged
Access gate Clinician credential verification (NPI or licence) Open sign-up, no credential verification
Regulatory status Software intended to inform diagnosis or treatment falls under EU MDR as a medical device Research tool, outside that scope by design

BioSkepsis is not a point-of-care product and is not sold as one. It has no guideline corpus, no drug labels, no dosing or interaction data, no clinical calculators and no credential gate, and it does not answer in seconds. Those are the requirements of a different category. What it does is the appraisal step that sits behind a clinical recommendation. Why that screening step matters in clinical evidence specifically is covered in the BioSkepsis analysis of zombie trials and their persistence in the clinical literature; the natural pairing, for a clinician who also does research, is set out in BioSkepsis vs OpenEvidence.

One disclosure in the spirit of the argument above. The research thread behind this post scored 30 out of 100 on the BioSkepsis Trust Index, capped by grounding: 11 of 25 cited supports were verified against full text, 10 were available from abstract only, and 4 were struck after failing verification. One cited randomised trial carries a published correction, which the provenance facet flagged and which is listed in the sources below. Checking the brief by hand then removed one more figure: a "91% valid citation rate" for the Almanac framework, widely repeated but absent from the peer-reviewed report, where 91% is a factuality score in the cardiology subset. We are publishing all of that rather than only the parts that survived, because a tool that reports nothing but its own confidence is the thing this post is warning about.

Frequently asked questions about AI tools for clinicians

What is the best AI tool for clinicians at the point of care in 2026?

There is no single best tool, because access is gated by jurisdiction. In the United States, OpenEvidence and Doximity Ask are free to verified clinicians. In the EU and UK, where OpenEvidence withdrew in April 2026, Vera Health, iatroX, Medwise AI, Praxis Medicine and ClariMed are the active options. Whichever you use, the load-bearing feature is statement-level citation linking, because it is the only thing that lets you check the answer.

How often do medical AI tools fabricate references?

In unaugmented general-purpose models, frequently. Across 636 citations in 84 generated papers, 55% of GPT-3.5 and 18% of GPT-4 citations were completely fabricated (PMID 37679503). In 115 references in ChatGPT-generated medical content, 47% were fabricated and only 7% were authentic and fully accurate (PMID 37337480). Retrieval-grounded systems do better: the Almanac framework outperformed base ChatGPT-4, Bing and Bard on factuality, completeness, user preference and adversarial safety, though the peer-reviewed report gives the direction rather than a citation-validity percentage (PMID 38343631). This is why a tool that retrieves before it writes is categorically different from one that writes then cites.

Why did OpenEvidence leave the EU and UK?

OpenEvidence withdrew from the European Union and the United Kingdom in April 2026, citing regulatory uncertainty over how AI systems are treated in those jurisdictions, including the EU Artificial Intelligence Act. The practical effect is that European clinicians lost access to the category's largest product and the European vendors moved to fill the gap.

Do AI clinical decision support tools actually improve physician diagnosis?

The randomised evidence is split. In 50 physicians working six diagnostic vignettes, GPT-4 access produced no significant improvement over conventional resources (76% vs 74%, p = 0.60), while the model alone scored 92% (PMID 39466245). In 92 physicians producing 400 case-responses across five management vignettes, GPT-4 access did improve management reasoning (43.0% vs 35.7%, p < 0.001) but added an adjusted 119.3 seconds per case (PMID 39910272). The gain, where it exists, appears to come from structured reflection rather than from the model answering for you.

Can I trust the benchmark scores that clinical AI companies publish?

Treat them as marketing until a third party reproduces them. USMLE and NEJM-AI figures in this category are almost entirely self-reported, run on datasets the vendor selected, with prompts and model versions that are usually not disclosed. Most of the comparative ranking content that appears when you search for clinical AI tools is also published by vendors competing in the same category. Guideline-concordance studies are more informative: GPT-3.5 produced at least one non-concordant recommendation in 34.3% of NCCN treatment prompts (PMID 37615976).

Is BioSkepsis a clinical decision support tool?

No. BioSkepsis is a biomedical research tool built for appraising published literature, not a point-of-care product and not a regulated medical device. It reads full text across 40 million or more curated papers, scores every brief on a Trust Index, and blocks retracted and hijacked-journal sources before weighing evidence. It does not carry guideline corpora, drug labels, dosing or interaction data, and it answers in research-session time rather than in seconds. If you need a dose at the bedside, use a clinical tool. If you need to know whether the literature behind a recommendation holds up, that is the different job BioSkepsis does.

What should a clinical AI tool do about retracted papers?

Screen them out before synthesis, not flag them afterwards. None of the point-of-care products surveyed here publishes a retraction-screening policy, which matters because retracted and questioned trials continue to be cited for years after withdrawal. Medical models are also demonstrably vulnerable to corrupted inputs: poisoning as little as 0.001% of pretraining tokens increased harmful medical outputs, and knowledge-graph verification of extracted triplets caught 91.9% of that harmful content (PMID 39779928).

Check the clinical evidence behind the answer

Point-of-care tools tell you what to do. BioSkepsis tells you how well the literature supports it: full-text analysis across 40 million or more biomedical papers, retracted and hijacked-journal sources blocked before evidence is weighed, and a Trust Index on every brief, including this one.

Start free

Sources: PubMed citations and clinical AI vendor material

  1. Singhal K, et al. Large language models encode clinical knowledge. PMID 37438534
  2. Diagnostic accuracy of a large language model in rheumatology: comparison of physician and ChatGPT-4. Rheumatology International 2023. PMID 37742280. A re-analysis of an existing rheumatology dataset rather than a prospective head-to-head; the headline nulls mask divergent subgroups.
  3. ChatGPT with GPT-4 outperforms emergency department physicians in diagnostic accuracy: retrospective analysis. JMIR 2024. PMID 38976865. The human comparator group was resident physicians, not attendings, despite the paper's title.
  4. Hager P, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine 2024. PMID 38965432
  5. Evaluating the use of large language models to provide clinical recommendations in the emergency department. PMID 39379357
  6. Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 2023. PMID 37679503
  7. Learning to fake it: limited responses and fabricated references provided by ChatGPT for medical questions. PMID 40206627
  8. High rates of fabricated and inaccurate references in ChatGPT-generated medical content. PMID 37337480
  9. Chen S, et al. Use of artificial intelligence chatbots for cancer treatment information. JAMA Oncology 2023. PMID 37615976
  10. Prompt engineering in consistency and reliability with the evidence-based guideline for LLMs. PMID 38378899
  11. Leveraging large language models for decision support in personalized oncology. PMID 37976064
  12. Goh E, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Network Open 2024. PMID 39466245
  13. Goh E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nature Medicine 2025;31(4):1233-1238. PMID 39910272. Publisher correction: Nature Medicine 2025;31(4):1370.
  14. Zakka C, et al. Almanac: retrieval-augmented language models for clinical medicine. NEJM AI 2024;1(2). PMID 38343631. The peer-reviewed report states improvement in factuality, completeness, user preference and adversarial safety without publishing a citation-validity percentage.
  15. Medical large language models are vulnerable to data-poisoning attacks. PMID 39779928
  16. A framework for human evaluation of large language models in healthcare derived from literature review (QUEST). PMID 39333376

Non-literature sources. Product functionality, market, access model, corpus size, user counts, funding and benchmark figures in the landscape tables are taken from vendor material and contemporaneous press reporting, and are self-reported unless otherwise stated. OpenEvidence's EU and UK withdrawal and its stated reasons were reported in April 2026. The Doximity acquisition of Pathway Medical, at $63M in cash and equity, was announced in 2025 and reported by CNBC, Fierce Healthcare and Doximity's own investor materials. Most comparative ranking material in this category is published by vendors competing in it.

Related reading. BioSkepsis vs OpenEvidence · Zombie trials in clinical research · How general-purpose LLMs are deepening the reproducibility crisis