AI Tools Oncologists Use in 2026: Evidence and the Tumour Board Gap

September 03, 2026

Reviewed

AI Tools Oncologists Use in 2026: What the Evidence Supports at the Molecular Tumour Board

The AI tools oncologists use in cancer care now have real trial evidence behind them: AI-supported mammography reading raised screening detection from 74% to 81% of cancers in a randomised trial of over 100,000 women, and ambient scribes are cutting documentation time in oncology clinics. None of that evidence touches the step where a molecular tumour board decides what a variant means. There, 134 specialists agreed with consensus on both tier and clinical significance only 59% of the time.

Five categories of clinical AI oncologists run today

Ask an oncologist which AI tools they use and the honest answer is usually a stack rather than a product. As of April 2026 more than 1,500 AI-enabled medical devices hold FDA marketing authorisation, and roughly 76% of them are radiology devices, with cardiovascular at 10% and neurology at 4%. Most cleared through the 510(k) pathway, which asks for substantial equivalence to a predicate device rather than evidence of clinical benefit.

That distinction matters for how oncologists read vendor claims. Authorisation is a floor, not a result. The table below separates what each category has actually been shown to do from what it is often assumed to do.

Categories of oncology AI by strongest published evidence
Category Strongest evidence to date What it does not do
Screening imaging AI MASAI randomised trial, over 100,000 women: 81% vs 74% of cancers detected at screening, 12% fewer interval cancers, 44% less screen-reading workload at interim Does not stage, does not select systemic therapy
Digital pathology AI First FDA authorisation for an AI pathology product, Paige Prostate, 2021; assistive detection on digitised slides Assistive only; does not replace pathologist sign-out or interpret genomics
Radiotherapy auto-contouring Multiple single-institution and multi-vendor evaluations showing contouring time reduction with physician edit Requires clinician review of every structure; no outcome trials
Ambient documentation scribes Randomised and prospective studies showing reduced documentation time and burnout scores Transcribes and drafts; contributes nothing to the treatment decision
General-purpose LLMs 34-study review of LLM use in oncology; single-centre tumour board comparisons Unstable across runs, cites weak evidence, no verifiable source trail
Literature and evidence retrieval No cleared device, no category owner, no trial The step the molecular tumour board actually performs by hand

Note the last row. Five categories have vendors, clearances and in one case randomised evidence. The sixth, the literature reconciliation a molecular tumour board does between a sequencing report and a treatment decision, has none of that, and it is the step this post is ultimately about.

Imaging AI in cancer screening: what the MASAI randomised trial proved

MASAI is the strongest evidence any AI tool in oncology has produced. Gommers, Hernström and colleagues, with Kristina Lång as senior author at Lund University, randomised more than 100,000 Swedish women between April 2021 and December 2022 to AI-supported mammography screening or standard double reading. Final results published in The Lancet in 2026 (407:505-514, PMID 41620232).

In the AI-supported arm, 338 of 420 cancers (81%) were detected at screening, against 262 of 355 (74%) under double reading. Interval cancers fell from 1.76 per 1,000 screened (93/52,872) to 1.55 (82/53,043). The AI arm found 16% fewer invasive cancers presenting later, 21% fewer large tumours and 27% fewer aggressive subtypes. False positive rates were 1.5% and 1.4%. The earlier interim safety analysis reported a 44% reduction in screen-reading workload.

This is what a validated oncology AI claim looks like: randomised, powered on interval cancer, published with denominators. It is also narrow. MASAI tells a breast unit how to read screening mammograms. It says nothing about the patient whose tumour has already been sequenced.

Digital pathology AI in prostate and breast cancer diagnosis

Paige Prostate received the first FDA authorisation for an AI product in digital pathology in September 2021, as an assistive tool flagging areas of a digitised prostate biopsy slide suspicious for cancer. The regulatory framing has held: the pathologist signs out, the software points.

Adoption follows scanner deployment rather than clinical enthusiasm, which is why uptake is uneven between academic centres and community laboratories. For the tumour board, digital pathology AI changes when a diagnosis is confirmed. It does not change what happens when the diagnosis is confirmed and standard therapy has already failed.

Ambient AI scribes in the oncology clinic: relief that is real and bounded

Ambient documentation is the fastest-spreading AI category in oncology, and the one with the clearest positioning from the profession itself. Debra Patt, MD, PhD, MBA, Executive Vice President of Texas Oncology, President of the Community Oncology Alliance and Chair of ASCO's AI Task Force, framed the current boundary plainly in an April 2026 interview: "Where we are today is that we want administrative burdens to be relieved, and we want our people to be able to work at the top of their license." She added that AI used appropriately "doesn't threaten [the doctor-patient relationship], it can enhance it, because doctors and nurses can have more meaningful interactions with patients when they are less burdened by administrative tasks."

Her first question about any oncology AI deployment is the right one to carry into the rest of this post: "What is the use case?" Scribes have a clean answer. They give back clinic time. They do not read a paper, weigh an ERBB2 exon 20 insertion against a phase I signal, or tell you whether the evidence for a repurposed agent holds up.

The boundary, stated by the people setting it

Douglas Flora, MD, Executive Medical Director of the Yung Family Cancer Center at St. Elizabeth Healthcare and Editor-in-Chief of AI in Precision Oncology, put the governance problem this way in September 2025: "We really have to take care as doctors to make sure that we're being responsible about this, and we are at the wheel this time, so that AI isn't thrust upon us." His warning about the alternative is specific: "We can bury our head in the sand and act concerned and let it happen to us like we did the medical record 15 years ago with the arrival of the EMR."

General-purpose LLMs at the molecular tumour board: what ChatGPT got wrong

This is where most oncologists have quietly experimented, and where the published evidence is least reassuring.

Schmutz and colleagues ran ChatGPT 4.0 against 20 consecutive real molecular tumour board cases from the Comprehensive Cancer Center Augsburg, published in The Oncologist in 2025 (PMID 40973166). The model produced a median of 3 therapeutic recommendations per case against 1 from the human board (P = .005). It drew on level 3 to 4 evidence in roughly 15% of cases where the human board used none (P = .0019). Internal consistency across triplicate runs was moderate, median Fleiss kappa 0.51 with a range of 0.12 to 1.0. Run the same case three times, get materially different boards.

The broader picture is the same. Carl, Schramm, Kather and colleagues reviewed 34 studies of LLM use in clinical oncology in npj Precision Oncology in 2024 (PMID 39443582). Treatment recommendations were the dominant test case (30 of 34 studies), but only 12 of 34 assessed test-retest reliability and only 9 of 34 reported prompt strategy. Performance variance across studies was substantial and largely methodological.

The tumour board literature step: general-purpose LLM versus citation-grounded retrieval
Axis OncoSkepsis General-purpose LLM
Source trail Every claim resolves to an openable paper Citations generated from parameters, not retrieved
Evidence tier shown Graded per source before you read Level 3 to 4 evidence used in ~15% of MTB cases without flagging it
Retracted sources Screened out No retraction awareness
Run-to-run stability Same corpus, same retrieval Median Fleiss kappa 0.51 across triplicate runs
Preprints bioRxiv and medRxiv as first-class content Bounded by training cutoff
Output type Cited evidence for the board to weigh Therapeutic recommendations the board did not ask for

More recommendations is not better recommendations

A board that produces three options where the specialists produced one has not been more thorough. It has widened the differential using weaker evidence, and it cannot show you which source carried which claim. For a decision that will be defended in a chart and, in some systems, in a reimbursement dossier, an unverifiable extra option is a liability.

The unclosed gap: somatic variant interpretation and the 14-day window

A molecular tumour board convenes when standard treatment has failed, to decide whether a mutation points to another therapy. Three findings define why that decision is still hard.

Classification is inconsistent between trained specialists. In the Association for Molecular Pathology multi-laboratory assessment, 134 participants classified 11 variants across four cancer cases against the AMP/ASCO/CAP guidelines. 86% classified variants correctly overall, range 54% to 94%. But only 59% of responses matched the working group consensus on both tier and category of clinical significance, range 39% to 84%.

The knowledge bases do not close the gap. Pallarz, Benary and colleagues at Charité and Humboldt University compared the public precision oncology knowledge bases and found substantial divergence in gene and variant coverage between CIViC, ClinVar, OncoKB and their peers, with no single resource complete on its own (PMID 32914021). Reconciliation is therefore manual, and it happens under time pressure.

Most of what sequencing returns is uninterpretable at the point of decision. Across the TCGA PanCancer Atlas, 10,967 samples over 32 studies, the proportion of identified mutations already classified as variants of uncertain significance runs at 62% for ATM (474/766), 68% for CHEK2 (101/148), 70% for BRCA1 (215/307) and 75% for BRCA2 (513/683) (PMID 38847368).

Against that, the 2025 IASLC consensus statement on molecular tumour boards, authored by Aldea, Rotow, Arcila and 30 colleagues, advises that regular meetings be held "to avoid delays beyond 14 days from result availability to discussion."

What the 14 days are actually spent on

Searching PubMed, checking three or four knowledge bases against each other, chasing a case report that may or may not exist, and reconciling a variant call against a population dataset that may not include the patient's population. None of the five validated AI categories above does any of this.

Where OncoSkepsis fits, and what it does not claim

OncoSkepsis is the free clinical tier of BioSkepsis, built for verified oncologists and molecular tumour boards. It searches the primary biomedical literature and returns the relevant evidence with graded citations, retraction screening and a trust index on every answer. Board briefs, variant evidence summaries, country-specific literature checks, and literature-backed reimbursement dossiers, which are needed for roughly one in three patients according to the clinicians who asked for the feature.

Three of those features came directly from oncologists in Cyprus, Greece, Portugal and the United Kingdom, not from a product roadmap. Country-specific literature was the most insistent request: a variant treated as benign on the basis of one population dataset may not be benign in another, and the databases do not carry that nuance.

What OncoSkepsis does not claim: it does not classify variants, it does not make treatment recommendations, and it is not a device. It does the literature step, with every claim traceable to a source you can open. The board decides. That boundary is deliberate, and it is the one Patt and Flora are both describing from inside the profession.

OncoSkepsisMolecular tumour board chairs and coordinators

Cited board briefs assembled inside the 14-day window, with the evidence tier and the retraction status visible on each source before the meeting rather than after it.

OncoSkepsisOncologists outside large academic centres

The reconciliation work that a dedicated molecular pathology service does at a comprehensive cancer centre, without the dedicated service. Free, and not dependent on which knowledge base your institution licenses.

OncoSkepsisMolecular pathologists and clinical geneticists

Full-text literature behind a VUS call, including preprints from bioRxiv and medRxiv, so the source trail for a reclassification is documented rather than remembered.

The engine underneath is BioSkepsis: over 40 million curated biomedical papers from 1931 onwards, analysed full text rather than by abstract, with preprints from bioRxiv and medRxiv treated as first-class content.

Why the clinical tier is free, and what leaves your tumour board

Free tools in oncology usually mean one of two things: a trial that will end, or a product whose real customer is not you. Neither is the case here, and the honest version is more reassuring than a mission statement, so here it is in full.

One engine, three surfaces. BioSkepsis is a single citation-grounded literature engine. What differs is the account type it serves. Academic and biotech researchers get the Research Workspace. Verified clinicians and tumour boards get OncoSkepsis, free. Pharmaceutical Medical Affairs teams get Insights, the paid dashboard sold under the name LeaderSkepsis. OncoSkepsis is not a stripped-down clinical edition of a research tool. It is the same engine, exposed to a different account type.

What pharma actually buys. Not your searches. Every existing expert-intelligence platform infers influence from artifacts an expert has already produced: publications, trials, grants. That is retrospective by construction, and the lag is long. Across 165,135 registered trials, the median time from first patient enrolled to published paper is 4.8 years, and only 53% are ever published in full (PMID 39601300). Bibliometric scoring is least accurate for early-career researchers (PMID 24165898), which is exactly where identifying someone first has value. Sustained, specific enquiry into a mechanism is an earlier marker of expertise than a paper about it. Insights receives aggregated, derived signals of that kind. It does not receive queries, cases or users.

The boundary is architectural, not a policy promise. The oncology workspace operates under an institutional licence, and clinicians opt in explicitly after seeing what is recorded, what is not, and what reaches pharmaceutical customers. Four guarantees follow from how the system is built rather than from what we undertake to do:

What the separation guarantees

No patient data leaves the institution; the system records the scientific question and its field, nothing else. Clinicians and researchers never see Insights. Pharma never sees clinical or research activity, only aggregated derived signals. Consent is explicit and withdrawable, and OncoSkepsis continues to work in full for any clinician who declines it.

Saying this plainly is a deliberate choice. A free clinical tier presented as charity invites the question of what is really being funded, and the answer is better than the suspicion.

Oncologists shaping AI in cancer care

The people setting the terms for clinical AI in oncology are not vendors. Debra Patt chairs ASCO's AI Task Force while running a community oncology network. Douglas Flora edits AI in Precision Oncology from a community cancer centre. Kristina Lång built the randomised evidence base for screening AI at Lund. Sanjay Juneja, a triple board-certified haematologist-oncologist, now leads clinical AI strategy at Tempus. Aldea, Rotow, Arcila, Sholl, Rolfo, Subbiah, Leighl, Drilon and their co-authors wrote the IASLC consensus that defines what a molecular tumour board owes a patient.

The pattern is consistent. Every functional piece of oncology AI was specified by clinicians who described the failure mode first. The literature step at the tumour board has not had that treatment yet, which is the reason for the invitation below.

If you sit on a molecular tumour board, OncoSkepsis is free and we want the criticism. Specifically: where the retrieved evidence misses a paper you knew about, where the country-specific check fails for your population, and where a board brief would not survive your own board's scrutiny. Those three failure reports are worth more to us than any feature request.

Questions oncologists ask about clinical AI tools

Which AI tools do oncologists actually use in clinical practice?

In routine use today: AI-supported reading in cancer screening imaging, assistive digital pathology for prostate and breast specimens, AI auto-contouring in radiotherapy planning, ambient documentation scribes in clinic, and general-purpose large language models used informally for drafting and literature orientation. Radiology accounts for roughly 76% of the more than 1,500 AI-enabled devices with FDA marketing authorisation as of April 2026.

Is ChatGPT reliable for molecular tumour board recommendations?

Not as an unsupervised source. In a critical evaluation on 20 consecutive real molecular tumour board cases, ChatGPT 4.0 produced a median of 3 therapeutic recommendations against 1 from the human board, drew on level 3 to 4 evidence in around 15% of cases where the human board used none, and showed only moderate consistency across replicate runs (median Fleiss kappa 0.51). The authors conclude that human oversight remains necessary.

How often do experts disagree on somatic variant classification?

In an Association for Molecular Pathology multi-laboratory exercise, 134 participants classified 11 variants across four cancer cases. 86% classified variants correctly overall, but only 59% of responses matched the working group consensus on both tier and category of clinical significance, with a range of 39% to 84% across variants.

How quickly should a molecular tumour board discuss a sequencing result?

The 2025 IASLC consensus statement on molecular tumour boards advises regular meetings to avoid delays beyond 14 days from the availability of the molecular result to its discussion.

Does AI-supported mammography screening detect more cancers?

Yes, in the MASAI randomised trial of more than 100,000 Swedish women. 81% of cancers in the AI-supported arm were detected at screening versus 74% under standard double reading, interval cancers fell from 1.76 to 1.55 per 1,000 screened, false positive rates were comparable at 1.5% versus 1.4%, and the interim safety analysis reported a 44% reduction in screen-reading workload.

What is OncoSkepsis and what does it cost?

OncoSkepsis is the free clinical tier of BioSkepsis for verified oncologists and molecular tumour boards. It searches the primary literature and returns cited board briefs, variant evidence summaries, country-specific literature checks and literature-backed reimbursement dossiers. Tumour boards pay nothing for it.

Does OncoSkepsis share my tumour board searches with pharmaceutical companies?

No. The oncology workspace runs under an institutional licence with explicit, withdrawable opt-in. No patient data leaves the institution, and the system records the scientific question and its field, nothing else. The paid Medical Affairs dashboard, sold as LeaderSkepsis Insights, receives only aggregated derived signals about which mechanisms and indications are drawing sustained enquiry. It does not receive queries, cases or user identities. Clinicians never see Insights, pharma never sees clinical activity, and OncoSkepsis works in full for any clinician who declines to opt in.

What is the difference between BioSkepsis, OncoSkepsis and LeaderSkepsis?

One engine, three surfaces determined by account type. BioSkepsis is the citation-grounded literature engine and the Research Workspace used by academic and biotech researchers. OncoSkepsis is the free clinical surface for verified oncologists and molecular tumour boards. LeaderSkepsis is the paid expert-intelligence product for pharmaceutical Medical Affairs teams, whose dashboard is called Insights. OncoSkepsis is not a reduced clinical edition of the research tool; it is the same engine serving a different account type.

Why does country-specific literature matter for variant interpretation?

Allele frequency and reported penetrance differ between populations, so a variant treated as benign on the basis of one population dataset may not be benign in another. Clinicians in Cyprus, Greece, Portugal and the United Kingdom raised this directly as a gap in the databases they use.

Bring the literature step to your next tumour board

OncoSkepsis is free for verified oncologists and molecular tumour boards. Cited board briefs, graded evidence, retraction screening, and country-specific literature checks. Bring a case that took you a week and tell us where it falls short.

Start free

Sources cited in this analysis of oncology AI evidence

  1. Li MM, Cottrell CE, Pullambhatla M, Roy S, Temple-Smolkin RL, Turner SA, Wang K, Zhou Y, Vnencak-Jones CL. Assessments of Somatic Variant Classification Using the AMP/ASCO/CAP Guidelines: A Report from the Association for Molecular Pathology. J Mol Diagn. 2023;25(2):69-86. PMID 36503149. DOI 10.1016/j.jmoldx.2022.11.002
  2. Aldea M, Rotow JK, Arcila M, et al. Molecular Tumor Boards: A Consensus Statement From the International Association for the Study of Lung Cancer. J Thorac Oncol. 2025;20(11):1594-1614. PMID 40633839.
  3. Mellgard GS, Atabek Z, LaRose M, Kastrinos F, Bates SE. Variants of uncertain significance in precision oncology: nuance or nuisance? Oncologist. 2024;29(8):641-644. PMID 38847368. DOI 10.1093/oncolo/oyae135
  4. Schmutz M, Sommer S, Sander J, et al. Large language model processing capabilities of ChatGPT 4.0 to generate molecular tumor board recommendations: a critical evaluation on real world data. Oncologist. 2025;30(10):oyaf293. PMID 40973166. DOI 10.1093/oncolo/oyaf293
  5. Carl N, Schramm F, Haggenmüller S, Kather JN, et al. Large language model use in clinical oncology. NPJ Precis Oncol. 2024;8:240. PMID 39443582. DOI 10.1038/s41698-024-00733-4
  6. Gommers J, Hernström V, Josefsson V, et al; Lång K. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trial. Lancet. 2026;407:505-514. PMID 41620232. DOI 10.1016/S0140-6736(25)02464-X
  7. Hernström V, Josefsson V, Sartor H, et al; Lång K. Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI). Lancet Digit Health. 2025;7(3):e175-e183. PMID 39904652. DOI 10.1016/S2589-7500(24)00267-X
  8. Pallarz S, Benary M, Lamping M, Rieke D, Starlinger J, Sers C, Wiegandt DL, Seibert M, Ševa J, Schäfer R, Keilholz U, Leser U. Comparative Analysis of Public Knowledge Bases for Precision Oncology. JCO Precis Oncol. 2019;3:PO.18.00371. PMID 32914021. DOI 10.1200/PO.18.00371
  9. Showell MG, Cole S, Clarke MJ, DeVito NJ, Farquhar C, Jordan V. Time to publication for results of clinical trials. Cochrane Database Syst Rev. 2024;(11):MR000011. PMID 39601300. DOI 10.1002/14651858.MR000011.pub3
  10. Penner O, Pan RK, Petersen AM, Kaski K, Fortunato S. On the predictability of future impact in science. Sci Rep. 2013;3:3052. PMID 24165898. DOI 10.1038/srep03052
  11. US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices list, updated April 2026.
  12. Paige. FDA De Novo authorisation of Paige Prostate, the first FDA-authorised AI product in digital pathology. September 2021.
  13. Patt D. Working at Top of License and Improving Care Delivery: A Q&A on AI's Practical Promise in Oncology. ASCO AI in Oncology, 7 April 2026.
  14. Flora D. Flora Charts the Future of AI in Community Cancer Care. OncLive, 27 September 2025.