AI Tools That Molecular Tumour Boards Use for Variant Interpretation

September 08, 2026

Reviewed

AI Tools That Molecular Tumour Boards Use for Variant Interpretation

Molecular tumour boards now run on four categories of software: curated knowledge bases, MTB portals, general large language models, and citation-grounded literature engines. They are not interchangeable, and the published concordance data show exactly where each one breaks. In the largest multi-laboratory test of variant classification, only 59% of assessments matched consensus on both tier and clinical significance (PMID 36503149).

Why molecular tumour boards need software: variant interpretation is inconsistent

A molecular tumour board convenes when standard treatment has failed, to decide whether a molecular alteration points to another therapy. The decision is a judgement about evidence, and the evidence is inconsistent before any software is involved.

The Association for Molecular Pathology ran a multi-laboratory classification exercise in which 134 participants assessed 11 variants across four cancer cases. Only 59% of classifications matched consensus on both tier and clinical significance (PMID 36503149). These were specialists applying a published standard to a small, curated set. Real boards face larger panels, rarer alterations and less time.

Time is the second constraint. IASLC recommends that the board discussion take place within 14 days of the sequencing result (PMID 40633839). A 2025 review of digital tumour board support found literature searches of 30 to 90 minutes for a single complex case, alongside manual data entry, media discontinuities between systems, and an absence of standardised case preparation (PMID 39819625).

The third constraint is that the whole exercise has a modest yield to protect. In a meta-analysis of 34 studies and 12,176 patients across 26 cancer entities, 20.8% received an MTB-recommended therapy, with an objective response rate of 21% and a disease control rate of 45% (PMID 40175535). Software that saves an hour of curation is worth having. Software that puts a wrong option in front of the board costs more than it saves.

The baseline problem, in one number

134 specialists, 11 variants, four cases, 59% concordance with consensus on tier and clinical significance (Li MM et al., J Mol Diagn 2023, PMID 36503149). Every tool below is an attempt to raise that number or to reach it faster.

Precision oncology knowledge bases: OncoKB, CIViC, ClinVar and COSMIC

Knowledge bases are curated assertions linking a variant, a tumour type and a therapy, usually with an evidence level attached. OncoKB and CIViC are the two most used in tumour board workflows; ClinVar carries germline and somatic submissions; COSMIC provides somatic mutation frequency context.

They are the fastest way to answer one specific question: has this alteration already been graded as actionable by someone. For a BRAF V600E in melanoma, that question is closed and the base answers it in seconds.

The limitation is coverage, and it is structural rather than a quality complaint. Comparative analyses of public precision oncology knowledge bases have consistently found that no single base holds all the relevant information for a case series, and that each holds assertions the others lack. Gene and variant coverage between CIViC, ClinVar and OncoKB differ substantially. A board that queries one base is making a coverage decision without knowing it.

The second limitation is latency. A knowledge base contains what curators have processed. The preclinical paper describing sensitivity of a rare fusion to an available inhibitor exists in PubMed months before it exists as a curated assertion, and for rare alterations it may never become one.

MTB portals and NGS visualisation software: MTBP, cBioPortal and MTPpilot

Portals handle the workflow rather than the biology. They ingest sequencing output, annotate against aggregated knowledge bases, rank findings by an actionability scale, flag trial eligibility and generate the report the board discusses.

The Molecular Tumor Board Portal is the best documented example. Across 500 consecutive advanced solid tumours evaluated between January 2019 and January 2021 within the Basket of Baskets trial, it automated variant classification, ranked biomarker matches on the ESMO-ESCAT scale, flagged 49 germline variants in 48 patients, detected trial eligibility across the Cancer Core Europe network and produced interactive reports in under 14 days, with average discussion time falling below 3 minutes per case after a learning curve of about 25 cases (DOI 10.1038/s43018-022-00332-x).

cBioPortal, including MTB-adapted deployments, covers visualisation and cohort context. MTPpilot provides interactive NGS visualisation built for the board meeting itself. MTB-Report, cbpManager, AMBAR and MIRACUM-Pipe cover annotation, data integration and pipeline workflow; a 2025 review of the digital tumour board field describes most institutions moving toward this kind of automation, with reported discussion times of 8 to 10 minutes per case at Heidelberg and written reports within 15 days in Italian boards (PMID 39609355).

What automation actually bought

500 consecutive tumours, reports inside the 14-day window, and discussion time under 3 minutes per case once the board had learned the report format. The gain is in preparation and standardisation, not in resolving contested biology.

The limitation is inherited. A portal is only as complete as the knowledge bases it aggregates, so it reproduces their coverage gaps and their latency. The reported barriers to adoption are also organisational: poor workflow integration and a lack of clinical trust in decision support output (PMID 39819625).

Large language models in the tumour board: what GPT-4 actually produced

General LLMs entered tumour board workflows informally, as the thing a registrar opens at 23:00 before a Thursday board. There is now direct evidence on what they produce.

Schmutz and colleagues ran ChatGPT 4.0 against 20 anonymised real molecular tumour board cases spanning breast, colorectal, glioblastoma and rare tumours, and compared its output with the board's own (The Oncologist 2025, PMID 40973166). The model generated a median of 3 therapeutic recommendations against 1 from the human board (P=.005). Information density was statistically comparable, at a median of 0.67 against 0.75 (P=.084). Two findings matter more than either of those.

Evidence overreach and run-to-run variance

Level of evidence 3 to 4 appeared in 15% of ChatGPT cases and 0% of human board cases (P=.0019). Across triplicate runs of the same case, agreement reached a median Fleiss kappa of 0.51, which is moderate. The same case, asked three times, prioritised evidence differently (PMID 40973166).

Read together, those two numbers describe a preparation aid, not a decision instrument. More options are useful when a board is stuck and costly when they arrive with preclinical support presented in the same register as a registrational trial. A 2026 systematic review of LLM applications across tumour boards reaches the same structural conclusion (PMID 42245825).

The reason is architectural. A general model reproduces language patterns from its training corpus. It has no citation layer, no retraction screening and no evidence tier, so it cannot mark the difference between a phase III readout and a cell line experiment unless that difference happens to be reflected in how the text was written. The BioSkepsis engine inverts that order: retrieve the papers first, grade them, screen for retractions, then answer, so a claim that rests on a single cell line arrives labelled as such. The same failure mode appears across oncology practice more broadly, which we covered in AI tools oncologists use in 2026.

Citation-grounded literature engines: how BioSkepsis grounds tumour board evidence

The fourth category exists because of the gap the first three leave: the primary literature on the specific alteration in front of the board, retrieved and graded rather than curated in advance or generated from memory.

OncoSkepsis is the clinical tier of BioSkepsis, free for verified oncologists and molecular tumour boards. It runs on the same BioSkepsis literature engine that academic and biotech researchers use, with the interface rebuilt around board work rather than around a research question. It searches the peer-reviewed literature and returns evidence with graded citations, a Trust Index and screening of retracted and questioned sources, in minutes rather than the 30 to 90 minutes a complex manual search takes. It supports four board-specific tasks: tumour board case questions, variant interpretation queries, country-specific literature checks, and literature-backed reimbursement dossiers, which boards report needing for roughly one in three patients.

Country-specific evidence is the request that came directly from clinicians in Cyprus, Greece, Portugal and the UK. A variant considered benign in one reference population is not necessarily benign in another, and default annotations rarely make that distinction visible. The retrieval method is the same one described in our review of AI tools for life-science literature review, applied to a clinical question rather than a research one.

Why the clinical tier is free, stated plainly

OncoSkepsis is free by design, not as charity. It is the clinical entry point that produces a consented, aggregated signal of what questions the field is asking, and that signal is what pharma Medical Affairs pays for in a separate product, LeaderSkepsis. The boundary is architectural: clinicians never see the pharma dashboard, pharma users never see clinical activity, and the dashboard receives only aggregated derived signals. No patient data leaves the institution, only scientific questions and their field are recorded, consent is explicit and withdrawable, and clinical access continues unchanged if a clinician declines.

Comparing the four tool categories on molecular tumour board tasks

What each category does on the jobs a board actually has
Board task Knowledge base MTB portal General LLM Citation-grounded engine (OncoSkepsis)
Is this variant already graded actionable Yes, its core function Yes, aggregated Sometimes, unverifiable Yes, with the source paper
Evidence published since the last curation cycle No No Only if in training data Yes
Traceable citation for every claim Yes, assertion level Yes, inherited No Yes, graded and retraction screened
Country or population-specific literature Rarely Rarely No Yes, as a query type
Report generation and trial flagging No Yes, its core function Drafting only No
Reimbursement dossier support No Partial Drafting only Yes, literature-backed
Main failure mode Coverage gaps and latency Inherits base coverage Evidence overreach, run-to-run variance Depends on published literature existing

No column wins. A board that has a portal and a knowledge base still lacks the literature layer; a board with only a literature engine still has to prepare and report the case by hand.

Building a molecular tumour board stack: which tool for which clinical question

MTB portalBoard coordinator preparing the weekly case list

Ingest the variant call file, annotate, rank by ESCAT, flag trial eligibility, generate the report. This is the workflow job, and portals are the only category that does it. Published performance is under 3 minutes of discussion per case after the board has learned the report format (DOI 10.1038/s43018-022-00332-x).

Knowledge baseMolecular pathologist checking known actionability

Query more than one base. OncoKB, CIViC and ClinVar differ in gene and variant coverage, so a single-base check is a silent coverage decision. Use them to close the question quickly when the alteration is established, not to conclude that an alteration is uninformative.

OncoSkepsisOncologist facing a variant of uncertain significance

When the bases return nothing graded, the question becomes what the primary literature says about this alteration, this pathway and this drug class, including recent work and negative results. Graded citations with retraction screening let the board see the evidence tier rather than infer it from confident phrasing.

OncoSkepsisTeam writing a reimbursement dossier or a population-specific check

Roughly one in three patients needs a literature-backed dossier for the payer, and boards outside the large reference cohorts need to know whether an allele frequency claim holds in their own population. Both are literature retrieval tasks with an audit trail requirement, which is the one thing a general LLM cannot supply.

What none of these AI tools solve in precision oncology yet

No tool in any of the four categories has been tested against a patient outcome in a randomised design. The pooled evidence for molecular tumour boards themselves gives an objective response rate of 21% and median overall survival of 13.5 months across 12,176 patients, with only 62% of studies reporting patient-level outcomes and no standardised endpoint definitions (PMID 40175535). Software vendors claiming survival benefit are extrapolating.

Adoption failures are also not technical. The reported barriers are workflow integration, manual data entry, media discontinuities between hospital systems, and clinical distrust of decision support output (PMID 39819625). A tool that returns the right answer into a system nobody opens during the meeting has not helped.

Finally, retrieval does not resolve contested biology. When two well-conducted studies disagree about a mechanism, the useful output is that the evidence is contested and where the disagreement sits, not a single confident recommendation. That is the standard a board should hold any of these tools to.

Molecular tumour board AI tools: frequently asked questions

Can ChatGPT replace a molecular tumour board?

No. In a retrospective evaluation of 20 real molecular tumour board cases, ChatGPT 4.0 produced a median of 3 therapeutic recommendations against 1 from the human board (P=.005), but drew on level of evidence 3 to 4 in 15% of cases where the board used none, and agreement across triplicate runs of the same case reached only a median Fleiss kappa of 0.51 (PMID 40973166). It generates more options with less consistency and weaker evidence, which is a preparation aid, not a decision.

Which knowledge base should a molecular tumour board use for variant actionability?

More than one. OncoKB, CIViC, ClinVar and COSMIC differ substantially in gene and variant coverage, and comparative work on precision oncology knowledge bases has repeatedly found that no single base contains all the relevant information for a case series while each holds information the others lack. Boards that query one base only will systematically miss annotations that exist elsewhere.

How long should a molecular tumour board take to return a recommendation?

IASLC recommends the board discussion happen within 14 days of the sequencing result (PMID 40633839). Published workflows meet this with automation: the Molecular Tumor Board Portal reported reports inside 14 days and under 3 minutes of discussion per case after a learning curve of roughly 25 cases across 500 consecutive advanced solid tumours (DOI 10.1038/s43018-022-00332-x).

Why do tumour boards still search the literature by hand?

Because the annotation layer stops at what has been curated. Knowledge bases hold curated variant to drug assertions; they do not hold the preclinical paper published last month on the specific fusion in front of the board. A 2025 review of digital tumour board support reported literature searches of 30 to 90 minutes for complex cases and named poor workflow integration and lack of trust as the main barriers to decision support adoption (PMID 39819625).

Do AI tumour board tools improve patient outcomes?

That has not been shown. The pooled evidence base for molecular tumour boards themselves is a meta-analysis of 34 studies and 12,176 patients in which 20.8% received an MTB-recommended therapy, with an objective response rate of 21% and disease control of 45% (PMID 40175535). No tool in any category has been tested against that endpoint in a randomised design. Claims of improved survival from software vendors are, at present, unevidenced.

Is a variant classified as benign in one population always benign in another?

No, and this is a recurring complaint from boards outside the large North American and Western European reference cohorts. Population frequency and founder effects change the interpretation of the same allele, so a board in Cyprus, Greece or Portugal often needs the country-specific literature rather than the default annotation. Few tools surface that layer, which is why OncoSkepsis was built to run country-specific literature checks as a distinct query type.

What does OncoSkepsis cost for a molecular tumour board?

Nothing. OncoSkepsis is the clinical tier of BioSkepsis, and it is free for verified clinicians by design, not by discount. It is the entry point that produces the consented, aggregated query signal behind LeaderSkepsis, the paid product sold to pharma Medical Affairs. No patient data leaves the institution, only scientific questions and their field are recorded, clinicians opt in explicitly after seeing what is captured, and clinical use continues unchanged if a clinician declines.

Add the literature layer your tumour board is missing

OncoSkepsis, the clinical tier of BioSkepsis, is free for verified oncologists and molecular tumour boards: graded citations, retraction screening, country-specific literature checks and literature-backed reimbursement dossiers, on tumour board cases and variant interpretation questions.

Start free

Sources for the molecular tumour board evidence cited above

  1. Li MM, et al. Reference standards for the interpretation and reporting of sequence variants in cancer: multi-laboratory concordance assessment. J Mol Diagn. 2023;25(2):69-86. PMID 36503149. Cited for: 134 participants, 11 variants, four cases, 59% concordance on tier and clinical significance.
  2. Aldea M, et al. IASLC recommendations on molecular tumour boards in thoracic oncology. J Thorac Oncol. 2025;20(11):1594-1614. PMID 40633839. Cited for: board discussion recommended within 14 days of the sequencing result.
  3. Gladstone BP, et al. Systematic review and meta-analysis of molecular tumor board data on clinical effectiveness and evaluation gaps. npj Precis Oncol. 2025;9:96. PMID 40175535. Cited for: 34 studies, 12,176 patients, 20.8% receiving MTB-recommended therapy, 21% objective response, 45% disease control, 13.5 months median overall survival, 62% reporting patient-level outcomes.
  4. Tamborero D, et al. The Molecular Tumor Board Portal supports clinical decisions and automated reporting for precision oncology. Nat Cancer. 2022;3(2):251-261. DOI 10.1038/s43018-022-00332-x (author correction PMID 35449310). Cited for: 500 consecutive advanced solid tumours, ESCAT ranking, 49 germline variants in 48 patients, reports within 14 days, under 3 minutes discussion per case after roughly 25 cases.
  5. Strantz C, et al. Empowering personalized oncology: evolution of digital support and visualization tools for molecular tumor boards. BMC Med Inform Decis Mak. 2025;25:29. PMID 39819625. Cited for: 30 to 90 minute literature searches for complex cases, workflow integration and trust as adoption barriers, manual data entry and media discontinuities.
  6. Lutz S, et al. Unveiling the digital evolution of molecular tumor boards. Target Oncol. 2025;20(1):27-43. PMID 39609355. Cited for: named support tools including cBioPortal and MIRACUM-Pipe, 8 to 10 minutes per case at Heidelberg, 15-day written reports in Italian boards.
  7. Schmutz M, et al. Large language model processing capabilities of ChatGPT 4.0 to generate molecular tumor board recommendations: a critical evaluation on real world data. Oncologist. 2025;30(10):oyaf293. PMID 40973166. Cited for: 20 cases, median 3 vs 1 recommendations (P=.005), information density 0.67 vs 0.75 (P=.084), level of evidence 3 to 4 in 15% vs 0% (P=.0019), median Fleiss kappa 0.51 across triplicates.
  8. Applications of large language models in tumor boards: a systematic review. PMID 42245825. Cited for: the structural conclusion that current LLM performance supports preparation rather than decision-making in tumour boards.
  9. Related reading on the BioSkepsis blog: AI tools oncologists use in 2026: evidence and the tumour board gap and 10 best AI tools for life-science literature review.