Image

Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations

Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations

Recruiting
18 years and older
All
Phase N/A

Powered by AI

Overview

BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.

Description

BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.

Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.

Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up \& biomarkers before any comparison.

Eligibility

Inclusion Criteria:

  • Synthetic oncology case within one of the five predefined localisations (breast, lung, urological, digestive, gynaecological).
  • Complete structured schema: UICC 8th-edition stage, biomarkers, ECOG performance status, comorbidities and a standardised clinical question.
  • A clinically answerable treatment-planning question that is mappable to the locked guideline matrix.

Exclusion Criteria:

  • Case outside the five predefined localisations.
  • Incomplete, internally inconsistent or ambiguous schema.
  • Duplicate or near-duplicate of an existing case in the set.
  • Question not resolvable by current guidelines.

Study details
    Breast Neoplasms
    Lung Neoplasms
    Urologic Neoplasms
    Prostatic Neoplasms
    Urinary Bladder Neoplasms
    Kidney Neoplasms
    Digestive System Neoplasms
    Genital Neoplasms
    Artifical Intelligence
    Large Language Models
    Decision Making
    Decision Support Systems
    Clinical

NCT07739121

Assistance Publique - Hôpitaux de Paris

1 August 2026

Step 1 Get in touch with the nearest study center
We have submitted the contact information you provided to the research team at {{SITE_NAME}}. A copy of the message has been sent to your email for your records.
Would you like to be notified about other trials? Sign up for Patient Notification Services.
Sign up

Send a message

Enter your contact details to connect with study team

Investigator Avatar

Primary Contact

  Other languages supported:

First name*
Last name*
Email*
Phone number*
Other language

FAQs

Learn more about clinical trials

What is a clinical trial?

A clinical trial is a study designed to test specific interventions or treatments' effectiveness and safety, paving the way for new, innovative healthcare solutions.

Why should I take part in a clinical trial?

Participating in a clinical trial provides early access to potentially effective treatments and directly contributes to the healthcare advancements that benefit us all.

How long does a clinical trial take place?

The duration of clinical trials varies. Some trials last weeks, some years, depending on the phase and intention of the trial.

Do I get compensated for taking part in clinical trials?

Compensation varies per trial. Some offer payment or reimbursement for time and travel, while others may not.

How safe are clinical trials?

Clinical trials follow strict ethical guidelines and protocols to safeguard participants' health. They are closely monitored and safety reviewed regularly.
Add a private note
  • abc Select a piece of text.
  • Add notes visible only to you.
  • Send it to people through a passcode protected link.