Clinical trial · Observational
Evaluating AI and Human Expert Decisions in Colorectal Cancer
Comparison of Large Language Models and Expert Multidisciplinary Team Decisions in Colorectal Cancer
- Source
- ClinicalTrials.gov
- Retrieved
- Sep 8, 2026
- Layer
- normalized (units and labels harmonized; values unchanged)
- Run
- ING-CLINICALTRIALS-20260908-000001
Summary
Brief summary (as posted)
The goal of this observational study is to evaluate the decision-making consistency between large language models (LLMs) and expert multidisciplinary teams (MDTs) in adult patients diagnosed with colorectal cancer who underwent MDT consultation between January 2023 and December 2024. The main questions it aims to answer are: How consistent are the treatment decisions generated by LLMs compared to actual MDT decisions? Do different LLMs (e.g., ChatGPT, DeepSeek) show varying levels of agreement with expert recommendations? What clinical factors contribute to differences between AI-generated and human expert decisions? Researchers will compare the AI-generated treatment recommendations with real-world MDT decisions using anonymized patient records to see if LLMs can reliably support clinical decision-making in oncology. Participants will: Have their de-identified clinical data (e.g., imaging, pathology, MDT notes) processed through several LLMs Not be contacted or receive any interventions, as this is a retrospective study using existing clinical records only.
Conditions
Conditions (1)
Free-text conditions as registered, with the CancerIndex entity they were reconciled to and the match type.
| Condition (as posted) | Mapped entity | Match | Confidence |
|---|---|---|---|
| Colorectal Cancer | Malignant Colorectal Neoplasm | CURATED_BROADER | 0.80 |
Interventions
Interventions (1)
| Intervention | Type | Mapped drug | Match |
|---|---|---|---|
| LLM-MDT | Other | — | UNRESOLVED |
Design
Arms and outcomes
Arms (0)
[]Primary outcomes (1)
- measure
- Agreement Between AI-Generated and MDT Treatment Decisions
- timeFrame
- January 1, 2023 to December 31, 2024 (based on MDT consultation date)
- description
- Description: The primary outcome is the consistency between treatment recommendations generated by large language models (LLMs) and those made by expert multidisciplinary teams (MDTs) for colorectal cancer cases. Consistency will be quantified using Cohen's Kappa coefficient. Higher Kappa values indicate stronger agreement
Secondary outcomes (3)
- measure
- Comparison of Agreement Across Different AI Models
- timeFrame
- January 1, 2023 to December 31, 2024
- description
- Description: To compare the consistency of treatment decisions generated by different large language models (e.g., ChatGPT, DeepSeek, Baichuan, Qwen) with expert MDT decisions using Cohen's Kappa and chi-squared tests. This outcome assesses whether performance varies across AI models.
- measure
- Output Stability of AI Models on Repeated InputsDescription
Eligibility
Eligibility (as posted)
- Sex
- All
Show eligibility criteria text
Inclusion Criteria: * Patients with a histologically confirmed diagnosis of colorectal cancer * Patients who received multidisciplinary team (MDT) consultation at Peking University Cancer Hospital between January 1, 2023 and December 31, 2024 * Availability of complete clinical records, including(MDT consultation notes, CT or MRI imaging reports, Pathology reports, Outpatient or inpatient medical summaries) Exclusion Criteria: * Incomplete or missing medical records related to MDT decision-making * MDT consultations conducted for non-oncologic purposes (e.g., hernia evaluation, stoma planning) * Missing critical clinical data such as imaging or pathology reports * Duplicate or conflicting records that prevent reliable data analysis
References
Publications (0)
Data not yet available