Nine frontier large language models, evaluated across 2,005 oncology decision points built from NCCN guidelines and colorectal cancer cases, fail collectively on almost half the items that require a therapeutic commitment. This is not a medical knowledge gap: in 3 to 9 percent of cases, the models state the correct next clinical step but do not commit to it. The benchmark, introduced as the Oncology Decision Boundary Benchmark, used a fully deterministic scorer with no LLM in the loop to judge outputs, and two oncologists independently validated a stratified sample of 225 items.
Treating the nine models as a pooled super-model, 42.1 percent of all items received no correct answer from any model: 35.7 percent of the 1,586 NCCN items and 66.4 percent of the 419 colorectal cancer cases. The gap is shared across closed-source and open-weight families, so combining models does not close it. Failures concentrate in choosing between guideline pathways before reasoning within any single pathway. The authors describe this as a blind spot in clinical meta-judgment that likely requires architectural intervention rather than more training data or fine-tuning.
Two models tuned for decisiveness, GPT-5.5 and Gemini 3.1 Pro Preview, made unsafe commitments three to five times more often than the seven cautious models without scoring higher. It is a warning against treating assertiveness as clinical utility: in oncology, a wrong commitment creates heavier review and liability costs than a cautious answer.
For teams looking at on-premise deployment, this reframes the discussion. Adding VRAM or choosing a model with more parameters does not address the observed limit, because the defect is not raw capacity but decision architecture. A self-hosted system can keep escalation rules, audit logs, and deterministic checks under local control, avoiding any transfer of clinical data to cloud APIs. But it shifts the cost toward the TCO of human supervision and hospital workflow integration, rather than just inference hardware. AI-RADAR offers analytical frameworks for evaluating these trade-offs, but the point here is more radical: the variable to measure is no longer average model accuracy but the quality of the boundary the system can recognize.
This changes incentives for model vendors. Selling decisive LLMs for healthcare risks offering apparent automation: a system that commits more often but unsafely increases clinical review workload and legal exposure instead of reducing it. The structural signal is that the next generation of clinical systems will not be won on parameter scale but on components that can detect when a model reaches its competence boundary and route the decision to a clinician.
If nine frontier models share the same blind spot, the question is no longer which LLM to choose, but how to design a system that knows when to stop before deciding.
💬 Comments (0)
🔒 Log in or register to comment on articles.
No comments yet. Be the first to comment!