NPCBench: A Clinical Apprenticeship Benchmark for Guideline-Constrained Care-Pathway Reasoning in Nasopharyngeal Carcinoma

ABSTRACT

Current medical LLM benchmarks typically decompose clinical competence into isolated factual questions, single-image interpretation, or brief case scenarios. This fragmented paradigm primarily measures local correctness but provides limited insight into whether models can progressively integrate subspecialty knowledge, multimodal evidence, and evolving patient states into safe, guideline-constrained longitudinal reasoning. To address this gap, we introduce NPCBench, a multimodal clinical apprenticeship benchmark for nasopharyngeal carcinoma (NPC). NPCBench operationalizes specialist training as a staged evaluation of (M)LLMs across guideline recall, atomic rule application, multimodal evidence grounding, longitudinal care planning, state revision, and expert-style consultation. It comprises 4,077 staged evaluation problems and 25 multimodal full-care path episodes, covering pathology diagnosis, MRI staging, and 146 downstream clinical decisions. The benchmark explicitly links stepwise tasks to complete patient-level care trajectories, enabling systematic evaluation of all clinical pathways. Beyond isolated precision, NPCBench evaluates whether the models ground decisions in multimodal evidence, maintain temporally consistent patient states, and provide safe expert-level consultation across complete care trajectories. Across 30 evaluated models, we observe a persistent composition gap: strong performance in local case-level tasks does not translate into episode-level clinical competence required for end-to-end patient management. Overall, NPCBench advances medical AI evaluation from static question answering toward apprenticeship-style workflow evaluation, providing a rigorous testbed for evidence-grounded and temporally coherent reasoning in complex oncology care.

MedicineBenchmarkNasopharyngeal Carcinoma
2026年5月13日
Pengkai Wang, Wei-Wei Zhang, Yan Li, Min Tang, Zhitian Hou, Zeyu Liu, Guanghao Zhu, Yuanyi Wang, Yanggan Gu, Wenjun Wang, Minheng Ni, Congkai Xie, Zhijie Sang, Jianmin Wu, Ying Sun, Hongxia Yang

Abstract

Current medical LLM benchmarks typically decompose clinical competence into isolated factual questions, single-image interpretation, or brief case scenarios. This fragmented paradigm primarily measures local correctness but provides limited insight into whether models can progressively integrate subspecialty knowledge, multimodal evidence, and evolving patient states into safe, guideline-constrained longitudinal reasoning. To address this gap, we introduce NPCBench, a multimodal clinical apprenticeship benchmark for nasopharyngeal carcinoma (NPC). NPCBench operationalizes specialist training as a staged evaluation of (M)LLMs across guideline recall, atomic rule application, multimodal evidence grounding, longitudinal care planning, state revision, and expert-style consultation. It comprises 4,077 staged evaluation problems and 25 multimodal full-care path episodes, covering pathology diagnosis, MRI staging, and 146 downstream clinical decisions. The benchmark explicitly links stepwise tasks to complete patient-level care trajectories, enabling systematic evaluation of all clinical pathways. Beyond isolated precision, NPCBench evaluates whether the models ground decisions in multimodal evidence, maintain temporally consistent patient states, and provide safe expert-level consultation across complete care trajectories. Across 30 evaluated models, we observe a persistent composition gap: strong performance in local case-level tasks does not translate into episode-level clinical competence required for end-to-end patient management. Overall, NPCBench advances medical AI evaluation from static question answering toward apprenticeship-style workflow evaluation, providing a rigorous testbed for evidence-grounded and temporally coherent reasoning in complex oncology care.

1. Introduction

Medical large language models (LLMs) and multimodal large language models (MLLMs) now achieve strong performance on medical examinations, visual question answering, and broad static benchmark suites. In clinical practice, however, the challenge is fundamentally different. Real-world clinical scenarios are inherently characterized by high specialization, strict guideline dependence, and the integration of multi-source evidence.

True specialist competence is a progressively integrated capability rather than an isolated skill. In clinical training, physicians recall guideline knowledge and apply it to specific contexts, progressively learn to ground decisions in multimodal evidence, formulate case-level plans, and dynamically revise management as patient states evolve. This hierarchical integration ultimately ensures global pathway consistency across an entire longitudinally coherent care pathway. However, existing medical evaluations predominantly rely on static, de-contextualized questions. This obscures where complex clinical reasoning breaks down and fails to reveal whether atomic capabilities can successfully compose into pathway-level care, leaving a construct-validity gap between static answer accuracy and true specialist-care competence..

Nasopharyngeal carcinoma (NPC) provides an informative clinical domain for evaluating guideline-constrained, multimodal, and longitudinal clinical reasoning. As a regionally salient cancer concentrated in Southern China and Southeast, NPC is biologically and clinically distinctive. Its management is inherently state-dependent, strictly relying on the continuous fusion of anatomical and staging evidence (e.g., MRI), histopathologic confirmation, and longitudinal biomarker trajectories (e.g., EBV DNA). Guideline-concordant care spans the entire disease trajectory, including induction chemotherapy, concurrent chemoradiotherapy, adjuvant or maintenance therapy, salvage local treatment, toxicity management, recurrence assessment, and surveillance.. Consequently, guideline-concordant care requires dynamic decision revisions across the evolving disease trajectory —making NPC an ideal crucible for testing pathway correctness over isolated diagnostic accuracy.

Figure 1. Overview of NPCBench.

NPCBench\ is a clinical apprenticeship benchmark for NPC following a specialist-training trajectory. To our knowledge, it is the first benchmark for a single oncologic disease to evaluate (M)LLMs within a unified specialist-care framework, spanning guideline knowledge, multimodal pathology and MRI interpretation, longitudinal patient-state revision, CarePath simulation, and rubric-based assessment of open-ended consultations. As shown in Fig. 1, it structures subspecialty care competence into a hierarchical progression comprising L1 guideline memory, L2 atomic guideline application, L3 multimodal evidence grounding, L4 case-level planning, L5 longitudinal revision, and L6 full CarePath episodes; O1 ConsultRubrics further evaluates the quality of open-ended specialist consultations. Lower levels assess individual capability components, whereas higher levels evaluate their integration across multi-stage clinical decision-making to form a coherent and consistent care pathway.

This study is organized around four research questions. RQ1: Do current (M)LLMs successfully acquire and apply guideline-constrained rules within isolated clinical decision contexts? RQ2: Do multimodal models accurately ground tumor assessment in histopathologic evidence and anatomical and staging evidence, or do they generate plausible predictions without strict evidence grounding? RQ3: Does high accuracy on isolated capability components translate into accurate decision revisions driven by patient-state transitions and the successful completion of full care episodes? RQ4: Do rubrics based evaluation in open-ended specialist consultation settings reveal clinically meaningful safety, clinical justification, and evidence-grounding limitations that are obscured by metrics based solely on static answer correctness?

We summarize the main contributions of this work as follows.

  1. [****] Pathway-sensitive evaluation paradigm. We shift the evaluation focus from static answer correctness to pathway correctness. NPCBench assesses whether models can maintain coherent clinical reasoning under patient-state transitions, seamlessly integrating guideline constraints, multimodal evidence grounding, and longitudinal disease evolution.
  2. [****] Disease-specific apprenticeship capability framework. We operationalize specialist competence into a progressive clinical training ladder (L1–L6) supplemented by cross-level consult rubrics (O1). This structured hierarchy decomposes NPC care into guideline memory, atomic guideline application, multimodal evidence grounding, case-level planning, longitudinal decision revisions, and full care episodes.
  3. [****] Expert-validated multimodal and longitudinal benchmark. We construct a rigorously validated evaluation framework that anchors clinical decisions in histopathologic evidence, anatomical and staging evidence, and biomarker trajectories. Expert verification of state transitions and evidence traceability ensures reliable measurement of step accuracy, transition accuracy, and overall episode success.

From exam QA to clinically structured benchmarks. Medical LLM evaluation began with static question-answering and examination-style datasets, including MedQA, MedMCQA, PubMedQA, MMLU medical subsets, MultiMedQA, MedBench, CMB, and CMExam. These benchmarks are useful for measuring medical knowledge and local reasoning, but they do not test whether a model can accumulate patient evidence, revise decisions, and preserve care-path coherence. Recent broad suites and construct-validity arguments, such as MedHELM, MedXpertQA, AnesBench, and AnesSuite, move evaluation toward clinically meaningful capabilities and specialty-aware reasoning. NPCBench\ follows this direction but makes the clinical apprenticeship ladder explicit: staged component abilities are measured first, and then tested under longitudinal and episode-level care constraints.

Multimodal, longitudinal, and rubric-based clinical evaluation.

Medical MLLM benchmarks have expanded from image QA datasets such as VQA-RAD, SLAKE, and PathVQA to broader modality, trustworthiness, endoscopy, clinical reasoning, and sequence-grounding evaluations. In parallel, HealthBench, LiveMedBench, LLMEval-Med, MedDialogRubrics, and ClinConsensus use expert-validated rubrics or temporally grounded cases to evaluate open-ended clinical responses. Workflow-oriented resources such as MedJourney, MedAgentBench, MedAgentBoard, and MTBBench further shift evaluation from static snapshots toward state transitions. NPCBench\ brings these directions into one disease setting: pathology and MRI evidence grounding, real-patient longitudinal revision, strict CarePath episodes, and criterion-level consult rubrics. Appendix B gives a fuller comparison with recent medical benchmark families.

3. NPCBench

3.1 Overview

NPCBench\ operationalizes the clinical apprenticeship framework via a progressive capability ladder (L1–L6) and cross-level consult rubrics (O1)—spanning guideline memory, atomic application, multimodal evidence grounding, case-level planning, longitudinal decision revisions, and full care episodes. Consequently, lower levels (L1–L4) diagnose failures in isolated capability components, while higher levels (L5–L6, O1) evaluate their composition into a coherent care pathway driven by patient-state transitions. The evaluated construct is strictly defined as: guideline-constrained, multimodal evidence-grounded, and longitudinally coherent specialist care.

Four design principles ensure rigorous construct validity: Competence isolation requires each level in the progressive capability ladder (L1–L6) to measure a distinct, clinically meaningful component of specialist competence. Evidence-boundary control strictly freezes the patient state at the intended clinical decision point, removing future outcomes and treatment, and answer-leaking cues. Clinical determinacy ensures a single optimal path supported by guidelines or expert consensus, given the provided clinical evidence. Auditable open-ended scoring evaluates cross-level consult rubrics (O1) to independently assess clinical justification, evidence grounding, conditional reasoning, and safety. All real-word patient materials are rigorously de-identified. NPC oncology experts verified capability component validity, decision determinacy, guideline/consensus consistency, reliability of multimodal evidence grounding and consult rubrics, and the fidelity of patient-state transitions. Figure 2 summarizes the scale of the seven evaluation levels, while Fig. 3 illustrates how this progressive clinical training ladder is operationalized.

Figure 2. Dataset scale and units distribution of NPCBench.

3.2 Construction and Validation

The construction pipeline transforms expert-curated clinical materials into rigorously validated evaluation assets. Structured around the progressive principle of the capability ladder, L1–L2 are constructed from CSCO and NCCN guidelines; L3 is built from histopathologic and anatomical evidence; L4–L6 are derived from de-identified real-world patient-state transitions; and O1 employs patient-specific clinical scenarios.

During validation, experts strictly delineate evidence boundaries across all levels based on these evaluation dimensions mentioned above. Specifically, for L1–L2, they verify optimal decision paths using solely the precise evidence snapshot. For L3, decoupling structure identification, invasion status, and T-stage assessment into independent steps forces models to explicitly reconstruct the diagnostic evidence chain—preventing high accuracy via spurious correlations without reliable multimodal evidence grounding. For L5–L6, anchoring each decision node to a patient-state transition enables independent measurement of step accuracy, transition accuracy, and overall episode success. Finally, for O1, decomposing cross-level consult rubrics into positive anchors and negative error criteria guarantees auditable open-ended judgments.

Figure 3. Data curation and validation pipeline of NPCBench.

3.3 Hierarchical Evaluation Levels

L1.G-Mem: Guideline memory. This level measures the retention of core NPC guideline knowledge requisite for downstream clinical application. It comprises 250 expert-validated evaluation** units** derived from authoritative guidelines—153 dually supported by CSCO and NCCN, 80 CSCO-specific, and 17 NCCN-specific. These assessments systematically probe treatment indications, staging implications, recommendation boundaries, surveillance strategies, and contraindication principles.

L2.A-App: Atomic guideline application. This level evaluates the translation of core guideline knowledge into concrete clinical actions. It comprises 185 expert-validated evaluation units constructed from standardized cases, systematically spanning initial definitive management, post-treatment recurrence, metastatic or systemic management, and toxicity or contraindication management. By isolating the clinical decision context, experts rigorously enforce clinical determinacy—ensuring that the provided patient state dictates a single, guideline-concordant optimal path.

L3.MM-EG: Multimodal evidence grounding. This level evaluates whether models can assess histopathologic and anatomical evidence. The pathology task comprises 69 expert-validated evaluation units (39 NPC-positive, 30 negative) for slide interpretation. Derived from 100 real-world patient scans, the MRI task includes 200 sequence-recognition units and 2,700 assessment units across 18 anatomical structures—rigorously evaluating structure identification and tumor invasion status. Furthermore, it incorporates 100 patient-level T-staging evaluations strictly defined by the 8th edition AJCC/UICC TNM system. Building upon direct T-stage accuracy as the primary metric, the appendix further presents rigorous structure-consistent staging analyses to probe the model's capability in reconstructing the diagnostic evidence chain.

L4.C-Plan: Case-level planning. This level evaluates the capability to formulate single-case management plans under a fixed, real-world patient state. It comprises 165 expert-validated evaluation units constructed from de-identified real NPC cases. Each unit integrates a comprehensive clinical snapshot—spanning diagnosis, TNM staging, risk features, prior treatment history, response status, toxicity, organ function, and patient-specific constraints—to determine the optimal next-step intervention. To amplify decision complexity under guideline-constrained settings, these units incorporate plausible but suboptimal clinical alternatives, forcing models to critically differentiate the optimal strategy.

L5.L-Rev: Longitudinal revision. This level evaluates whether models can assimilate newly arrived clinical evidence, systematically review historical treatment trajectories, and formulate the optimal next-step management. It comprises 305 expert-validated evaluation units anchored to clinically meaningful transition points across 100 real-world patient trajectories, systematically covering treatment transitions, response reassessments, toxicity management, recurrence or metastatic evaluations, and follow-up planning. To capture robustness against cascading errors over time, we report step accuracy, transition accuracy, and overall episode success—rigorously measuring whether models maintain globally consistent longitudinal reasoning.

L6.CP-Epi: CarePath episodes. This level serves as the integrative endpoint of NPCBench, evaluating the execution of full-course NPC care pathways under multimodal and longitudinal constraints. It comprises 146 evaluation units derived from 25 multimodal-complete, real-world patient trajectories. Each episode spans the entire clinical continuum, encompassing initial workup, diagnosis and staging, treatment planning, management, and long-term follow-up. Accurately inferring across these trajectories requires models to synthesize multimodal evidence (e.g., pathology, MRI, labs, EBV-DNA, and clinical note) and maintain cross-stage reasoning throughout longitudinal state evolution. Emphasizing pathway-sensitive evaluation, success demands both local step accuracy and global pathway consistency. Ultimately, we report overall episode success as a measure of pathway integrity, evaluating whether models can approximate specialist-level longitudinal care under strict guideline constraints.

O1.C-Rub: ConsultRubrics. This level evaluates the quality of open-ended specialist consultations in NPC care. It comprises 103 expert-validated, rubric-based evaluation units derived from 42 real-world patient records, covering initial workup, diagnosis, treatment, and adverse-event management within authentic consultation scenarios. The rubric system integrates 1,035 positive clinical anchors and 657 negative safety and error criteria, yielding 1,692 criterion-level assessments per evaluated model. During evaluation, major errors incur a two-point deduction and minor errors a one-point deduction, with scores clipped at zero at the criterion level.

4. Experimental Evaluation

4.1 Experimental Setup

Evaluation Models. We evaluate 30 model (M)LLMs comprising medical LLMs such as HuatuoGPT-o1, medical MLLMs such as Lingshu, open-source models such as Qwen and GPT-OSS, and proprietary models from the GPT and Gemini families. More selected (M)LLM details are shown in App. C.1. Text LLMs are evaluated on guideline memory, atomic guideline application, case planning, longitudinal revision, and open-ended consultation, while MLLMs are additionally evaluated on pathology, MRI, and CarePath streams that require visual evidence.

Evaluation Metrics. NPCBench is reported as a staged profile, not a single averaged score. For closed-choice and structured outputs, we parse responses with fixed task-specific rules and compute exact-match accuracy. When a subset spans multiple capability groups or repeated items from one patient, we report macro scores over groups or patients to avoid over-weighting large categories or long records. For L3, we separately report pathology recognition, MRI sequence recognition, anatomical-structure identification, direct T-stage accuracy, and evidence-grounded staging, so final-stage labels are not conflated with the visual evidence supporting them. For L5 and L6, we report both turn-level accuracy and strict patient- or episode-level success, distinguishing local decision correctness from trajectory-level coherence. For O1, ConsultRubrics scores free-text consultations using positive anchors and major/minor error criteria: anchors add credit, while triggered errors subtract credit. Full metric definitions and relative details are shown in App.C.2.

More Evaluation Details. All models are evaluated with identical prompt templates, identical clinical evidence at each decision point, and identical scoring rules. For O1, we use the same saved responses for primary scoring, judge-sensitivity analysis, and stratified physician audit. The full generation and criterion-level judging prompts for L1–L6 and O1 appear in App.C.4. Local open-weight models and rubric-based judges run on a GPU server with eight NVIDIA A800-SXM4-80GB GPUs. Additional details on the experimental environment are in App.C.

4.2 Overall Results

Tab. 1 presents the main NPCBench results as a staged capability profile, spanning guideline knowledge, multimodal evidence grounding, case planning, longitudinal revision, full CarePath episodes, and open-ended consultation. We do not aggregate these heterogeneous metrics into a single score. Instead, we report them side by side to show each model's modality coverage, local strengths, and whether those strengths carry through to evidence-grounded longitudinal care.

Table 1. Main NPCBench\ results.

MM-EGL-Rev.CP-Epi.
ModelG-Mem.A-App.Path.T-Stg.Seq.Struct.C-PlanStepEpi.StepEpi.C-Rub.
HuatuoGPT-o1-72B88.474.1--------48.561.623.0----13.7
HuatuoGPT-o1-7B84.851.4--------40.655.118.0----8.0
Baichuan-M2-32B74.428.1--------24.862.327.0----37.9
Fleming-R1-32B85.647.0--------46.773.440.0----32.8
Fleming-R1-7B82.457.8--------35.856.720.0----27.0
HealthGPT-Pro-4B80.054.656.539.054.012.747.937.715.025.30.018.3
HealthGPT-Pro-8B81.258.956.533.050.010.847.345.913.020.60.021.5
Hulu-Med-4B78.053.556.535.030.08.452.760.325.046.60.012.2
Hulu-Med-7B82.450.856.515.058.09.849.747.218.037.70.012.7
Hulu-Med-32B86.071.460.916.075.55.954.569.535.057.50.010.4
Lingshu-7B78.454.656.535.047.09.342.449.816.056.90.04.4
Lingshu-32B85.658.959.435.075.011.452.769.534.068.50.010.7
Qwen3-30B-A3B83.261.6--------51.561.023.0----30.1
GPT-OSS-20B81.660.5--------36.447.916.0----30.0
GPT-OSS-120B86.873.0--------49.161.325.0----46.6
DeepSeek-V3.190.079.5--------52.772.536.0----32.3
Qwen3-VL-8B-Instruct83.660.556.535.029.54.753.367.934.068.50.027.2
Qwen3-VL-30B-A3B-Instruct86.064.356.535.040.036.252.768.933.064.40.029.9
Qwen3.5-9B85.268.172.526.065.538.954.561.324.071.20.039.4
Qwen3.5-27B89.273.571.034.066.047.658.278.749.082.20.046.8
Qwen3.5-35B-A3B86.474.166.743.064.550.355.874.139.080.80.043.0
Qwen3.6-27B87.671.471.035.075.541.455.278.750.085.60.045.2
Qwen3.6-35B-A3B85.671.466.735.059.544.856.479.751.084.34.041.9
InternVL3.5-38B88.069.459.433.048.03.750.962.028.054.80.013.2
GPT-5.190.877.862.339.047.026.361.870.238.083.64.059.1
GPT-4.189.675.171.030.048.528.355.277.442.080.80.036.8
Gemini-3.1-Pro-Preview95.290.363.834.077.072.073.985.665.087.016.048.6
Gemini-2.5-Pro92.083.260.933.052.048.666.783.055.078.14.039.7
Gemini-2.5-Flash87.269.265.234.059.524.558.865.634.057.50.038.0
Claude-Sonnet-4.693.283.256.533.057.027.967.983.660.084.324.048.1

Staged evaluation elucidates clinically critical failure modes that would be obscured by a single aggregate metric. First, robust performance on localized reasoning tasks does not guarantee competence across the full end-to-end care pathway. Although Gemini-3.1-Pro-Preview demonstrates high guideline adherence (G-Mem. 95.2) and acceptable treatment appropriateness (A-App. 90.3), its strict CarePath success rate declines precipitously to 16.0; the best-performing model achieves only 24.0. This pronounced degradation indicates that models frequently generate intermediate outputs that appear locally correct yet fail to preserve a coherent, clinically consistent patient trajectory. Second, plausible label prediction can conceal a lack of genuine visual grounding. For example, while anatomical structure recognition attains 72.0, direct patient-level T-stage classification accuracy reaches only 43.0, and evidence-grounded staging accuracy is even lower. These findings suggest that high performance on aggregate metrics may be driven by linguistic heuristics or dataset-specific biases rather than authentic image-based reasoning.

4.3 Diagnostic Analysis

The primary evaluation framework conceptualizes assessment as a multidimensional profile of specialist-care capabilities rather than as a competitive ranking. The subsequent analyses investigate the points at which this capability profile may fail, including: (i) whether constituent skills reliably compose into coherent, pathway-level clinical care; (ii) whether direct staging decisions faithfully reflect image-grounded evidence extraction; (iii) whether locally correct decisions remain robust when aggregated to patient-level and episode-level performance metrics; and (iv) whether model-family designations are predictive of the intended clinical use case.

(1) Local correctness does not aggregate into globally successful care pathways. Gemini-3.1-Pro-Preview scores 95.2 on G-Mem., 90.3 on A-App., and 73.9 on C-Plan, yet its best strict CP-Epi. score is only 24.0, despite care-turn accuracies up to 87.0. The full capability-profile heatmap in 6 shows that medical-domain–specialized models, open-source MLLMs, and proprietary systems each fail in distinct ways at different pathway stages. NPCBench quantifies this gap: high-quality, locally correct responses do not ensure guideline-concordant, evidence-based performance across end-to-end specialist care pathways.

(2) Direct T-stage accuracy obscures the MRI evidence-grounding bottleneck. Direct T-stage accuracy is 43.0, while full 18-structure invasion extraction and evidence-grounded T-stage prediction are each capped at 4.0 in Table 19. In contrast, deterministic or GPT-OSS-120B staging given gold 18-structure evidence reaches 99.0 in Tables 21 and 22. Thus, most errors arise upstream: models can output a clinically plausible T-stage label but fail to consistently extract, structure, and align the MRI-derived evidence that should support it.

Figure 4. Local correctness collapses under patient-level and episode-level success criteria.

(3) Longitudinal decisions collapse under full-patient and full-episode scoring. In L5, the top model achieves 85.6 step accuracy but only 65.0 full-patient success. In L6, Gemini-3.1-Pro-Preview reaches 87.0 care-turn accuracy but just 16.0 strict full-episode success, while Claude-Sonnet-4.6 attains the best strict episode score of 24.0. 4 explains this gap: an episode counts as successful only if pathology evidence, MRI stage, treatment-state tracking, phase transitions, and downstream care decisions are all correct simultaneously.

(4) Model family is not a reliable proxy for NPC-care competence. Medical specialization, multimodal access, model scale, and proprietary serving each help in some areas but none excels overall. Medical MLLMs add image interfaces but still lag on strict CarePath, and proprietary systems differ by whether the task is staging, evidence extraction, longitudinal consistency, or open-ended rubric satisfaction. Model choice for NPC support must therefore be matched to clinical phase and evidence interface, not just whether it is medical, multimodal, or large.

4.4 Expert Calibration of Rubric Judging

Open-ended consultation needs a separate validity check because O1 uses criterion-level rubric judging, not a fixed answer key. Under the primary GPT-OSS-120B judge, the best C-Rub. score is 59.1. Top proprietary models still perform poorly: in the zero-truncation rerun, GPT-5.1 and Gemini-3.1-Pro-Preview each have a 13.6% major-error rate and >40% minor-error rates (Tab. 28 in the Appendix). To calibrate the automated O1 scorer, a clinical NPC expert annotated a stratified audit set of 548 criterion-level labels from 48 frozen model responses. The main set includes rubric labels for GPT-5.1, Qwen3.6-27B-Instruct, GPT-OSS-120B, and the non-truncated Gemini-3.1-Pro-Preview response. Tab. 2 supports GPT-OSS-120B as the primary judge: it shows the best agreement with expert labels (0.814 accuracy, κ=0.610\kappa=0.610) and the highest specificity on expert-negative criteria (0.890). Qwen3-30B-A3B-Instruct has higher sensitivity on expert-positive criteria (0.834) but lower specificity (0.658), indicating more permissive scoring; GPT-4.1 is intermediate. We therefore use GPT-OSS-120B as the primary ConsultRubrics judge and Qwen3-30B-A3B-Instruct as a permissiveness-sensitivity probe.

Figure 5. Criterion-level ConsultRubrics reveals safety errors and judge sensitivity.

Table 2. Human-expert calibration of O1 model judges.

Judge modelnnAcc.κ\kappaSens.Spec.
GPT-4.15480.7770.5530.8170.749
GPT-OSS-120B5480.8140.6100.7070.890
Qwen3-30B-A3B-Instruct5480.7320.4720.8340.658

5. Conclusion

NPCBench\ reframes medical benchmark design around the staged acquisition and execution of specialist clinical competence. Modeled after the clinical apprenticeship trajectory of NPC care, it integrates guideline memory, atomic application, multimodal evidence grounding, case-level planning, longitudinal revision, and full CarePath episodes within a unified, disease-specific framework. Across the 30 evaluated models, we demonstrate that high local step accuracy does not necessarily translate into robust multimodal grounding, global pathway consistency, or overall episode success. Ultimately,NPCBench\ establishes a rigorous capability-profile evaluation for clinical AI systems operating under strict guideline constraints, multimodal evidence, and evolving patient states.

Limitation. First, NPCBench\ relies on simulated retrospective trajectories for a single disease; future work should validate this paradigm in prospective real-world deployments across broader oncologic domains. Second, it uses standardized 2D images and static interfaces. Future work should advance toward interactive 3D-native evaluations to simulate authentic clinical workflows. Third, despite being specialist-calibrated, automated evaluation of open-ended clinical reasoning inherently carries LLM-as-a-judge bias risks.

BibTeX

@misc{npcbench-2026,
      title={NPCBench: A Clinical Apprenticeship Benchmark for Guideline-Constrained Care-Pathway Reasoning in Nasopharyngeal Carcinoma},
      author={Pengkai Wang and Wei-Wei Zhang and Yan Li and Min Tang and Zhitian Hou and Zeyu Liu and Guanghao Zhu and Yuanyi Wang and Yanggan Gu and Wenjun Wang and Minheng Ni and Congkai Xie and Zhijie Sang and Jianmin Wu and Ying Sun and Hongxia Yang},
      year={2026},
}