首页 > AI前沿 > A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses

arXiv自然语言 2026-09-29 23:19 6 阅读 查看原文

We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work:

  • medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE)
  • clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench)
  • and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith.

We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes.

It includes provisional behavioural anchors, applicability rules and a separate flag for case-specific safety-critical errors.

General-domain frameworks inform its design but are not treated as validated clinical instruments.

The rubric does not replace case-specific reference criteria or the task-specific metrics of existing benchmarks.

It has not yet been tested for inter-rater reliability, construct validity or clinical utility.

Its immediate purpose is to make evaluation decisions explicit and open to scrutiny before empirical testing.