首页 > AI前沿 > Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

Classifying Interpretive Canons at the Sentence Level: A Benchmark from the German Federal Constitutional Court

arXiv自然语言 2026-09-23 02:34 6 阅读 查看原文

Judicial reasoning remains challenging for large language models (LLMs) to analyze.

This paper contributes a sentence-level benchmark for evaluating the ability of LLMs to classify interpretive canons as articulated by Larenz in the tradition of Savigny.

Our contributions are threefold.

  • First, we operationalize this conception of interpretation as classification criteria.
  • Second, we provide a dataset of decisions of the German Federal Constitutional Court annotated at the sentence level.
  • Third, we report baseline evaluations of four LLMs from three model families under expert hand-written prompts, compared against prompts optimized with Genetic-Pareto (GEPA).

Mean F1 over the seven binary subtasks clusters between 70.4 and 79.2 across models, with grammatical interpretation usually the easiest canon to identify and systematic interpretation usually the hardest;

under the tested configuration, GEPA-optimized prompts do not systematically outperform the hand-written ones, suggesting that the expert prompts provide a meaningful baseline.