首页 > AI前沿 > Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi

Knowing the Rules, Applying the Rules: Evaluating Language Models on Traditional Chinese Bazi

arXiv自然语言 2026-10-05 09:50 5 阅读 查看原文

Knowing domain rules does not guarantee applying them to a case.

We study this distinction in traditional Chinese Bazi through 3,000 Chinese multiple-choice questions spanning 14 Theory and 11 Case categories.

Six endpoint systems are evaluated, with primary results reported on a 2,492-item model-informed refinement.

Theory accuracy exceeds Case accuracy for every system, and gaps of 16.60-29.56 percentage points remain when invalid responses are excluded.

The contrast is more specific than a general case-reasoning deficit.

Across six systems, Twelve Stages and Nayin reach mean accuracies of 89.10% and 88.62%, while Shensha Basics reaches 75.96%.

Within Case, Luck Pillars averages 84.62%, but Career and Family Relations average only 36.98% and 38.19%.

Overall rankings also conceal different category strengths.

On the original 3,000 items, paired DeepSeek native/disabled comparisons associate native configurations with Theory gains of 6.53 and 12.20 points for Flash and Pro, respectively; Case changes are -3.67 and +1.27 points.

These are provider-configuration associations, not isolated causal effects of reasoning.

The results motivate task-specific evaluation of cultural-domain applications rather than reliance on aggregate knowledge scores.

The benchmark measures agreement with a model-generated, model-verified answer key, not real-world predictive validity.

Final-set results are post-selection descriptions, and incomplete provenance and expert validation constrain their interpretation.