首页 > AI前沿 > The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations

The Canonical Order Problem: When Large Language Models Are Unreliable Knowledge Bases for Multi-Valued Relations

arXiv自然语言 2026-09-29 04:09 6 阅读 查看原文

Large language models (LLMs) are increasingly used as knowledge bases (KBs) due to the vast amount of knowledge they acquire during pre-training.

While many works focus on extracting single relational triples, most real-world relations are multi-valued and require generating sets of entities.

In this paper, we investigate how LLMs represent and generate multi-valued relations.

We identify the canonical order problem: The probabilistic distributions inside LLMs organize many multi-valued relations according to a canonical ordering (e.g., alphabetical or chronological).

Through mechanistic analysis, we show that set generation in LLMs can be thought of in terms of three phases:

  • (1) retrieval of candidate entities
  • (2) internal sorting
  • (3) selection of the next element

As a result, prompts aiming to construct KBs that deviate from this internal canonical ordering lead to a markedly reduced reliability of LLMs when aiming to generate complete sets for multi-valued relations.