首页 > AI前沿 > TPBench: A Turning-Point Benchmark for Dialogue Compression

TPBench: A Turning-Point Benchmark for Dialogue Compression

arXiv自然语言 2026-10-02 11:08 5 阅读 查看原文

A compressor can keep the facts of a dialogue and still drop the turn that changed them.

A user corrects a price, reverses a choice, or adds a constraint.

We call this failure turning-point eviction.

One overall retention score hides it, because that score mixes what the user first wanted with what the user wants now.

We introduce TPBench

We introduce TPBench, which evaluates three complementary information targets at shared nominal retention budgets.

P1 asks for the user's initial goal.

P2 asks for the current value of a slot the user revised.

P3 asks for both, in dialogues with a late annotated slot update.

The current-value answers come from the human dialogue-state annotations of MultiWOZ and SGD.

The initial-goal answer is the first sentence of the first user turn.

Neither requires new crowdsourcing.

Probe-specific evaluations

The probe-specific evaluations rank compression methods differently.

On the joint probe at a retained fraction of 0.30, every tested compressed method remains below full context with the main Llama reader.

Deleting the turn that carries the update sharply lowers current-value accuracy, while deleting one matched irrelevant turn leaves it unchanged.

A Mistral reader

A Mistral reader repeats the P2/P3 rankings and the joint-probe gap.

Current-value recovery

Current-value recovery is tested on an additional corpus, LongMemEval-KU, and on Chinese RiSAWOZ:

full context has the highest accuracy, and recency has the highest compressed-method mean in both evaluations.