首页 > AI前沿 > Word-Level Text Unmixing via Evidence-Preserving Ownership Routing with Language Models

Word-Level Text Unmixing via Evidence-Preserving Ownership Routing with Language Models

arXiv自然语言 2026-10-06 00:12 3 阅读 查看原文

Text from multiple sources can become interleaved into a single sequence when attribution metadata is lost, such as overlapping speech transcripts, document reading flows, or concurrent agent streams.

We formalize this challenge as Word-Level Text Unmixing: given an interleaved lexical stream and source count K, recover the original source sequences while preserving every word occurrence and its within-source order exactly.

Directly generating separated texts with LLMs can omit, duplicate, or hallucinate words, violating this exact-reconstruction objective.

We therefore propose Evidence-Preserving Ownership Routing (EPOR), which decouples source-ownership prediction from reconstruction.

EPOR adapts a causal LLM to predict canonical ownership routes conditioned on the mixed stream and prior routing decisions.

At inference, completion-safe constrained decoding is combined with deterministic indexed reconstruction, yielding structurally valid K-source partitions that preserve every observed occurrence exactly once.

We also introduce UNMIXBENCH, covering controlled synthetic mixtures, timestamp-derived speech from AMI and ICSI, layout-derived document streams from ReadingBank, and simulated concurrent digital outputs.

Across five evaluation tracks, a 4B EPOR model achieves the lowest mean minimum-permutation word error rate among finetuned baselines, reducing the five-track mean by 22.3% relative to compact source-array generation and remaining competitive with zero-shot frontier LLMs.

These results show that when lexical evidence is fully observed, separating ownership inference from lexical regeneration provides a reliable alternative to direct generation.