首页 > AI前沿 > Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery

Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery

arXiv自然语言 2026-09-15 06:15 6 阅读 查看原文

Korean dialect corpora are available but not redistributable: weights may be released, while reproducing training and evaluation from the underlying data cannot be.

We ask how much of that supervision synthetic data recovers, and whether that recovery can be measured independently of the synthesis pipeline.

We contribute KoDialectBench

We contribute KoDialectBench, 1,000 items across five regions on three axes, released as identifier hashes and scoring code so users reconstruct the items from their own licensed copy.

Recovery is strongly axis-dependent: our best synthetic arm reaches 91.2% of the real-data gain on region identification but 63.7% on comprehension.

On generation

On generation the answer depends on the metric: the deployed marker lexicon reports 92.3% on dialectness and 119.3% on region match, the latter exceeding the real-data reference, whereas reference-based generation reaches 72.4%.

We find the marker metrics' scoring inventory is entirely contained in the inventory our transformation rules can emit.

Test the effect of construction access

We test the effect of construction access directly with an exact-form construction-disjoint arm that withholds 20% of marker types from the rules.

At exactly matched training size (8,600 examples) it reduces dialectness recovery from 91.8% to 8.1% and region-match recovery from 101.9% to 25.6% on the held-out marker inventory, while the three pipeline-independent measurements do not fall at all.

A complementary evaluator sweep

A complementary evaluator sweep defines metric-construction coverage (MCC) and finds measured dialectness recovery increasing monotonically as overlap rises from MCC=0 to MCC=1.

Shared construction and evaluation inventories can therefore substantially inflate estimates of synthetic-data recovery.