Verifiable instruction-following benchmarks often express each constraint through one fixed template.
We test whether scores remain stable when the operational requirement is unchanged but its wording varies.
We introduce WISE, a matched evaluation suite and reporting protocol instantiated on exact word count, keyword inclusion exactly once, and an inclusive 8--12 word range.
实验结果
Across 100 matched tasks, up to thirteen models from seven providers, and repeated generations scored over the complete visible output, wording alone produces substantial compliance shifts.
In an avoidance-family panel, five avoidance and exclusion forms fall below the positive baseline, while constructional controls also shift compliance substantially:
- in the nine-model control panel, compliance is 54.9% for the original positive form,
- 48.2% for a longer positive form,
- 36.7% when the target appears later,
- and 33.8% for AVOID1.
A strict JSON-structure probe shows wording sensitivity beyond counting, with a different direction of effect.
Effect sizes, failure directions, weakest forms, and model rankings vary across realizations.
Under the most disruptive exclusion form, the top-ranked model changes and 24.1% of strictly ordered model pairs reverse.
Human validation further shows that unanimous agreement on an exact-count interpretation can coexist with substantially different model behavior.
WISE补充
WISE supplements conventional scores with mean and worst-form compliance, wording gaps, failure profiles, and ranking stability.