首页 > 资讯 > WDCD Run #316: Grok 4 Leads Multi-Turn Commitment Benchmark as Claude Sonnet 4.6 Sets Decay Resistance Record

WDCD Run #316: Grok 4 Leads Multi-Turn Commitment Benchmark as Claude Sonnet 4.6 Sets Decay Resistance Record

赢政天下 2026-09-09 05:12 3 阅读 查看原文

The Winzheng Dynamic Contextual Decay (WDCD) benchmark measures how AI models' commitment to user instructions decays across multi-turn dialogue. In Run #316, executed on 2026-09-09 across 11 frontier models, Grok 4 finished first with 93.8 points, followed closely by GPT-o3 (89.8) and Claude Sonnet 4.6 (87.2).

Top 3 Ranking:

  • Grok 4 — 93.8 pts (decay: -50%)
  • GPT-o3 — 89.8 pts (decay: -50%)
  • Claude Sonnet 4.6 — 87.2 pts (decay: -100%)

WDCD scoring is fully rule-based with zero AI judges. Each run consists of 30 questions distributed across five real-world scenarios: data_boundary, resource_limit, business_rule, security, and engineering. Models are evaluated across three rounds: R1 confirms initial instruction acknowledgment, R2 introduces distractor material in the form of 2,000–5,000 word professional documents, and R3 performs a final constraint integrity check.

Decay patterns. The aggregate average commitment decay across the 11-model cohort was 0% from Round 1 to Round 3, indicating that the field as a whole held its ground between the opening and closing checks — though this aggregate figure masks significant variance at the individual model level. Both Grok 4 and GPT-o3 registered -50% decay, meaning meaningful mid-dialogue slippage even as their absolute scores remained high. Claude Sonnet 4.6, despite recording a -100% decay figure, produced the strongest decay resistance profile in the run under WDCD's scoring methodology.

Worst performer on decay. Qwen3 Max recorded the deepest instruction decay of the cohort at -100%, marking it as the model most susceptible to constraint erosion once distractor documents were introduced in R2.

Interpreting the results. WDCD is designed to isolate multi-turn commitment from single-turn instruction-following ability. A high R1 score paired with a steep R2–R3 drop indicates a model that accepts constraints but fails to retain them under contextual load. Conversely, models that resist instruction decay across all three rounds demonstrate durable adherence — a property increasingly relevant to agentic and long-context deployments where user rules must survive extended reasoning chains.

Run #316 confirms that top-line leaderboard position and decay resistance are distinct axes: Grok 4 leads on total score, while Claude Sonnet 4.6 leads on decay-profile stability within the top tier.

Methodology: https://www.winzheng.com/yz-index/methodology
Data API: https://www.winzheng.com/yz-index/api/v1/dcd