DoGBench: The first user-facing docs generation benchmark. No model scores >50%
AI agents can write polished documentation, but they often miss what readers need to complete a task. DoGBench evaluates whether agents can recognize when documentation needs updating and produce changes that meet expert review standards. The benchmark contains 292 tasks drawn from real open-source projects, including Helm, PostHog, Mautic, and Doc Detective. Of these, 205 require a documentation change; 87 require leaving the documentation unchanged. Tasks include creating new documentation, updating existing guidance, and deleting stale or redundant content. Our primary evaluation uses a random, stratified 117-item held-out split: 82 requiring updates and 35 requiring no change. For reproducibility, we are publishing 735 complete agent trajectories from seven non-cloud agents. Agents struggle to decide when documentation needs updating Documentation work begins with a judgment: does this change affect users, and do they need new guidance? Agents make mistakes in both directions. Some document internal refactors or performance optimizations that require no new user guidance. Others overlook necessary updates because they find no existing page about the feature, even when that missing coverage is the problem they need to solve. Among the seven model-and-harness lanes reported in the paper, the highest combined score on the 117-item held-out split was 47.3 out of 100, achieved by Qwen3.8 Max with OpenCode. The leaderboard also includes three cloud-agent systems evaluated on that same split. Polished writing can still leave readers unable to finish a task A patch can sound clear and contain correct facts while omitting a prerequisite, decision point, or essential step. It can also appear on a page readers are unlikely to find or contradict guidance elsewhere. In a separate, broader audit of 1,267 submissions: 45.5% had a task-completion gap, such as a missing prerequisite, procedural step, verification step, or recovery path. 36.6% contained technical inaccuracies. 32.5% omitted part of the central concept or reference information. These categories overlap. Fabricated content, such as invented classes, flags, or endpoints, appeared in 6.1% of submissions. With frontier agents, hallucinations can also take subtler forms: distorting real functionality by extending its scope, misstating defaults, or presenting behavior that holds only under specific conditions as universally true. Our audit classified these distortions as technical inaccuracies. The trajectory analysis also found that agents frequently missed decisive evidence or stopped investigating after finding the first plausible page to edit. Expert review remains necessary A patch is P0-clean when it has no critical failures under the rubric. Highest P0-clean delivery rate among the paper's seven lanes: 39.0%. GPT-5.6 Sol with Codex delivered a P0-clean patch for 32 of the 82 held-out tasks requiring an update. Documentation owners bring context that a code change alone may not reveal: who the reader is, what they are trying to accomplish, and how the change affects their workflow. That context helps determine whether an update is needed and what it must cover. Review an agent’s draft against that workflow. Check that the reader can find the guidance, complete the steps, and verify the result. Confirm the conditions behind each technical claim and check related pages for contradictions or stale instructions. How we evaluated agents Each agent received a repository as it existed before the change and a triggering signal, such as a code pull request or a reported documentation gap. The merged documentation was withheld. We evaluated seven agent lanes on the same task set. Task-specific rubrics assessed accuracy, completeness, reader guidance, placement, style, and repository conventions. Across the 205 documentation-needed items, the current rubrics contain 3,273 criteria, including 798 P0 criteria. DoGBench reports update decisions and patch quality separately. The scoring rules capture both: Delivered patch quality averages quality across all 82 held-out tasks requiring an update. Missed or empty patches receive zero. Correct abstention measures how often the agent leaves documentation unchanged across the 35 held-out tasks requiring no update. The combined score uses the harmonic mean of delivered patch quality and abstention recall: 2DN / (D + N). Strong performance on one cannot fully compensate for weak performance on the other. A critical failure caps a patch’s score at 60 out of 100. Each rubric criterion has equal weight, but failing any P0 criterion triggers the cap. Strengths on secondary criteria cannot average away a critical defect.