首页 > AI前沿 > ORCA-bench: How Ready Are Language Model Agents for Oncall?

ORCA-bench: How Ready Are Language Model Agents for Oncall?

arXiv自然语言 2026-07-31 01:14 6 阅读 查看原文

Large language models can write, patch, and search code, but oncall root cause analysis (RCA) demands something different: reasoning over noisy metrics, logs, traces, and source code, starting from ambiguous user-facing reports, often hours after the incident began.

We introduce ORCA-bench, a benchmark that puts general-purpose coding agents in a production-fidelity oncall setting.

ORCA-bench pairs 1,079 RCA tasks with six days of metrics, logs, and traces collected from an OpenTelemetry-instrumented microservice system under continuous simulated user load.

Agents investigate this recorded history through real observability interfaces---Prometheus, Jaeger, and OpenSearch via Grafana---with full access to application source code.

Tasks systematically vary report specificity, time-to-detection, and co-occurring fault scenarios.

Ground-truth symptoms are curated and signed off by expert SREs, and our LLM-as-judge is independently re-scored by humans (Cohen's $κ_w = 0.91$).

Across five frontier agents, the best RCA Accuracy is 25.3% on Medium-difficulty tasks (the realistic-input setting) and 10.0% on Hard---a gap that remains even with Claude Fable 5.

The weakest model hallucinates an implausible root cause in 40% of incident reports, and removing source-code access reduces RCA accuracy and increases the hallucination rate for every evaluated model.

These results come from a curated 50 GB / six-day testbed of standalone tasks on a system whose code and instrumentation are public.

Since real production systems are orders of magnitude larger, more dynamic, and more idiosyncratic, the gap we report underscores the engineering work still needed before agents can be entrusted with production reliability.

We release the public set at https://hub.harborframework.com/datasets/orca-bench/orca-bench.