Workshop paper

Context or Capability? Debugging Agentic Workflows

Abstract

When a multi-step LLM agent fails, the terminal accuracy score saysthatsomething1 went wrong but is silent onwhereandwhy. We address this diagnostic problem with2 a structural causal framework that localizes faults to specific nodes in the workflow3 DAG and distinguishes failures of insufficient context from failures of inadequate4 decision-making, mapping each diagnosis onto a well-defined intervention. We5 validate on IT-Bench, a Kubernetes incident diagnosis benchmark, across three6 models. For GPT-OSS-120B, the framework traces failures to a long-context7 bottleneck mid-workflow; addressing it lifts end-to-end task accuracy from 59% to8 74% — +15pp absolute / 27% relative improvement.