Estimating Avoidable Harm for Test-Time Compute Allocation in Long-Horizon Tool-Using Language Agents: A Controlled Computational Study and External Validation Protocol

agentic AI, test-time compute, tool-using agents, risk-sensitive decision making, adaptive inference, AI safety, long-horizon agents, counterfactual value of computation

Authors

September 8, 2026

Downloads

Long-horizon tool-using language agents must decide not only what action to take, but how much inference-time computation to spend before an action whose consequences may be difficult to reverse. Existing adaptive test-time methods allocate compute using uncertainty, planning need, budget state, or task consequence, while recent counterfactual work shows that additional computation can help in one state and harm in another. This paper studies a narrower quantity: avoidable harm, defined as action consequence multiplied by the estimated reduction in failure risk obtainable from additional computation. We formulate discrete compute allocation as a budget-constrained sequential decision problem and estimate counterfactual failure risk from randomized calibration data. In a controlled stochastic-agent study with 2,500 calibration trajectories and 6,000 paired test trajectories of horizon 20, a learned avoidable-harm allocator at the same 40-unit compute budget reduced catastrophic trajectory failures from 27.55% under uniform scaling to 19.95% (difference -7.60 percentage points; 95% paired-bootstrap CI [-8.33, -6.87]) and reduced consequence-weighted failure from 6.618 to 5.887. Relative to consequence-only routing, the method further reduced catastrophic failure by 0.62 percentage points and weighted failure by 0.274 while increasing task success by 1.10 percentage points. Across 20 independent replications, the safety and weighted-loss improvements remained significant after Holm correction. The result is not an accuracy win: task success remained lower than uniform scaling (7.85% versus 10.55%), exposing a safety-utility trade-off rather than a universal dominance claim. Distribution-shift and ablation studies further show that persistent risk state and verifier signals are not uniformly beneficial. The study therefore supports avoidable harm as a useful allocation target while explicitly limiting its claims to controlled computation; we provide a pre-registered external-validation protocol for WebArena, ToolHaystack, LongCLI-Bench, OS-Harm, and related environments.