Estimating Avoidable Harm for Test-Time Compute Allocation in Long-Horizon Tool-Using Language Agents: A Controlled Computational Study and External Validation Protocol
Downloads
Long-horizon tool-using language agents must decide not only what action to take, but how much inference-time computation to spend before an action whose consequences may be difficult to reverse. Existing adaptive test-time methods allocate compute using uncertainty, planning need, budget state, or task consequence, while recent counterfactual work shows that additional computation can help in one state and harm in another. This paper studies a narrower quantity: avoidable harm, defined as action consequence multiplied by the estimated reduction in failure risk obtainable from additional computation. We formulate discrete compute allocation as a budget-constrained sequential decision problem and estimate counterfactual failure risk from randomized calibration data. In a controlled stochastic-agent study with 2,500 calibration trajectories and 6,000 paired test trajectories of horizon 20, a learned avoidable-harm allocator at the same 40-unit compute budget reduced catastrophic trajectory failures from 27.55% under uniform scaling to 19.95% (difference -7.60 percentage points; 95% paired-bootstrap CI [-8.33, -6.87]) and reduced consequence-weighted failure from 6.618 to 5.887. Relative to consequence-only routing, the method further reduced catastrophic failure by 0.62 percentage points and weighted failure by 0.274 while increasing task success by 1.10 percentage points. Across 20 independent replications, the safety and weighted-loss improvements remained significant after Holm correction. The result is not an accuracy win: task success remained lower than uniform scaling (7.85% versus 10.55%), exposing a safety-utility trade-off rather than a universal dominance claim. Distribution-shift and ablation studies further show that persistent risk state and verifier signals are not uniformly beneficial. The study therefore supports avoidable harm as a useful allocation target while explicitly limiting its claims to controlled computation; we provide a pre-registered external-validation protocol for WebArena, ToolHaystack, LongCLI-Bench, OS-Harm, and related environments.
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, "ReAct: Synergizing Reasoning and Acting in Language Models," International Conference on Learning Representations (ICLR), 2023. arXiv:2210.03629.
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou, "Self-Consistency Improves Chain of Thought Reasoning in Language Models," International Conference on Learning Representations (ICLR), 2023. arXiv:2203.11171.
C. Snell, J. Lee, K. Xu, and A. Kumar, "Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters," arXiv:2408.03314, 2024.
N. Lee, L. E. Erdogan, C. J. John, S. Krishnapillai, M. W. Mahoney, K. Keutzer, and A. Gholami, "Agentic Test-Time Scaling for WebAgents," arXiv:2602.12276, 2026.
D. Paglieri et al., "Learning When to Plan: Efficiently Allocating Test-Time Compute for LLM Agents," arXiv:2509.03581, 2025.
T. Liu et al., "Budget-Aware Tool-Use Enables Effective Agent Scaling," Conference on Language Modeling (COLM), 2026. arXiv:2511.17006.
J. Wen, L. He, and Z. He, "Not All Errors Are Equal: Consequence-Aware Reasoning Compute Allocation," arXiv:2606.04402, 2026. DOI: 10.48550/arXiv.2606.04402.
Z. Li, J. Huang, X. Guo, G. Wang, and C. Zhang, "Same Signal, Opposite Meaning: Direction-Informed Adaptive Learning for LLM Agents," arXiv:2605.06908, 2026.
DOI: 10.48550/arXiv.2605.06908.
W. Si, S. Jang, I. Lee, and O. Bastani, "Conformal Constrained Policy Optimization for Cost-Effective LLM Agents," Proceedings of the AAAI Conference on Artificial Intelligence, vol. 40, no. 30, pp. 25446-25453, 2026. DOI: 10.1609/aaai.v40i30.39739.
A. N. Angelopoulos, S. Bates, A. Fisch, L. Lei, and T. Schuster, "Conformal Risk Control," International Conference on Learning Representations (ICLR), 2024.
S. Sadhuka et al., "E-valuator: Reliable Agent Verifiers with Sequential Hypothesis Testing," arXiv:2512.03109, 2025, rev. 2026.
DOI: 10.48550/arXiv.2512.03109.
D. Prinster et al., "Conformal Policy Control," arXiv:2603.02196, 2026.
DOI: 10.48550/arXiv.2603.02196.
J. Opoku and D. Banahene, "ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift," arXiv:2606.18467, 2026. DOI: 10.48550/arXiv.2606.18467.
Z. Chen et al., "Cordon: Semantic Transactions for Tool-Using LLM Agents," arXiv:2606.17573, 2026. DOI: 10.48550/arXiv.2606.17573.
C. Wu et al., "Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents," arXiv:2608.27141, 2026.
DOI: 10.48550/arXiv.2608.27141.
S. Zhou et al., "WebArena: A Realistic Web Environment for Building Autonomous Agents," International Conference on Learning Representations (ICLR), 2024. arXiv:2307.13854.
B.-W. Kwak et al., "ToolHaystack: Stress-Testing Tool-Augmented Language Models in Realistic Long-Term Interactions," Findings of the Association for Computational Linguistics: EMNLP 2025, 2025.
DOI: 10.18653/v1/2025.findings-emnlp.1344.
Y. Feng et al., "LongCLI-Bench: A Preliminary Benchmark and Study for Long-horizon Agentic Programming in Command-Line Interfaces," Findings of the Association for Computational Linguistics: ACL 2026, 2026.
DOI: 10.18653/v1/2026.findings-acl.1497.
T. Kuntz et al., "OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents," Advances in Neural Information Processing Systems 38, Datasets and Benchmarks Track, 2025. DOI: 10.52202/085713-1502.
J. Yang, S. Shao, D. Liu, and J. Shao, "RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents," Advances in Neural Information Processing Systems 38, 2025.
DOI: 10.52202/085713-0286.
T. Sadhu, Y. Chen, and A. Pesaranghader, "VestaBench: An Embodied Benchmark for Safe Long-Horizon Planning Under Multi-Constraint and Adversarial Settings," Proceedings of EMNLP 2025: Industry Track, 2025.
DOI: 10.18653/v1/2025.emnlp-industry.149.
P. S. Ponduru, "AgentMesh-MCP: A Secure and Governed Framework for Agentic AI Systems Using LLM Agents and Model Context Protocol Servers," International Journal of Scientific Research in Engineering and Management, 2026.
DOI: 10.55041/IJSREM62689.
J. Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models," Advances in Neural Information Processing Systems 35, 2022. arXiv:2201.11903.
B. Brown, J. Juravsky, R. Ehrlich, R. Clark, Q. V. Le, C. Re, and A. Mirhoseini, "Large Language Monkeys: Scaling Inference Compute with Repeated Sampling," arXiv:2407.21787, 2024.
S. Yao et al., "Tree of Thoughts: Deliberate Problem Solving with Large Language Models," Advances in Neural Information Processing Systems 36, 2023. arXiv:2305.10601.
M. Besta et al., "Graph of Thoughts: Solving Elaborate Problems with Large Language Models," Proceedings of the AAAI Conference on Artificial Intelligence, vol. 38, no. 16, pp. 17682-17690, 2024. DOI: 10.1609/aaai.v38i16.29720.
A. Madaan et al., "Self-Refine: Iterative Refinement with Self-Feedback," Advances in Neural Information Processing Systems 36, 2023. DOI: 10.52202/075280-2019.
N. Shinn, F. Cassano, A. Gopinath, K. Narasimhan, and S. Yao, "Reflexion: Language Agents with Verbal Reinforcement Learning," Advances in Neural Information Processing Systems 36, 2023. arXiv:2303.11366.
S. Hao, Y. Gu, H. Ma, J. Hong, Z. Wang, D. Wang, and Z. Hu, "Reasoning with Language Model is Planning with World Model," Proceedings of EMNLP 2023, pp. 8154-8173, 2023.
DOI: 10.18653/v1/2023.emnlp-main.507.
A. Zhou, K. Yan, M. Shlapentokh-Rothman, H. Wang, and Y.-X. Wang, "Language Agent Tree Search Unifies Reasoning, Acting, and Planning in Language Models," Proceedings of the 41st International Conference on Machine Learning, PMLR 235, pp. 62138-62160, 2024.
K. Cobbe et al., "Training Verifiers to Solve Math Word Problems," arXiv:2110.14168, 2021.
H. Lightman et al., "Let's Verify Step by Step," arXiv:2305.20050, 2023.
T. Schick et al., "Toolformer: Language Models Can Teach Themselves to Use Tools," Advances in Neural Information Processing Systems 36, 2023. DOI: 10.52202/075280-2997.
E. Karpas et al., "MRKL Systems: A Modular, Neuro-Symbolic Architecture That Combines Large Language Models, External Knowledge Sources and Discrete Reasoning," arXiv:2205.00445, 2022.
Y. Qin et al., "ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs," International Conference on Learning Representations (ICLR), 2024. arXiv:2307.16789.
M. Li et al., "API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs," Proceedings of EMNLP 2023, pp. 3102-3116, 2023. DOI: 10.18653/v1/2023.emnlp-main.187.
Y. Song et al., "RestGPT: Connecting Large Language Models with Real-World RESTful APIs," arXiv:2306.06624, 2023.
S. G. Patil et al., "Gorilla: Large Language Model Connected with Massive APIs," arXiv:2305.15334, 2023.
S. Hao, T. Liu, Z. Wang, and Z. Hu, "ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings," Advances in Neural Information Processing Systems 36, 2023. DOI: 10.52202/075280-1988.
P. Lu et al., "Chameleon: Plug-and-Play Compositional Reasoning with Large Language Models," Advances in Neural Information Processing Systems 36, 2023.
DOI: 10.52202/075280-1882.
Y. Zhuang, Y. Yu, K. Wang, H. Sun, and C. Zhang, "ToolQA: A Dataset for LLM Question Answering with External Tools," Advances in Neural Information Processing Systems 36, 2023. DOI: 10.52202/075280-2180.
S. Yao et al., "WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents," Advances in Neural Information Processing Systems 35, 2022.
DOI: 10.52202/068431-1508.
X. Liu et al., "AgentBench: Evaluating LLMs as Agents," International Conference on Learning Representations (ICLR), 2024. arXiv:2308.03688.
T. Xie et al., "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments," Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track, 2024.
A. Drouin et al., "WorkArena: How Capable are Web Agents at Solving Common Knowledge Work Tasks?" Proceedings of the 41st International Conference on Machine Learning, PMLR 235, pp. 11642-11662, 2024.
J. Y. Koh et al., "VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks," Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, pp. 881-905, 2024. DOI: 10.18653/v1/2024.acl-long.50.
C. Ma et al., "AgentBoard: An Analytical Evaluation Board of Multi-turn LLM Agents," Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track, 2024. DOI: 10.52202/079017-2365.
X. Deng et al., "Mind2Web: Towards a Generalist Agent for the Web," arXiv:2306.06070, 2023.
C. E. Jimenez et al., "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" International Conference on Learning Representations (ICLR), 2024. arXiv:2310.06770.
J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press, "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering," Advances in Neural Information Processing Systems 37, 2024. arXiv:2405.15793.
R. Wang et al., "ScienceWorld: Is your Agent Smarter than a 5th Grader?" Proceedings of EMNLP 2022, 2022. DOI: 10.18653/v1/2022.emnlp-main.775.
Y. Ruan et al., "ToolEmu: Identifying the Risks of LM Agents with an LM-Emulated Sandbox," arXiv:2309.15817, 2023.
E. Debenedetti et al., "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents," Advances in Neural Information Processing Systems 37, Datasets and Benchmarks Track, 2024. DOI: 10.52202/079017-2636.
Q. Zhan, Z. Liang, Z. Ying, and D. Kang, "InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents," Findings of the Association for Computational Linguistics: ACL 2024, pp. 10471-10506, 2024. DOI: 10.18653/v1/2024.findings-acl.624.
J. Ye et al., "ToolSword: Unveiling Safety Issues of Large Language Models in Tool Learning Across Three Stages," Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, 2024. DOI: 10.18653/v1/2024.acl-long.119.
K. Greshake et al., "Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection," arXiv:2302.12173, 2023.
C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger, "On Calibration of Modern Neural Networks," Proceedings of the 34th International Conference on Machine Learning, PMLR 70, pp. 1321-1330, 2017.
S. Kadavath et al., "Language Models (Mostly) Know What They Know," arXiv:2207.05221, 2022.
L. Kuhn, Y. Gal, and S. Farquhar, "Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation," International Conference on Learning Representations (ICLR), 2023.
S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal, "Detecting Hallucinations in Large Language Models Using Semantic Entropy," Nature, vol. 630, pp. 625-630, 2024. DOI: 10.1038/s41586-024-07421-0.
P. Manakul, A. Liusie, and M. J. F. Gales, "SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models," Proceedings of EMNLP 2023, 2023. DOI: 10.18653/v1/2023.emnlp-main.557.
Y. Geifman and R. El-Yaniv, "Selective Classification for Deep Neural Networks," Advances in Neural Information Processing Systems 30, 2017.
Y. Geifman and R. El-Yaniv, "SelectiveNet: A Deep Neural Network with an Integrated Reject Option," Proceedings of the 36th International Conference on Machine Learning, PMLR 97, pp. 2151-2159, 2019.
A. N. Angelopoulos and S. Bates, "A Gentle Introduction to Conformal Prediction and Distribution-Free Uncertainty Quantification," arXiv:2107.07511, 2021.
I. Gibbs and E. Candes, "Adaptive Conformal Inference Under Distribution Shift," Advances in Neural Information Processing Systems 34, 2021.
R. F. Barber, E. J. Candes, A. Ramdas, and R. J. Tibshirani, "Conformal Prediction Beyond Exchangeability," Annals of Statistics, vol. 51, no. 2, 2023. arXiv:2202.13415.
A. N. Angelopoulos, S. Bates, E. J. Candes, M. I. Jordan, and L. Lei, "Learn then Test: Calibrating Predictive Algorithms to Achieve Risk Control," Annals of Applied Statistics, 2025. arXiv:2110.01052.
S. Bates, A. N. Angelopoulos, L. Lei, J. Malik, and M. I. Jordan, "Distribution-Free, Risk-Controlling Prediction Sets," Journal of the ACM, vol. 68, no. 6, pp. 1-34, 2021.
R. J. Tibshirani, R. F. Barber, E. J. Candes, and A. Ramdas, "Conformal Prediction Under Covariate Shift," arXiv:1904.06019, 2019.
M. Campos, A. Farinhas, C. Zerva, M. A. T. Figueiredo, and A. F. T. Martins, "Conformal Prediction for Natural Language Processing: A Survey," Transactions of the Association for Computational Linguistics, vol. 12, pp. 1497-1516, 2024. DOI: 10.1162/tacl_a_00715.
J. Achiam, D. Held, A. Tamar, and P. Abbeel, "Constrained Policy Optimization," Proceedings of the 34th International Conference on Machine Learning, PMLR 70, pp. 22-31, 2017.
J. Garcia and F. Fernandez, "A Comprehensive Survey on Safe Reinforcement Learning," Journal of Machine Learning Research, vol. 16, no. 42, pp. 1437-1480, 2015.
M. Alshiekh et al., "Safe Reinforcement Learning via Shielding," Proceedings of the AAAI Conference on Artificial Intelligence, vol. 32, no. 1, 2018. DOI: 10.1609/aaai.v32i1.11797.
Y. Chow, A. Tamar, S. Mannor, and M. Pavone, "Risk-Sensitive and Robust Decision-Making: A CVaR Optimization Approach," Advances in Neural Information Processing Systems 28, 2015.
E. Altman, Constrained Markov Decision Processes. Boca Raton, FL, USA: Chapman & Hall/CRC, 1999. ISBN: 9780849303821.
G. Dulac-Arnold et al., "Challenges of Real-World Reinforcement Learning: Definitions, Benchmarks and Analysis," Machine Learning, vol. 110, pp. 2419-2468, 2021. DOI: 10.1007/s10994-021-05961-4.
D. B. Rubin, "Estimating Causal Effects of Treatments in Randomized and Nonrandomized Studies," Journal of Educational Psychology, vol. 66, no. 5, pp. 688-701, 1974. DOI: 10.1037/h0037350.
P. W. Holland, "Statistics and Causal Inference," Journal of the American Statistical Association, vol. 81, no. 396, pp. 945-960, 1986. DOI: 10.1080/01621459.1986.10478354.
P. R. Rosenbaum and D. B. Rubin, "The Central Role of the Propensity Score in Observational Studies for Causal Effects," Biometrika, vol. 70, no. 1, pp. 41-55, 1983. DOI: 10.1093/biomet/70.1.41.
B. Efron, "Bootstrap Methods: Another Look at the Jackknife," Annals of Statistics, vol. 7, no. 1, pp. 1-26, 1979. DOI: 10.1214/aos/1176344552.
F. Wilcoxon, "Individual Comparisons by Ranking Methods," Biometrics Bulletin, vol. 1, no. 6, pp. 80-83, 1945. DOI: 10.2307/3001968.
S. Holm, "A Simple Sequentially Rejective Multiple Test Procedure," Scandinavian Journal of Statistics, vol. 6, no. 2, pp. 65-70, 1979.
Q. McNemar, "Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages," Psychometrika, vol. 12, no. 2, pp. 153-157, 1947. DOI: 10.1007/BF02295996.
R. D. Wright and J. B. Ramsay, "On the Effectiveness of Common Random Numbers," Management Science, vol. 25, no. 7, pp. 649-656, 1979. DOI: 10.1287/mnsc.25.7.649.
J. P. C. Kleijnen, "Analyzing Simulation Experiments with Common Random Numbers," Management Science, vol. 34, no. 1, pp. 65-74, 1988. DOI: 10.1287/mnsc.34.1.65.
