Why Enterprise Agent Deployments Fail: A Field Taxonomy
DOI:
https://doi.org/10.15680/IJCTECE.2025.0805037Keywords:
AI agents, agentic AI, enterprise AI, human-AI systems, reliability, AI governance, AgentOps, sociotechnical systems, evaluation, incident analysisAbstract
Enterprise agents retrieve privileged data, select workflow steps, invoke tools, and modify organizational state, yet public evidence conflates canceled pilots, incorrect runs, unsafe actions, and uneconomic services. This paper develops an evidence-grounded field taxonomy of enterprise-agent deployment failure from a structured synthesis of academic research, technical benchmarks, enterprise surveys, standards, and public incidents available by June 30, 2025. Failure is defined at four nested levels (run, service, workflow, program) and attributed across eight causal domains spanning mission and value logic, control specification, state integrity, deliberation, tool execution, coordination and supervision, runtime reliability, and organizational absorption. The evidence does not support a universal failure rate, but it shows that short-task model competence is insufficient evidence for reliable, repeated, policy-constrained, tool-using work, and that consequential failures emerge at system boundaries that lack proportional assurance. The paper contributes a causal cascade, incident codebook, testable propositions, and deployment gates linking evaluation to authority, reversibility, observability, economics, and organizational learning.
References
[1] Gartner, "Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027," Gartner Newsroom, Jun. 25, 2025. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
[2] S&P Global Market Intelligence, "Generative AI experiences rapid adoption, but with mixed outcomes: Highlights from VotE: AI & Machine Learning," May 30, 2025. [Online]. Available: https://www.spglobal.com/market-intelligence/en/news-insights/research/ai-experiences-rapid-adoption-but-with-mixed-outcomes-highlights-from-vote-ai-machine-learning
[3] McKinsey & Company, "The state of AI: How organizations are rewiring to capture value," Mar. 12, 2025. [Online]. Available: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value
[4] Boston Consulting Group, "Where's the value in AI?" Oct. 24, 2024. [Online]. Available: https://www.bcg.com/publications/2024/wheres-value-in-ai
[5] L. Boisvert et al., "WorkArena++: Towards compositional planning and reasoning-based common knowledge work tasks," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, 2024.
[6] T. Xie et al., "OSWorld: Benchmarking multimodal agents for open-ended tasks in real computer environments," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, 2024. [Online]. Available: https://arxiv.org/abs/2404.07972
[7] S. Yao, N. Shinn, P. Razavi, and K. Narasimhan, "Tau-bench: A benchmark for tool-agent-user interaction in real-world domains," arXiv preprint arXiv:2406.12045, 2024.
[8] Anthropic, "Building effective agents," Dec. 19, 2024. [Online]. Available: https://www.anthropic.com/engineering/building-effective-agents
[9] OpenAI, "A practical guide to building agents," Apr. 2025. [Online]. Available: https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf
[10] L. Wang et al., "A survey on large language model based autonomous agents," Front. Comput. Sci., vol. 18, art. no. 186345, 2024.
[11] L. Bainbridge, "Ironies of automation," Automatica, vol. 19, no. 6, pp. 775-779, 1983.
[12] D. L. Goodhue and R. L. Thompson, "Task-technology fit and individual performance," MIS Quart., vol. 19, no. 2, pp. 213-236, 1995.
[13] N. G. Leveson, Engineering a Safer World: Systems Thinking Applied to Safety. Cambridge, MA, USA: MIT Press, 2011.
[14] J. Rasmussen, "Risk management in a dynamic society: A modelling problem," Safety Sci., vol. 27, no. 2-3, pp. 183-213, 1997.
[15] D. Sculley et al., "Hidden technical debt in machine learning systems," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 28, 2015.
[16] Deloitte, "The state of generative AI in the enterprise, Q3 report," Aug. 27, 2024. [Online]. Available: https://www.deloitte.com/uk/en/about/press-room/deloitte-ai-institute-state-of-generative-ai-in-the-enterprise-report.html
[17] Gartner, "Gartner survey finds generative AI is now the most frequently deployed AI solution in organizations," Gartner Newsroom, May 7, 2024. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2024-05-07-gartner-survey-finds-generative-ai-is-now-the-most-frequently-deployed-ai-solution-in-organizations
[18] Gartner, "Gartner survey finds 45% of organizations with high AI maturity keep AI projects operational for at least three years," Gartner Newsroom, Jun. 30, 2025. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2025-06-30-gartner-survey-finds-forty-five-percent-of-organizations-with-high-artificial-intelligence-maturity-keep-artificial-intelligence-projects-operational-for-at-least-three-years
[19] H. Mayer, L. Yee, M. Chui, and R. Roberts, "Superagency in the workplace: Empowering people to unlock AI's full potential," McKinsey & Company, Jan. 28, 2025. [Online]. Available: https://www.mckinsey.com/capabilities/tech-and-ai/our-insights/superagency-in-the-workplace-empowering-people-to-unlock-ais-full-potential-at-work/
[20] IBM, "Data suggests growth in enterprise adoption of AI is due to widespread deployment by early adopters," IBM Newsroom, Jan. 10, 2024. [Online]. Available: https://newsroom.ibm.com/2024-01-10-Data-Suggests-Growth-in-Enterprise-Adoption-of-AI-is-Due-to-Widespread-Deployment-by-Early-Adopters
[21] Gartner, "Lack of AI-ready data puts AI projects at risk," Gartner Newsroom, Feb. 26, 2025. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk
[22] J. Ryseff, B. De Bruhl, and S. Newberry, "The root causes of failure for artificial intelligence projects and how they can succeed," RAND Corporation, Santa Monica, CA, USA, Rep. RRA2680-1, 2024. [Online]. Available: https://www.rand.org/pubs/research_reports/RRA2680-1.html
[23] McKinsey & Company, "The state of AI in early 2024: Gen AI adoption spikes and starts to generate value," May 30, 2024. [Online]. Available: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-2024
[24] LangChain, "State of AI agents," Nov. 14, 2024. [Online]. Available: https://www.langchain.com/resources/state-of-ai-agents-report
[25] G. Mialon, C. Fourrier, C. Swift, T. Wolf, Y. LeCun, and T. Scialom, "GAIA: A benchmark for general AI assistants," in Proc. Int. Conf. Learn. Represent. (ICLR), 2024.
[26] K.-H. Huang et al., "CRMArena-Pro: Holistic assessment of LLM agents across diverse business scenarios and interactions," arXiv preprint arXiv:2505.18878v1, 2025.
[27] M. Cemri et al., "Why do multi-agent LLM systems fail?" arXiv preprint arXiv:2503.13657v2, 2025.
[28] F. F. Xu et al., "TheAgentCompany: Benchmarking LLM agents on consequential real world tasks," arXiv preprint arXiv:2412.14161v2, 2025.
[29] A. Drouin et al., "WorkArena: How capable are web agents at solving common knowledge work tasks?" in Proc. 41st Int. Conf. Mach. Learn. (ICML), vol. 235, 2024, pp. 11642-11662.
[30] C. Ma et al., "AgentBoard: An analytical evaluation board of multi-turn LLM agents," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, 2024.
[31] E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramer, "AgentDojo: A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, 2024.
[32] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, "Not what you've signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection," in Proc. 16th ACM Workshop Artif. Intell. Secur. (AISec), 2023, pp. 79-90.
[33] Q. Zhan, Z. Liang, Z. Ying, and D. Kang, "InjecAgent: Benchmarking indirect prompt injections in tool-integrated large language model agents," in Findings Assoc. Comput. Linguistics: ACL 2024, 2024, pp. 10471-10506.
[34] Y. Ruan et al., "Identifying the risks of LM agents with an LM-emulated sandbox," in Proc. Int. Conf. Learn. Represent. (ICLR), 2024.
[35] Moffatt v. Air Canada, 2024 BCCRT 149, Civil Resolution Tribunal of British Columbia, 2024. [Online]. Available: https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
[36] Slack, "Slack security update," Aug. 21, 2024. [Online]. Available: https://slack.com/blog/news/slack-security-update-082124
[37] National Vulnerability Database, "CVE-2025-32711," National Institute of Standards and Technology, 2025. [Online]. Available: https://nvd.nist.gov/vuln/detail/CVE-2025-32711
[38] J. Reason, "Human error: Models and management," BMJ, vol. 320, no. 7237, pp. 768-770, 2000.
[39] J. D. Lee and K. A. See, "Trust in automation: Designing for appropriate reliance," Human Factors, vol. 46, no. 1, pp. 50-80, 2004.
[40] R. Parasuraman, T. B. Sheridan, and C. D. Wickens, "A model for types and levels of human interaction with automation," IEEE Trans. Syst., Man, Cybern. A, Syst. Humans, vol. 30, no. 3, pp. 286-297, 2000.
[41] H.-P. Lee et al., "The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers," in Proc. 2025 CHI Conf. Human Factors Comput. Syst. (CHI), 2025, art. no. 1121, pp. 1-22.
[42] S. Amershi et al., "Software engineering for machine learning: A case study," in Proc. 41st Int. Conf. Softw. Eng.: Softw. Eng. Pract. (ICSE-SEIP), 2019, pp. 291-300.
[43] N. Sambasivan, S. Kapania, H. Highfill, D. Akrong, P. Paritosh, and L. M. Aroyo, "'Everyone wants to do the model work, not the data work': Data cascades in high-stakes AI," in Proc. 2021 CHI Conf. Human Factors Comput. Syst. (CHI), 2021, art. no. 39.
[44] Y. Dong, Q. Lu, and L. Zhu, "AgentOps: Enabling observability of LLM agents," arXiv preprint arXiv:2411.05285, 2024.
[45] OWASP Foundation, "OWASP Top 10 for LLM applications 2025," Nov. 18, 2024. [Online]. Available: https://owasp.org/www-project-top-10-for-large-language-model-applications/assets/PDF/OWASP-Top-10-for-LLMs-v2025.pdf
[46] N. F. Liu et al., "Lost in the middle: How language models use long contexts," Trans. Assoc. Comput. Linguistics, vol. 12, pp. 157-173, 2024.
[47] D. Ru et al., "RAGChecker: A fine-grained framework for diagnosing retrieval-augmented generation," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, 2024.
[48] X. Liu et al., "AgentBench: Evaluating LLMs as agents," in Proc. Int. Conf. Learn. Represent. (ICLR), 2024.
[49] J. Huang et al., "Large language models cannot self-correct reasoning yet," in Proc. Int. Conf. Learn. Represent. (ICLR), 2024.
[50] K. Zhu et al., "MultiAgentBench: Evaluating the collaboration and competition of LLM agents," arXiv preprint arXiv:2503.01935v1, 2025.
[51] C. Autio et al., "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile," National Institute of Standards and Technology, Gaithersburg, MD, USA, NIST AI 600-1, 2024.
[52] Information Technology, Artificial Intelligence, Management System, ISO/IEC 42001:2023, International Organization for Standardization, Geneva, Switzerland, 2023.
[53] European Commission, "AI Act enters into force," Aug. 1, 2024. [Online]. Available: https://commission.europa.eu/news-and-media/news/ai-act-enters-force-2024-08-01_en
[54] European Commission, "First rules of the Artificial Intelligence Act are now applicable," Feb. 2, 2025. [Online]. Available: https://ai-watch.ec.europa.eu/news/first-rules-artificial-intelligence-act-are-now-applicable-2025-02-02_en
[55] S. Kapoor, B. Stroebl, Z. S. Siegel, N. Nadgir, and A. Narayanan, "AI agents that matter," arXiv preprint arXiv:2407.01502v1, 2024.
[56] L. Shi, C. Ma, W. Liang, X. Diao, W. Ma, and S. Vosoughi, "Judging the judges: A systematic study of position bias in LLM-as-a-judge," arXiv preprint arXiv:2406.07791v8, 2024.
[57] L. Zheng et al., "Judging LLM-as-a-judge with MT-Bench and Chatbot Arena," in Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, 2023.

