Agentic AI Failure Modes: What Government Tech Buyers Must Know

The News

Ilia Razvin, founder of IOSYA, a bootstrapped AI workflow automation platform, is available for on-record commentary on the real-world failure modes of agentic AI in high-stakes industrial and enterprise environments. Razvin builds AI agents for chemical plants, industrial operations, and engineering-document pipelines, contexts where a hallucinated output has physical consequences. His commentary covers topics including context drift, silent tool-call failures, RAG poisoning, multi-modal document ingestion, and the conditions under which small models outperform frontier models on cost and latency grounds. IOSYA has grown to a profitable enterprise platform without venture funding.

Analyst Take

The agentic AI conversation in the public sector and enterprise technology markets is dominated by demos, launch events, and capability announcements. Razvin’s framing describes agentic coding as writing “like a senior engineer with dementia.” That line is uncomfortable precisely because it’s accurate. Agents that perform brilliantly in controlled demonstrations frequently degrade within days of deployment in live environments because the conditions that made the demo work, clean inputs, stable context windows, reliable tool responses, simply don’t exist in production.

Why Government Technologists Should Pay Attention

This is not an abstract engineering concern. The ECI Research Google GovTech Survey found that 49.6% of respondents estimated that 26% to 50% of their organization’s code will be assisted or generated by AI within the next 12 months. That is a substantial deployment footprint taking shape quickly. At the same time, the single largest blocker to AI adoption cited by survey respondents was FedRAMP and compliance approval friction, at 31.8%. What that combination means in practice is that agencies are being asked to absorb meaningful AI-generated code volumes while operating under approval frameworks that have not kept pace with the technology and without the production-hardening experience that operators like Razvin have accumulated through failure.

The failure modes Razvin identifies are also directly relevant to the cognitive load problem government developers already carry. According to ECI Research’s Google GovTech Survey, 48.0% of respondents selected “Navigating compliance documentation and audit evidence collection” as the greatest source of cognitive load for their developers today. An agent that fails silently, returns a confidently wrong document, or drifts in context doesn’t reduce that burden. It compounds it, because a developer now has to audit AI outputs on top of the compliance scaffolding they were already managing.

The Technical Reality Behind “It Works in the Demo”

For developers, Razvin’s specific failure taxonomy deserves close attention. Context drift is the slow degradation of agent coherence across long task chains, a particular risk in document-heavy workflows like engineering spec review or policy analysis, where a 400-page technical document is not a clean input. Silent tool-call failures are worse: the agent proceeds, produces output, and nothing in the workflow signals that the underlying call returned garbage. RAG poisoning in live enterprise systems is not a theoretical threat when the retrieval corpus includes legacy documents, inconsistently formatted records, or data that was never designed to be machine-readable.

The model-size question Razvin raises is also practically important. The instinct in government AI procurement tends toward frontier models, in part because “best available” is a defensible position in an acquisition. But Razvin’s point that there is an exact crossover where cost and latency make a smaller model the correct choice for a given agent task is an engineering decision that procurement frameworks are not currently built to accommodate. The ECI Research survey found that 31.8% of respondents cited FedRAMP friction as the top AI adoption blocker, meaning the approval cycle is already long. Adding the complexity of model-selection tradeoffs into that cycle creates a real gap between what engineers know works and what they can actually deploy.

Who Wins and Who Doesn’t

Vendors and system integrators that can demonstrate production-hardened agentic implementations, not just POC-grade demos, will have a meaningful advantage in the next wave of government AI contracts. The gap between demo performance and week-two reliability is where incumbent integrators are most exposed, particularly those whose AI practice is thin and conference-driven. Operators like Razvin who built under industrial constraints, where a wrong answer has physical consequences, have accumulated a body of failure knowledge that is genuinely scarce and commercially valuable. The bootstrapped, profitable trajectory of IOSYA also signals something worth noting: the margin structure of AI workflow automation, when built without the overhead of venture-driven growth, can be compelling.

Looking Ahead

The government technology market is heading toward a reckoning on agentic AI that will look a lot like the one the commercial enterprise market experienced with RPA a decade ago: widespread initial adoption, followed by a wave of failed deployments, followed by a more disciplined second generation of implementations built around the lessons of the first. The organizations that reach that second generation fastest will be the ones that invested early in understanding production failure modes rather than demo success rates. Razvin’s work represents exactly the kind of practitioner knowledge that should be informing acquisition criteria, not just vendor roadmaps.

For ITDMs evaluating AI investments, the immediate question is whether the vendors and integrators they are considering have genuine production failure experience to draw on. For developers, the near-term priority is instrumentation: agents that fail silently are not acceptable in compliance-heavy environments, and building observable, auditable agent pipelines is the engineering challenge that will define the next 18 months of this market.

Authors

  • Paul Nashawaty

    Paul Nashawaty, Practice Leader and Lead Principal Analyst, specializes in application modernization across build, release and operations. With a wealth of expertise in digital transformation initiatives spanning front-end and back-end systems, he also possesses comprehensive knowledge of the underlying infrastructure ecosystem crucial for supporting modernization endeavors. With over 25 years of experience, Paul has a proven track record in implementing effective go-to-market strategies, including the identification of new market channels, the growth and cultivation of partner ecosystems, and the successful execution of strategic plans resulting in positive business outcomes for his clients.

    View all posts
  • With over 15 years of hands-on experience in operations roles across legal, financial, and technology sectors, Sam Weston brings deep expertise in the systems that power modern enterprises such as ERP, CRM, HCM, CX, and beyond. Her career has spanned the full spectrum of enterprise applications, from optimizing business processes and managing platforms to leading digital transformation initiatives.

    Sam has transitioned her expertise into the analyst arena, focusing on enterprise applications and the evolving role they play in business productivity and transformation. She provides independent insights that bridge technology capabilities with business outcomes, helping organizations and vendors alike navigate a changing enterprise software landscape.

    View all posts