Skip to content

UPDATED 09:00 EDT / OCTOBER 11 2026

AI

Agentic workloads break assumptions about software testing. Here’s how to cope

Most traditional enterprise systems were built around three assumptions: Jobs finish quickly, retrying one is free and the same input always produces the same output.

Agentic workloads break all three of these expectations, which is why pilots that perform well turn into operational problems once they run unattended. The difficulty is rarely the model; it’s the surrounding infrastructure and management practices that assume properties these workloads no longer have.

What has changed

Agents work by different rules than conventional software. Here are three significant differences:

Work takes minutes, not milliseconds. An agentic task can run long enough to exceed timeout thresholds nothing in your stack has encountered before. Systems that seemed stable start failing in ways that look mysterious until someone checks the duration.

Retries now cost money. Traditionally, retry logic has been nearly free, so teams retry generously. But every attempt against a metered model consumes compute resources whether the result is usable or not. The combination of loose quality thresholds and automatic retries creates a budget event. Because cloud and model application programming interface usage is billed asynchronously on monthly cycles, these compounded retry costs quietly accumulate and become visible only when the invoice arrives weeks later.

Failures cannot be reproduced. This is the biggest change. When an engineer investigates a bad result, the standard practice is to rerun it and watch it fail again. That doesn’t work when agents are involved. Without a step-by-step trace of agent decisions, tool calls and application programming interface runs, there is nothing to investigate and reviews become speculation.

Why pilots hide all of this

Experienced teams get caught because the evaluation environment conceals all of this.

When working interactively, humans are the error handlers. They read each result, notice problems and try again. Costs are visible because attempts are counted and fixes applied by hand. There’s no need for an audit record or a way to reproduce the exact failure.

In contrast, automated agent workflows run headlessly, with the potential for unrecorded failures to break downstream systems. A workflow that behaved reliably when a person would chat with AI and check each response behaves differently when a scheduler fires it 400 times overnight with nobody watching. The error was always there, but it was invisible because a human operator absorbed and corrected it one response at a time.

Before moving to automated production with agents, engineering leaders must identify every task the human was performing manually and specify which automated check or system will take over that responsibility.

Three risks worth your attention

Invisible costs. Spending is no longer a function of usage volume alone. Autonomous AI loops automatically retry failed tasks, regenerate responses and hit APIs repeatedly without human intervention or approval. A threshold set by a developer can move the monthly bill more than a procurement negotiation. Ask for the cost per completed unit rather than cost per API call, since the second number excludes discarded attempts.

Failures that report success. The most expensive defects are not errors but results that appear structurally valid but are substantively wrong. Probabilistic agents can’t recognize their own logical errors, so they produce correctly formatted outputs with bad data that appears structurally valid. Downstream automated systems accept and process it without triggering any alerts. Conventional monitoring misses these errors because nothing failed. AI output quality must to be measured using predefined programmatic criteria (such as assertion checks, LLM-as-a-judge rules or semantic benchmarks) rather than relying on standard system uptime or error logs.

Incidents nobody can explain. If your team cannot say which model version, inputs and settings produced a specific output, they cannot investigate it and neither can an auditor. That information is cheap to capture while the work runs and close to impossible to reconstruct later.

What to ask your team

Five questions address most of the scenarios just described:

  • What defines acceptable output, and is it written down before the work runs? Automation requires clear, reproducible acceptance criteria. If success criteria change based on individual human opinion during review, software cannot validate the output automatically.
  • How many attempts does a typical completed unit take, and is that capped? An unbounded retry loop against a vague standard is the most common source of overspending
  • What do we record for every run, and could we produce that record on request in six months? Job data should include the exact input payload, prompt/model version, timestamp, execution latency, retry counts, token cost and the final output artifact.
  • Who approves various classes of output, by name? Not a team, an individual. When something goes out wrong, that is the first question that will be asked.
  • What happens when the provider updates the model? Model updates can subtly alter response formatting, accuracy or reasoning logic. These downstream shifts can break automated pipelines, which is why model version changes must be tested in a staging environment before being deployed to production.

Where consolidation helps

Agentic work tends to be scattered because teams manage agentic tools using traditional software practices. Prompts live in one team’s repository, artifacts are in various buckets, approvals happen in chat threads and nobody can produce an execution record without an afternoon of archaeology.

The desired AI workspace is one place where runs, inputs, outputs, quality results and approvals are recorded together, functioning as an operational surface that allows teams to actively manage execution, rerun inputs, trace logs and control agent workflows in real time. This is a discipline more than a product decision. Reporting and logging must be continuous across every execution run rather than periodic, so teams can trace anomalies and maintain auditability in real time.

These same questions decide which platform anchors your AI workspace. Can the platform report on what produced an output? Can it re-execute from a saved input? Does it distinguish a failure from a refusal from a degraded result? Vendors serving production pipelines, including platforms such as ImagineArt, increasingly expose this. A polished product demo shows curated, best-case results, but it doesn’t reveal how the AI model handles edge cases, unexpected errors or continuous long-running tasks when operating automatically.

The skills conversation

One pattern is easy to miss in planning. Teams building agentic features often come from application development, where synchronous request-and-response patterns are the norm. These workloads behave like data pipelines: long-running, partially failing and expensive to re-execute.

If nobody on the team has run these kinds of systems before, the skills gap shows up as a series of surprising incidents rather than an indication that skills need to be upgraded. The skills conversation is cheaper.

Agentic AI did not create a new class of operational problem. It removed assumptions that have existed for decades, and most incidents trace back to those process gaps rather than to model quality.

Define what correct means before the work runs. Measure cost per completed unit. Record enough to investigate what you cannot reproduce. Keep a named person accountable for what ships. None of it is exotic, but none of it is what your current stack was built to do.

M. Touheed is a growth specialist at Imagine Art. He wrote this article for SiliconANGLE.

Image: SiliconANGLE/ChatGPT

A message from John Furrier, co-founder of SiliconANGLE:

Support our mission to keep content open and free by engaging with theCUBE community. Join theCUBE’s Alumni Trust Network, where technology leaders connect, share intelligence and create opportunities.

  • 15M+ viewers of theCUBE videos, powering conversations across AI, cloud, cybersecurity and more
  • 11.4k+ theCUBE alumni — Connect with more than 11,400 tech and business leaders shaping the future through a unique trusted-based network

Are you an AWS customer?  Support SiliconANGLE financially by buying your AWS services from our Marketplace portal page and links: https://siliconangle.com/aws-marketplace/

 

About SiliconANGLE Media
SiliconANGLE Media is a recognized leader in digital media innovation, uniting breakthrough technology, strategic insights and real-time audience engagement. As the parent company of SiliconANGLE, theCUBE Network, theCUBE Research, CUBE365, theCUBE AI and theCUBE SuperStudios — with flagship locations in Silicon Valley and the New York Stock Exchange — SiliconANGLE Media operates at the intersection of media, technology and AI.

Founded by tech visionaries John Furrier and Dave Vellante, SiliconANGLE Media has built a dynamic ecosystem of industry-leading digital media brands that reach 15+ million elite tech professionals. Our new proprietary theCUBE AI Video Cloud is breaking ground in audience interaction, leveraging theCUBEai.com neural network to help technology companies make data-driven decisions and stay at the forefront of industry conversations.

Send us a news tip

Send us a News Tip

  • This field is for validation purposes and should be left unchanged.
  • Max. file size: 244 MB.

Sign in

SIGN IN

Bio

Ethics statement

Extract the signal from the noise

Get SiliconANGLE updates and analysis.

Contact us

Partner with us

Contact us

Guest inquiry