Skip to content
56NorthExperts
All insights

Run & maintain

Maintaining AI agents in production: what degrades, and how to keep control

Pascal Mennesson
By Pascal Mennesson

Founder, 56North · 1 October 2026 · 3 min read

Most AI projects on enterprise platforms are judged on the go-live date. Then the team moves on to the next use case, and the first agent is left to run. Within a few months, answers get worse, costs go up and nobody can say exactly why. An agent built on Agentforce, Copilot Studio, Joule or Now Assist needs maintenance like any production system, with a few failure modes of its own.

What degrades after go-live

  • Knowledge goes stale. The agent answers from the content it is grounded on: knowledge articles, product data, policies. When a price, a procedure or an offer changes and the source is not updated, the agent keeps giving the old answer, confidently.
  • The platform changes under you. Vendors update their AI features frequently, sometimes including the underlying model or the way the agent plans its steps. Salesforce, for example, ships three major releases a year. A change you did not make can alter how your agent behaves.
  • Users find new uses. People ask the agent things it was not designed for. Some of those questions fall outside its instructions, and that is where answers become improvised, or wrong.
  • Permissions drift. New data sources are connected, new actions are added, access rights are widened "temporarily". Each change widens what the agent can read or do.
  • Costs drift. Longer conversations, more actions per request and more users all raise consumption. Without a cost per conversation tracked monthly, the bill becomes the first alert.
  • Integrations break. An API changes, a field is renamed, a system is migrated. The agent may not fail visibly: it may simply stop using the data and answer without it.

The risks that come with drift

Two risks deserve special attention because they grow over time:

  • Data exposure. An assistant that searches company content shows users everything they are technically allowed to see. If permissions on shared drives or sites are too broad, the assistant makes that oversharing visible in seconds.
  • Prompt injection. Text hidden in a document, an email or a web page can carry instructions that the agent follows. OWASP ranks it as the first risk for applications built on language models. The more data sources and actions an agent has, the larger the exposure.

A monthly routine that keeps control

Frequency What to do Who
After every change, including vendor releases Replay a fixed set of test conversations and compare the answers Platform team
Weekly Read a sample of real conversations, especially escalations and negative feedback Agent owner
Monthly Review four indicators: resolution or deflection rate, escalation rate, user feedback, cost per conversation Agent owner and business sponsor
Quarterly Review permissions, data sources and actions; remove what is no longer needed Platform team and security

Keep a change log for every agent: what changed, when, why, and who approved it. It is the first document an auditor will ask for, and the fastest way to understand a sudden change in behaviour.

Who owns the agent

Give each agent a named owner, a person rather than a team, who answers three questions at any time: what does the agent do, how well is it doing it, and what changed last. For agents used in sensitive processes, the EU AI Act makes this monitoring an obligation: deployers of high-risk systems must monitor their operation and report serious incidents.

Questions and answers

How often should an AI agent be tested after go-live?

After every change, including vendor releases, replay a fixed set of test conversations. Add a weekly review of a sample of real conversations and a monthly review of resolution, escalation, feedback and cost indicators.

Why does an AI agent get worse over time?

Its knowledge sources age, the vendor updates the platform and sometimes the model, users bring new kinds of questions, and integrations change. None of these shows up as an error message.

Who should own an AI agent in production?

A named person who can say at any time what the agent does, how well it performs and what changed last. A team is not an owner.

Sources

The 56North platform

Measure the AI you run. Prove you control it.

The 56North Cockpit lists the AI systems in service across your company, tracks them on five dials (reliability, costs, AI Act evidence, usage, reference data) and gathers the dated evidence the regulation requires.

Discover the 56North Cockpit

Need this expertise on your project?

Free to brief. A practice lead replies within one business day.

Request experts