The implementation team finishes the pilot. The agent is live. The demo went well, the project plan is closed, and the people who built the workflow move to their next engagement.
Then a connector starts timing out. The agent keeps creating work, but the CRM is not receiving it. Operations calls IT. IT calls the integrator. The integrator points to the platform vendor. The vendor says its service is up.
Everyone may be telling the truth. The workflow is still broken.
This is why support cannot be a line item labeled "ongoing assistance." Before production, decide who detects a failure, who owns the business queue, who can change the workflow, who contacts each vendor, and who has authority to stop the agent.
The agent is a chain of responsibilities
A production agent is usually several services tied together: a model, orchestration logic, business rules, data sources, connectors, identities, approval queues, monitoring, and one or more destination systems. One company rarely controls all of them.
The model provider can report healthy service while your connector is failing. The integrator can repair custom logic but may not own the source data. IT can restore access but should not decide whether a customer exception is acceptable. Operations can spot a bad outcome but may not have the evidence needed to diagnose it.
Do not ask one team to "own AI" as if that settles the issue. Assign ownership by work. If the implementation still lacks a clear component boundary, start with the build, buy, or integrate decision.
Assign six production jobs
A workable AI agent support model separates six jobs. A small organization may give several jobs to one person or provider. That is fine. The names, access, response expectations, and backup coverage still need to exist.
| Production job | What the owner does | Likely owner |
|---|---|---|
| Business outcome | Defines acceptable results, handles policy exceptions, and decides when the workflow should pause | Business operations or process owner |
| First response | Receives alerts and user reports, finds the affected transaction, classifies impact, and starts the runbook | Service desk, managed service, or internal support |
| Technical workflow | Diagnoses orchestration, prompts, rules, state, custom code, and failed tool calls | Internal engineering, integrator, or managed AI team |
| Applications and data | Owns source records, permissions, APIs, field mappings, and downstream application behavior | Application and data owners |
| Risk and security | Handles unauthorized action, data exposure, credential misuse, and control failures | Security, privacy, risk, or compliance |
| Supplier escalation | Opens vendor cases, tracks service commitments, coordinates evidence, and manages commercial follow-through | Vendor manager, IT, procurement, or integrator |
The integrator can fill several of these roles under a support agreement. It should not quietly become the business owner or risk authority. Your company still decides what outcomes are acceptable and when the agent may continue operating.
Route incidents by symptom, not vendor logo
Users will report what they can see: a missing order, wrong account update, slow response, duplicate ticket, blocked approval, or customer complaint. They will not know whether the model, connector, policy, identity, or destination caused it.
Give users one intake route. The first responder should be able to search by transaction ID, confirm the final business state, identify the failed step, and see recent changes. The AI agent audit trail checklist shows the evidence that makes this possible.
After triage, route the case to the component owner. Keep one incident owner responsible for the whole business outcome. Without that role, five suppliers can close five technically correct tickets while the customer's work remains stuck.
Set severity using business impact. A delayed internal summary does not need the same response as an agent sending unauthorized customer messages. Define which conditions require immediate containment, credential revocation, manual fallback, security review, or executive notice.
Put the handoff in the implementation scope
Do not wait until the final project meeting to ask for documentation. Make the support handoff an acceptance requirement with named artifacts and a practical test.
Require an architecture and dependency map, inventory of production identities, current configuration and version record, alert list, runbooks, escalation contacts, known limitations, recovery procedures, change process, and support coverage. Record which items your team can inspect and change without calling the integrator.
Then run a handoff drill. Break a test connector, revoke a test credential, send an out-of-policy request, and create an uncertain downstream write. Have the future support team detect, triage, contain, recover, and document each event. If the builder has to drive every step, the handoff has not happened.
This drill belongs beside the vendor pilot acceptance test. Output quality matters, but supportability is also production evidence.
Buy support by responsibility and response
Vendor support packages often describe channels, hours, initial response, and product boundaries. An integrator's managed service may cover workflow logic and connectors. Your internal teams may still own data correction, business exceptions, user communication, access approval, and incident decisions.
Put that split into a responsibility matrix. For each failure class, name the accountable owner, first responder, consulted specialists, evidence required, response target, escalation path, and authority to stop or restart the workflow.
Ask direct commercial questions:
- Which components and failure types are included in the support fee?
- Does response time mean acknowledgement, diagnosis, workaround, or restoration?
- Who manages model, prompt, policy, and connector changes after launch?
- What coverage exists outside business hours, and which incidents qualify?
- How are third-party vendor cases coordinated when root cause is uncertain?
- What logs, configurations, runbooks, and test assets can the buyer retain?
- What happens when the named specialist leaves either company?
Price the internal work too. A low-cost support plan can be expensive if your team performs every investigation, vendor chase, exception review, and recovery. Add observed support effort to the full AI agent cost model.
Separate routine support from controlled change
Fixing a failed credential is not the same as changing the prompt, swapping a model, remapping a field, or adding a new action. The first restores approved behavior. The second may change what the agent does.
Define which support actions are pre-authorized and which require testing and release approval. Emergency changes need an owner, time limit, evidence record, and follow-up review. Permanent changes should follow the AI agent release test.
Keep business policy decisions with the business. A support engineer should not rewrite a pricing exception, eligibility rule, or customer promise simply because it appears in a prompt.
NIST supports clear ownership after deployment
The NIST AI RMF Playbook says unclear responsibilities and chains of command limit effective risk management. Its GOVERN 1.5 guidance calls for ongoing monitoring and periodic review with clearly defined organizational roles and review frequency. It also describes incident response and human override as common IT management processes.
The NIST Generative AI Profile gets more specific. It recommends defining responsibilities for incident monitoring, establishing after-action reviews, keeping relevant test and validation history, and recording human oversight roles in the system inventory.
NIST does not prescribe your org chart or require a managed service. It does give buyers a sensible standard: production monitoring, response, review, and third-party risk need assigned owners. "Call the vendor" is not an operating model.
Use this support-readiness gate
Before the project closes, require a written yes to these questions:
- Does every workflow have a business owner and a technical owner?
- Can users report one failed business outcome without choosing a vendor?
- Can the first responder find the affected transaction and current system state?
- Are component owners, backups, vendor contacts, and escalation paths current?
- Are containment, manual fallback, recovery, and restart authorities documented?
- Has the future support team completed a handoff drill without the builder driving?
- Does the support agreement state what is covered, when, by whom, and with what response?
- Are routine repair and production change handled through different controls?
If those answers are vague, the agent is not ready to become an operational dependency. Keep the scope narrow, extend the handoff, or buy the support coverage the workflow requires.
The integrator may leave. The agent still has to work on Monday morning.