What should an AI agent own?
Four questions that separate the work worth an agent from the work that will cost you
Work that is continuous, cross-system and cheap to check. Four tests that separate the workflows worth an agent from the ones that cost you.
This is the question every discovery call turns into, usually about forty minutes in, once the demo enthusiasm has worn off. Not “can an agent do this” — the answer is almost always yes, and it’s not useful. What an agent should own is a question about your business rather than about the technology.
Four tests do most of the work.
1. Is the work continuous?
An agent earns its keep on work that never stops arriving and that nobody has the capacity to keep up with. The intake queue that runs three days behind. The renewal that lapses because the reminder was a person’s memory. The document check that happens when somebody gets to it.
Work that happens twice a month isn’t an agent problem. It’s a calendar problem, and automating it produces a maintenance burden with no offsetting return. The rule of thumb: if nobody is currently behind on it, an agent won’t give you anything you don’t already have.
2. Does it cross systems a person has to bridge by hand?
The most reliable agent wins involve work where a human is currently the integration layer — reading something in one system, deciding and typing it into another. That re-keying is pure cost. It’s where errors enter, and it’s exactly what an agent with access to both sides removes.
If the work happens entirely inside one screen and one record, the honest answer is often that you want a better form, not an agent.
3. Is “correct” checkable after the fact?
This is the test people skip, and it’s the one that decides whether the deployment survives. An agent’s output has to be something a person can look at and tell whether it was right — a routed record, a drafted reply, a flagged discrepancy, a populated field. Then when it’s wrong you can find out why, fix the cause, and show the trail.
Work whose correctness only becomes apparent months later, or that depends on context nobody wrote down, doesn’t fail loudly. It fails quietly, and for a long time. Those workflows need a person, and a good implementer will say so.
4. Would you be comfortable explaining the decision rule out loud?
If the work involves a judgment you’d struggle to articulate to a colleague — whose account to prioritize, whether a client relationship can take bad news this week, how hard to push on a renewal — an agent will produce a confident answer to a question you never actually defined. Confidence without a defined rule is the failure mode, and it looks like success right up until it doesn’t.
Judgment work is not a gap waiting for a better model. It is the part of the job that is the job.
What passing all four looks like
Run the four tests over a real list and the shape falls out quickly. Most organizations find three or four workflows that pass all four, another handful that pass with a human approval gate in the middle, and a long tail that shouldn’t be touched.
The middle group is the interesting one, because it’s where hybrid workflows live: the agent prepares the work and a person releases it. Anything a customer sees, and anything with money attached, generally belongs there rather than in full autonomy — not because the agent can’t do it, but because the review is cheap and the mistake is not. The other half of this decision — which steps stay with a person, and why — is its own piece.
The tail matters too. Telling a client that a workflow should stay human shrinks the engagement. Saying it anyway is the difference between a deployment that’s still running in two years and one that got switched off in six months.
The one thing not to do
Don’t decide this from a demo. A demo shows you the technology working on a scenario chosen because it works. The decision you actually need is which of your workflows clears the four tests, and in what order. That’s a couple of weeks of looking at how work moves through your business, not an afternoon of watching software.
We run the four tests on your own queue before we scope an agent for any of it.
Want this applied to your business?
Thirty minutes with a principal turns the general case into your specific list: which workflows, what they’re worth, and in what order. The article can only tell you what we would look at; the call tells you what we found.