The word "AI" now covers so much that it no longer helps anyone decide anything. The useful question is much narrower: which task, repeated every day in your company, consists of reading something, understanding it, and then acting? That is where an agent earns its cost, and almost nowhere else.
Three families come up consistently. Answering the questions customers ask on a loop — hours, availability, pricing, order status. Extracting data from documents: supplier invoices, delivery notes, identity papers, with no retyping. And triaging inbound requests so they reach the right person at the right urgency.
The trap is the agent that invents. A language model is built to produce a plausible answer, not a true one; left unconstrained it will quote a price that does not exist or promise availability you cannot honour. The fix is not a better model but a tighter construction: it answers only from your documents, it cites its source, and it has an explicit exit — hand over to a human — when it finds nothing.
That handover is the part demos skip and reality punishes. An agent that leaves a customer waiting because nobody receives the notification costs more than having no agent at all. Build the escalation path before the conversation flow.
Finally, one metric matters: conversations genuinely closed with no human intervention, measured over a full month. Not messages handled, not response time. Start with a single task, measure it, and widen only if the number holds. An AI project that fails in a month costs a month.