Introduce AI into support in three stages: operational automation such as tagging and routing, then agent assistance, then customer-facing resolution on documented topics. Measure cost per resolution with reopen rate and satisfaction as guardrails, and route judgement calls to people by explicit rule.
Key takeaways
- Sequence by volume times repetitiveness, not by what demos best.
- Operational automation is low risk and produces the data for later decisions.
- Customer-facing resolution is the only stage that changes capacity.
- Every automation needs an owner and a monthly review or it drifts.
Most advice about AI in customer support is either a vendor pitch or a warning. Here is the practical version: a sequence that works, what to skip, and how to tell whether any of it is helping.
Start by mapping your actual volume
Before any tooling decision, categorise your last thousand tickets by topic and record the average handling time for each. Multiply volume by handling time and rank the result.
This takes an afternoon and it changes everything downstream. Almost every team that does this finds the ranking surprising: the topic they assumed dominated does not, and something unglamorous is consuming a third of the queue.
Without this, you will automate whatever the loudest person remembers, which is how automation projects end up demonstrably working and delivering nothing.
Stage one: operational automation
Start with the steps that carry no customer-facing risk.
Classification and tagging. Consistent topic labels on every conversation, applied automatically. Manual tagging degrades under time pressure because it competes directly with handle time, so the data you plan with is systematically wrong. Automated classification is more consistent even when it is occasionally less nuanced, and consistency is what reporting needs.
Routing and prioritisation. Send tickets to the right queue based on topic, customer tier, and urgency. This removes a decision from every ticket and applies your commercial priorities at three in the morning as reliably as at midday.
Context assembly. Pull the account details, order status, or plan information an agent needs into the ticket, so the first minutes are not spent searching three systems.
None of this is exciting and all of it pays back immediately. It also produces the clean data you need for stage three.
Stage two: agent assistance
AI that helps the human rather than replacing them. Suggested replies, retrieved documentation, conversation summarisation, and drafted notes.
The honest assessment: agent assistance compresses new-agent ramp time more than it compresses experienced-agent handle time. Summarisation at handoff is usually the highest-value feature and the least demoed one, because reading a three-line summary instead of twelve messages is where the real minutes go.
What it does not do is change capacity. Every contact still needs a human, so the ceiling is efficiency rather than headcount. Be careful about paying an AI add-on price for what is essentially better search, and measure how often agents edit suggestions before sending. A high edit rate means the suggestions are wrong, not that agents are being difficult.
Stage three: customer-facing resolution
This is the stage that changes the maths, and the one that needs the most care.
Take your ranked topic list. Find the entries that are high volume, well documented, and low risk. Those are your candidates. For most software companies that means access questions, plan and feature explanations, billing policy questions, and setup or integration how-tos.
Three things determine whether this works:
Grounding. The system must answer from your own documentation rather than from general model knowledge. An ungrounded model will answer a policy question with something plausible, and a customer will act on it. Grounding converts an unbounded risk into a content management problem you can solve.
An honest refusal path. A system that always answers is not grounded, it is confident. It must be able to say it cannot confirm something and route to a person. This lowers your resolution rate and it is the feature that makes the whole thing safe.
Escalation design. Customers judge AI support almost entirely on what happens when it cannot help. A handoff that carries the full conversation, so nobody repeats themselves, is what separates an acceptable experience from a resented one. Design this before the answers.
What to leave alone
Some categories should route to a person by explicit rule, not by confidence score:
- Complaints of any kind. A complaint handled by automation becomes a bigger complaint.
- Cancellations. These are retention conversations, and they are the highest-value moment you get.
- Commercial exceptions and disputes.
- Anything involving a vulnerable customer or financial difficulty.
- Any irreversible action with a large blast radius.
Write these down as rules. Leaving them to a confidence threshold means they will occasionally not fire, and those occasions are exactly the ones that matter.
Fix the content first
Whatever you deploy will faithfully reproduce your documentation, including its errors.
Before launch, audit for contradictions on the topics customers act on: refunds, pricing, cancellation terms, and security. Two articles disagreeing about a refund window is worse than having neither, because a grounded system will pick one and sound certain about it.
Then check coverage against your top thirty ticket topics. The gaps on that list are your content roadmap, ranked by real demand rather than by intuition. We wrote up how to run a knowledge gap analysis.
How to measure it
One metric, two guardrails.
Cost per resolution is the metric. Total fully loaded support cost divided by issues genuinely resolved. It captures both the spend and the outcome, which is why it is more defensible than any volume claim.
Reopen rate is the first guardrail. If resolutions rise and reopens rise with them, you are automating failure at scale.
Satisfaction is the second. A cost reduction that comes with falling satisfaction is deferred cost, not saved cost.
Do not report deflection or containment as success metrics. Both count customers who gave up as wins, and any programme built on them will eventually be contradicted by churn data.
Ownership and review
Every automation needs a named owner and a monthly review, or it drifts silently out of line with your product. Rules that were correct when written stop being correct when the product changes, and nobody notices because nothing throws an error.
Sample twenty AI answers weekly and check each claim against the source documentation. AI errors are systematic rather than random, which means one unnoticed pattern reaches hundreds of customers before any aggregate metric moves. That weekly half hour is the highest-return quality activity available to you.
A realistic expectation
If a large share of your volume is repetitive and documented, expect meaningful capacity recovery within a month and a changed cost curve within a quarter.
If your distribution is flat and every ticket is genuinely different, expect much less, and know that before you start. Automation removes repeated work. It cannot remove work that was never repeated.
Frequently asked questions
What should support automate first with AI?
What should never be automated in support?
Will AI replace support agents?
How do I know if AI support is working?
Resolve, don't deflect.
See Fidiora resolve a ticket, capture a lead, and keep the bill predictable.