Customers judge AI support on what happens when it cannot help. A good handoff is available on request without negotiation, carries the full transcript and what the AI already tried, and fires before the customer is frustrated rather than after three failed attempts.
Key takeaways
- The escalation path should be designed before the answers.
- Carry the transcript, not a summary. Summaries lose the wording that matters.
- Escalate on frustration signals, not only on low confidence.
- Billing for handoffs means paying a vendor for the product failing.
Teams building AI support spend their time on answer quality. Customers judge it almost entirely on what happens when the answer is not there.
That asymmetry is worth taking seriously, because it means the most important design work is the part usually left until last.
What customers actually experience
A customer with a question your documentation covers gets an instant answer and forgets the interaction. That is a success and it produces no feeling at all.
A customer with a question your documentation does not cover experiences one of two things. Either a clean transition to a person who already knows what they said, or a loop, a repeated menu, and eventually explaining everything again to someone who has no context.
The second experience is what people describe when they say they hate chatbots. It is not caused by the answering, it is caused by the failing.
Rule one: the human option is always available
A customer who asks for a person should get one, immediately, without negotiating.
The objection is that everyone will use it and deflection will collapse. In practice, customers use self-service when it works. The ones who ask for a human immediately are the ones already trained not to trust the alternative, and blocking them confirms the training.
Hiding the escalation path raises deflection and generates the complaints. It is the single most common design mistake in this category and it is entirely self-inflicted.
Rule two: carry everything
A handoff should include:
- The full transcript. Not a summary. Summaries lose the exact wording, and the exact wording is frequently where the actual problem is.
- Customer identity and account context. Plan, tier, order, whatever is relevant.
- What the AI already tried. So the agent does not repeat a suggestion that already failed.
- Why it stopped. Low confidence, explicit request, topic rule, or frustration signal.
The last two are the ones usually missing, and they are the difference between an agent picking up mid-conversation and an agent starting an investigation.
Rule three: escalate early
The instinct is to try harder before giving up. It is wrong.
A customer who gets three unhelpful answers before reaching a person is more annoyed than one who gets a single “I cannot help with that, let me get someone who can”. The failed attempts do not earn credit, they spend patience.
Escalate on:
- Explicit request. Immediately, no friction.
- Low confidence. Before answering, not after.
- Topic rules. Complaints, cancellations, exceptions, vulnerability signals. These fire regardless of confidence.
- Frustration signals. Repeated rephrasing, short replies, or sentiment shift. The customer is telling you something.
- A second failed attempt. Not a third or fourth.
Rule four: route by rule for judgement topics
Some things should never depend on a confidence score:
- Complaints of any kind
- Cancellations, which are retention conversations
- Commercial exceptions and disputes
- Any signal of financial difficulty or vulnerability
- Safety, security, and legal matters
Write these as explicit rules. A confidence threshold will occasionally not fire, and these are precisely the occasions where that matters.
Rule five: set the expectation at the handoff
Tell the customer what happens next. “I am passing this to the team, they will reply within four hours” is enormously better than a silent transfer.
If nobody is available right now, say so and say when. An honest wait is tolerable. An unexplained silence after being transferred is the thing people complain about publicly.
What this looks like when it works
The customer asks something your documentation does not cover. The assistant says it cannot confirm that and is getting a person. Within the stated window, an agent replies referencing what the customer already said, without asking them to repeat anything.
The customer’s experience is that they asked a question and got help. The automation is invisible, which is the goal.
The commercial version of this argument
If a vendor bills for handoffs, they are paid identically whether or not they resolved the issue. That removes the alignment that outcome pricing is supposed to create, and it quietly discourages escalating early.
Ask directly whether handoffs are billed. Fidiora does not bill them, along with abandoned chats and timeouts, because charging for a failure and calling it outcome pricing is not outcome pricing.
Testing the path before launch
Do this deliberately, not incidentally:
- Ask for a human as your first message. Confirm you get one without friction.
- Ask something your documentation does not cover. Confirm the system declines rather than invents.
- Check the agent view. Is the full transcript there? The context? What was tried?
- Ask a complaint-shaped question. Confirm it routes by rule.
- Rephrase the same question three times. Confirm frustration triggers a handoff.
Five tests, fifteen minutes, and they catch the failures that generate the reviews.
Frequently asked questions
When should AI hand off to a human?
What should a handoff include?
Should handoffs be billable?
How do I stop customers being frustrated by AI support?
Resolve, don't deflect.
See Fidiora resolve a ticket, capture a lead, and keep the bill predictable.