Run an AI support pilot on one channel and a defined topic set, using your own real questions and a resolution definition you wrote first. Measure genuine resolution rate, reopen rate, and satisfaction for four weeks, and read transcripts yourself rather than relying on a vendor report.
Key takeaways
- Write the success criteria before the vendor sees anything.
- Use real past questions, never a curated demo set.
- Four weeks minimum, because reopens take two to appear.
- Read transcripts. Aggregate scores hide the failures that matter.
Most AI support pilots produce a feeling rather than a decision. Here is a design that produces a number you can act on.
Write the success criteria first
Before any vendor sees anything, write down:
Your resolution definition. Ours would be: the customer received a correct and complete answer, the issue closed, and they did not contact us about it again within fourteen days on any channel.
Your thresholds. What resolution rate would make this worth continuing? What reopen rate would make it a failure regardless of resolution rate? What satisfaction floor is non-negotiable?
Your guardrails. What must never happen? An invented policy, a customer unable to reach a human, a complaint handled by automation.
Doing this first is the single most important step. Criteria written afterwards get shaped by the result, unconsciously and inevitably.
Scope it narrowly
One channel. Chat is usually the right one: high volume, immediate feedback, and easy to switch off.
A defined topic set. Take your top three to five documented topics from your queue analysis. Everything else escalates immediately.
Real traffic. Not a test environment, not a sandbox. Real customers asking real questions, with a visible human option.
Narrow scope produces a clear answer. Broad scope produces noise and a lot of caveats.
Prepare the content first
The pilot will test your documentation as much as the vendor. That is fine as long as you know it.
Audit the topics in scope for contradictions and gaps before starting. Otherwise you will conclude the product is poor when the actual finding is that your refund policy exists in two conflicting versions. We wrote the audit checklist separately.
Run for four weeks
Two weeks measures first impressions. Repeat contacts take one to two weeks to appear, and reopen rate is the metric that makes everything else trustworthy.
Four weeks also covers a full billing cycle for most businesses, which surfaces the billing question wave that a shorter window misses.
Measure three things
Genuine resolution rate against your written definition. Not the vendor’s number, yours.
Reopen rate across channels. Match by customer and topic, not by ticket thread. A resolved chat followed by an email about the same issue is a reopen.
Satisfaction, segmented by topic and by whether the conversation was resolved or escalated.
Do not measure deflection or containment. Both count customers who gave up as successes, which means a pilot measured on them will succeed regardless of what the product actually did.
Read the transcripts
This is where the real information is and it is the step people skip.
Read twenty a week yourself. You will find things no dashboard surfaces: the customer who accepted an answer that was wrong, the phrasing that consistently confuses the system, the topic where the escalation fires too late.
An hour of reading teaches more than any report, and it also gives you the specific examples you will need when presenting the result.
The traps
Vendor-supplied test questions. They will be questions the product handles well. Use your own, from real tickets, including the awkward ones.
A curated content set. Do not build documentation specially for the pilot. Test against what you actually have, because that is what you will actually deploy against.
Running too short. Two weeks will look great and tell you nothing about reopens.
Only reading the summary. Aggregate scores hide the failure modes that determine whether this works in production.
Comparing to the wrong baseline. If your helpdesk already has AI capability, measure what it resolves first. That is the number a specialist has to beat, and most teams never establish it.
Making the decision
At four weeks you should be able to state:
- Genuine resolution rate on the piloted topics, against your definition.
- Reopen rate, compared to your human baseline.
- Satisfaction, compared to your baseline.
- The categories of failure you found, and whether they were content problems or product problems.
- What it would cost at full volume, as cost per resolution.
If resolution is good and reopens are flat, expand. If resolution is good and reopens rose, the answers are shallow and you should investigate before expanding. If most failures were content gaps, that is a documentation project rather than a vendor rejection.
Expanding afterwards
Add topics before adding channels. A new topic tests the same configuration against different content. A new channel tests different customer expectations at the same time, which confounds the result.
And keep the weekly transcript reading going. The pilot discipline is the production discipline, and teams that drop it after go-live are the ones surprised six months later.
Frequently asked questions
How long should an AI support pilot run?
What should I measure in a pilot?
How do I scope an AI support pilot?
What makes a pilot meaningless?
Resolve, don't deflect.
See Fidiora resolve a ticket, capture a lead, and keep the bill predictable.