AI Support

How to Run an AI Support Pilot That Tells You Something

Short answer

Run an AI support pilot on one channel and a defined topic set, using your own real questions and a resolution definition you wrote first. Measure genuine resolution rate, reopen rate, and satisfaction for four weeks, and read transcripts yourself rather than relying on a vendor report.

Key takeaways

  • Write the success criteria before the vendor sees anything.
  • Use real past questions, never a curated demo set.
  • Four weeks minimum, because reopens take two to appear.
  • Read transcripts. Aggregate scores hide the failures that matter.

Most AI support pilots produce a feeling rather than a decision. Here is a design that produces a number you can act on.

Write the success criteria first

Before any vendor sees anything, write down:

Your resolution definition. Ours would be: the customer received a correct and complete answer, the issue closed, and they did not contact us about it again within fourteen days on any channel.

Your thresholds. What resolution rate would make this worth continuing? What reopen rate would make it a failure regardless of resolution rate? What satisfaction floor is non-negotiable?

Your guardrails. What must never happen? An invented policy, a customer unable to reach a human, a complaint handled by automation.

Doing this first is the single most important step. Criteria written afterwards get shaped by the result, unconsciously and inevitably.

Scope it narrowly

One channel. Chat is usually the right one: high volume, immediate feedback, and easy to switch off.

A defined topic set. Take your top three to five documented topics from your queue analysis. Everything else escalates immediately.

Real traffic. Not a test environment, not a sandbox. Real customers asking real questions, with a visible human option.

Narrow scope produces a clear answer. Broad scope produces noise and a lot of caveats.

Prepare the content first

The pilot will test your documentation as much as the vendor. That is fine as long as you know it.

Audit the topics in scope for contradictions and gaps before starting. Otherwise you will conclude the product is poor when the actual finding is that your refund policy exists in two conflicting versions. We wrote the audit checklist separately.

Run for four weeks

Two weeks measures first impressions. Repeat contacts take one to two weeks to appear, and reopen rate is the metric that makes everything else trustworthy.

Four weeks also covers a full billing cycle for most businesses, which surfaces the billing question wave that a shorter window misses.

Measure three things

Genuine resolution rate against your written definition. Not the vendor’s number, yours.

Reopen rate across channels. Match by customer and topic, not by ticket thread. A resolved chat followed by an email about the same issue is a reopen.

Satisfaction, segmented by topic and by whether the conversation was resolved or escalated.

Do not measure deflection or containment. Both count customers who gave up as successes, which means a pilot measured on them will succeed regardless of what the product actually did.

Read the transcripts

This is where the real information is and it is the step people skip.

Read twenty a week yourself. You will find things no dashboard surfaces: the customer who accepted an answer that was wrong, the phrasing that consistently confuses the system, the topic where the escalation fires too late.

An hour of reading teaches more than any report, and it also gives you the specific examples you will need when presenting the result.

The traps

Vendor-supplied test questions. They will be questions the product handles well. Use your own, from real tickets, including the awkward ones.

A curated content set. Do not build documentation specially for the pilot. Test against what you actually have, because that is what you will actually deploy against.

Running too short. Two weeks will look great and tell you nothing about reopens.

Only reading the summary. Aggregate scores hide the failure modes that determine whether this works in production.

Comparing to the wrong baseline. If your helpdesk already has AI capability, measure what it resolves first. That is the number a specialist has to beat, and most teams never establish it.

Making the decision

At four weeks you should be able to state:

  • Genuine resolution rate on the piloted topics, against your definition.
  • Reopen rate, compared to your human baseline.
  • Satisfaction, compared to your baseline.
  • The categories of failure you found, and whether they were content problems or product problems.
  • What it would cost at full volume, as cost per resolution.

If resolution is good and reopens are flat, expand. If resolution is good and reopens rose, the answers are shallow and you should investigate before expanding. If most failures were content gaps, that is a documentation project rather than a vendor rejection.

Expanding afterwards

Add topics before adding channels. A new topic tests the same configuration against different content. A new channel tests different customer expectations at the same time, which confounds the result.

And keep the weekly transcript reading going. The pilot discipline is the production discipline, and teams that drop it after go-live are the ones surprised six months later.

Frequently asked questions

How long should an AI support pilot run?
At least four weeks. Repeat contacts take one to two weeks to appear, so anything shorter measures the first impression rather than the outcome.
What should I measure in a pilot?
Genuine resolution rate against a definition you wrote first, reopen rate across channels, and satisfaction segmented by topic. Ignore deflection and containment, which count abandonment as success.
How do I scope an AI support pilot?
One channel, a defined set of topics from your highest-volume documented categories, and your real traffic rather than a test environment. Narrow scope produces a clear answer.
What makes a pilot meaningless?
Using vendor-supplied questions, running for two weeks, measuring deflection, and reading only the summary report. Any one of those will produce a positive result regardless of the product's actual quality.
ai support pilotvendor evaluationsupport automationproof of concept

Resolve, don't deflect.

See Fidiora resolve a ticket, capture a lead, and keep the bill predictable.

See Pricing