AI support tools fall into four categories: scripted chat widgets, helpdesk built-in AI, specialist automation layers, and enterprise agent platforms. Comparing them on feature lists is unproductive. Compare them on your own hardest questions, with your own definition of a resolution, and track reopens afterwards.
Key takeaways
- The category label tells you almost nothing about what a tool can resolve.
- Test on your real questions, never on the vendor's demo set.
- Ask how a resolution is defined before comparing any quoted rate.
- Escalation design matters more to customers than answer quality.
Every tool in this category is described as AI customer support, which makes the label useless for choosing between them. Here is a more practical map, followed by the questions that separate options faster than any feature comparison.
The four categories
Scripted chat widgets. Decision trees with a chat interface. Cheap, fast to deploy, and they break the moment a customer phrases something outside the script. Products like Tidio, Crisp, and Tawk.to live here, some with language-model features layered on. Good for small sites with a narrow question set. Poor for anything open-ended.
Helpdesk built-in AI. Capability included in or added to a helpdesk you already run: Zendesk, Freshdesk, Intercom, HubSpot Service Hub. The advantage is one vendor and no integration. The disadvantage is that depth varies and the pricing is frequently a separate add-on stacked on your seat licence.
Specialist automation layers. Products that sit on top of an existing helpdesk and add resolution: Ada, Forethought, Ultimate. These add a budget line rather than replacing one, and typically involve an implementation programme.
Enterprise agent platforms. Decagon, Sierra, and similar. Language-model-native agents that answer and take actions, sold through quoted agreements to large organisations with a dedicated owner.
Fidiora sits between the third and fourth categories deliberately: grounded resolution with published pricing and no implementation programme.
Why feature lists do not help
Every vendor in every category will tell you they resolve customer questions, ground answers in your content, and escalate cleanly. Most of them are describing something true and incomparable.
The variance that matters is not in the feature list. It is in how each system behaves on your specific questions, with your specific documentation, when the customer phrases things badly.
The six questions
Ask these of every vendor, and ask for written answers.
1. What happens with an unanticipated phrasing? This single question separates scripted from generative systems faster than anything else. Ask them to demonstrate with a question you write during the call.
2. Where does the answer come from? Scripts, an intent library, or retrieved passages from your documentation. Three different risk profiles. Ask whether the system can answer from general model knowledge when retrieval fails, because that is where invented policies come from.
3. What happens when it does not know? A system that always answers is not grounded, it is confident. An honest refusal path lowers the resolution rate and is the feature that makes deployment safe.
4. How does it escalate? Customers judge AI support almost entirely on this. Does the handoff carry the full transcript, the customer context, and what the AI already tried? Can a customer reach a human on request without negotiating?
5. How is a resolution defined and billed? Does abandonment count? Are handoffs billed? Does a reopen retract the charge? Can you audit a sample? A vendor quoting seventy percent and a vendor quoting thirty-five may be describing identical performance under different definitions.
6. How long until we are live? Not the demo, the deployment. Ask for a reference timeline from a customer of your size. Enterprise platforms measured in weeks are a poor fit for a five-person team, regardless of quality.
How to run the actual evaluation
The vendor comparison is not the evaluation. This is:
- Take fifty real questions from your last month of tickets. Include the awkwardly phrased ones, the ones that span two topics, and the ones your own agents found hard.
- Write your resolution definition first, before any vendor sees it. Ours would be: the customer got a correct answer, the issue closed, and they did not come back about it within fourteen days.
- Run each candidate on the same set. Score against your definition, not theirs.
- Track repeat contacts for two weeks after any live pilot. This is the honesty check on everything else.
- Read twenty transcripts yourself. Aggregate scores hide the failure modes that matter, and you will learn more in an hour of reading than from any report.
Fix your content before you evaluate anything
Whatever you deploy will faithfully reproduce your documentation, including its errors.
Audit for contradictions on the topics customers act on: refunds, pricing, cancellation, security. Then check coverage against your top thirty ticket topics. Doing this first makes every vendor look better and makes the comparison meaningful, because you are testing the system rather than testing your content.
We wrote a full pre-deployment knowledge base audit if you want the checklist.
What good looks like after three months
- Cost per resolution has fallen.
- Reopen rate is flat or better.
- Satisfaction is flat or better.
- Your team spends less time on the top three repetitive topics.
- You have a list of documentation gaps generated from real customer questions.
If deflection is up and satisfaction is down, the deployment is failing regardless of what the dashboard says. That combination means customers are giving up rather than being helped, and we wrote about why deflection is the wrong metric.
Frequently asked questions
What is the best AI customer support tool?
How do I compare AI support vendors fairly?
Are Ada competitors worth evaluating?
Should I use my helpdesk's built-in AI?
Resolve, don't deflect.
See Fidiora resolve a ticket, capture a lead, and keep the bill predictable.