Skip to content

How to evaluate an AI receptionist: what it should do, what it costs, and where it fails

The questions that separate an AI receptionist from a chatbot with a nicer greeting: what it is allowed to say, what happens when it does not know, what it actually costs once messaging is counted, and the failures worth planning for.

Frontiva · · 8 min read

An AI receptionist answers the messages your front desk cannot get to, from facts you approved, and hands over to a person when it should not answer. Everything else in a demo is decoration. This guide is the set of questions that tells you which one you are looking at, in the order they matter, including the ones vendors would rather you asked last.

We build one of these. That is a reason to read the section on where they fail more carefully than the rest, not a reason to skip it: the failures below are ours as much as anybody's.

What it actually is

A receptionist does four things: answers what is asked, takes the details, books the time, and knows when to fetch somebody. An AI receptionist that does the first and none of the other three is an FAQ widget. One that does all four but invents answers is worse than nothing, because a confident wrong answer in your practice's name is a complaint rather than a missed call.

The distinction worth holding: a chatbot generates a plausible reply, an AI receptionist retrieves an approved one. Both sound fluent. Only one of them can be held to what you said it could say.

Question one: what is it allowed to say?

Ask where the answers come from. There are three honest answers and one bad one.

From facts you entered and approved. Prices, hours, policies, services. The narrowest and the safest. The AI can only ever repeat something a person in your practice wrote down.

From documents you uploaded, after review. The vendor extracts candidate answers from your PDFs or your website and shows them to you before any of it is live. Slower to set up, much faster than typing everything.

From a general model, with your documents as context. Fluent, broad, and it will occasionally produce a sentence nobody in your practice has ever said. Whether that is acceptable depends entirely on what you do. For a restaurant, a slightly wrong description of a dish is survivable. For a dental practice answering a question about a child's anaesthetic reaction, it is not.

The bad answer is "it learns from your business". Ask what that means. Learns from what, reviewed by whom, live from when. If the vendor cannot say, the answer is that nobody knows what it will say.

Then ask the follow-up that matters more: can I see, before it sends, which of my answers it used? A reply you can trace to an approved fact is a reply you can defend. One you cannot is a reply you are hoping about.

Question two: what happens when it does not know?

This is the question that separates the products, and it is the one demos skip, because a demo is built from questions the AI can answer.

There are only two behaviours, and you should insist on knowing which you are buying:

Ask to see the second one happen. A vendor who cannot show you their product declining to answer has either not built it or does not want you to see how often it fires.

Ask also what the customer is told at that moment. Silence is the wrong answer: somebody who asked for a human and watched nothing happen for two hours has been failed twice.

Question three: approval mode, and whether you can leave it on

Most products offer a mode where the AI drafts and a person approves before anything sends. It is the right setting for the first few weeks whatever the vendor's confidence.

The questions are: can you stay in it indefinitely, and does the product still work if you do? A drafts queue nobody opens is not a front desk. Ask what the review screen looks like on a phone at 8pm, because that is when the drafts pile up.

And ask what going autonomous requires. A product that lets you switch it on by flicking a toggle on day one is a product that will eventually be switched on by somebody who should not have.

Question four: what does it cost, once messaging is counted

Software pricing is the smaller half and vendors quote it alone. The pieces:

The subscription. Usually per location per month. Straightforward.

Messaging. Texts cost money per segment and somebody pays. Ask whether you bring your own carrier account or the vendor resells you one, and get the per-message rate either way. A product with a low subscription and a marked-up message rate is more expensive at volume, which is exactly when you will be using it.

Carrier registration. In the US, business texting requires A2P 10DLC registration, which has its own one-off and monthly fees and is not the software vendor's charge. Anybody who tells you texting has no carrier cost has not registered a brand.

The included allowance, and what happens past it. Ask for the overage rate in cents, in writing, before you sign. "We will talk about it" means you will talk about it after you are dependent.

Setup time. Counted in hours of your staff's time, not the vendor's. A product that needs a week of configuration costs a week.

Question five: what happens to your customers' data

Short version of a longer list: where it is stored, who at the vendor can read a conversation and under what conditions, whether any of it trains anybody's model, how long it is kept, and how you get it out if you leave.

For a medical, dental or legal practice, add: does the vendor sign a BAA, and does the model provider behind them. A vendor who has a BAA but whose AI provider does not has given you a piece of paper that does not cover the part you were worried about.

Where these products actually fail

Five failures worth planning for, in rough order of how often they happen.

Setup was never finished. The commonest by a distance. Hours are half-entered, two services have no prices, nobody linked a staff member to the service. The AI then answers "I will check on that" to a third of questions and everyone concludes it does not work. Ask the vendor what their setup completion rate looks like and what the product does about an incomplete one.

The knowledge went stale. A price rose in March, the AI quoted the old one in June. Ask whether facts can carry an expiry, whether a promotion can be given an end date, and whether anything tells you when a source has drifted. Most products have no answer to this at all.

It was too eager. Booked something it should have escalated, or texted somebody who had asked not to be contacted. Ask what consent checks run before an outbound message, and whether they run on every path or only the obvious one.

The hand-off was a dead end. The AI escalated correctly, the conversation was flagged, and nobody was looking at the flag. Escalation without a queue somebody actually works is a thoughtful way to lose a customer.

Nobody could explain a reply afterwards. A customer says "your system told me X". Ask whether you can look up that conversation, see the reply, and see what it was based on. If the answer is no, every complaint becomes your word against theirs.

What a sensible evaluation looks like

Two weeks, not two months, and structured:

  1. Set it up with your real information. Not the demo data. The setup is the product; evaluating a pre-filled account tells you nothing about the one you would own.
  2. Run it in approval mode on a real channel. Read every draft. Keep a note of the ones you would not have sent, and why.
  3. Ask it ten things it should refuse. A question about a medical outcome, a price you never set, something a competitor does. Count how many it guesses at.
  4. Break the setup on purpose. Remove a price, unlink a staff member. See whether the product tells you, or just gets quieter.
  5. Read the outbound log. Every message it sent, why, and to whom. If there is no such log, stop here.

If it clears that, the remaining question is commercial rather than technical, and you will know what to negotiate about.

The honest summary

An AI receptionist is worth having when your front desk is genuinely losing messages, your answers are stable enough to write down, and somebody will own the queue it escalates into. It is not worth having if your prices change weekly and nobody will maintain them, or if there is no person behind it when it steps back.

The products differ far less in how well they write than in what they do when they should not write at all. Evaluate that part.

What Frontiva's AI agents will and will not answer is where our own version of the above lives, and Knowledge is where the approved facts come from.

Start free. Be live this week.

Built for dental, med spa, home services, and the businesses that live on the phone.