Skip to content

Auditing what your AI said: the log, the weekly read, and what to look for

You are accountable for what the AI says to customers. What a proper log contains, how to read a sample of conversations each week without it becoming a job, and the patterns that mean a rule or a fact needs changing.

Frontiva · · 4 min read

You are responsible for what an AI receptionist says in your name, so you need to be able to read it. A proper log records every message the AI sent, what facts it based each on, which guardrails fired, when it handed off and why, and what a person changed in approval mode. A weekly read of twenty conversations, chosen partly at random and partly from the flagged ones, takes twenty minutes and finds most problems before customers do.

What the log must contain

For every AI reply: the customer's message, the reply as sent, the facts retrieved and used, the guardrails that ran and their result, whether the reply was approved or edited by a person and what changed, and the channel and time. For every handoff: the trigger, the wording sent to the customer, who it was assigned to and when they first replied. For every action: the booking, the payment link, the template sent, with the parameters.

If a vendor cannot show this for a sample conversation, the product does not have it, and you cannot audit what you cannot see.

The weekly read

Twenty conversations: ten at random, five from the handoffs, five from the edited-in-approval-mode list. Read them as the customer. For each, four questions: was every fact correct, was the tone yours, did it hand off when it should have, and did it say anything you would not want screenshotted. Twenty minutes. Note anything that fails and fix the cause (a fact, a rule, the tone paragraph), not the reply.

Patterns that mean a fact is wrong or missing

The same question answered "I am not sure" repeatedly (a gap; see the gaps log). A price that differs from your current list (a stale fact). An answer that is correct for one branch and wrong for another (a fact that should be per location). A confident answer to something that should have been a gap (a grounding failure; raise it with the vendor).

Patterns that mean a rule needs changing

A handoff that should have been automatic and was not (add the trigger). A handoff that fired on something harmless every time (tune the trigger). Handoffs that sit for hours (an assignment problem, not an AI one). A person's edits in approval mode that keep changing the same phrase (a tone rule).

Patterns that mean stop and look properly

Any reply containing another customer's details. Any reply that gave medical, legal or financial advice. Any action the AI took that it should not have had a tool for. Any sign that a customer's instructions changed the AI's behaviour. These are rare in a well-built system and serious when they happen. Pull the full log for the conversation, tell the vendor, and fix the cause before the AI answers again in that category.

Retention and access

Logs are records of customer conversations and are kept with the customer's record, subject to your retention policy. Access to them is limited to people who need it, and the log of who viewed what is itself kept, especially for medical and legal practices. Exports should include the audit detail, not just the messages.

Sharing the audit

A one-page monthly summary: conversations handled, share handed off, share edited in approval mode, gaps filled, anything serious found and fixed. It is the document that lets an owner say, honestly, that they know what the AI is doing in their name.

Frequently asked questions

Is twenty conversations a week enough?

For a single-location business, yes, once the first month's supervised period is done. Multi-location businesses read twenty per branch, by the branch manager, with the owner reading the summaries.

Should the AI audit itself?

It can flag conversations worth a person's look (a guardrail fired, a customer seemed unhappy, a reply was edited). It should not be the only reader. The point of the audit is a human judgement about the business's voice.

What do I do with what I find?

Fix causes: facts, rules, tone, assignment. Keep a short log of the fixes. Over a year it is a record of the AI getting better at your business, which is what an auditor, a regulator or a buyer would want to see.

What Frontiva does here

In Frontiva every AI reply shows the source it answered from, analytics include an AI report of what the agent answered, what it escalated and what it could not answer at all, and every conversation stays on the contact's timeline for the weekly read. Branch managers see their own branch's figures. See approval mode for the first month.

Start free. Be live this week.

Built for dental, med spa, home services, and the businesses that live on the phone.