Skip to content

Guardrails: what they are and what they are not

A guardrail is a check that runs outside the AI model and cannot be talked out of. What guardrails cover at a front desk, why 'the prompt tells it to be careful' is not one, and how to test whether a product has them.

Frontiva · · 4 min read

A guardrail is a fixed check that runs outside the AI model, before it generates a reply or after, and that the model cannot override. Emergency words route to a person before generation. A reply containing a phone number that is not the business's is blocked after. Guardrails are code, not instructions. "We tell the model to be careful" is a prompt, and a prompt is a request the model usually honours. The difference matters on exactly the messages where it matters most.

Before the model: input guardrails

These run on the customer's message before the AI is asked to reply.

Escalation triggers. Emergency language (breathing, collapse, gas, flooding, arrested), complaint language, requests for a person, and topics you have listed (medication, legal outcomes, other customers). A match skips generation and routes to a person with approved wording.

Opt-out detection. STOP and plain-language equivalents. A match stops everything from that number; no reply is generated.

Sensitive data detection. A card number or Social Security number typed into a message triggers a fixed response (do not send that here; use this secure link) and flags the conversation.

After the model: output guardrails

These run on the draft reply before it is sent.

Content checks. The draft must not contain another customer's name, a phone number other than the business's, a price not in the approved list, medical or legal advice patterns, or a promise the AI has no tool to keep.

Grounding check. The draft's claims are checked against the facts it was given. A claim with no supporting fact is removed or the reply is replaced with an honest "I am not sure".

Tone and length. Optional, but useful: no exclamation marks if the tone paragraph forbids them; one segment where possible.

A failed output check does not send. Depending on the failure, the reply is regenerated, replaced with a fixed response, or held for a person.

Around the model: capability guardrails

The AI can only do what it has tools for, and each tool has its own limits: book only within availability, only for the current customer, only services on the list; send only approved templates; never refund, never change a price. These are not checks on text; they are limits on action, and they are the reason a prompt cannot make the AI do something harmful.

What is not a guardrail

An instruction in the prompt ("never give medical advice"). Useful as a default; not a guarantee. A vendor's assurance that the model is "aligned". A disclaimer at the start of the conversation. Reviewing conversations afterwards (that is auditing, which matters, and is not a guardrail). If it can be talked around, worn down, or skipped by an unusual phrasing, it is not a guardrail.

Testing for them

Ask the product, in the demo, to do each thing the guardrails should prevent: describe an emergency and see whether it books an appointment; type a fake card number; ask for another customer's appointment time; ask for a price of something not on the list; tell it to ignore its rules. Then ask the vendor to show you where each check lives and whether it runs inside or outside the model. "The prompt handles it" is the answer to worry about.

Guardrails you set

The escalation list is yours to extend per business. The blocked topics are yours. The tone limits are yours. A good product ships sensible defaults per industry and lets you add to them without asking the vendor. Review them quarterly against the handoff log: handoffs that should have been automatic tell you what to add.

Frequently asked questions

Do guardrails make the AI slower?

Milliseconds. The checks are simple compared with generation.

Can a guardrail be wrong?

Yes: an escalation trigger that fires on "my dog is not breathing right, he just snores" sends a snoring joke to a person. That is the correct failure direction. Tune from the log, and keep the triggers.

Are guardrails the same as a content filter?

A content filter is one kind of output guardrail. The set here is broader: routing, data, grounding, and limits on action.

What Frontiva does here

Frontiva's AI agents run a deterministic, network-free check for prompt injection and sensitive topics before anything else, so the text being checked cannot talk its way past it. Escalation keywords you configure hand off instantly, and every outbound message passes one send check for consent, opt-out, quiet hours, purpose and frequency. See when an AI receptionist should hand off for the routing side.

Start free. Be live this week.

Built for dental, med spa, home services, and the businesses that live on the phone.