No sensible owner puts a new front-desk hire on live calls on day one. They shadow someone, they get a script, a manager listens in, and they earn more responsibility as they prove they can handle it. Yet the same owners will switch on an AI chat or phone assistant on a Tuesday afternoon and let it talk to real prospects that evening — with no test, no script review, and nobody listening. The tool isn't the risk. The missing onboarding is. Here's a practical way to QA an AI assistant before and after it goes live, borrowed from how you'd manage a person.
Write the failure cases before the happy path
Every vendor demo shows the happy path: a polite prospect asks about pricing, the bot answers, a booking appears. Real inbound is messier. A roofer's inbox gets storm-panicked homeowners asking whether insurance will cover the damage. A med spa gets people asking whether a treatment is safe with a medical condition. An HVAC company gets a furious customer whose "fixed" unit died again, a wrong-number caller, and a price-shopper collecting quotes from five companies at once.
Before launch, write down 15 to 20 realistic conversations — and make at least half of them awkward ones. Include the angry customer, the out-of-area request, the question the bot shouldn't answer, and the message that's really a complaint dressed up as a question. Then have someone on your team play each one against the assistant and log what it said. You are not checking whether the bot is smart. You're checking whether it does something sensible when the conversation stops being easy. The happy path will take care of itself; the failure cases are where your reputation lives.
Give it a lane and a hand-off rule
Decide in writing what the assistant is allowed to do and what it must never do. A reasonable lane for most service businesses: answer questions about services, hours, and service area; explain how pricing works in general terms; collect name, contact details, and job details; and book or offer appointment times. Off-limits: quoting firm prices on jobs that need eyes on site, giving anything resembling medical advice, promising completion dates, and arguing with a complaint.
The second half of the lane is the hand-off rule, and this is where most setups quietly fail. Define the situations that must go to a human — complaints, safety issues, anything medical, a job outside the lane — and define what the hand-off actually looks like. "Please call us during business hours" is not a hand-off; it's a polite dead end. A real hand-off means the bot takes the details, tells the person a human will follow up within a stated window, and a notification actually reaches someone who owns that follow-up. If nobody is on the other end of the escalation, you haven't built a safety valve — you've built a place where your hardest conversations go to die.
Read the transcripts like a manager, not an engineer
QA doesn't end at launch, because the questions your customers ask will drift with seasons, promotions, and whatever happened on the local news this week. The habit that keeps an assistant honest is the same one that keeps a front desk sharp: someone reads the conversations. Fifteen minutes a week is enough. Pull ten transcripts and ask four questions. Did it say anything wrong? Did any conversation hit a dead end? What question came up that it couldn't answer well? And did it lose a lead a human would have saved?
When you find a bad answer, resist the urge to bolt on another instruction and move on. Most bad answers trace back to missing or stale information — a price range that changed, a service you stopped offering, a policy that only exists in the owner's head. Fix the knowledge, not just the wording. And track one number over time: the share of conversations that end with contact details captured or an appointment booked. If that number drifts down, something changed — in the bot, or in what people are asking it.
What to automate, and what stays human judgement
Automation belongs in the plumbing of this process. The assistant answers around the clock — that's the point of having it. Transcripts should be collected and logged automatically. Flagging is automatable too: conversations containing complaint language, safety or medical keywords, or an abrupt ending can be tagged for review so your weekly fifteen minutes goes to the conversations that matter. And when you change the bot's instructions or knowledge, re-run your original test conversations — a regression test, in plain terms — so a fix in one place doesn't quietly break another.
The judgement stays human. You write the test cases, because you know how your customers actually talk. You set the lane, because you carry the liability when the bot steps out of it. You decide which complaints need the owner's phone call rather than a templated apology. The AI can flag; a person decides.
Treat the assistant the way you'd treat a promising new hire: a narrow job on day one, someone checking the work, and a wider lane earned over weeks — not granted at setup. The businesses that get burned by AI assistants aren't the ones that used them. They're the ones that never checked the work.
Want this built into your business, properly?
Start with the $250 Growth & AI Audit — a concrete teardown of where your leads leak, credited to setup if you start within 30 days.
Book a 10-min call →