2026-08-01

How to Design a Voice AI Pilot That Tells You Something

A two-week pilot with no success criteria teaches you nothing. How to scope, instrument, and judge a restaurant voice AI trial before you sign anything longer.

Most voice AI pilots fail to produce a decision. They run for two weeks, everyone forms a vague impression, and the choice gets made on gut feel or on whichever vendor followed up more persistently. That's not a pilot, it's an extended demo.

A pilot that earns its name has three things a demo doesn't: a defined scope, a small number of pre-committed success criteria, and a mechanism for seeing what actually happened on the calls. Set those up in advance and two weeks will give you a real answer. Skip them and two months won't.

Decide what the pilot is testing

Be specific, because different questions need different setups.

If you're testing whether AI can handle your menu, you want maximum exposure to ordering calls, including complicated ones, and you should be reading transcripts closely.

If you're testing whether it reduces staff interruption, the measurement is on your floor, not in the software: how many times per shift does someone drop what they're doing for the phone, before and after.

If you're testing whether it recovers missed calls, you need a baseline miss rate first. Our missed-call framework walks through pulling that from your existing call log. Without it you have nothing to compare against.

If you're testing whether customers accept it, you're watching hang-up rates and requests for a human, and you should be asking regulars directly.

Most operators care about all four, which is fine, but rank them. The top one determines how you configure the pilot.

Scope the traffic

Three practical patterns, in increasing order of risk and information:

After-hours only. The agent answers when you're closed. Nearly zero risk, and it tests information handling and order-ahead but not peak-rush behavior. Good first step for a nervous operator, insufficient on its own.

Overflow only. The agent picks up after your staff have had a set number of rings. This is the setup we'd recommend for most pilots. The traffic is real, the calls are exactly the ones you're currently losing, and your staff remain the first line, so a bad agent response can't cost you a call you were already going to answer.

Full front line. The agent answers everything. Most information, most risk. Reasonable for a second pilot phase once overflow has gone well, not as a starting point.

Whichever you pick, keep it consistent for the duration. Changing the routing mid-pilot destroys your ability to compare week one to week two.

Write down your criteria before day one

This is the step people skip, and it's the one that turns impressions into a decision. Three to five criteria, written down, with a number or a clear yes/no.

Reasonable examples:

Set the thresholds at a level you'd genuinely sign at. Setting them impossibly high isn't rigor, it's a way of avoiding the decision.

Instrument it so you can see what happened

You need access to transcripts, not just a summary dashboard. A vendor that shows you aggregate numbers but not the underlying calls is asking you to trust their scoring of their own product. Confirm transcript access before the pilot starts, not after.

On your side, keep a simple log. A clipboard by the pass works. Every time a staff member has to fix something the agent did, or a customer mentions it, one line. Twenty lines over two weeks tells you more than any dashboard.

Seed the hard cases deliberately

Real traffic won't cover your edge cases in two weeks, so create them. Have three different people — different voices, different accents, different speaking speeds — call in and place orders. Include:

Do this in week one, fix what breaks, and repeat the same set in week two. That repeat run is how you learn whether the vendor's fixes actually stick, which is a better signal about the vendor than the initial results were. Our guide to what to test in a demo covers the same scenarios in a shorter format.

Tell your staff, and tell them what to do

Staff who discover the AI from a confused customer will undermine it, reasonably. Brief them: what it is, why you're trying it, what to do if a caller asks for a person, and where to report anything that went wrong. Ask for their honest read at the end — they're closer to the phone than you are.

Don't ask them to hide it from customers. If a regular asks, "yes, we're trying an AI system for calls we can't get to, tell me if it's annoying" is a fine answer and often a well-received one.

Judge it against the criteria, not the impression

At the end, sit down with the written criteria and mark each one. Then, separately, note your impression. If they disagree, take it seriously in both directions: sometimes the numbers are fine and the experience is wrong for your brand, and sometimes a single bad call has colored an otherwise good result.

Also weigh how the vendor behaved during the pilot. How fast did they fix things? Did they explain failures honestly or deflect? You're evaluating a relationship, not just software, and the pilot is the only sample of that relationship you'll get before committing. That's the same lens we apply in evaluating voice AI vendors.

The bottom line

A useful pilot is small, scoped, instrumented, and judged against criteria you wrote before you started. Route overflow traffic rather than everything, run it two full weeks including your worst shift, read the actual transcripts, seed the hard cases twice, and hold the vendor to how they respond when something breaks. Do that and you'll end with a decision you can defend — including, legitimately, the decision not to buy.

More on buying guides

All buying guides articles

Frequently asked questions

Hear it answer a real call.

Call the demo line and order like a customer would, or book time and we'll walk your team through it.