Direct answer: do not evaluate an AI receptionist by listening to the vendor's demo. Evaluate it by calling it ten times yourself, from a script, and scoring what actually happened — whether it understood you, captured your details correctly, booked the right slot, handled the question it was not prepared for, and knew when to hand off. This article gives you the ten calls, a scoring sheet, and the failure modes that separate a receptionist that works from one that merely sounds good. Run it on every vendor you are considering. Run it on ours. We built the checklist expecting that.
Why test calls instead of demos?
A demo is a call the vendor chose, on a topic the assistant was tuned for, with a caller who knows the script. Your customers are none of those things. They mumble, they call from a truck with the window down, they ask two questions at once, they change their mind mid-sentence, and may ask something nobody anticipated. The only way to know how a receptionist handles that is to be that caller.
There is a second reason. The thing that matters — did a real conversation happen and did the right outcome result — is invisible in the metrics vendors usually show. A call log that says "answered, 2 min 14 s" tells you the line connected and the audio ran. It does not tell you whether the caller said a word, whether the assistant heard it, or whether the booking landed in the right calendar. We learned this operating our own phone systems: a carrier-completed call is not a conversation, a greeting playing is not comprehension, and a transcript with no caller speech in it is a failed call no matter what the connection log says. Your test calls are how you look past the log.
What should be set up before you test?
Test the configuration you would actually run, not a blank one. That means:
- Your real hours, services, and the common questions you want answered. On our system this is the information the assistant uses to "answer common questions using your hours, services, and business information." Give any vendor the same inputs.
- A calendar the assistant can book into — ideally a test calendar you can watch. Ours books "through an approved calendar connection." If the vendor's booking needs a connection they have not set up for the trial, note that; an unbookable trial cannot pass the booking tests below.
- Transfer rules: who gets what, and when. Ours transfers "according to the rules and availability you agree with us." A trial with no transfer rule cannot pass the escalation test.
- A phone you can call from that is not on file, so the assistant treats you as a new caller.
Ask the vendor to configure the trial before testing. Our published setup process includes connecting approved services and testing; the ten calls below are your own evaluation in addition to that setup.
Ask for a trial or supervised walkthrough, and confirm its price and limits before placing test calls. Obtain permission for your test and use your own test identity, phone number and calendar. The checklist does not authorize calls to customers or unannounced testing of someone else's service.
The ten test calls
Make each call from the script. Do not help the assistant. Score each one on the sheet in the next section.
Call 1 — The plain booking. "Hi, I'd like to book a [your most common service] sometime next week, afternoons are best." Give your name and number when asked. What you are testing: basic comprehension, slot offering, capture of name and number, confirmation.
Call 2 — The hours question. "Are you open Saturday? What time do you close?" Testing: whether it answers from your actual hours rather than a generic guess.
Call 3 — The price question. "How much is a [service]?" Testing: whether it gives the answer you configured, says it will have someone call back, or invents a number. An invented price is an automatic fail.
Call 4 — The mumbler. Same as Call 1, but speak quickly, use a nickname, and give your number as "five oh five, three one nine... five eight six six." Testing: capture accuracy under realistic speech. Check the captured number digit by digit afterward.
Call 5 — The interruption. Start booking, then cut in mid-sentence: "Actually, wait — do you do [a service you don't offer]?" Testing: whether it handles the change of direction and correctly says it does not offer that, rather than booking it anyway.
Call 6 — The reschedule. Call back and say "I booked for Tuesday, I need to move it." Testing: whether it can find the booking and change it, or hands off cleanly.
Call 7 — The off-script question. Ask something reasonable that you did not configure: "Do you charge extra for a second-floor unit?" Testing: what happens at the edge. The right answers are "I'll have someone confirm that and call you back" or a transfer. The wrong answer is a confident fabrication.
Call 8 — The urgent one. "I've got water coming through the ceiling right now." (Adapt to your trade.) Testing: whether it recognizes urgency and follows your transfer rule instead of offering a slot next Thursday.
Call 9 — The after-hours call. Repeat Call 1 outside your stated hours, within the agreed test window. Test whether the assistant follows the vendor's after-hours commitment: answering, capturing a message, offering permitted booking or using the agreed fallback. AlgoEthos includes after-hours answering in its published receptionist scope; test that commitment instead of assuming a connected line is enough.
Call 10 — The text confirmation and STOP. Use your own opted-in test number. After a successful test booking, check that the confirmation reaches the right number with the correct details. Then test the configured opt-out keyword, such as STOP. Verify the suppression record and check that scheduled reminders are canceled for that number. If the platform sends an opt-out acknowledgment, check that it is not duplicated by your application. This is an operational test of the chosen platform's documented behavior, not a claim that one keyword test establishes legal compliance. Twilio's Advanced Opt-Out documentation illustrates why platform configuration and application handling both matter.
The scoring sheet
Score every call on five lines. Be strict; the point is to find problems before your customers do.
| Line | Question | Pass condition |
|---|---|---|
| S1 Heard | Did the assistant respond to what I actually said? | No repeated "sorry, I didn't catch that" loops; no answering a different question |
| S2 Captured | Are my name, number, and reason recorded correctly? | Check the record, not the recap. Number must be digit-perfect |
| S3 Outcome | Did the call end in the right result — booked, transferred, callback promised, or correctly declined? | The result matches what a competent human would have done |
| S4 Honest | Did it avoid inventing anything — prices, availability, services? | Any fabrication fails the call outright |
| S5 Handoff | When it could not help, did it hand off cleanly? | Transfer or callback, stated clearly; no dead end |
For each call, mark each line pass, fail or not applicable, with a short reason. A call passes only if every applicable line passes and there is no fabrication. Count passes out of ten. In a hypothetical result, seven calls pass and Calls 4, 7 and 10 fail. That identifies three separate problems to investigate: number capture, handling an unknown question and text opt-out behavior. That is a made-up example. Score your own calls.
Which failures matter most?
Not all fails are equal. Rank them:
- Fabrication (S4). An assistant that invents a price or promises a service you do not offer is creating liabilities in your name. This is disqualifying at any frequency.
- Wrong capture (S2). A misheard phone number is a lost customer who thinks you never called back. Check numbers digit by digit on every test call.
- Silent dead ends (S5). The caller asks something, the assistant cannot answer, and the call just ends. Worse than a busy signal, because the caller thinks they were heard.
- Booking to the wrong slot (S3). Annoying and fixable, but it erodes trust fast.
- Comprehension loops (S1). Frustrating, and worth checking in the recording rather than guessing at the cause.
A vendor that fails on items 1 or 2 is not ready regardless of how natural the voice sounds.
What should you ask the vendor after the calls?
Bring your scoring sheet and ask:
- Show me the transcript for Call 4. Where did the number get captured, and can I see the raw audio?
- On Call 7, what did the assistant do with the question it could not answer — where does that land for a human to follow up?
- On Call 10, show me the opt-out record. What stops the system from texting that number again?
- Who reviews calls, how often, and what changes as a result? On our side the answer is a real person and monthly reviews — "A real human to call" and "Setup + monthly reviews" are on our page as commitments, and you should hold us to them.
- Does the assistant identify itself as automated? Decide what you want your customers told, and make sure the vendor can do it.
- What does it cost, in writing, at my volume, and what is the notice period? Ours is $200 per month, no setup fee, month-to-month.
Checklist: running the evaluation
- [ ] Configured the trial with real hours, services, common questions, a test calendar, and transfer rules
- [ ] Calling from a number the system does not know
- [ ] Made all ten calls from the script without helping the assistant
- [ ] Scored each call on all five lines; number capture checked digit by digit
- [ ] Ranked the failures by severity, not by count
- [ ] Verified the text confirmation and the STOP path end to end
- [ ] Reviewed transcripts, not just call logs, for at least three calls
- [ ] Asked the six post-call questions and recorded the answers
- [ ] Repeated the protocol on every vendor under consideration, including AlgoEthos
Where to go from here
If you want to run this protocol on our receptionist, ask for a walkthrough at (505) 319-5866 and tell us you are going to make the ten calls. We would rather you found a problem in a test than a customer found it during a real appointment request. And if you want the broader context on when an AI receptionist is the right option at all — versus voicemail, forwarding, or a live service — the after-hours comparison helps you decide before you spend time on test calls.
Service details and limits
AlgoEthos AI Receptionist and pricing, checked September 14, 2026. The scenarios and scoring sheet are our suggested test method, not an industry benchmark or a claim of measured vendor performance. Ten successful tests do not prove reliability across every caller, outage or edge case. Keep reviewing actual outcomes after launch.