Case study · · 5 min read · Lurto

We put an AI receptionist on our own phone line first. Here is what broke.

We put an AI receptionist on our own phone line first. Here is what broke.

Quick answer

We built a voice receptionist and pointed our own business number at it before selling one to anyone, so that every failure landed on us rather than a client. It answers at any hour, books into a real calendar, writes the caller into the CRM and leaves a summary in the morning. Two things it taught us: in week two it double-booked a slot because availability was checked ninety seconds before the write, and our eval harness once graded the agent on a transcript it never heard, so a change that scored better made the calls worse.

A plumber's phone rings at 6:40pm. He is under a sink. It goes to voicemail, the caller does not leave one, they ring the next number on the list. That call was worth more than the rest of his evening and he never knew it happened. It is the most common problem we are asked about by trade and service businesses, and before we would sell a fix for it we wanted to run one on a line where the mistakes would cost us, not a client.

So we built a receptionist and pointed our own business number at it. It has answered that line every day since, and it is the voice product we now sell to US small businesses as Rinqly. This is the build story, with the failures kept in.

What it does

It answers the phone at any hour, talks with the caller, books the appointment into a real calendar, writes the caller into the CRM, and leaves a summary waiting the next morning. It also runs the website chat, places outbound calls where that is wanted, and connects to the tools a small business already has: Zapier, HubSpot, Pipedrive, GoHighLevel and ActiveCampaign among them. Billing is metered through Stripe. None of that is the interesting part.

Why we built it on our own line first

A voice demo is easy. A recorded call in a quiet room, a friendly script, a caller who behaves. Real telephony is nothing like it. Callers interrupt, change their mind mid-sentence, phone from a van with the window open, ask for a slot that filled up while they were spelling their name, and hang up if the line goes quiet for two seconds. A system that survives that is a different thing from a system that survives a demo, and the only way we know to find the difference is to take real calls and read what went wrong.

Running it on our own number meant every failure landed on us. That was the point.

How it is built

The model hears the caller and speaks back directly, with no transcription and synthesis steps between, which is what makes interruptions and corrections feel natural. Transport is LiveKit, telephony is Twilio, and the model is OpenAI's realtime audio. The application around it is Next.js. We chose speech-to-speech over the cheaper cascaded pipeline on purpose, and we accepted paying roughly ten dollars a month more for it at our volume, because a pipeline that needs an extra 300 milliseconds to notice a caller has started talking produces exactly the stilted call that makes people hang up. The full arithmetic of what each part costs per minute is in what an AI voice agent costs.

There is a warm replica running whether or not the phone is ringing. Letting the agent scale to zero saves money and buys a cold start on a call that is already ringing, which a caller tolerates for about two seconds. That is a fixed monthly cost the per-minute maths never shows and we pay it.

Week two: it double-booked

In its second week, two calls hit the calendar within the same second. Both callers were told the slot was free, both were booked into it, and two people expected the same appointment. Nothing in the prompt was wrong. The system checked availability at the start of the booking flow, spent ninety seconds collecting a name and a number, then wrote to the slot it had checked earlier. The write was never locked.

The fix was to verify the slot again at the moment of writing, treat "slot taken" as a normal conversational turn, and take a tentative hold when the slot is offered so it can be confirmed or released at the end. That incident became a permanent test case, and the calendar-locking check has run before every deploy since. The other four ways voice systems fail on live calls are written up in why a voice agent drops calls and double-books.

The eval that lied to us

The receptionist has an evaluation set: a fixed list of real call scenarios, each with the outcome a competent person would expect, run by a script before any change goes live. It is the thing we would most insist on for any voice system, and it still caught us out.

We shipped a change to the agent because the harness said it made the calls better. It made them worse, and we reverted the same day. The mistake was whose transcript we were scoring. We had been transcribing the recordings ourselves, with a different speech-to-text model than the one the agent hears through on the call, so we were grading the agent on words it never heard. Once the harness scored against the agent's own transcript, the improvement disappeared and the regression showed up. Impressions lie, and so does a second-hand transcript. How to build an eval set that does not is in evals for business automations.

What we build client voice work on now

Every voice agent we build for a client runs on this infrastructure: the same transport, the same telephony, the same eval harness, the same calendar locking, the same warm replica. The client owns every account, the meters run at the provider's price with no markup, and bugs in the first thirty days are ours. A client build is usually one calendar plus message taking at the bottom of the range and multi-location routing, CRM writes and outbound callbacks at the top; the bands are on the pricing page.

Two honest limits. If your business takes a handful of calls a day, a good voicemail greeting with a same-day callback will serve those callers better than anything in this post. And a template subscription at $49 to $299 a month is the right purchase when the calls are simple and you do not yet know what your callers ask; sixty days on a template tells you your real volume and where the script breaks. We would rather say that here than sell a build to someone who needs a subscription.

Questions people ask about this

Why speech-to-speech rather than the cheaper pipeline?

The model hears the caller and speaks back directly, with no transcription and synthesis steps in between, which is what makes interruptions and mid-sentence corrections feel natural. We pay roughly ten dollars a month more at our volume for it, because a pipeline that needs an extra 300 milliseconds to notice a caller has started talking produces exactly the stilted call that makes people hang up.

How did it double-book a slot?

Check-then-book. The system confirmed availability at the start of the booking flow, spent ninety seconds collecting a name and a number, then wrote to the slot it had checked earlier, and the write was never locked. The fix was to verify the slot again at the moment of writing, treat slot taken as a normal conversational turn, and take a tentative hold when the slot is offered. That incident became a permanent test case and the check has run before every deploy since.

How can an eval set mislead you?

We shipped a change because the harness said it made the calls better. It made them worse and we reverted the same day. The mistake was whose transcript we were scoring: we had been transcribing the recordings ourselves with a different speech-to-text model than the one the agent hears through on the call, so we were grading the agent on words it never heard. Once the harness scored against the agent's own transcript, the improvement disappeared and the regression showed up.

Which cost does the per-minute arithmetic never show?

A warm replica running whether or not the phone is ringing. Letting the agent scale to zero saves money and buys a cold start on a call that is already ringing, which a caller tolerates for about two seconds. It is a fixed monthly cost, and we pay it.

Is a voice agent right for a small business?

Two honest limits. If your business takes a handful of calls a day, a good voicemail greeting with a same-day callback will serve those callers better than anything we could build. And a template subscription at $49 to $299 a month is the right purchase when the calls are simple and you do not yet know what your callers ask, sixty days on a template tells you your real volume and where the script breaks.

Have a system that needs this treatment?