CRYPTOITDATACRYPTOITDATA

Artificial intelligence

AI voice agent on your phone line: what it can and cannot do

What an AI voice agent on a small company phone line can and cannot do: the duty to say it is AI, the limits outside English and the real cost per minute.

9 min read

Think of the last morning when the company phone rang three times in a row while you were with a customer. You answered none of them and you probably still do not know who was calling. An AI voice agent — a program that picks up, holds a conversation and writes down what needs writing down — is exactly what gets sold to you for that situation, with the promise that nobody will realise they are talking to a robot. The part that does not get sold: since 2 August 2026, that program is required to tell the caller, in the first seconds, that it is an AI system.

The obligation comes from Regulation (EU) 2024/1689 — the AI Act — and it turns around the question almost everyone who asks us about voice agents starts with. They ask “can you tell it is a robot?”. Since August, the answer is: yes, you can, because the law requires it. The useful question is a different one: which of your calls can be handled correctly by a robot that declared itself a robot in the first second. Not “all of them”, but not “none” either — and the gap between those two lists decides whether the project is worth doing.

The call nobody answers shows up in no report

A missed call is the only loss in a business that leaves no trace anywhere: it appears in no report and it never reaches you as a complaint. The patient who could not get through booked somewhere else, the customer with the broken car called the next garage on the list, and the one who wanted a table for four on Saturday evening will never tell you he tried. The cost is real and invisible at the same time — which is why the decision gets postponed indefinitely.

There are convenient figures circulating on this subject: that 62% of calls to small businesses go unanswered, that 85% of the people who get no answer never call again. We traced them to the source and we do not use them: they appear almost exclusively on sites that themselves sell solutions for missed calls, and the citation chain ends at a study of 85 companies with an unpublished methodology. The argument does not need them: the owner of a car workshop already knows how often the phone rings when his hands are inside an engine.

What can be supported with data is more modest, but it is enough: the reminder call works. A randomised study in the Journal of General Internal Medicine (2016) measured the effect of a call placed seven days before the appointment, on patients at high risk of not showing up — the rate fell from 29.2% in the control group to 22.8% in the called group. The caveats are mandatory: the call was placed by a human, not by an AI agent, and the population was US primary-care patients selected by risk. The figure shows that the call matters; not that a voice agent produces the same effect.

The alternatives anyone tries before arriving at a voice agent each fail in their own way:

  • Voicemail — it moves the effort onto the caller, exactly when their patience is lowest. An observational comparison published in Psychiatric Services (Teo et al., 2017), on 250 primary-care patients with depression, found 3% no-shows when the reminder was answered live, 24% when a message was left and 39% when nobody answered. A non-randomised study on a narrow population: it shows a direction, not a ratio of eight to one.
  • The touch-tone menu (“press 1 for appointments”) — it does not lose the call, but it forces the caller to navigate a tree they have never seen; many hang up at the second branch.
  • A receptionist who only answers during office hours — evening and weekend calls simply do not exist for the business.
  • Unconfirmed appointments — the slot stays blocked in the calendar and empty in reality, and nobody has time for ten confirmation calls a day.

What an AI voice agent can and cannot do

Every voice agent demo has something in common: the other person speaks clearly, in turn, in a quiet room. The best public evidence about what happens when that is no longer true comes from a paper on Romanian speech recognition (Pîrlogeanu, Georgescu, Cucu), posted on arXiv in November 2025, with a model trained on more than 2,600 hours of Romanian. On read speech, the word error rate is 3.09% for RoWhisper-large-v2 and 1.73% for the proposed model. On spontaneous speech, the same models reach 25.05% and 8.12% on one corpus, and 61.46% and 10.75% on another.

The difference between read and spontaneous speech is an order of magnitude, and a phone call is, without exception, spontaneous speech: people correct themselves mid-sentence, talk over you, call from the car, give you the number in groups of two digits. Two warnings, without which the figures become manipulation: the paper is a preprint and it does not measure the commercial models that production agents are built on. What stays valid is the shape of the problem, not the percentage: quality depends more on how predictable what the caller says is than on how expensive the model is — and that holds for every language that is not English.

From there comes the only useful selection criterion. A voice agent is feasible when the call is, in essence, a short form read out loud — a name, a number, a time slot, one service from a list of five. It becomes unreliable as the call looks less and less like a form:

  • Works: taking a new booking, confirming or moving an existing one, opening hours and address, frequently asked questions with a fixed answer, taking a message with a name and a number, screening calls after 6 p.m.
  • Works with care: price estimates with conditions (“it depends on the part”), long service lists, reservations with many variables — these need a tight script and a clear giving-up threshold.
  • Does not work: an unhappy customer’s complaint, any substantive medical discussion, negotiating a price, the problem that fits none of the prepared categories — not because the technology is weak, but because here a mistake costs more than the missed call.
It no longer matters whether you can tell it is a robot — since August 2026 it has to say so itself. All that matters is which calls deserve a robot.

“75 milliseconds” is not the pause the caller hears

The second limit nobody explains is latency, because the vendor figures sound excellent. The ElevenLabs documentation states roughly 75 ms (Flash v2.5), around 100 ms median (v4 Turbo), around 280 ms (v3 Conversational) and 150 ms for real-time transcription with Scribe v2 Realtime. The problem is that, in speech synthesis, latency measures strictly the time to the first byte of generated audio — not the pause the caller hears.

The real pause adds up telephone transport, speech recognition, language-model inference, synthesis and the network buffer. We did not find a credible independent end-to-end latency benchmark for Romanian, so we are not giving a figure — but we do not accept 75 ms as the answer to “how long until it answers me” either. Anyone can run the check before signing: ask for a test number, call from a mobile, from the car, and interrupt the agent halfway through a sentence.

Romanian is on the list of supported languages — included in the multilingual model with 29 languages, and the newest ones list it explicitly. What does not exist, as far as we found, is any published evaluation of how natural the synthesis sounds in Romanian compared with English. So the person who says it sounds identical has as little to stand on as the one who says it sounds bad. You check it with your own ear, in your own language, where synthesis stumbles: proper names, addresses, abbreviations and numbers.

Three obligations you inherit the day you switch the agent on

The legal part is not a warning at the end of the article, it is part of the specification: two of the three obligations below turn into sentences the agent speaks in the first seconds.

1. The agent has to say it is AI

Article 50(1) of the AI Act requires providers to ensure that AI systems intended to interact directly with natural persons are designed so that those persons are informed that they are interacting with an AI system — unless this is obvious to a reasonably well-informed, observant and circumspect person. On the voice channel, with good synthesis, that is exactly what stops being obvious. Paragraph (5) adds the condition that settles the matter for the phone: the information is given in a clear and distinguishable manner, at the latest at the time of the first interaction. It is not a line in the privacy notice on your website — it is a sentence spoken at the start of the call. The text binds the provider, so for you it is an acceptance condition: whoever builds your agent has to deliver it with the disclosure spoken, not leave it for you to configure.

Two useful nuances. The obligation to mark generated audio in a machine-readable format, in paragraph (2), falls on the provider of the generative system, not on the clinic that uses it; for systems placed on the market before 2 August 2026 it has grace until 2 December 2026. And if the agent runs sentiment analysis on the caller’s voice or identifies them by voice, you also come under paragraph (3) — a good reason not to switch on features you do not need.

As for consequences: breaching Article 50 falls under Article 99(4) — up to EUR 15,000,000 or 3% of total worldwide annual turnover, whichever is higher — but for SMEs Article 99(6) reverses the rule and applies the lower of the two. The EUR 35 million figure that circulates in the press refers to the prohibited practices in Article 5, not to transparency. Enforcement is national: every member state designates its own market surveillance authority and sets the inspection procedure and the penalties itself. In Romania, by government memorandum of 12 March 2026, ANCOM — the national communications regulator — was designated market surveillance authority and single point of contact for the AI Act. In mid-August 2026 the national law establishing the inspection procedure and the penalty regime had not been adopted, and ANCOM was publicly confirming that the act was still being drafted; up to the publication date of this article we found no indication that this had changed. The delay negotiated in 2026 concerned high-risk systems, not transparency. The obligation is in force, the national enforcement instrument was under construction — a window for getting ready, not a suspension.

2. An agent that answers is one thing, an agent that calls is something else entirely

This is the only fine in this article that has actually been issued. In June 2026 the Romanian data protection authority, ANSPDCP, closed an investigation into AMATO BESTSELLER S.R.L. and fined the controller, among other things, RON 50,000 for breaching Article 12(1) of Law 506/2004: it had made commercial communications through automated calling and communication systems that do not require human intervention, without the prior express consent of the recipients. The same case added RON 78,465 for Article 32(4) GDPR and RON 104,620 for Articles 5(1)(c) and 9 GDPR. The rule behind it is not Romanian: Article 13 of the ePrivacy Directive 2002/58/EC requires prior consent for automated calling machines and is transposed in every member state.

The pattern that was penalised is exactly that of a voice agent that calls instead of answering. The practical dividing line: an agent that takes incoming calls is not unsolicited commercial communication, because the customer initiated the call; one that dials numbers to pitch offers needs prior express consent, documented before the first call. The third case — the automated call that confirms an appointment the customer asked for — is not marketing in the usual logic, but we did not find an explicit public position from the authority on that situation, so we are not giving a verdict.

3. If you keep the transcript, you inform the caller

A useful voice agent comes with transcripts — without them you cannot check what it told the customer and you cannot fix the script. But the transcript and any recording are personal data, and in a medical practice they can be health data. That means informing the caller at the start of the call, a legal basis you have thought through, a retention period set down in writing, limited access and deletion on request — handled together with the rest of your data protection obligations, not improvised in a file on the receptionist’s laptop.

The order of operations, if you want it to work

The steps below are useful whether you build it yourself or never build it at all. The first three are decisions, not settings, and getting them wrong at the start is paid for in month three.

  1. 1Pick one measurable use case — bookings or frequently asked questions, not “a full reception desk”. An agent that does one thing well is a project; one that does three things badly is a demo.
  2. 2Write the opening sentence before anything else: the disclosure that the caller is talking to an AI system, the notice about recording and the option to ask for a human. If it does not fit into 10-12 seconds, rewrite it until it does.
  3. 3Define the handover point to a human explicitly: at the second consecutive misunderstanding, at any direct request, at any medical, payment or complaint subject. A Gartner survey of 3,566 customers, from February-March 2026, shows that 87% consider it essential to be able to reach a human agent when a company uses generative AI in customer relations. The door to a human is not a concession, it is the acceptance condition.
  4. 4Decide whether you also need outbound calls. If you do, prior express consent is obtained and documented before the first campaign, not after.
  5. 5Choose how the call comes in: a new number or integration into your existing telephony, with written rules for out-of-hours, the busy line and simultaneous calls.
  6. 6Run the pilot in parallel with the manual process for two to four weeks, the way we work on any AI automation, and measure three things: how many calls it took, how many it escalated, how many bookings landed correctly in the calendar.
  7. 7Listen to 20 real calls in the dashboard before you expand. Not the reports — the calls. That is where you see the script stumble, and that is where you fix it in 20 minutes.

What it costs and why it is billed per minute

With us, an agent on a single use case — bookings or frequently asked questions — is built in 3-4 weeks, at RON 12,500 for implementation, and is operated for RON 450 a month for up to 1,000 minutes. The multi-function version (reception, bookings and support) takes 5-7 weeks and RON 24,500, with RON 1,250 a month for up to 5,000 minutes. The deliverables are the same: the agent with a custom voice, integration with Twilio or your existing telephony, your own workflow for bookings, FAQ and escalation, plus the dashboard with transcripts. The operating subscription can be bought straight from packages; implementation is sized after a conversation, because the script differs between a clinic and a car workshop.

The minute cap is not a commercial trick, and it is worth explaining because it changes the way you design the call. Platform cost is per minute spoken: ElevenLabs bills agents at USD 0.08 per minute on all plans, with USD 0.16 above the concurrent-call limit, while Twilio telephony for Romania adds USD 0.0100 per minute for an inbound call on a local number, USD 3 a month for the number and USD 0.0125 and 0.0320 per minute respectively for outbound calls to landline and mobile — prices as displayed on the publication date. The practical conclusion: an agent that settles a booking in 90 seconds costs several times less than one that holds six-minute conversations, so a short script shows up on the invoice, not only in the quality.

The expense that hurts is not the subscription, it is the project started without a clear use case. We have written separately about why AI projects fail inside companies, in the analysis of AI ROI for SMBs.

Conclusion

An AI voice agent is not a digital secretary that settles every call, and anyone who promises you that has never listened to 20 real calls. It is a tool that handles a narrow, predictable category of calls well — bookings, confirmations, questions with a fixed answer, the calls after closing time — and that has to say what it is in the first second. Everything that decides whether it works happens before implementation: which calls you give it, what it says at the start and when it hands the call to a human. If you want us to look together at which of your calls fall into that category, let us talk for 30 minutes, with no obligation.

Sources

  • Regulation (EU) 2024/1689 (AI Act) — Article 50 (transparency obligations) and Article 99 (penalties), via the AI Act Service Desk, European Commission (ai-act-service-desk.ec.europa.eu)
  • Analyses of the Digital Omnibus package on AI and the postponement of deadlines for high-risk systems, with Article 50 kept at 2 August 2026 (gibsondunn.com)
  • Alexandra Crăsoveanu — “The AI Act applies in Romania, but cannot yet be enforced”, JURIDICE.ro, 13 August 2026 (juridice.ro)
  • ANSPDCP — press release of 6 August 2026, penalty for commercial communications through automated calling systems without prior consent (dataprotection.ro); Law 506/2004, Article 12(1); ePrivacy Directive 2002/58/EC, Article 13
  • Gabriel Pîrlogeanu, Alexandru-Lucian Georgescu, Horia Cucu — “Open Source State-Of-the-Art Solution for Romanian Speech Recognition”, arXiv preprint 2511.03361, November 2025 (arxiv.org)
  • ElevenLabs — documentation on models and supported languages, the blog post on latency optimisation in conversational AI and the agents pricing page (elevenlabs.io)
  • Twilio — Voice pricing for Romania (twilio.com/en-us/voice/pricing/ro)
  • Gartner — press release of 4 August 2026, survey of 3,566 B2B and B2C customers on access to a human agent (gartner.com)
  • Journal of General Internal Medicine, 2016 — randomised study on the effect of the reminder call on no-shows (link.springer.com)
  • Alan R. Teo, Christopher W. Forsberg, Heather E. Marsh, Somnath Saha, Steven K. Dobscha — “No-Show Rates When Phone Appointment Reminders Are Not Directly Delivered”, Psychiatric Services, 68(11), 2017 — observational comparison of appointment reminder types (psychiatryonline.org)

Frequently asked questions

Does the voice agent have to say it is a robot?+

Yes. Article 50(1) of the AI Act requires that persons be informed they are interacting with an AI system, and paragraph (5) requires the information to be clear and distinguishable, at the latest at the time of the first interaction. On the phone that translates into a sentence spoken at the start of the call, not a line on the website. The obligation applies from 2 August 2026 and was not postponed by the package that moved the deadlines for high-risk systems.

Can I use an AI agent to call customers with offers?+

Only with the prior express consent of the recipients, obtained and documented before the first call. Article 13 of the ePrivacy Directive — transposed in Romania by Article 12(1) of Law 506/2004 — prohibits commercial communications through automated calling systems that do not require human intervention without that consent, and in 2026 ANSPDCP fined a controller RON 50,000 on exactly this pattern. An agent that answers incoming calls is a completely different situation, because the customer initiated the call.

Do I need the caller’s consent to record the conversation?+

The recording and the transcript are personal data, and in a medical practice they can be health data, so the caller has to be informed at the start of the call, before recording. In practice you need a legal basis you have thought through, a retention period set down in writing, limited access and a deletion mechanism on request. It is part of the general data protection obligations, not a separate formality of the voice agent.

How well does a voice agent understand a language other than English?+

It depends decisively on whether the speech is predictable or free. A November 2025 paper on Romanian speech recognition measures, on read speech, word error rates of 3.09% and 1.73% for two models, while on spontaneous speech the same models climb to 25.05% and 8.12% on one corpus and 61.46% and 10.75% on another. A phone call is spontaneous speech, so an agent that collects a name, a number and a time slot is feasible, while a free conversation is not. The paper is a preprint and does not measure the commercial production models, so the figures show the shape of the problem, not a vendor’s performance.

Have a concrete question?

30 minutes, free. We discuss exactly your situation.

Book a consultation