Open voice AI: what has already happened in practice, and how not to turn it into an expensive toy for attackers
A voice AI consultant on a website is convenient: a visitor talks to the company right in the browser instead of hunting for a phone number. But an open button is also an open paid resource. The company pays for speech recognition, synthesis, the model’s work and sometimes a live employee’s time. The person who pressed the button out of curiosity or malice pays nothing.
We are deploying a voice agent on our own site. That is why we think it is right not to hide the limits of this technology, and to explain in advance which abuses have already occurred in practice and which safeguards we are building in before a public launch. For us voice AI is not an experiment without rules but a public service that has to be bounded in its powers, its spending and its handling of data.
What follows is neither scare stories nor an advertisement for security products, but a few telling cases and the conclusions that follow from them.
1. The bot that started swearing at its own company
In January 2024 a user got the chatbot of the delivery company DPD to swear and to write a poem criticising the company. The recording spread quickly: one post gathered around 800 thousand views in a day. DPD said it had disabled the problematic AI feature after a system update. The Guardian, 20 January 2024.
That was a text bot, not a voice one. For the customer the difference is slight: if a voice agent says something on behalf of the company, the recording is easy to clip, caption and publish. The protection here is not a “magic ban in the prompt” but a limited set of topics, testing answers against provocation, and a fast way to switch off a bad version.
2. A company was held responsible for what its bot said
On 14 February 2024 a British Columbia tribunal heard a passenger’s dispute with Air Canada. The bot on the website gave incorrect information about compensation terms, the passenger relied on it, and the airline later refused. The tribunal ruled that a company is responsible for the information on its own website regardless of whether a page or a chatbot displayed it, and awarded C$812.02. Moffatt v. Air Canada, 2024 BCCRT 149.
The lesson is not the size of the award. For a public agent the more dangerous thing is the casual “yes, that is possible” — about a price, a deadline, a discount, the terms of a service. Where an answer may amount to a promise, it is better to give an approved fact from a database or to hand the conversation to a human.
3. An official bot advised businesses to break the law
In March 2024 an investigation by The Markup showed that the New York City MyCity chatbot gave business owners incorrect advice on rent, tips, cash payments and employment rules. In some answers it effectively suggested actions that were illegal. The Markup, 29 March 2024.
This is not a “hacker attack” but a dangerous combination of an authoritative tone and a wrong answer. For a voice assistant it is especially important not to feign confidence where it cannot verify a fact. A useful principle: the agent explains general information, while prices, contracts, legal and personal matters go into a verifiable process.
4. Millions of logs and audio recordings were exposed through a storage mistake
In February 2026 the security researcher Jeremiah Fowler found three publicly accessible databases connected to the voice AI agent of Sears Home Services. According to him they held 3.7 million chat logs, 1.4 million audio files and text transcripts of calls from 2024–2026 — with names, phone numbers, addresses and order details. Some recordings lasted for hours: the call stayed open after the customer considered the conversation finished, and everything happening nearby went into the recording. The databases were closed after the researcher’s notification; whether anyone else had gained access is unknown. WIRED, March 2026.
This is the clearest argument against keeping all the audio “just in case”. Recordings need a defined retention period, role-based access, encryption and regular checks that the storage has not been left open by accident.
5. Telephone infrastructure is already being used as a weapon
Telephony denial of service — TDoS — works simply: a huge number of calls occupies the lines and the queue while real people cannot get through. The US cybersecurity agency CISA published a review of a real case: an emergency communications centre fought off daily TDoS attacks for years. CISA, 2022. In another documented case the FCC established that the participants in the ScammerBlaster scheme made almost 10 million automated calls to toll-free numbers, earned money from the fee the recipient pays for such calls, and spent that money on TDoS attacks. In September 2023 the regulator imposed a $116 million fine. FCC, 2023.
In a browser voice widget the transport is different but the goal is the same. Instead of thousands of phone calls one can try to open thousands of browser sessions, burn model tokens or hold a microphone open in silence. That is why the limit has to come before an expensive conversation starts, not only after it.
6. A synthetic voice has already become a tool of telephone fraud
In January 2024 residents of New Hampshire received robocalls with a spoofed number and an AI clone of President Biden’s voice, urging them not to vote. In August 2024 the carrier Lingo Telecom, whose network carried the calls, reached a $1 million settlement with the FCC. In September 2024 the FCC finally imposed a separate $6 million fine on the organiser of the calls. FCC, 21 August 2024.
This case is not about a web bot and does not mean that every artificial voice is criminal. It shows something else: once ordinary telephone lines are connected, neither the number on the screen nor a familiar voice can be treated as sufficient proof of identity. Outbound calls, callbacks and transfers to external numbers need separate fraud protection.
7. The “expensive button” is a recognised class of risk, even if public figures are scarce
In its list of top risks for applications built on language models (2025), the application security community OWASP singled out “unbounded consumption”: an attacker issues too many or too heavy requests to the model and turns pay-per-use into an economic denial of service. OWASP recommends input validation, rate limiting, quotas, timeouts and resource monitoring. OWASP, 2025.
So far there is no widely published, reliable case with a named loss figure caused specifically by a free browser-based voice AI widget. That is an important honest caveat. But the absence of a public number does not make the mechanism safe: a voice interface pays for several layers at once — audio, the model and sometimes an operator.
What anyone opening a voice AI agent should do
- Give the browser a short-lived session rather than a permanent key to the services.
- Limit the number of starts, concurrent conversations, minutes, turns and tokens; keep a daily and hourly spending cut-off.
- Close the conversation after a reasonable silence instead of paying for a forgotten microphone.
- Hand over to a human only into an internal queue, and never let the agent call or message an arbitrary number.
- Answer questions about prices and terms only from an approved source; when in doubt, promise nothing and pass the caller to a human.
- Test every update against provocation: requests to reveal instructions, abuse, wrong prices, repeated transfers, synthetic audio and long pauses.
- Keep recordings to a minimum, protect them, and name in advance the person who can switch the widget off quickly.
The honest conclusion
An open voice AI does not become safe by being “polite” or by recognising the topic of a conversation. It has to be designed as a public paid service and as an employee whose words may be recorded. You cannot fully tell a human from a good machine, and you cannot reliably stop every manipulation of the model’s instructions with a single filter. What you can do is cap the maximum spend in advance, withhold dangerous powers from the model, stop an attacker from occupying the queue indefinitely, and never leave recordings unprotected.
None of this kills the convenience of the service. It simply makes it fit for open access.
What we are working on now
We build voice AI agents ourselves and know well how synthesised speech sounds. So our next step is a fake-voice detector: a module that tells a live person from a recording and from synthesis during the conversation. It will work inside our own voice engine. We will publish the details when there is something to show.