Issue 01 · Tuesday, September 22, 2026
iLand Daily

News, opinion & art from the AI / human frontier

Opinion · Analysis

Stolen voices, licensed voices

Voice cloning has split into two lanes: licensed voices under contract, and cloned voices under nobody's authority. The difference between them is a piece of paper.

By · · Issue 01

In 2023, voice actors began finding their own voices speaking in places they had never recorded for. Some discovered synthetic clones of themselves narrating audiobooks; others found their voices in phone systems and video games they had never touched. The recordings that trained those clones had often been scraped or taken from past contracts that said nothing about artificial intelligence.

Two years on, the industry has split into two lanes, and the difference between them is a piece of paper.

The licensed lane

The clearer lane is consent by contract. Voice actors now negotiate AI-specific clauses: a per-project license, a defined scope, a renewal date, and — critically — a right to revoke. Some talent agencies report that standard sessions now include a "synthetic use" schedule, priced separately from the original recording. The actor gets paid for the clone; the studio gets a defensible paper trail.

The market for licensed voice work has grown accordingly. Marketplaces that connect actors with game studios and advertisers report that voice-cloning licensing is among the fastest-growing categories of remote talent work. The buyer is not paying for a recording session. The buyer is paying for permission.

The unlicensed lane

The other lane runs on no permission at all. Open-source text-to-speech systems can now produce a convincing replica of a voice from a sample as short as a few seconds of clean audio — a podcast, a YouTube video, a conference talk. The tools are free, the inputs are everywhere, and the output is a voice that says whatever the operator types.

The result is a consent gap that enforcement has not caught up with. The Right of Publicity statutes of states like California and New York were extended in 2023 and 2024 to cover digital replicas of a performer's voice, and the federal NO FAKES Act has been repeatedly introduced to create a national standard. But a law against counterfeiting is not the same as a licensing system. It tells you what happens after a violation. It does not tell you who checks permission before a voice is spun up.

Where does the burden fall?

That is the uncomfortable question underneath the technology. A text description — "middle-aged British man, warm, unhurried" — produces a generic synthetic voice and raises no one's rights. A reference clip of a specific person produces something legally and ethically heavier. The current state of platform practice is mixed: some commercial voice services require the uploader to warrant they have consent to clone the reference, a declaration checked by no one unless a complaint arrives. Others simply host the tools.

Voice professionals argue the burden belongs at the point of creation, not the point of complaint. Their industry bodies have spent two years pushing for consent, credit, and compensation as the floor for any synthetic voice derived from a real performer. The counterargument from developers is that identity verification at the tool level is technically fragile and would push cloning to less scrupulous corners of the internet. Both things are probably true, which is why the dispute has not resolved.

A quiet precedent

There is a precedent worth watching, and it is not legal but commercial. It is the slow emergence of "voice licensing" as a normal line item — the idea that a performer's voice is a licensed asset with a term, a scope, and a renewal, the way music publishing has treated compositions for a century. Musicians did not stop copyright infringement with lawsuits alone; they built a licensing system so ordinary that asking permission became the path of least resistance. Voice is at the beginning of that road.

For now, the burden sits in practice with whoever presses the button. If you are a platform generating a voice from a description, no one's rights are implicated — check nothing. If you are generating from a person's audio, someone's rights almost certainly are, and the fact that no one will check does not mean no one is entitled to object.

Two lanes, then: licensed voices under contract, and cloned voices under nobody's authority. The technology does not care which lane it serves. The contract, or the absence of one, is the only thing that separates a synthetic voice from a stolen one.

← Front page