← Back to blog
NewsAbout 8 min read

Shanghai's First AI Voice Infringement Case Decided on Appeal: Identifiability Becomes the Yardstick, Platform Ordered to Pay 50,000 Yuan

Published Oct 6, 2026
Shanghai's First AI Voice Infringement Case Decided on Appeal: Identifiability Becomes the Yardstick, Platform Ordered to Pay 50,000 Yuan

--- title: Shanghai's First AI Voice Infringement Case Decided on Appeal: Identifiability Becomes the Yardstick, Platform Ordered to Pay 50,000 Yuan meta_title: How Are AI Voice Infringement Cases Decided? Shanghai's First Second-Instance Ruling Sets the Yardstick meta_description: The Shanghai No. 1 Intermediate People's Court upheld the original judgment on appeal: a voice actor's voice was synthesized by AI for promotional use, and the platform was ordered to pay 50,000 yuan. The court established identifiability as the core standard for voice infringement. ---

Voice actor Ms. Wang only discovered the problem after a friend tipped her off. Someone told her that the voice used in an app's user-acquisition campaign sounded a lot like hers.

She looked into it and concluded that the audio was most likely AI-synthesized. She then had the evidence notarized, sued the operator, and sought 300,000 yuan in damages. The case went all the way to the Shanghai No. 1 Intermediate People's Court, where the final judgment upheld the original ruling: the platform constituted voice infringement and was ordered to pay 50,000 yuan.

This is Shanghai's first dispute over the protection of voice rights arising from AI-synthesized speech. It deserves to be examined on its own, and the reason has nothing to do with the amount of damages: the court clarified one question—under what conditions does AI imitating a person's voice count as infringement.

Three key facts in the case

First, the defendant admitted the audio was AI-generated but could not explain where the source material came from.

Company A's defense was that the audio in question was indeed AI-generated by the company, but the specific source and the training material fed into it could not be confirmed because the former employee had left. The company also said it had not previously known that Ms. Wang was a voice actor, and that it had not collected or used her voice.

Second, the appraisal opinion provided quantitative results.

Ms. Wang applied for a judicial appraisal, comparing the notarized audio with her own voice collected on the spot. Of 28 formant acoustic indicators, 24 had a deviation of less than 10%, of which 16 were below 5.36%. The relatively similar and highly similar portions together reached 90%.

Third, the court placed the burden of proof on the platform that held the training data.

The Shanghai No. 1 Intermediate People's Court held that Company A, as the party controlling the original training data, had failed to adequately prove the lawfulness of the source of the fed material, that it had not used the voice characteristics of a natural person, and that it had not caused public confusion or misidentification, and should therefore bear the adverse consequences of failing to meet its burden of proof.

The rule the court established: look at identifiability, not the technical source

The core holding of this case comes down to one sentence: using a natural person's voice as training corpus without that person's consent, imitating that person's timbre, intonation, and pronunciation style, and generating a synthesized voice that can identify that person, should be found to infringe the natural person's voice rights.

There are two layers here worth unpacking.

One is that a defense AI voice developers often rely on no longer holds. Some argue that the voice is an algorithm's reconstruction of a voiceprint, which differs from directly copying an original recording, and therefore does not constitute infringement. The court's response is that the judgment rests on the synthesized result: whether it can lead ordinary members of the public to associate it with a specific identity.

The other is that the recognition threshold is set at the level of the ordinary public, not professionals. In this case, the appraisal result was a high degree of similarity, so identifiability was established. But this also means that once similarity approaches this level, the line is crossed—it is not only when the imitation is indistinguishable from the real thing that it counts.

This precedent connects to a document from the Supreme People's Court

If you look at the timeline, this case is not an isolated judgment.

On September 7, the Supreme People's Court issued the Opinions on the Lawful Adjudication of Cases Involving Artificial Intelligence, Article 4 of which clarifies two things: without consent, using artificial intelligence to process another person's name or portrait to generate an identifiable virtual digital likeness and then using or publicly disclosing it constitutes infringement of personality rights such as the right to name and the right to portrait; and without consent, using another person's voice for training, imitating their timbre, intonation, and pronunciation style, and generating an identifiable synthesized voice constitutes infringement of voice rights.

This is China's first document of judicial rules on AI-related cases issued by the country's highest judicial body. The most significant point is that, for the first time, voice receives the same protection as portrait, and identifiability becomes the unified yardstick for determining infringement.

The Shanghai case came after these Opinions, effectively putting the written rule into a concrete factual setting: the appraisal gave a 90% similarity, the court found identifiability established on that basis, the platform could not produce evidence of a lawful source for the material, and the damages were thus established.

There was also another typical case cited for reference by the Supreme People's Court. In 2024, a cultural media company in Xuzhou commissioned a livestream host to produce a promotional video; the host used public speech clips of Ms. Li, a scholar in the field of family education, paired with AI-synthesized audio highly similar to her voice in timbre, intonation, and pronunciation style. In July 2025, the Beijing Internet Court ordered the defendant to apologize and pay 120,000 yuan in economic losses and reasonable costs of safeguarding rights. Both cases point to the same logic: unauthorized use of another person's portrait, compounded by a highly matching AI-synthesized voice, constitutes a dual infringement of portrait rights and voice rights.

Market reaction: face-swapping tightened up, voice cloning still on sale

Once the rules landed, the reaction on the transaction side was asymmetric.

According to media reports, within a week of the Supreme People's Court's Opinions being issued, searches on some e-commerce platforms showed a strong blocking of face-swapping services, with no related products appearing in search results; but when searching for voice cloning, products were still widely on sale. Some speech-synthesis products were priced at 33.88 yuan, with over 700 sold; some voice model training services were listed at 20 yuan; and on second-hand trading platforms, voice cloning services were listed at 8.85 yuan, showing over 770 sold.

This asymmetry itself reflects a reality: the barrier to entry for voice cloning is much lower than for face-swapping, and the demand for source material is more discreet. A publicly available voice-over work, an episode of a podcast, or a short video can constitute sufficiently good training corpus.

For users, the signal from these cases is clear. First, don't expect to be exempt from liability just because you can't explain where the material came from—the party holding the training data is at a disadvantage in terms of proof. Second, the determination of identifiability depends on technical appraisal; once the synthesized result is close enough to be associated by the public with a specific individual, commercial use constitutes infringement. Third, the amount of damages will be determined by reference to the rights holder's licensing fees for comparable uses—the costs saved by using AI to bypass authorization may ultimately be paid back in the form of damages.

What this means for teams building AI voice products

The reasoning in this judgment offers several direct takeaways for the product side.

Proof of the source of training data needs to be kept on file. The most fatal point in the case was that Company A could not explain where the fed material came from. For teams working on voice cloning, speech synthesis, and audio content generation, the authorization chain for corpus must have traceable records, rather than relying on verbal confirmation.

Imitation of voice characteristics must be decoupled from specific individuals. If a product is to generate a specific timbre, authorization is a precondition, not an after-the-fact fix. Conversely, if you just want a generic timbre, the training data must be ensured to come from commercially usable, traceable sources.

The identifiability line should be treated as a red line, not a reference value. The 90% similarity in the appraisal conclusion is the basis for the finding in this case, but the court did not set a fixed numerical threshold. This means any generated result that approaches the voice characteristics of a specific natural person needs to be carefully assessed.

Ms. Wang's case resulted in 50,000 yuan in damages. Compared with the 300,000 yuan claimed, this figure is not high. But its significance lies not in this case, but in the fact that the default assumptions for everyone doing AI voice work afterward have changed.

Related articles