ByteDance Could Not Shake a Lawsuit Over How It Got the Videos

A group of YouTubers sued ByteDance, alleging the company scraped millions of videos to train its text-to-video model. ByteDance asked a federal court in California to throw the case out. On October 2, Judge Jacqueline Scott Corley declined.
The ruling turns on a section of copyright law that had been sitting to one side of the AI training debate. Everyone has been arguing about fair use, which is section 107. The YouTubers are arguing about section 1201, which forbids circumventing technical measures that control access to a copyrighted work.
That is a different case, and ByteDance's defence depends on the difference.
The argument ByteDance made
The company's position was that YouTube's protective measures regulate downloading, not access. The videos in question are public. Anyone can watch them. If the content is already available to everyone, ByteDance argued, a measure that slows bulk downloading is not an access control in the sense the statute means, and section 1201 does not apply.
Judge Corley rejected it. She held that the YouTube safeguards ByteDance allegedly worked around are the kind of access control the DMCA prohibits circumventing. What the platform was restricting, in her reading, was the manner and volume of access, and that counts.
The distinction she drew is the one that matters going forward. A paywall is an access control. So is a login. So, apparently, is a rate limiter or an anti-bot measure guarding content that is otherwise free to view. The statute does not say the work has to be secret.
Why this is a problem for the training pipeline
The practical effect is that a data collection method can be unlawful even when the data itself was publicly posted.
If a court agrees with that reading, then the compliance question for a model trainer shifts. It is no longer only "did I have the right to learn from this work." It also becomes "did I get the work in a way the platform permitted." Those are separate questions with separate answers, and a company can be right on the first and wrong on the second.
The second question is also easier for a plaintiff to litigate. Fair use requires a four-factor analysis about purpose, nature, amount and market effect, and the answers are contestable. Whether a scraper defeated a technical measure is closer to a factual determination. Either the tool worked around the control or it did not.

Section 1201 has been circling AI cases for a while
This is not the first time that section has appeared in a training dispute, and the earlier cases went the other way.
In Doe v. GitHub, developers argued that Copilot reproduced their code without licence, and the court's treatment of the section 1202 claim, which covers removal of copyright management information, was narrower than the plaintiffs hoped. In the Roblox case, an artist named Austin Beaulier sued over the removal of NoAI tags from 3D models used in a training dataset. Judge Beth Labson Freeman dismissed it on October 2 as well, holding that the NoAI tag does count as content management information but that the plaintiff had not shown Roblox intentionally stripped it.
Both of those are about section 1202, which concerns information attached to a work. ByteDance's case is about section 1201, which concerns the gate in front of it. The Roblox dismissal and the ByteDance survival on the same day make the split visible: a claim about metadata that got dropped is hard to prove, and a claim about a gate that got bypassed is easier.
The fair use picture is not settled either
The same week brought the unredacted opinion in Thomson Reuters v. ROSS Intelligence, where the Third Circuit held that training a legal research tool on Westlaw headnotes was not fair use. That case is about section 107 and has no technical measures in it at all. ROSS copied editorial summaries and used them to build a competing product.
The three decisions together describe a setting where the legal exposure of an AI training pipeline depends on mechanics rather than on principle. How the data was obtained, whether metadata survived the pipeline, and whether the output competed with the source each carry their own rule.
For a company collecting training data at scale, the operational consequence is that the collection layer is now a compliance surface in its own right. Logging how a crawler behaved, honouring rate limits and robots directives, and refusing to build around platform defences have stopped being courtesies. They now decide whether a case can be dismissed at the pleading stage.
Why the platform's interest matters as much as the creator's
A section 1201 claim belongs to the person who built the gate as well as to the person whose work sits behind it.
YouTube's protective measures exist because the platform has its own interest in controlling how its catalogue is accessed. Bulk downloading is expensive for a service that serves video at scale, and scraping has consequences for bandwidth, for advertising measurement and for the licensing deals the platform negotiates with rights holders. A ruling that treats those measures as access controls hands platforms a legal tool against automated collection of any kind, and they did not have one before.
That reaches beyond AI training. It covers price monitoring, academic research datasets, competitive analysis and anything else that reads a site at volume. The anti-circumvention provision predates all of it by decades, and it is now being asked to govern a category of activity its drafters did not imagine.
For rights holders, the consequence is a second way into court. A creator who cannot win a fair use argument against a model trainer may still have a claim about how the work was obtained, and that claim does not require proving economic harm in the same way. It requires proving that a technical measure was defeated.
Where the discovery is likely to go
If the case survives to discovery, the interesting material will be operational rather than legal. Which crawling tools were used, whether they were configured to respect platform directives, how many requests were made and over what period, and what the company's own engineers wrote about the risks at the time.
Internal documents have driven the outcome of several recent AI cases, and a training pipeline at ByteDance's scale leaves a substantial trail. The question of whether anyone flagged the legal exposure before the scraping began is the one that tends to decide settlement value.
What comes next
The ruling does not decide whether ByteDance infringed. It keeps the case alive at the pleading stage, which means the YouTubers get discovery. That is where the specifics become public: which videos, which tools, which measures were defeated, and how many times.
The likeliest path is a settlement before trial. Suits of this shape usually end that way, and ByteDance has no interest in a public record of how a training pipeline operated. But the legal question is now in the system, and the answer will shape how every future scraper confronts a platform's defences.
There is a wider point. The AI industry's data problem has been treated as a licensing problem: pay the rights holder and everything is fine. This line of cases suggests that paying is not sufficient, because the platform that hosts the work has its own interests, codified in technical measures and backed by a statute that does not care whether the work was free to read. Getting the content legally and getting it cleanly are two separate obligations now.
Related articles
A Renewable Energy Company's Founders Just Ordered 20,000 Nvidia Rubin GPUs
The chip order is the easy part. The financing structures underneath it are where the risk sits.
The White House Handed AI Policy to Its Intelligence Chief, and the Enforcement Is Coming From Somewhere Else
The federal government's AI agenda is being set in at least two places at once, and the two are not obviously aligned.
Australia Ordered Every Federal Agency to Audit Its Legacy Systems After the Medicare Breach
The audit will surface the exposure. Funding the fix is a separate exercise with a separate budget.
A Wuhan Court Priced AI Compute Into a Copyright Ruling, and Set a New Kind of Precedent
The amount is small. The method treats an API invoice as evidence of what the work cost to make.