Twitch streamer Warren Pandiscia has sued Twitch and its parent company, Amazon, alleging that they used his streams to develop AI products without obtaining permission or paying licensing fees.
Pandiscia claims Amazon began collecting video streams before Twitch updated its terms of service to address AI data use and introduced an opt-out option for creators.
According to the complaint, the alleged data collection had been taking place for at least two years. Pandiscia cited a previous statement from former Twitch Chief Monetization Officer Mike Minton, who said video data had been used “in a prototyping capacity.”
“Because Amazon AI products are commercialized, Amazon had an overwhelming incentive to acquire training data on an unprecedented scale. Rather than negotiate for lawful licenses or seek permission, defendants accessed the Twitch streams and videos to utilize them as a massive dataset necessary to fuel Amazon’s AI products,” Pandiscia alleged in his complaint filed in the U.S. District Court for the Northern District of California.
It is not clear how much control Amazon has over Twitch streamers’ content, given that they sign a variety of waivers in return for access to the platform and a share of revenue. YouTube has been scraped by almost every major AI model maker as a key resource for training, with its owner Google continuing to rely heavily on the platform to train its AI models, but there has yet to be a successful lawsuit against YouTube over the practice.
Amazon has not been as successful as other tech giants in the frontier model field, with its Nova models far behind Anthropic, OpenAI, and Google in usage. It recently pivoted its strategy to consolidate several smaller models into one frontier model.
The copyright lawsuits are piling up
Amazon is not the only tech giant facing lawsuits over the use of user data to train AI models. Google has at least two lawsuits ongoing, one concerning access to copyrighted books without permission, an issue several AI model makers have been sued over, and another alleging it failed to offer clear opt-out settings for Gmail users.
Music publishing groups have also fought back against the wholesale harvesting of content for AI training. Round Hill sued Suno and Anthropic for up to $1 billion, alleging that the two companies had accessed up to 500 compositions without proper licensing agreements.
The litigation is unlikely to end soon, as courts continue to consider whether different uses of copyrighted material for AI training qualify as fair use and what licensing obligations apply.
Hunting for more data to feed the models
There seems to be an insatiable appetite for almost any kind of data, but as its value has increased, more sites have created licensing agreements and commercial plans for AI model makers. Reddit and Wikipedia are two major examples, with licensing agreements generating hundreds of millions of dollars for the companies. Some music publishers have also signed agreements with Suno and other AI companies.
Proprietary information is also becoming more valuable as AI companies look for material their competitors cannot easily obtain. Google’s reported $20 million purchase of Spirit Airlines corporate records during the carrier’s bankruptcy proceedings illustrates that demand.
For Twitch creators, the immediate concern is whether platform terms clearly explain how uploaded streams may be used beyond hosting and distribution. Regardless of how Pandiscia’s case is decided, it could increase pressure on platforms to disclose their AI-training practices and give creators more meaningful control over their content.