What Happened
The Seattle Times filed a complaint in the U.S. District Court for the Western District of Washington, accusing OpenAI and Microsoft of using automated crawlers to copy articles behind its paywall and incorporate them into training data for GPT models. Newsday followed with a similar suit in the Eastern District of New York, claiming the same methodology was applied to its archives. Both plaintiffs cite internal logs showing millions of requests from IP ranges linked to Azure and OpenAI infrastructure, alleging violation of the Copyright Act and breach of terms of service.
The lawsuits come as the EU’s voluntary General‑Purpose AI Code of Practice, signed by over 150 companies including Microsoft, commits signatories to respect copyright and avoid illicit data harvesting. Plaintiffs argue the defendants’ actions contradict those commitments, seeking injunctive relief, destruction of infringing datasets, and monetary damages potentially exceeding $150 million combined.
Why It Matters
A ruling that scraping paywalled content for AI training constitutes infringement would reshape the data acquisition strategies of generative AI firms across Europe, where regulators are already scrutinizing compliance with the AI Act. It would likely compel companies to negotiate licensing deals with publishers or invest in synthetic data pipelines, raising operating costs and slowing model iteration.
Second‑order effects include a possible shift toward collective licensing models among European news conglomerates, strengthening their bargaining power. Conversely, AI startups lacking the resources to secure broad licenses may face barriers to entry, consolidating advantage among incumbents that can afford premium data deals.
Who Wins & Loses
Winners are the Seattle Times, Newsday, and other European publishers who could secure licensing revenue or legal precedent protecting their content. Losers include OpenAI and Microsoft, which may face costly settlements, mandatory data purges, and reputational harm; smaller AI ventures that rely on scraped web data could also be disadvantaged if licensing becomes the norm.
What to Watch
Watch for the courts’ decisions on fair use defenses and the potential for a consolidated multidistrict litigation. Monitor any settlement talks that could set a benchmark royalty rate for news content used in AI training. In Brussels, observe whether the European Commission uses the case to tighten enforcement of the General‑Purpose AI Code or to propose amendments to the AI Act that explicitly ban paywall‑bypassing scraping.
Social PulseRedditHackerNews
Engineers express unease about the legal risk of using web‑scraped data, while founders debate whether licensing costs will stifle innovation. Publishers welcome the lawsuits as a long‑overdue defense of their intellectual property, signaling a broader shift toward valuing journalistic content in the AI economy.
Sources
- Two more newsrooms join the case against OpenAI and Microsoft