
The long-running legal battle between news publishers and OpenAI is finally heading into its next big phase. After consolidating separate copyright infringement lawsuits from major outlets like The New York Times, Ziff Davis, and the New York Daily News, the parties presented their core arguments against OpenAI to a federal court this month, asking a judge to rule on key questions around AI training data.
At the heart of the dispute is whether training ChatGPT on copyrighted news articles counts as “fair use” under US copyright law, or if AI developers must pay to license the written material that powers their products.
“A rounding error” for tech giants
The federal government submitted a statement suggesting that licensing burdens could hinder American AI development. However, publishing leaders are pushing back against that premise. Writing in Fortune, Ziff Davis CEO Vivek Shah pointed out a striking financial contrast.
Shah noted that paying annual royalties across the news industry would cost roughly $20 billion. This figure pales in comparison to the $750 billion OpenAI expects to spend building out computing infrastructure by 2030. In Shah’s view, paying publishers for quality content amounts to little more than a “rounding error” for AI firms, yet it would create a sustainable ecosystem for journalism.
Publishers also argue that unauthorized AI scraping hurts web traffic and replaces original reporting with low-quality, AI-generated summaries (via CNET).
Implied licenses and expressed consent
OpenAI, however, is fully committed to its legal defense. In court filings, the company argued that using publicly available internet data constitutes fair use because pretraining causes no direct market harm and transforms facts into new tools.
OpenAI also claimed it held an implicit license to scrape publisher sites because media outlets did not initially block AI crawlers using standard web protocols like robots.txt. The ChatGPT maker pointed to search engines like Google, which have relied on automated web indexing for decades to build internet search engines.
However, things have changed quickly on the internet. Many media outlets are now actively blocking AI bots and erecting more formidable paywalls to safeguard their work. As discovery wraps up, the court could soon rule on whether scraping published articles crosses the legal line into copyright infringement. The final word will set a massive precedent for the entire tech industry.
The post Publishers Push OpenAI to Pay Up for Their Scraped Content in Landmark Copyright Case appeared first on Android Headlines.