WikiHow Sues OpenAI Over AI Training Data

0
50

WikiHow, the world’s largest collaborative “how-to” guide encyclopedia, has filed a lawsuit against OpenAI, alleging that the AI giant unlawfully scraped over 11,000 of its tutorial articles to train its AI models, including ChatGPT and GPT-4.

The lawsuit was filed on August 21st in the U.S. District Court for the Southern District of Manhattan. WikiHow claims that OpenAI’s actions constitute copyright infringement, specifically violating at least 1,200 registered copyrights.

WikiHow, known for its mission to help anyone learn how to do anything, stated in its filing that the scraped content spans a wide range of topics, from everyday tasks to specialized skills. The platform asserts that after OpenAI ingested this vast amount of its content, ChatGPT became capable of replicating WikiHow’s tutorial formats and generating similar content in response to user prompts.

The models, after absorbing wikiHow articles, are now capable of producing competing tutorial content in a fraction of the time, effort, and cost required to research, write, and edit wikiHow articles. This substitution reduces wikiHow’s revenue, and over time, it has even eroded the incentive to continue producing articles.

The legal document highlights that this alleged unauthorized use of its content allows OpenAI to generate comparable content at a significantly lower cost and faster pace than the original creators. WikiHow argues that this directly impacts its revenue and diminishes its motivation to continue producing high-quality articles.

WikiHow is seeking damages from OpenAI, though the specific amount has not been detailed in the lawsuit. Additionally, the company is requesting a court injunction to prevent OpenAI from further infringing on its copyrighted material.

In response to the allegations, an OpenAI spokesperson stated on August 24th that the company’s models are trained using publicly available data and that such use is considered fair use. Representatives for WikiHow and its legal counsel were not immediately available for comment when contacted by the press.

This legal battle is the latest in a series of challenges faced by AI companies regarding the data used for training their powerful language models. The core issue revolves around whether using publicly accessible web content for AI training constitutes fair use or copyright infringement, a question that courts are increasingly being called upon to decide.

The outcome of this lawsuit could have significant implications for the AI industry and content creators alike, potentially setting precedents for how copyrighted material can be utilized in the development of artificial intelligence.

Source: https://www.ithome.com/0/993/824.htm

LEAVE A REPLY

Please enter your comment!
Please enter your name here