Microsoft and OpenAI have been sued again, this time by a group of local news media that rarely attract the spotlight. Led by the Emmerich Media Group in Mississippi, several local news organizations in the United States recently filed a lawsuit against the two tech giants, accusing them of long-term, systematic unauthorized use of thousands of copyrighted news articles in large model training processes, generating significant commercial value from these AI products, while the original content providers have received no compensation at all.

This lawsuit is not an isolated case. In recent years, similar tensions have persisted. Large media outlets such as The New York Times have already taken Microsoft and OpenAI to court on the same grounds. However, the involvement of local media has extended the conflict from major newspapers to community newspapers. The plaintiffs clearly stated in the complaint: Microsoft and OpenAI have created and will continue to generate billions of dollars in revenue, while content creators have "received nothing"; more strikingly, some articles that were originally behind paywalls or had access restrictions are also accused of being bypassed and used for training.

The most technical details of the lawsuit focus on copyright notices. The plaintiffs point out that the articles in question originally contained copyright management information—author names, publication institution names, copyright notices, and usage terms—equivalent to the identity cards of the works. They accuse Microsoft and OpenAI of removing these identifiers before feeding the content into training, secretly tearing off the works' identities in family photo albums. Another piece of evidence comes from the output side: the plaintiffs claim that their news content has been repeatedly spit out by AI systems almost verbatim in recent years, which effectively proves that the materials had already been deeply absorbed into the models.

For local media, this is not just about copyright, but also about survival. Local newspapers in the United States have been struggling with declining print sales and shrinking advertising revenues for years, and AI can directly deliver answers to users, further draining their already thin traffic. Therefore, the plaintiffs have called the tech companies' actions "the death knell of local journalism"—in many of the markets they serve, the Emmerich Media Group is the only source of news, and if such media is forced to shrink or close down, the information channels of the neighborhood will be cut off. The joint lawsuit includes not only Emmerich, but also Ojai Media and several publishing and media companies under Coopwood, all independent newspapers and magazine publishers.

In legal terms, the plaintiffs believe that both companies have violated the U.S. Copyright Act and the Digital Millennium Copyright Act (DMCA), constituting copyright infringement and related violations. The complaint also highlights a clear double standard: while Microsoft and OpenAI are accused of ignoring publishers' copyrights, they actively protect their code, models, and systems through licensing agreements, paid services, and legal means; the plaintiffs also specifically brought up a previous issue—OpenAI once publicly complained that competitors used its generated content to train models, harming its own interests, a stance that sharply contrasts with its current alleged behavior.

The demands of the local media are clear: they want the court to award damages and compensation, and also a injunction to force Microsoft and OpenAI to completely remove all content covered by the plaintiffs' copyright from the model training data. As more and more news organizations, publishers, and content creators take up legal weapons, whether AI training data is legal and how content should be licensed and compensated are steadily becoming the top legal issues in the global artificial intelligence industry.