Big Tech leaders are spending millions of dollars ā and pushing dubious national security concerns ā to try to prevent federal regulators from forcing them to pay for the copyrighted works their companies are using to train their artificial intelligence (AI) systems.
At issue is a new effort by the US Copyright Office to consider how to apply US copyright law to the nascent AI industry. The matter has triggered impassioned pushback from powerful tech interests who say they must have access to peopleās hard work for free, or the future of their industry will be jeopardized.
The fight comes asĀ artists,Ā actors,Ā news organizations, and others have sued AI companies using their work to train the emergent technology on how toĀ create images in the style of certain artists,Ā replicate voicesĀ of singers,Ā write new literature based on copyrighted works, and many other instances in which original work is being harvested off the internet free of charge.
As the AI industry is buffeted byĀ executive shake-upsĀ and mounting concerns that AI systems areĀ growing too powerful, Google, Microsoft, Meta Platforms, and Big Tech venture capital firm Andreessen Horowitz have spent over $30 million lobbying lawmakers and regulators on AI and other tech-related issues.
Andreessen Horowitz ā whichĀ providedĀ funding for Airbnb and Facebook, andĀ helpedĀ financeĀ Elon Muskās takeover of Twitter ā has even claimed that if the Copyright Office were to enforce its existing laws protecting copyrighted works from exploitation, investment dollars could be lost and US national security could be threatened.
āOver the last decade or more, there has been an enormous amount of investment ā billions and billions of dollars ā in the development of AI technologies, premised on an understanding that, under current copyright law, any copying necessary to extract statistical facts is permitted,ā Andreessen HorowitzĀ wrote in a commentĀ to the Copyright Office.
āA change in this regime will significantly disrupt settled expectations in this area,ā the firm continued. āThose expectations have been a critical factor in the enormous investment of private capital into US-based AI companies which, in turn, has made the United States a global leader in AI. Undermining those expectations will jeopardize future investment, along with US economic competitiveness and national security.ā
TheĀ New York Times, the Screen Actors GuildāAmerican Federation of Television and Radio Artists (SAG-AFTRA), theĀ News Media Alliance,Ā Getty Images, and other organizations and trade groups representing artists, musicians, and journalists have complained that AI companies are violating copyright law by copying their material and using it to train AI. The use of AI was a core concern during this yearās historicĀ writersāĀ andĀ actorsāĀ strikes.
āAlmost all of these AI companies are ingesting copyrighted works in order to train their AI, and in most instances they are not licensing, theyāre not getting the permission, and theyāre not compensating the copyright owners for using those works,ā Keith Kupferschmid, CEO of the Copyright Alliance, told us.
The Copyright Alliance, which represents over two million copyright holders and over fifteen thousand organizations,Ā said inĀ a comment to the Copyright Office that other than online piracy, āno copyright issue has drawn more interest from the Copyright Alliance membership than generative AI.ā
āOverzealous Enforcement of Copyrightā
The Copyright Office has roughly 440 employees tasked with examining hundreds of thousands of copyright registrations each year. RoughlyĀ five hundred thousand copyrights registrations are issued, which provide protections for original works, ideas, concepts, art, music, and other works.
Last month, President Joe Biden issued anĀ executive orderĀ requiring the US Patent and Trade Office to work with the Copyright Office on recommendations governing AI, including how copyrighted material is used to train AI.
The Copyright Office isĀ currently conducting a studyĀ examining a potential mandate that would require AI developers to disclose training materials and compensation for copyright holders whose works were used to train AI.
The Copyright Office began soliciting public comments on August 20 and has received over ten thousand comments from rightsholders, trade groups, AI developers, and venture capital firms.
Andreessen Horowitz, a Silicon Valley-based firm thatĀ spentĀ more thanĀ $800,000Ā inĀ 2023Ā lobbying the White House, lawmakers, and federal agencies on AI, cryptocurrencies, and other matters,Ā issued a comment with the Copyright OfficeĀ on October 31.
In the comment letter, the firm claims that AI can revolutionize the fields of medicine, education, technology, and warfare, but companies need free access to copyrighted material to do so. Especially for an AI technology called āLarge Language Models,ā which is ātrained on something approaching the entire corpus of the written word,ā Andreessen Horowitz wrote.
The firm claimed that if the Copyright Office were to make AI developers pay to use copyrighted material, it would risk billions of dollars in investments and threaten national security.
āThe United States is currently at the vanguard of the AI industry as a direct result of these expectations and investments,āĀ wroteĀ Andreessen Horowitz. āThere is a very real risk that the overzealous enforcement of copyright when it comes to AI training. . . could cost the United States the battle for global AI dominance.ā
Andreessen Horowitz and its Big Tech brethren believe that the fair use doctrine of copyright law allows them to hoover up information and use it to train AI. The fair use doctrine allows the use of copyrighted material for news, commentary and criticism, research, and when the use of the material produces a new concept or body of work that is different from the original version.
MicrosoftĀ spentĀ $6.8 millionĀ lobbyingĀ Congress and a slew of federal departments on AI, facial recognition technology, and other issues. Microsoft is aĀ partial owner of OpenAI, which operatesĀ DALL-EĀ andĀ ChatGPT, two of the leading image- and text-based AI technologies currently in use.
OpenAI, whichĀ recently registeredĀ to be its own lobbying firm, claims that its technology does not store exact copies of text and images and that ChatGPT doesnāt provide āverbatim repetition or āmemorizationā of training data,ā according to itsĀ comment filed with the Copyright Office.
OpenAI said its AI technology is trained on information publicly available on the internet, information obtained through licensing agreements, and information āthat our users or human trainers create and provide,ā OpenAI wrote to the Copyright Office.
OpenAI went on to say that given the vast amount of information on the internet, having to pay to use it would be impractical.
āThe diversity and scale of the information available on the internet is thus both necessary to training a āwell-educatedā model (which, again, does not contain copyrighted expression) and also makes licensing every copyrightable work contained therein effectively impossible,ā OpenAI wrote.
OpenAI said fair use is central to its training process and that a ārestrictive interpretation . . . could drive massive investments in AI research and supercomputing infrastructure overseas.ā
Meta, Facebookās parent company, hasĀ spentĀ $14.6 millionĀ this yearĀ lobbying Congress and the Biden administration on AI and other tech-related issues. In comments filed with the Copyright Office,Ā Meta claimedĀ that it is only extracting āunprotectable facts, ideas, and conceptsā from copyrighted work, all of which are not protected by copyright law.
But even if they were protected, Meta argued, the widespread extraction and use of those works would fall under the fair use doctrine. Meta compared using the extracted material to train AI to teaching a child how to speak.
āJust as a child learns language . . . by hearing everyday speech, bedtime stories, songs on the radio, and so on, a model ālearnsā language by being exposed ā through training ā to massive amounts of text from various sources,ā Meta wrote.
Google spentĀ $9.2 millionĀ this yearĀ lobbyingĀ lawmakers on intellectual property enforcement and a slew of issues pertaining to AI and other tech-related matters. Google also believes the fair use doctrine protects AI from copyright infringement, according to aĀ comment the company filed with the Copyright Office.
Big Techās interpretation of fair use doesnāt sit well with copyright advocates.
āWhen you use a copyrighted work without permission of the copyright owner . . . you are an infringer, thereās no question about it,ā Kupferschmid of the Copyright Alliance said.
Kupferschmid said he doesnāt agree with most of Big Techās fair use arguments. Many tech companies pointed to aĀ Supreme Court caseĀ that found Google was allowed to copy copyrighted work and use it on its website for search purposes, which is fundamentally different from what is happening with AI, Kupferschmid said.
āWhatās going on here is that AI is copying works to create works that could be a substitute in the market for the works that are being copied,ā he added. āWe expect AI companies to license copyrighted works that they are ingesting to train their AI engines and their AI models.
āForced to Pay for It, or Go Bankruptā
Bidenās Federal Trade Commission (FTC), which oversees economic competitiveness and enforces monopoly laws,Ā wrote to the Copyright OfficeĀ that the commission has concerns about AIās potential harm to consumers, workers, and small businesses.
The FTC provided a short list of AI usage it sees as potential copyright violations, which includes training AI on protected works without the creatorās consent, selling work that mimics a creatorās āstyle, vocal or instrumental performance,ā or actions that devalue the work of creators.
Devaluation of work and imitating copyrighted material is especially concerning for some news organizations.
āPublishers invest in producing high-quality content that is taken without permission to train the AI systems . . . that then compete directly with publisher content, reducing publisher revenues and employment, tarnishing their brands, and undermining their relationships with readers,ā the News Media AllianceĀ wrote to the Copyright Office.
The Thomson Reuters Enterprise Centre, which owns Reuters News and a legal research platform called Westlaw, is suing Ross Intelligence, Inc., a legal research company, for allegedly mining Westlawās content and using it to train Rossās AI.Ā Ross shut down in 2021, citing financial issues after being sued by Reuters, but the case is still headed to a jury trial, a federalĀ judge ruledĀ on September 25.
Ross is a direct competitor to Westlaw, and the case could determine how AI companies will operate in the future, Scott Hervey, an entertainment, intellectual property, and business attorney, told us.
ā[The case] will certainly have a significant impact on the way courts look at whether or not the use of third-party content in training and AI is fair use,ā Hervey said.
Hervey doesnāt foresee any federal legislation on copyright and AI coming anytime soon, given other, more pressing issues facing Congress. However, he believes AIās extraction of copyright material will most likely be settled in the courts and result in licensing deals similar toĀ arrangements worked out by music streaming platformsĀ and musicians.
The Associated PressĀ signed a deal with OpenAIĀ to give the tech company access to the APās vast archive of stories and to train AI technology on it.
Hervey added that it is disingenuous for tech companies to say their investments are at risk if they canāt have unlimited access to peopleās hard work.
āJust because the technology company hasnāt figured out a way to make money doesnāt mean that they should get away with infringing work and not paying for it,ā Hervey said. āThere will eventually be a judgment and [AI companies] will either be forced to pay for it or go bankrupt. But weāll see ā this is a quickly moving space.ā
ZNetwork is funded solely through the generosity of its readers.
Donate
