19.7 C
London
Tuesday, July 28, 2026
Home AI “Google and Reddit do not own the Internet,” web scraper says after...
“google-and-reddit-do-not-own-the-internet,”-web-scraper-says-after-court-win
“Google and Reddit do not own the Internet,” web scraper says after court win

“Google and Reddit do not own the Internet,” web scraper says after court win

3
0

After a big court loss last week, Google has confirmed that it won’t give up its fight to block AI bots from scraping its search results. And Reddit is weirdly along for the ride.

Curiously invoking the Digital Millennium Copyright Act (DMCA), Google sued SerpApi last December. The search giant accused the web scraper of circumventing its anti-scraping technology and then selling content scraped from Google search results through an unauthorized “Google Search API” software service.

According to Google, the anti-scraping tech was in place to protect copyrighted content in search results. Allegedly, SerpApi’s circumvention threatened to disrupt Google’s relationships with rights holders, including some who license content to Google to appear in so-called “knowledge panels” that are displayed in some search results for well-known people or entities.

It was an odd use of the DMCA, since Google search results can’t be copyrighted. But Google was apparently emboldened to explore the legal theory after Reddit filed a very similar lawsuit in October, accusing SerpApi and Google-rival Perplexity of scraping Reddit content that appears in Google results.

In a blog, Google cited Reddit’s lawsuit when announcing its own challenge, which it said it filed as a “last resort” to block “malicious scraping” that violates rights holders’ choices over who can access their content.

Specifically, Google alleged that SerpApi’s circumvention violated its terms and made it impossible to profit from—or offset the cost of—“billions” of bot searches. And before it, Reddit claimed that SerpApi was evading two levels of security: Reddit’s own controls blocking scraping on its platform and Google controls blocking scraping of Reddit content in search results.

Meredith Rose, a senior policy counsel with expertise in the DMCA for a nonprofit public interest group called Public Knowledge, told Ars that Google and Reddit seem to be “sort of grasping at whatever tool is available” in the face of the sudden, continuous rise of AI scraping over the past three years. And while the way they’re using the DMCA is “bizarre”—and “not what the law had sort of contemplated as a use case”—she says it’s not “surprising.” Historically, the DMCA has been an effective tool to quickly stop disfavored content uses and force discussions around licensing, so turning to it may have been an obvious starting point, given Google’s goals.

But Google’s and Reddit’s unusual DMCA arguments don’t seem to be winning ones. Last week, a court took the somewhat rare step of granting SerpApi’s motion to dismiss very early on in Google’s lawsuit. In that case, the judge found that Google had no DMCA standing to sue SerpApi, since it didn’t own any of the content in the search results and has not shown that it’s acting on behalf of any rights holders.

“That does not happen terribly often,” Rose told Ars. “It really boiled down to Google didn’t allege enough about what it was protecting that was copyrighted.”

Likely the timing of that decision wasn’t great for Reddit, which faced a hearing on SerpApi’s motion to dismiss its lawsuit last Thursday. It’s unclear which way the court will rule in that case, but Rose told Ars that the Google ruling doesn’t bode well for Reddit since Reddit can’t claim that it is the content owner or exclusive licensee of content in search results.

“The judge in the Google case said, ‘Well, in order to have standing to bring a lawsuit under the DMCA, you can be the copyright owner or the exclusive licensee or the person who is deploying and manufacturing the technological protection measure at issue,’” Rose told Ars. “Reddit is none of those things.”

SerpApi is hoping that the fight will be over soon, telling Ars that the costly legal battle is worth sticking it out to defend the open web.

“The bottom line is that both Google and Reddit appear to be engaged in attempts to use the DMCA to wall off the open Internet by retroactively claiming control over content that they didn’t author and don’t own,” SerpApi told Ars.

Google’s last chance to keep fight alive

Although Rose agreed with SerpApi that, in granting the motion to dismiss, the court gave SerpApi a big win, the fight is not over yet, as Google has a narrow path forward to keep its war against web scraping alive.

Google acknowledged that search results can’t be copyrighted but argued that “knowledge panels” sometimes include copyrighted content that Google licenses from rights holders. If Google can amend its complaint to argue that rights holders directly authorized Google to use its anti-scraping technology to prevent unauthorized access to content, then Google may be able to block a very limited amount of SerpApi’s scraping.

Google’s spokesperson, José Castañeda, told Ars that Google plans to amend the complaint and is “pleased to see that the Court rejected nearly all of SerpApi’s legal arguments” otherwise attempting to dispute Google’s standing.

“We look forward to filing an amended complaint, as the Court invited us to do, and we remain committed to protecting our services and partners from unauthorized access,” Castañeda said.

However, Rose told Ars that Google has somewhat “talked themselves into a little bit of a corner here, both in this litigation and historically.”

For Google, it could be “very dangerous” to argue that the knowledge panel is “chock full of copyrighted material,” Rose suggested. Since the search giant doesn’t license all the content in the knowledge box, Google could risk future lawsuits if the act of algorithmically creating the knowledge box without licenses suddenly becomes viewed as infringement, Rose said.

“They have to make an argument somehow that there are parts of that that are reproductions of copyrighted material, but the only parts that are reproductions of copyrighted material are the ones they’ve explicitly licensed,” Rose suggested. “Otherwise, they’re admitting that they have been reproducing stuff without licensing it, and that gets them into another fair use fight that they probably don’t want to have.”

Google was given 21 days to amend its complaint, at which point it will become clearer how it plans to thread the needle to keep its DMCA fight going.

SerpApi calls fight “misguided”

Reddit did not respond to Ars’ requests to comment but claimed in a filing ahead of last week’s hearing that it was prepared to discuss how Google’s court loss impacted its case.

Perhaps notably, SerpApi said that Reddit was not among attendees in the courtroom. SerpApi did not comment much on Reddit’s arguments at the hearing but said that the judge appeared to be focused on the nuances of the legal questions. Most particularly, the judge seemed interested in whether Reddit’s agreement with Google actually authorized Google to protect its copyrighted content.

It would appear then that both courts have somewhat narrowed the fight to this key question, but SerpApi seems confident that neither Google nor Reddit can show evidence that scraping public search results harms rights holders.

Asked for comment on Google’s plan to amend its complaint, SerpApi told Ars that there may be little point in continuing to argue over “snippets of text that appear in Google’s Knowledge Panels” after Google launched its attack to supposedly defend “hundreds of thousands of publishers” that appear in search results.

“We hope Google drops this misguided attack on the open Internet, but if necessary, we are prepared to defend SerpApi, our customers, and our principles,” SerpApi said.

SerpApi told Ars that its business has continued to grow while it has fought the DMCA lawsuits but that its customers, including tech giants like Nvidia, Uber, and Adobe, have faced uncertainty as both cases have dragged on.

They “rely on SerpApi every day to provide structured access to search data, which is and has always been a lawful and legitimate business,” SerpApi said.

SerpApi’s right to defend open web, expert says

Beyond its own customers, people invested in the open web have also pondered what the impact of a Google win could be, SerpApi suggested.

“Of course, litigation is incredibly expensive and disruptive,” SerpApi said. “However, we feel strongly that platforms such as Google and Reddit do not own the Internet and that their attempts to weaponize the DMCA threaten all users of the open web.”

Rose told Ars that while SerpApi hasn’t necessarily “endeared itself to a lot of people,” she thinks that “they’re in the right from a policy perspective here.”

Ever since the summer of 2023—when a sort of “API apocalypse” began—publishers have been cutting off access to the open web, Rose said, by attempting to find ways to quickly stop high volumes of AI scraping like SerpApi does.

“This mass reaction to scraping” has caused a “re-enclosure of a lot of the web,” Rose said. And concerningly, that doesn’t just gum up AI training efforts, it also harms research, archiving, journalism, public health reporting, and other important work that depends on anonymous crawling and automated scraping at scale, Rose said.

SerpApi told Ars that it has gotten “overwhelming” feedback from its customers, researchers, and SEO providers in support of its efforts to defend against the Google and Reddit lawsuits. The web scraper is optimistic that the court will grant its motion to dismiss against Reddit and last week celebrated that Google’s lawsuit may only proceed on very limited claims.

In a win, SerpApi warned that Reddit planned to act as a “toll collector,” requiring payments for scraping when it doesn’t even own the content.

“Reddit seeks to consolidate its control over its users’ content so that it will be in a better position to tax that content. Reddit is not acting to protect its users; Reddit is clearing a path to exploit them,” SerpApi warned.

Similarly, SerpApi accused Google of asking the court to ignore that it’s the “largest scraper on the planet,” while agreeing to cut off other web scrapers from information that is “100 percent public.”

For Rose, there’s no winning for online users if, at the end of the fight, public access to data is cut off.

“It’s a very fraught time,” Rose said, acknowledging that it’s not just Reddit and Google but many publishers across the web that are currently tempted to use “whatever tool is available in the toolbox to tamper down” AI scraping.

Some publishers may be financially motivated, like Google and Reddit, while others may be protecting their infrastructure or taking a moral stance, Rose said. Whatever the motivation is, “it’s leading to this kind of broader ecosystem-wide consequence of re-enclosure,” she warned.

And although SerpApi expects to win, the problem with DMCA cases, she suggested, is that it can be difficult to predict the outcome.

“I’m very curious how this is going to play out with Reddit,” Rose said. “I’m always a little bit cynical because to some extent, whenever you get into the copyright realm, a lot of judges just decide it’s all vibes.”

Advance Publications, which owns Ars Technica parent Condé Nast, is the largest shareholder in Reddit.