Perplexity AI's bid to dismiss Reddit data-scraping lawsuit rejected

Courts are willing to entertain that AI companies cannot simply harvest copyrighted material without consequence
A federal judge's decision to let Reddit's lawsuit proceed signals judicial skepticism of the AI industry's data practices.
Mark

Why does it matter that the judge didn't dismiss this case? Couldn't Perplexity just win later anyway?

Mimi

The dismissal stage is where weak cases die before they cost anyone real money. If the judge had thrown it out, Reddit would have lost without ever getting to prove its case. Now Perplexity has to actually defend itself in discovery, which means producing documents about how it scraped the data and what it did with it.

Mark

So the court is saying Reddit's argument is plausible?

Mimi

Exactly. The judge is saying: if what Reddit alleges is true—that Perplexity scraped their content without permission—that could be illegal. Now we find out if it actually happened the way Reddit claims.

Mark

What does Reddit stand to gain if it wins?

Mimi

Potentially damages, an injunction stopping the scraping, and precedent that AI companies can't just take whatever data they want. But more immediately, it sends a signal to other AI companies that this strategy might not be cost-free.

Mark

Is this about money or principle?

Mimi

Both. Reddit wants compensation for its data, but it's also about control. The platform spent years building a community that generates valuable content. If AI companies can harvest that for free, Reddit loses leverage over its own product.

Mark

What happens if Perplexity wins?

Mimi

Then the current model holds: scraping is legal, AI companies can train on whatever they find online, and content creators have no recourse. That's the world we've been living in. This case is about whether that changes.

  • A federal judge refused to let Perplexity AI escape Reddit's lawsuit early, rejecting the company's argument that its data scraping fell within legal bounds.
  • The case exposes a fault line running through the AI industry: companies have long treated publicly accessible web content as raw material, but that assumption is now being tested in court.
  • Reddit, which rewrote its terms of service and introduced paid data licensing specifically to block unauthorized AI harvesting, chose litigation when Perplexity continued scraping anyway.
  • The case now enters discovery, where the precise mechanics of how Perplexity collected Reddit's data and what protections it bypassed will face full legal scrutiny.
  • The outcome could force AI companies to license third-party content or negotiate with creators — or, if Perplexity prevails, entrench the current model where scraping remains largely unchallenged.

In late July, a federal court declined to dismiss Reddit's lawsuit against Perplexity AI, allowing claims of copyright infringement and DMCA violations to move forward over the AI company's unauthorized scraping of platform content for model training. The ruling places the case on a path toward discovery and potential trial, marking a meaningful moment in the slow reckoning between an industry built on borrowed data and the creators who generated it. Courts, it seems, are beginning to ask whether the age of large language models can coexist with the rights of those whose words made them possible.

A federal court has cleared the way for Reddit's copyright lawsuit against Perplexity AI to proceed, rejecting the AI company's bid to have the case dismissed before trial. The ruling, handed down in late July, means Reddit's claims under copyright law and the Digital Millennium Copyright Act will now advance into discovery and potentially to trial.

At the heart of the dispute is a familiar tension: Perplexity systematically scraped Reddit's content to train its AI models without permission, Reddit alleges. The platform had already moved to protect its data — revising its terms of service and introducing paid licensing to signal that its content was not freely available for AI training. When Perplexity continued scraping regardless, Reddit sued.

The court's decision to allow the case to proceed is not a verdict, but it carries real weight. A motion to dismiss clears only a low bar — the judge simply found that Reddit's allegations, taken as true, could constitute a legal violation. Still, it signals judicial willingness to scrutinize the AI industry's long-standing assumption that harvesting publicly available web content is either fair use or transformative enough to escape liability.

The stakes extend well beyond this single case. Reddit's lawsuit is part of a broader wave of legal action by publishers, creators, and platforms challenging AI companies' data practices. If Reddit ultimately prevails, it could compel the industry to license content or compensate creators — reshaping the economics of AI development. If Perplexity wins, it would reinforce the current model. For now, the question of who owns the right to use internet data for AI training will be litigated in earnest.

A federal court has rejected Perplexity AI's attempt to have Reddit's lawsuit thrown out before trial, clearing the way for the social media platform to pursue copyright and digital rights claims against the artificial intelligence company. The ruling, handed down in late July, means the case will proceed to discovery and potentially trial—a significant moment in the escalating legal battle over how AI companies source and use data from the internet.

Reddit's complaint centers on a straightforward allegation: Perplexity has systematically scraped content from the platform without permission to train its AI models. The company claims this violates both copyright law and the Digital Millennium Copyright Act, which protects against circumventing technological safeguards. Perplexity had asked the court to dismiss the case early, arguing that its use of Reddit's data fell within legal bounds. The judge disagreed, allowing Reddit's claims to move forward.

The decision carries weight beyond this single dispute. It signals that courts are willing to entertain arguments that AI companies cannot simply harvest vast quantities of copyrighted material from the web without consequence. For years, the AI industry has operated in a legal gray zone, with companies arguing that scraping publicly available data constitutes fair use or that training models on existing content is transformative enough to escape liability. Reddit's case—and now this court ruling—challenges that assumption.

The timing matters. Reddit has become increasingly vocal about protecting its data as AI companies have grown more aggressive in their collection practices. The platform explicitly revised its terms of service and pricing structure to make clear that its content is not freely available for AI training. When Perplexity continued scraping anyway, Reddit decided to fight. This lawsuit is part of a broader wave of legal action: other content creators, publishers, and platforms are pursuing similar claims against various AI companies, testing whether the internet's traditional norms of data access can survive the age of large language models.

Perplexity's loss in this motion to dismiss does not mean Reddit will ultimately win the case. Dismissal motions are a low bar—the court simply has to find that Reddit's allegations, taken as true, could constitute a legal violation. But it does mean the case will now move into the discovery phase, where both sides will exchange documents, depose witnesses, and build their full arguments. This is where the real litigation begins, and where the specific details of how Perplexity scraped Reddit's data, what safeguards it bypassed, and how it used the material will come under scrutiny.

The broader implications are significant. If Reddit prevails, it could establish that AI companies cannot freely use third-party content without permission or compensation. That would reshape the economics of AI development, potentially requiring companies to license data or negotiate with content creators. If Perplexity wins, it would reinforce the current model where scraping is largely permissible. The court's decision to let the case proceed suggests at least some judicial skepticism of that model, though the ultimate outcome remains uncertain.

For now, Reddit has cleared an important hurdle. The lawsuit will continue, and the question of who owns the right to use internet data for AI training—and under what terms—will be litigated in earnest.

Reddit's complaint centers on the allegation that Perplexity systematically scraped content without permission to train its AI models
— Court filing summary
Contattaci Domande frequenti