AI

Perplexity AI fails to dismiss Reddit lawsuit over user content scraping

A judge denied Perplexity AI's motion to dismiss Reddit's lawsuit accusing the AI company of illegally scraping user content to train its search engine.

By Tim Editorial

Perplexity AI fails to dismiss Reddit lawsuit over user content scraping
astintech.id

Perplexity AI has lost a motion to dismiss a lawsuit filed by Reddit, which accuses the AI powered search engine company of unlawfully scraping user content to train its systems. The development, first reported via the X account @Polymarket on Saturday, August 1, 2026, marks a significant step in a legal dispute centered on the use of social media platform user data for AI development. The lawsuit stems from Reddit's allegations that Perplexity AI, along with several other entities, harvested user conversations and comments on an industrial scale without permission. Reddit asserts that this scraping was done for commercial purposes, namely to build and train AI products that Perplexity then commercializes.

The discussion platform argues that such actions violate intellectual property rights and its terms of service. According to a Business Insider report published on October 22, 2025, Reddit claims the sued companies obtained information about Reddit posts through Google, rather than by signing data licensing agreements. Reddit views this strategy as an attempt to avoid paying for data that should be protected. Reddit also highlighted that such practices allowed Perplexity to build a company valuation of up to $20 billion from data allegedly obtained unlawfully. The chronology of the dispute began in October 2025 when Reddit formally filed a lawsuit against Perplexity AI and three other entities.

The Associated Press reported on October 22, 2025, that the lawsuit accused the parties of industrial scale scraping of user comments for commercial gain. Reddit contends that these actions harm the platform and its users because data generated by the community is used without fair compensation. The court's decision to deny the motion to dismiss means the case will proceed to trial. This sets an important precedent for the AI industry, which has long relied on public data from various platforms to train large language models and search systems. Perplexity AI is known as a search engine that uses large language models to provide direct answers to users, a departure from traditional search engines. The impact of this ruling could extend across the AI ecosystem.

Many AI companies rely on data from public platforms like Reddit to improve their models. If Reddit wins this lawsuit, other platform companies may be encouraged to take similar steps to protect their users' data. Conversely, AI companies may need to be more cautious in acquiring training data and consider formal licensing schemes. Reddit has long had a complex relationship with AI companies. On one hand, the platform is a valuable data source due to its diverse, constantly updated user generated content. On the other hand, Reddit began restricting third party data access after its API policy changed in 2023, which sparked widespread protests from the developer community. That policy was followed by data licensing agreements with several AI companies willing to pay.

Perplexity AI has not yet issued an official statement regarding the court's decision. However, the company has previously denied allegations of illegal scraping and maintained that its data collection practices comply with applicable regulations. The company also emphasized that it respects intellectual property rights and is committed to cooperating with content owners. This case is one of the biggest legal tests for the AI industry regarding the use of third party data. The trial's outcome could determine the legal boundaries of data scraping for AI model training. If the court rules that scraping user content without permission is a violation, AI companies may have to fundamentally change how they obtain training data.

On the other hand, this decision could also affect negotiation dynamics between content platforms and AI companies. Platforms like Reddit, X, and Facebook hold vast amounts of user data that are highly valuable for AI development. With the court allowing the lawsuit to proceed, content platforms' bargaining power in data licensing negotiations is likely to strengthen. Next developments to watch include the trial schedule and the possibility of an out of court settlement. Given Perplexity's reported $20 billion valuation, the case carries significant financial implications. Reddit has shown its seriousness by suing not only Perplexity but also other entities involved in similar data scraping practices. This court decision also signals to investors and industry players that unauthorized data harvesting carries real legal risks.

Rapidly growing AI companies now face pressure to ensure their training data sources are legal and ethical. This could drive a shift toward more transparent and mutually beneficial data licensing models between content platforms and AI developers.

Sources and references