the tech buzz

SUBSCRIBE
AIEnterpriseDealsSecurityCrypto
Newsletter

the tech buzz

Your premier source for technology news, insights, and analysis. Covering the latest in AI, startups, cybersecurity, and innovation.

FOLLOW US

THE DAILY

Get the latest technology updates delivered straight to your inbox.

Company

  • About Us
  • Editorial Team
  • Write For Usnew
  • Contact Us
  • Advertisenew

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Disclaimer
  • EULA
  • AI Code of Conduct

Resources

  • Newsletters
  • RSS Feeds
  • Subscribe
  • Pricing & Packages
  • Sitemap
  • Archives
  • TechBuzz Pressnew

PUBLISH WITH US

Reach 1.1M+ subscribers via TechBuzz Press.

TechBuzz Press

HAVE A TIP?

Send us a tip using our anonymous form.

Send a tip

HAVE QUESTIONS?

Reach out to us on any subject.

Ask Now

Browse by Category

AIBlockchainCloudSecurityDataDealsInvestmentsEnterpriseVenturesIoTMobileRoboticsSoftwareStartupsAppleMetaMicrosoftOpenAiGoogleTesla

© 2026 The Tech Buzz. All rights reserved.

the tech buzz

Reddit sues Perplexity for allegedly stealing data via scrapers

ArticlesNewsletters
ArticlesNewsletters
AI

Reddit sues Perplexity for allegedly stealing data via scrapers

Reddit files major lawsuit against AI search company for circumventing protections

by The Tech Buzz

PUBLISHED: Wed, Oct 22, 2025, 6:53 PM UTC | UPDATED: Fri, Sep 4, 2026, 9:13 PM UTC

Add as a preferred source on Google
Reddit sues Perplexity for allegedly stealing data via scrapers

Reddit just filed a bombshell lawsuit against AI search startup Perplexity, alleging the company used third-party data scrapers to steal Reddit content and circumvent the platform's protections. The lawsuit targets not just Perplexity but three data-scraping companies that Reddit says are fueling an "industrial-scale data laundering economy" in AI. This marks Reddit's most aggressive legal push yet to monetize its treasure trove of human conversation data that's become gold for AI training.

Reddit is declaring war on AI companies that won't pay up. The social platform just dropped a lawsuit against Perplexity AI and three data-scraping companies, accusing them of running an elaborate scheme to steal Reddit's most valuable asset - millions of authentic human conversations.

The complaint filed today reads like a heist movie. Reddit compares the defendants to "would-be bank robbers" who "knowing they cannot get into the bank vault, break into the armored truck carrying the cash instead." The armored truck in this case? Third-party scrapers SerpApi, Oxylabs, and AWMProxy that Reddit says Perplexity hired to do its dirty work.

Here's where it gets juicy. Reddit sent Perplexity a cease-and-desist letter back in May 2024, demanding the AI search company stop scraping Reddit data. Perplexity promised they'd play nice and respect Reddit's robots.txt file. But according to the lawsuit, the volume of Reddit citations on Perplexity actually went up after that conversation.

Reddit didn't stop there. They set a trap - creating a post that could only be crawled by Google. "Within hours," the lawsuit claims, Perplexity had somehow accessed and used that content in its answer engine. "The only way that Perplexity could have obtained that Reddit content," Reddit argues, "is if it and/or its co-defendants scraped Google search results."

Advertisement

This lawsuit hits at the heart of the AI industry's biggest tension right now. Reddit's treasure trove of human-written, community-ranked content is exactly what AI companies need to train better models. But Reddit learned its lesson from the 2023 API pricing controversy that sparked massive protests - if AI companies want the data, they need to pay for it.

The strategy's been working. Reddit has already struck lucrative deals with OpenAI for ChatGPT integration and Google for AI training access. The company reportedly wants even better terms as these partnerships come up for renewal.

"AI companies are locked in an arms race for quality human content," Ben Lee, Reddit's chief legal officer, said in a statement. "That pressure has fueled an industrial-scale 'data laundering' economy." Lee painted a picture of a shadowy ecosystem where scrapers "mask their identities, hide their locations, and disguise their web scrapers" to steal content from Google search results.

The defendants read like a cybercrime lineup. Oxylabs UAB is described as a Lithuanian data scraper, AWM Proxy as a "former Russian botnet," and SerpAPI as a company that "openly advertises its shady circumvention tactics."

Advertisement

Perplexity isn't backing down. "We will always fight vigorously for users' rights to freely and fairly access public knowledge," Jesse Dwyer, the company's head of communication, told The Verge. The response frames this as a battle over internet openness rather than copyright infringement.

But Reddit's legal team clearly did their homework. This isn't their first rodeo - they previously sued Anthropic over similar alleged scraping violations. They're building a pattern of aggressive enforcement that sends a clear message: pay up or face the courts.

The timing couldn't be more critical. As AI companies race to build better models, Reddit sits on one of the internet's largest collections of authentic human dialogue. Every upvoted comment, every community discussion, every niche subreddit conversation represents training data that's incredibly hard to replicate artificially.

This lawsuit represents more than just Reddit protecting its turf - it's a defining moment for how AI companies will access training data going forward. If Reddit wins, it could establish a legal precedent forcing AI companies to negotiate licensing deals rather than rely on scraped content. For Perplexity, this is an existential threat to their business model. The outcome will likely influence how every major AI company approaches data acquisition, potentially reshaping the entire industry's relationship with content platforms.

Advertisement

Advertisement

Trending Now

1

Does Gemini Have a Limit? How Google's Usage Caps Actually Work in 2026

2

Black Friday 2026: When It Is, and Why It Often Isn't the Cheapest Day

3

Nscale Eyes $3.5B Pre-IPO Round After Anthropic Deal

4

GoPro CEO Vows Cameras Stay Core After Starman Deal

5

Judge Splits Ruling in X vs. Twitter Rival Fight

People Also Ask

Reddit is suing Perplexity AI for allegedly using third-party data scrapers to steal Reddit content and circumvent the platform's protections. The lawsuit claims Perplexity continued scraping Reddit data even after receiving a May 2024 cease-and-desist letter, with citations actually increasing afterward.

Reddit's lawsuit targets Perplexity AI plus three data-scraping companies: SerpApi, Oxylabs UAB (Lithuanian data scraper), and AWMProxy (described as a former Russian botnet). Reddit alleges these scrapers helped Perplexity bypass Reddit's protections to access content illegally.

Reddit set a trap by creating test content that could only be crawled by Google. Within hours, Perplexity had accessed and used that content in its answer engine. Reddit argues this proves Perplexity was scraping Google search results to obtain Reddit data.

Reddit has struck lucrative licensing deals with OpenAI for ChatGPT integration and Google for AI training access. These partnerships allow the companies to legally access Reddit's content for AI model training, unlike alleged unauthorized scraping by Perplexity.

Reddit is aggressively enforcing paid licensing deals with AI companies rather than allowing free scraping. After the 2023 API pricing controversy, Reddit learned to monetize its valuable human conversation data by requiring AI companies to pay for access or face lawsuits.

Perplexity's head of communication Jesse Dwyer stated they will fight vigorously for users' rights to freely access public knowledge. The company frames the dispute as a battle over internet openness rather than acknowledging copyright infringement claims.

More in AI

Does Gemini Have a Limit? How Google's Usage Caps Actually Work in 2026

Does Gemini Have a Limit? How Google's Usage Caps Actually Work in 2026

Google's Lyria 3.5 Brings AI Music to Gemini

Google's Lyria 3.5 Brings AI Music to Gemini

Rogue OpenAI Agents Hijacked a German Wiki

Rogue OpenAI Agents Hijacked a German Wiki

Altman Apologizes for Messy GPT-6 Astra Rollout

Altman Apologizes for Messy GPT-6 Astra Rollout

Microsoft's Project Zenith Targets AI Developers

Microsoft's Project Zenith Targets AI Developers

Nvidia's $99B Bet: AI's Biggest Backer Emerges

Nvidia's $99B Bet: AI's Biggest Backer Emerges

More Articles

Samsung's AI Rally Reshapes Dating, TV and Majors

Samsung's AI Rally Reshapes Dating, TV and Majors

Sep 4

Accel Nears $1B Deal for Thinking Machines at $40B

Accel Nears $1B Deal for Thinking Machines at $40B

Sep 3

Utilities Race to Fusion Startups as AI Strains Grid

Utilities Race to Fusion Startups as AI Strains Grid

Sep 3

Meta Offers 95% AI Discount for Your Data

Meta Offers 95% AI Discount for Your Data

Sep 3

Abliteration.AI Sells Access to Uncensored Models

Abliteration.AI Sells Access to Uncensored Models

Sep 3

OpenAI Launches GPT-6 Astra, Claims 'AGI Era'

OpenAI Launches GPT-6 Astra, Claims 'AGI Era'

Sep 3