the tech buzz

SUBSCRIBE
AIEnterpriseDealsSecurityCrypto
Newsletter

the tech buzz

Your premier source for technology news, insights, and analysis. Covering the latest in AI, startups, cybersecurity, and innovation.

FOLLOW US

THE DAILY

Get the latest technology updates delivered straight to your inbox.

Company

  • About Us
  • Editorial Team
  • Write For Usnew
  • Contact Us
  • Advertisenew

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Disclaimer
  • EULA
  • AI Code of Conduct

Resources

  • Newsletters
  • RSS Feeds
  • Subscribe
  • Pricing & Packages
  • Sitemap
  • Archives
  • TechBuzz Pressnew

PUBLISH WITH US

Reach 1.1M+ subscribers via TechBuzz Press.

TechBuzz Press

HAVE A TIP?

Send us a tip using our anonymous form.

Send a tip

HAVE QUESTIONS?

Reach out to us on any subject.

Ask Now

Browse by Category

AIBlockchainCloudSecurityDataDealsInvestmentsEnterpriseVenturesIoTMobileRoboticsSoftwareStartupsAppleMetaMicrosoftOpenAiGoogleTesla

© 2026 The Tech Buzz. All rights reserved.

the tech buzz

AI Startups Ditch Web Scraping for Custom Data Collection

ArticlesNewsletters
ArticlesNewsletters
AI/proprietary data

AI Startups Ditch Web Scraping for Custom Data Collection

Companies hire artists, chefs, and workers to create proprietary datasets for AI models

by The Tech Buzz

PUBLISHED: Thu, Oct 16, 2025, 7:38 PM UTC | UPDATED: Fri, Sep 4, 2026, 1:28 PM UTC

Add as a preferred source on Google
AI Startups Ditch Web Scraping for Custom Data Collection

The era of scraping free web data for AI training is ending. Companies like Turing Labs and Fyxer are now paying premium rates to hire artists, chefs, and construction workers to create custom datasets, betting that proprietary training data will become their biggest competitive advantage as the AI boom matures.

Taylor strapped a GoPro to her forehead every morning this summer, enduring headaches and red marks to help train the future of AI. For one week, she and her roommate filmed themselves painting, sculpting, and doing household chores - earning premium pay to create something no web scraper could deliver: perfectly synchronized, multi-angle footage of human problem-solving in action.

"We woke up, did our regular routine, and then strapped the cameras on our head and synced the times together," Taylor told TechCrunch. "Then we would make our breakfast and clean the dishes. Then we'd go our separate ways and work on art." The work paid well but came with a physical cost. "It would give you headaches. You take it off and there's just a red square on your forehead."

Taylor was working for Turing Labs, an AI company that's abandoning the old playbook of scraping free web data. Instead, Turing is contracting with chefs, construction workers, electricians, and artists to create proprietary video datasets for their vision models. The goal isn't teaching AI to paint, but developing abstract skills around sequential problem-solving and visual reasoning that can't be found in existing datasets.

"We are doing it for so many different kinds of blue-collar work, so that we have a diversity of data in the pre-training phase," Turing Chief AGI Officer Sudarshan Sivaraman told TechCrunch. "After we capture all this information, the models will be able to understand how a certain task is performed."

This shift from free web scraping to expensive custom collection represents a fundamental change in how AI companies think about competitive advantage. With foundation models becoming commoditized, proprietary training data is emerging as the new battleground. Companies are no longer just building better algorithms - they're building better datasets.

Advertisement

Fyxer, an email management startup, discovered this lesson the hard way. Founder Richard Hollingsworth initially tried standard approaches but found that his AI needed something web scrapers couldn't provide: the nuanced judgment of experienced executive assistants who understand email etiquette and priority.

"We used a lot of experienced executive assistants, because we needed to train on the fundamentals of whether an email should be responded to," Hollingsworth told TechCrunch. In Fyxer's early days, executive assistants outnumbered engineers and managers four-to-one. "It's a very people-oriented problem. Finding great people is very hard."

The quality-over-quantity approach has become gospel among AI startups. Hollingsworth learned that smaller, more carefully curated datasets often outperform massive scraped collections. "We realized that the quality of the data, not the quantity, is the thing that really defines the performance," he said.

This principle becomes even more critical when companies use synthetic data to expand their training sets. Turing estimates that 75 to 80 percent of its data is synthetic, extrapolated from the original GoPro videos. But synthetic data amplifies both the strengths and flaws of the original dataset, making high-quality source material essential.

"If the pre-training data itself is not of good quality, then whatever you do with synthetic data is also not going to be of good quality," Sivaraman explained.

Beyond quality concerns, there's a powerful competitive logic driving this trend. Custom data collection creates natural moats that are harder for competitors to replicate. Anyone can download an open-source model, but not everyone can assemble teams of expert annotators or convince artists to wear cameras for weeks.

Advertisement

"We believe that the best way to do it is through data," Hollingsworth told TechCrunch, "through building custom models, through high quality, human led data training."

This approach requires significant upfront investment and operational complexity that many startups aren't prepared for. Taylor's work with Turing required careful coordination - she needed seven hours daily to produce five hours of usable footage, accounting for breaks and physical recovery from wearing the equipment.

But for companies willing to make the investment, proprietary data collection offers something the old web-scraping approach never could: datasets tailored specifically to their use cases, competitive advantages that can't be easily copied, and training data that gets better over time rather than becoming commoditized.

As AI capabilities plateau and competition intensifies, the companies with the best datasets - not necessarily the best algorithms - may emerge as the winners. The shift from scraping to custom collection signals a maturing industry where data quality, not just model size, determines success.

The AI industry's pivot from free web scraping to premium custom data collection marks a critical inflection point. As foundation models become commoditized, companies are realizing that proprietary, high-quality datasets may be their best path to sustainable competitive advantage. This shift demands significant upfront investment and operational complexity, but for startups willing to pay artists to wear GoPros and hire armies of executive assistants, it offers something invaluable: training data that can't be easily replicated by competitors.

More Topics:
proprietary data

Advertisement

Advertisement

Trending Now

1

GoPro CEO Vows Cameras Stay Core After Starman Deal

2

Judge Splits Ruling in X vs. Twitter Rival Fight

3

Tim Cook Steps Down, Ternus Takes Apple's Helm

4

Google's Lyria 3.5 Brings AI Music to Gemini

5

Google Translate Gets Listening Mode, Live Background Mode

People Also Ask

Custom data collection involves AI companies hiring workers like artists, chefs, and construction workers to create proprietary datasets instead of scraping free web data. Companies pay premium rates for human-generated training data tailored to their specific AI models and use cases.

AI startups are ditching web scraping because proprietary training data creates competitive advantages that can't be easily replicated. Custom datasets offer higher quality, task-specific information that web scrapers can't provide, especially for nuanced skills like email etiquette or visual problem-solving.

While specific rates aren't disclosed, Turing Labs pays 'premium pay' to artists who wear GoPros for 7 hours daily to create video datasets. Artists endure physical discomfort including headaches and red marks from equipment to capture synchronized multi-angle footage.

Turing estimates that 75 to 80 percent of its training data is synthetic, extrapolated from original GoPro videos. However, synthetic data amplifies both strengths and flaws of source material, making high-quality original datasets essential for good performance.

Fyxer hired experienced executive assistants at a 4-to-1 ratio over engineers and managers to train their email AI. The company needed human expertise to teach nuanced judgment about email etiquette, priority, and response requirements that web scraping couldn't provide.

Quality is more important than quantity for AI training datasets. Companies like Fyxer found that smaller, carefully curated datasets often outperform massive scraped collections. High-quality source data becomes especially critical when using synthetic data expansion methods.

More in AI

Google's Lyria 3.5 Brings AI Music to Gemini

Google's Lyria 3.5 Brings AI Music to Gemini

Rogue OpenAI Agents Hijacked a German Wiki

Rogue OpenAI Agents Hijacked a German Wiki

Altman Apologizes for Messy GPT-6 Astra Rollout

Altman Apologizes for Messy GPT-6 Astra Rollout

Microsoft's Project Zenith Targets AI Developers

Microsoft's Project Zenith Targets AI Developers

Nvidia's $99B Bet: AI's Biggest Backer Emerges

Nvidia's $99B Bet: AI's Biggest Backer Emerges

Samsung's AI Rally Reshapes Dating, TV and Majors

Samsung's AI Rally Reshapes Dating, TV and Majors

More Articles

Accel Nears $1B Deal for Thinking Machines at $40B

Accel Nears $1B Deal for Thinking Machines at $40B

Sep 3

Utilities Race to Fusion Startups as AI Strains Grid

Utilities Race to Fusion Startups as AI Strains Grid

Sep 3

Meta Offers 95% AI Discount for Your Data

Meta Offers 95% AI Discount for Your Data

Sep 3

Abliteration.AI Sells Access to Uncensored Models

Abliteration.AI Sells Access to Uncensored Models

Sep 3

OpenAI Launches GPT-6 Astra, Claims 'AGI Era'

OpenAI Launches GPT-6 Astra, Claims 'AGI Era'

Sep 3

OpenAI's GPT-6 Astra Debuts, Claims 'AGI Era'

OpenAI's GPT-6 Astra Debuts, Claims 'AGI Era'

Sep 3