the tech buzz

SUBSCRIBE
AIEnterpriseDealsSecurityCrypto
Newsletter

the tech buzz

Your premier source for technology news, insights, and analysis. Covering the latest in AI, startups, cybersecurity, and innovation.

FOLLOW US

THE DAILY

Get the latest technology updates delivered straight to your inbox.

Company

  • About Us
  • Editorial Team
  • Write For Usnew
  • Contact Us
  • Advertisenew

Legal

  • Privacy Policy
  • Terms of Service
  • Cookie Policy
  • Disclaimer
  • EULA
  • AI Code of Conduct

Resources

  • Newsletters
  • RSS Feeds
  • Subscribe
  • Pricing & Packages
  • Sitemap
  • Archives
  • TechBuzz Pressnew

PUBLISH WITH US

Reach 1.1M+ subscribers via TechBuzz Press.

TechBuzz Press

HAVE A TIP?

Send us a tip using our anonymous form.

Send a tip

HAVE QUESTIONS?

Reach out to us on any subject.

Ask Now

Browse by Category

AIBlockchainCloudSecurityDataDealsInvestmentsEnterpriseVenturesIoTMobileRoboticsSoftwareStartupsAppleMetaMicrosoftOpenAiGoogleTesla

© 2026 The Tech Buzz. All rights reserved.

the tech buzz

Meet the Billionaires Selling AI Its Training Data

ArticlesNewsletters
ArticlesNewsletters
AI/Handshake

Meet the Billionaires Selling AI Its Training Data

Meet Mercor, Handshake, and the $10B AI data startups now profiting where labs struggle

by The Tech Buzz

PUBLISHED: Mon, Dec 15, 2025, 12:04 PM UTC | UPDATED: Fri, Sep 4, 2026, 7:17 PM UTC

Add as a preferred source on Google
Meet the Billionaires Selling AI Its Training Data

A year after Mercor hit $500 million in annualized revenue, the 22-year-old founders of the data-labeling startup just became billionaires, joining an unexpected gold rush. While frontier AI labs like OpenAI and Anthropic burn billions chasing superintelligence, the only companies actually making money off AI right now are the ones selling them the raw material: human expertise. Roughly $10 billion is flowing annually into training data providers, turning a sleepy corner of the AI infrastructure world into the hottest startup category.

Mercor started as something almost boring. When Brendan Foody was 19, he and two high school friends launched it in 2023 to help their other startup-founding friends hire software engineers overseas. Language models screened resumes. Models did the interviews. By the time Scale AI came knocking with a request for 1,200 specialized coders in early 2024, the startup was already pulling in $1 million a month.

Then Scale called. At the time, Scale AI was nearly the only household name in AI training data, having grown to a $14 billion valuation by orchestrating hundreds of thousands of people worldwide labeling data for autonomous vehicles, e-commerce algorithms, and coding tasks. But when OpenAI and Anthropic started pushing their chatbots toward actual programming ability, Scale needed software engineers to produce the training data—the kind of work that requires real expertise, not just crowdsourced button-clicking.

Foody saw something bigger unfolding. When the engineers he recruited complained about missed payments and chaotic platform management at Scale, he pivoted. By September, Mercor announced $500 million in annualized revenue. Foody's most recent fundraising round valued the company at $10 billion. At 22 years old, he and his two cofounders are now the youngest self-made billionaires.

Mercor's success isn't an outlier. It's the clearest signal yet of a wholesale reshuffling in how frontier labs approach AI development. While everyone obsesses over data center buildouts and chip shortages, an analogous race is happening for something equally critical: training data that actually works.

Advertisement

Labs have exhausted the easy stuff. They've already fed their models centuries' worth of publicly available text. When that didn't produce the superintelligence investors were promised, the labs pivoted to something different: teaching models specific skills through reinforcement learning, a technique where models get rewarded for producing outputs that humans prefer. But unlike traditional crowdsourcing where you pay someone $3 to label images of dogs, this requires hiring lawyers, consultants, physicists, and surgeons to define what "good" means in their respective domains.

Surge AI, founded by data scientist Edwin Chen, figured this out first. After watching vendors cut corners with cheap labor and poor quality at past jobs at Google, Twitter, and Facebook, Chen built something different: smaller, higher-paid teams of actual experts. The company has been profitable since launch and pulled in more than $1 billion in revenue last year, surpassing Scale's reported $870 million. Reuters reported in July that Surge is now raising at a $15 billion valuation.

But something shifted in June when Meta hired Scale's CEO and took a 49 percent stake. Suddenly, rival labs panicked. Why would they trust a provider now partially owned by a competitor? The answer: they wouldn't. Demand for alternative data providers exploded overnight. Handshake, which started as a LinkedIn-for-college-students platform, had built a network of 20 million alumni, grad students, and PhDs. It launched an AI data arm in early 2025 and watched demand triple in the weeks after the Meta announcement. By November, Handshake had surpassed a $150 million run rate—exceeding the original decade-old business entirely.

Advertisement

The economics are brutal for the data companies themselves, though. Labs want "grading rubrics"—detailed specifications for what counts as correct output in every conceivable context. These aren't simple checklists. Joelle Pineau, chief AI officer at Cohere, explains the core problem: "There seems to be a belief that there's a single reward function, that if we can just specify what we want, we can train [models] to do it. But the reality is more varied." When success depends on context, goals, and audience—as it does for legal briefs or consulting analyses—defining that reward function requires humans spending hours refining rubrics with dozens of criteria each.

The market has exploded into a Cambrian explosion of competitors. Turing, Labelbox, Invisible Technologies, Snorkel AI, Micro1, even Uber—which started letting drivers annotate between rides—are now positioning themselves as data infrastructure for AI labs. Everyone's touting increasingly prestigious talent: Surge boasts Fields Medal mathematicians and Supreme Court litigators; Mercor advertises Goldman analysts; Handshake draws from 1,000+ universities.

But this concentration of demand creates fragility. When Appen, the Australian data annotation giant, dominated the market in 2020 with a $4.3 billion valuation, 80 percent of its revenue came from just five clients: Microsoft, Apple, Meta, Google, and Amazon. Today it's worth less than $130 million. The data industry is littered with former giants undone by training technique shifts or single customer departures.

The training data boom reveals something uncomfortable about AI's actual trajectory. If frontier labs truly believed they were months from artificial general intelligence, they wouldn't be spending billions on task-specific training data for accounting, law, and contact centers. Instead, what's unfolding looks more like the industrialization of AI—a future where companies need custom-trained models for their particular workflows, bought repeatedly as their needs shift. That's bad for the AGI-is-imminent narrative but great for the data startups. In the race to build superintelligence, these companies discovered the most reliable way to make money: selling the shovels, not digging the holes.

More Topics:
Handshakedata annotation

Advertisement

Advertisement

Trending Now

1

GoPro CEO Vows Cameras Stay Core After Starman Deal

2

Judge Splits Ruling in X vs. Twitter Rival Fight

3

Tim Cook Steps Down, Ternus Takes Apple's Helm

4

Google's Lyria 3.5 Brings AI Music to Gemini

5

Google Translate Gets Listening Mode, Live Background Mode

People Also Ask

AI training data consists of specialized expert-annotated examples used to teach models specific skills through reinforcement learning. Frontier labs like OpenAI and Anthropic exhausted public datasets and now pay billions for specialized data from lawyers, engineers, and doctors to define what "good" output means in each domain.

Mercor ($500M annualized revenue, $10B valuation), Surge AI ($1B+ revenue, $15B valuation), and Scale AI ($870M reported revenue) lead the market. Handshake, a newer entrant, hit a $150M run rate by November 2025, tripling demand after Meta's investment in Scale AI.

Approximately $10 billion flows annually into training data providers. This explosive demand increased dramatically after labs exhausted public datasets and shifted to specialized reinforcement learning, requiring custom data from subject matter experts across industries like law, medicine, and finance.

Brendan Foody and his two cofounders, all in their early 20s, built Mercor into a $10B valuation startup by capitalizing on frontier AI labs' desperate need for specialized coding and expert data. The company grew from $1M monthly revenue to $500M annualized in just over a year.

Meta took a 49% stake in Scale AI in June 2025, prompting rival labs like OpenAI and Anthropic to panic about data provider conflicts. Competitors like Handshake and Surge AI saw demand triple overnight as labs sought alternative training data providers independent from major tech companies.

These companies hire specialized experts—lawyers, engineers, doctors, mathematicians—to annotate training data and develop detailed grading rubrics defining successful outputs. Labs pay millions for this expertise-based data labeling. Unlike crowdsourced cheap labor models, the high-pay expert approach proved superior for complex AI tasks.

More in AI

Google's Lyria 3.5 Brings AI Music to Gemini

Google's Lyria 3.5 Brings AI Music to Gemini

Rogue OpenAI Agents Hijacked a German Wiki

Rogue OpenAI Agents Hijacked a German Wiki

Altman Apologizes for Messy GPT-6 Astra Rollout

Altman Apologizes for Messy GPT-6 Astra Rollout

Microsoft's Project Zenith Targets AI Developers

Microsoft's Project Zenith Targets AI Developers

Nvidia's $99B Bet: AI's Biggest Backer Emerges

Nvidia's $99B Bet: AI's Biggest Backer Emerges

Samsung's AI Rally Reshapes Dating, TV and Majors

Samsung's AI Rally Reshapes Dating, TV and Majors

More Articles

Accel Nears $1B Deal for Thinking Machines at $40B

Accel Nears $1B Deal for Thinking Machines at $40B

Sep 3

Utilities Race to Fusion Startups as AI Strains Grid

Utilities Race to Fusion Startups as AI Strains Grid

Sep 3

Meta Offers 95% AI Discount for Your Data

Meta Offers 95% AI Discount for Your Data

Sep 3

Abliteration.AI Sells Access to Uncensored Models

Abliteration.AI Sells Access to Uncensored Models

Sep 3

OpenAI Launches GPT-6 Astra, Claims 'AGI Era'

OpenAI Launches GPT-6 Astra, Claims 'AGI Era'

Sep 3

OpenAI's GPT-6 Astra Debuts, Claims 'AGI Era'

OpenAI's GPT-6 Astra Debuts, Claims 'AGI Era'

Sep 3