Snorkel AI Closes $350M Series E to Scale Training Data Operations
The Stanford-founded startup, which pivoted to providing ready-made training datasets, has reached a $3.5 billion valuation with backing from Insight and S32.

Snorkel AI Inc. announced a $350 million Series E funding round, valuing the training data provider at $3.5 billion. Insight and S32 led the investment, alongside more than six additional backers including Alphabet Inc.'s GV startup fund.
Founded in 2019 by Stanford AI Lab researchers, Snorkel AI initially launched Snorkel Flow, a software platform designed to streamline supervised learning projects by automating the creation of labeled datasets. Supervised learning trains neural networks using labeled datasets—files containing prompts paired with human-generated correct answers.
Labeling datasets manually is extremely labor-intensive. Snorkel Flow addressed this bottleneck by applying statistical methods developed by the company's founders at Stanford, solving accuracy problems that had plagued earlier automation solutions.
Last year, Snorkel AI shifted its business model away from selling software tools toward delivering pre-built training datasets. The company also expanded beyond supervised learning into reinforcement learning, a more complex AI training methodology.
Reinforcement Learning and Evaluation Infrastructure
Reinforcement learning datasets differ fundamentally from supervised learning datasets. Rather than containing prompts with correct answers, reinforcement learning datasets present unanswered questions. The AI model must generate responses independently, after which human reviewers or automated systems assess accuracy and provide feedback to refine the model's reasoning.
Snorkel AI maintains a network of tens of thousands of human experts who create reinforcement learning training tasks. Beyond datasets, the company supplies additional technical resources required for AI training operations.
When an AI model completes a reinforcement learning task, human reviewers evaluate the output using predefined evaluation criteria—sometimes spanning multiple pages. For programming tasks, evaluation guidance must encompass cybersecurity and performance requirements that AI-generated code must satisfy.
Snorkel AI develops customized AI evaluation rubrics for clients and refines them continuously based on reviewer feedback. Discrepancies between two human reviewers scoring the same response often reveal inconsistencies in the evaluation criteria itself.
AI models frequently train within specialized virtual environments. A code generation model, for instance, might need a simulated developer workstation. Snorkel AI provides training sandboxes alongside its datasets and evaluation rubrics.
Growth and Future Plans
Since launching our new data-as-a-service offering nearly a year ago, we've grown over 18 times, and this week crossed an annualized revenue run rate of $375 million
Alex Ratner, co-founder and Chief Executive Officer
The company plans to deploy the new capital toward expanding its engineering team. Snorkel AI will also invest in AI safety initiatives and support the creation of open-source model evaluation benchmarks.


