Ooak Data, a five-person Paris startup in Y Combinator's S26 batch, has built a platform that pays real companies for access to their operational guts — Slack channels, Jira boards, CRM records, the buried spreadsheet nobody dares refactor — then strips the identifiers and ships the result to frontier AI labs as reinforcement-learning sandboxes.
The product is called Alexandria, and the bet underneath it is simple: the next generation of capable agents won't be trained on bigger models but on better records of how work actually gets done. Co-founders Pierre-Louis Vouteau, Grégoire Lamy, and Thomas Aubry framed the thesis in direct terms in an August 2026 LinkedIn post: "We've recorded almost everything, our history, our science, our music, but the way we actually work was never recorded by anyone. The most valuable knowledge humanity produces is the only kind we've never collected. Until now."
That pitch lands in a market with a known problem. High-scoring agents in 2026 still beat benchmarks and lose in production because benchmarks miss the hard parts: the rule overridden in last Thursday's Slack DM, the legacy spreadsheet, the contract clause only three people know exists. "You can't synthesize it either," reads the YC listing. "Real org charts, permissions, and cross-tool dependencies don't come out of a generator." Ooak Data's wager is that the only way to teach an agent to handle that mess is to give it the mess, scrubbed of anything that could identify the company it came from.
The mechanics start with a counterintuitive sales motion. Ooak Data pays the companies it sources from — about $2 million in cumulative partner payments as of August 2026, according to Y Combinator's launch post — and frames those checks as runway extensions. "Money CEOs have used to extend their runway, keep the lights on and the team paid, or launch their next venture… all from an asset they didn't even know was sitting there."
The trajectory since founding has been steep by YC standards. Ooak Data joined STATION F's Founders Program in November 2025, entered YC's S26 batch in mid-2026, and by August 2026 reported working with "multiple frontier labs" and revenue "well above 7 figures," with ingestion-to-delivery time cut threefold across data from 20 source companies. The team is hiring — a first product manager, a lead data engineer, senior software engineers — and plans to open a fundraising round in November 2026.
The founders' earlier careers matter here. Vouteau was previously COO at the UK ed-tech firm Visely and worked at Gopuff and L'Oréal; Aubry was Head of Data at PayLead, an Applied ML Scientist at Samsung AI, and earlier co-founded Macro Vision on autonomous-vehicle engineering; Lamy's background is described as CPO/COO. The company's decision to pay data partners rather than scrape them reads as a deliberate departure from the ad-tech playbook that has shaped much of the AI data supply. Whether Alexandria can scale to the 300-plus enterprise ecosystems it targets over the next six months, and whether anonymization can preserve the messy edge cases that make workflows worth training on, are the questions that will decide if this bet pays off.
Product Mechanics: From Slack/Jira/CRM Data to Anonymized RL Environments
Ooak Data's pipeline runs in three stages: ingest from the enterprise tools a company already uses, anonymize at the schema and entity level, and package the result as reinforcement-learning environments with calibrated tasks. Each stage preserves the messiness that breaks agents in production, including multi-tool dependencies, permission boundaries, and buried rules in old threads, while stripping the identifiers that would make the raw data legally and commercially unusable.
Ingestion starts with read-only API connections to the standard SaaS stack: Gmail, Slack, Notion, Jira, Drive, SharePoint, and Microsoft Teams. The YC launch page says these connectors pull documents, communications, and tool activity "with full organizational context preserved," meaning Ooak captures not just individual messages and tickets but the relationships among them: who owns what, which channel rolls up to which project, which document feeds which decision. Source companies tend to have at least 20 full-time employees and span the operating lifecycle (active, winding down, or recently closed), a deliberately wide funnel meant to capture the long tail of how work actually unfolds.
The anonymization layer is where Ooak stakes its claim. Names, dates, companies, and proprietary content are transformed by an "automated multimodal anonymization pipeline," but structure, relationships, and complexity are preserved verbatim. The result is a digital twin that is "structurally identical, but privacy-safe." In plain terms: a thread that contained a buried pricing rule becomes a thread between anonymized personas, with the same subject line, the same attachment, and the same position in the inbox, but with no real names, no real customer, and no real date.
Packaging is the third move and the most consequential for buyers. Each anonymized company becomes a reinforcement-learning environment, a sandbox where an agent must complete multi-step, multi-tool workflows calibrated against the latest frontier models. Ooak describes the task design as "expert-level" and explicitly designed to "expose weaknesses, not confirm strengths." That is a deliberate inversion of benchmark culture, where static test sets reward memorization. A workflow environment instead requires an agent to navigate permission boundaries, retrieve context scattered across tools, and chain the right actions in the right order to clear a task.
The mechanics are not theoretical: the company is running, billing, and shipping today. The product's target, per the company's public materials, is to acquire and anonymize more than 300 full data ecosystems over the next six months.
Early Adoption: Labs Are Buying, Receipts Are Thin
Ooak Data's traction with frontier labs is, as of September 2026, the clearest commercial proof point for the anonymized-workflow-twins thesis — but the company has not disclosed which specific labs are under contract or what those labs are paying. The startup's own materials describe the customer profile in the abstract ("frontier AI labs for training") and confirm the licensing model, yet the buyer roster is not public. That gap is worth naming up front: the adoption story is real, but the receipts are thin.
What is documented is the demand environment those contracts land in. The biggest labs are signaling large commitments to reinforcement-learning environments and broader R&D compute in 2026. In a market where a vendor that delivers structurally identical, privacy-safe workflow twins at scale starts to look like procurement infrastructure, not a niche dataset play, Ooak Data's value proposition is that it pays businesses for an anonymized copy of how work actually happens inside their company, then licenses those digital twins to labs for agent training.
Ooak Data's economic case compounds with the compliance case. As EU AI Act data-governance obligations took effect in August 2026, a twin that is structurally identical to the source but free of personal data removes a layer of provenance work the lab would otherwise own. A lab tapping into even a fraction of Ooak Data's targeted pipeline, which is 300-plus full data ecosystems over six months, is effectively buying a year of enterprise process capture without running a single data-collection RFP.
The early-adoption signal also reads through the labor market. Scale AI (the closest incumbent in the RL-environment layer) added 10 roles in the past week on Zero G Talent, including a VP, Research slot listing $453,600–$567,000 and a Director of Engineering, Physical AI role at $302,400–$378,000. That hiring pattern is consistent with a vendor scaling capacity to meet the same frontier-lab demand Ooak Data is chasing. Labelbox, by contrast, has paused hiring (0 roles added in the past seven days) but lists a Forward Deployed Engineer, RL Environments position at $140,000–$200,000.
The honest read: Ooak Data has staked out a clear product wedge and a credible go-to-market, and the macro math supports the thesis that those contracts exist and are growing. What isn't yet public is the named-customer list, the contract values, or any lab-reported cost savings. Until the company or its customers disclose those, the adoption story is best framed as directional rather than quantified.
Competitive Response: Scale, Snorkel, and Labelbox Are Repositioning
The three names drawing the most attention as Ooak Data courts frontier AI labs are the incumbents it has to displace, not the startups it has to beat. Scale AI, Labelbox, and Snorkel AI each sit at a different layer of the same workflow (environment generation, programmatic labeling, and human-in-the-loop curation), and each has moved within the past several months to defend ground that Ooak Data's anonymous workflow twins are designed to absorb.
Scale AI's shift is the most direct. The San Francisco data infrastructure company has rebuilt a chunk of its revenue mix around RL environments that run on desktop VMs, MCP servers, and other workflow surfaces, per its product page. That rebalancing puts Scale squarely in the path of Ooak Data's pitch to labs that want pre-built, privacy-safe environments instead of paying for fresh human demonstrations. Scale's open job board reflects the bet: 10 roles posted in the past week, with the company's salary band topping out around $331,000 (median $249,000) across 156 salaried postings.
Labelbox is taking a different angle: buying and partnering its way into the agent-training stack. The company acquired Upcraft, an agentic sales automation startup, and used the deal to scale the human expertise powering frontier AI, per a June 2026 press release. Separately, Labelbox expanded its Google Cloud partnership so its labeling and curation tools run natively inside Vertex AI. The moves suggest Labelbox is positioning itself as the curation layer above raw environments; Ooak Data's twin environments would, in theory, slot into that layer.
Snorkel AI is the outlier. The Stanford-spinout programmatic-labeling shop, which encodes subject-matter expertise as noisy "labeling functions" that get statistically combined, laid off 13% of its workforce in 2026, per reporting. Programmatic labeling trades up-front labeling effort for scaled training labels rather than trading real data for anonymized twins, which puts it on a different value proposition than synthetic workflows. But Snorkel's pressure to find new growth vectors illustrates how broadly Ooak Data's pitch is being read across the market.
Mercor and Surge AI round out the set. Mercor, which works with the top five AI labs, acquired Sepal AI in February 2026 specifically to deepen its RL-environment capabilities at the intersection of human data and research. Surge AI runs CoreCraft, a large-scale enterprise environment suite, and counts OpenAI, Anthropic, Meta, and Google among its partners. None of these players are building exactly what Ooak Data builds (privacy-safe digital twins of entire enterprise tool stacks) but together they form a counter-movement that labs can route through instead.
The countervailing force is that most of these rivals are reacting to demand Ooak Data helped surface, not the other way around. Frontier labs have signaled they want more environments, not fewer, and they want provenance. That leaves room for an anonymous-workflow-twin vendor to grow alongside, rather than against, the Scale-and-Labelbox axis, as long as the rivals keep launching their own synthetic-data tools in parallel.
Regulatory Push: GDPR and the EU AI Act Reshape What Counts as Defensible Data
The compliance math for frontier AI labs changed on August 2, 2026. That's the day Article 50 of the EU AI Act (the world's first comprehensive AI framework, in force since August 2024) became fully applicable. Any provider of generative AI placed on the EU market must now mark AI-generated or manipulated audio, image, video, and text in a machine-readable format, with detectable signals that survive metadata stripping. A limited transition runs until December 2, 2026 for systems placed on the market before August 2. The fine schedule is severe: up to €15 million, or 3% of worldwide annual turnover, whichever is higher.
The law's reach is wider than the EU's borders. Article 50 binds all 27 member states plus the European Economic Area (Norway, Iceland, Liechtenstein) and Switzerland through bilateral agreements. Any enterprise worldwide that builds, sells, or deploys generative AI tools for European users falls inside it. Agentic AI systems land in scope where their actions generate outputs intended to be directly perceived by users.
The provenance demand that follows is reshaping what counts as defensible training data. Frontier labs now need to demonstrate, on demand, where their data came from, what was redacted, and whether the underlying consent or licensing chain holds. Anthropic responded in August 2026 by announcing that supported Claude models would embed watermarks in generated text and attach digitally signed C2PA Content Credentials to files wherever Claude was offered worldwide, a move one provenance executive called "an inflection point." That infrastructure layer is itself a market: Grand View Research's figures put the global digital provenance market at roughly $4.2 billion in 2026, growing to $16.9 billion by 2033, a 21.9% compound annual growth rate.
GDPR sits beneath Article 50 and pulls in the same direction. The original 2018 framework already required lawful bases for processing personal data, data minimization, and demonstrable safeguards on cross-border transfers. Layered onto that, Article 50 adds output-side obligations that can't be met if training data itself was contaminated by personal information that wasn't adequately anonymized. The harder the provenance audit gets on outputs, the more valuable upstream data that was anonymized at collection becomes.
That's the wedge Ooak Data drives into. A workflow twin built from a single consenting enterprise, stripped of identifiers, comes with a paper trail Article 50's traceability requirements demand. By contrast, a scraper-built corpus assembled from public web pages inherits every unresolved consent claim of every contributing site, claims a regulator in Frankfurt or Dublin is now empowered to chase.
In the United States, the regulatory direction is messier but moving in the same direction on the marketing and disclosure side. The FTC has finalized orders against firms including Cox Media Group and a marketing group tied to an "Active Listening" AI service, requiring payment of nearly $1 million to settle deceptive-practice allegations. With the establishment of a dedicated AI enforcement unit in early 2026, the maximum penalty for disclosure violations rose to $53,088 per individual violation, meaning a single automated campaign with unverified AI content across thousands of social posts can rack up millions. New York enacted a law, effective June 9, 2026, requiring conspicuous disclosure when AI-generated "synthetic performers" appear in commercial advertisements.
The practical consequence for Ooak Data's customers is straightforward: anonymized real-world workflow data is one of the few categories of training input where provenance can be evidenced at the dataset level rather than reconstructed post-hoc. Labs that license it sidestep the worst-case audit — having to defend an opaque supply chain in front of an EU regulator with turnover-percentage teeth. Labs that don't will spend 2026 and 2027 doing it anyway, which is why procurement teams at frontier labs now ask synthetic-data vendors two questions they didn't ask two years ago: where did this come from, and can you prove it?
Market Outlook: Synthetic Data Growth and Where the Money Is Going
The category Ooak Data is selling into is still small in absolute dollars but compounding fast, and the analyst forecasts disagree loudly about how big it gets and how fast. The spread among recent estimates tells the story: nobody has settled on a taxonomy, and the forecasters are measuring different baskets.
| Forecast source | Scope | Base year figure | Endpoint | CAGR |
|---|---|---|---|---|
| MarketsandMarkets | Global synthetic data (all uses) | $0.3B (2023) | $2.1B by 2028 | 45.7% |
| Technavio | Synthetic data for AI training only | $182M (2025) | — (through 2030) | 37.3% |
| KBV Research | Global synthetic data (narrower basket) | — | $880.2M by 2028 | 34.1% |
The bigger addressable market above it gives the opportunity room to scale. MarketsandMarkets reported the AI agents market at $7.84 billion in 2025, rising to $52.62 billion by 2030 at a 46.3% CAGR, with vertical AI agents (the closest analog to Ooak Data's workflow-specific environments) growing fastest at a projected 62.7% CAGR through 2030. BCG's separate infrastructure lens pegs data-center capital deployment at $1.8 trillion from 2024 through 2030 to keep up with AI compute demand, and global data-center electricity consumption is on track to roughly double from 415 TWh in 2024 to 945 TWh by 2030. Synthetic data sits one layer below the model and one layer above the raw compute, and that middle layer is where the spend is crystallizing.
Capital is already rotating into that layer. The pattern flagged across the market is the relevant one: investors are targeting the often-overlooked infrastructure of AI (compute, data pipelines, reliability engineering) rather than model wrappers. That bucket includes Ooak Data's category directly.
The tension worth flagging: the synthetic-data TAM forecasts above describe a market smaller than Ooak Data's own implied six-month goal of 300 anonymized enterprise ecosystems suggests the company believes it can capture, and the company has not disclosed pricing. If even a modest share of the projected synthetic-data pool by 2028 flows through workflow-level RL environments, the category supports multiple specialized vendors, not just one. Whether Ooak Data's anonymized-enterprise-data wedge ends up as the dominant form factor, or as one ingredient inside broader platforms from Scale AI, Snorkel AI, and Labelbox, is the bet the next funding round will price.
Why Defense, Space, and Internal Hiring Are Out of Scope
This article is deliberately narrow. The story of Ooak Data's anonymized enterprise-workflow environments sits inside a much larger agent-deployment landscape, and three adjacent territories warrant an explicit "not here."
The first is defense. The research base in this piece — Y Combinator's company page, Ooak Data's product materials, and the competitive-move coverage around Scale AI, Snorkel AI, and Labelbox — describes workflow environments sourced from commercial SaaS tools (Slack, Jira, Gmail, Notion, Drive, SharePoint) and from CRM, sales, financial, and product-analytics datasets inside customer businesses. None of those sources document a defense application, a Department of Defense contract, or an autonomous-systems testbed for Ooak Data's twins.
The second excluded territory is space. Mission operations at NASA, ESA, and the commercial launch providers generate a distinct category of training data (telemetry streams, ground-system logs, fault trees, crew-procedure documents) that overlap almost not at all with the Jira tickets, CRM records, and Slack conversations Ooak Data ingests.
The third exclusion is Ooak Data's internal recruiting beyond its founding team. Thomas Aubry, Pierre-Louis Vouteau, and Grégoire Lamy are named on the YC page; the company is listed with five employees and is hiring across engineering and operations. Those hiring details belong to a separate HR or company-profile piece, not to a story about the synthetic-RL-environment market, regulatory pressure on data provenance, and competitive dynamics among Scale AI, Snorkel AI, and Labelbox.
Working in frontier tech? Zero G Talent tracks the openings: see every open Scale AI role, browse frontier tech jobs, openings at Labelbox, and the people building the field.