The Hire That Unlocks Everything
Ooak Data has signed seven-figure contracts, Ooak Data's Y Combinator profile reported, with the two largest AI labs in the world while operating as a five-person, bootstrapped team, according to Ooak Data's Y Combinator profile, in Paris. That fact, confirmed on the company's Y Combinator profile, is the reason a recent job posting for a single Senior Operations Manager reads like a founding-engineer brief. The role demands someone who will "build the machine" — not maintain it, and the mandate spans data sourcing, pipeline automation, and the task-creation engine that will feed reinforcement-learning environments for those labs. By the time most early-stage startups are still figuring out payroll, Ooak Data is already hiring for the infrastructure that will let it ingest 300-plus full corporate data ecosystems, Ooak Data's Y Combinator jobs page's data shows, in six months.
The hiring push centers on that one senior operations position, but the scope reveals the company's actual headcount trajectory. Ooak Data lists seven open roles across operations, sales, marketing, engineering, and recruiting on its Y Combinator profile — breadth that signals a transition from founding team to functional organization. Co-founder Thomas Aubry put the current team at three co-founders plus one ML engineer, four people total. The Y Combinator directory shows five. Either way, adding a senior operations lead at this stage is the clearest signal that the bottlenecks have shifted from model research to data logistics.
The job description leaves no ambiguity about what "operations" means here. The hire will "drive complex data projects, scale our sourcing pipeline, automate repetitive workflows, and help build the task-creation pipeline powering real-world AI agent training." Those bullet points map directly to the Alexandria library build-out: acquiring company data, anonymizing it, structuring it into multi-step interaction trajectories, and turning those trajectories into evaluatable RL environments. The posting explicitly rejects "keep things running" maintenance work. The founders want an owner who treats the data pipeline as a product — versioned, monitored, and continuously improved.
This hire is the keystone. Everything else in Ooak Data's near-term roadmap — the Alexandria library expansion, the lab partnerships, and the early-2027 research lab launch, depends on turning raw corporate workflows into clean, licensed, evaluatable data at a rate the current team cannot sustain. The posting isn't a hiring signal. It's a capacity announcement.
Inside the Pipeline: Three Layers, One Bottleneck
The pipeline that feeds Alexandria is not a single ETL job. It is a multi-stage system that ingests raw enterprise data (documents, chat logs, project-management records, org charts) and outputs reinforcement-learning environments where an agent can take actions, observe outcomes, and adjust strategy across multiple tools. Ooak Data's research page describes the flow in three layers: an automated multimodal anonymization pipeline that transforms names, dates, and proprietary content while preserving structure and relationships; a multi-modal data enrichment and curation pipeline that adds expert-level tasks calibrated on frontier models; and the RL environment runtime itself, which serves those tasks to agents for training and evaluation.
The anonymization stage is where the hybrid design lives. Each "full data ecosystem" (the company's term for a complete organizational digital twin) arrives as heterogeneous data: Slack exports, Jira tickets, Notion pages, email threads. The pipeline must strip identifiers without flattening the relational graph that makes the workflow realistic. Ooak's job postings note that the anonymization pipeline combines algorithmic approaches with human verification. That hybrid design means every new ecosystem requires both compute scaling and reviewer coordination. A senior operations manager hired today will own the throughput target: that cadence.
Enrichment and curation add a second scaling dimension. Expert-level tasks are not generated by template; they are authored by domain specialists who map real workflows into multi-step, multi-tool trajectories that expose model weaknesses. The job description for data engineers lists "multi-modal data enrichment & curation pipeline" as a core contribution area, alongside the anonymization pipeline and the datasets feeding RL environments. That phrasing signals a single integrated platform, not three disconnected scripts. Versioning, lineage, and observability must span all three stages.
DataOps practices are explicit requirements. The same posting calls for "DataOps know-how: versioning, monitoring, data quality and testing" and asks candidates to "define our technical standards and take an active part in the structuring architecture decisions (tech choices, design patterns, scalability, DataOps)." In practice, that means the operations hire will set standards for pipeline latency, define data-quality gates, and build the observability that lets the research team debug a training run that stalled because a tool schema drifted.
The scale target — that cadence is not aspirational. It is the delivery cadence implied by the seven-figure agreements with them, which the company has publicly confirmed. Each lab receives a curated slice of Alexandria tailored to its model-evaluation roadmap. Miss a weekly delivery, and the lab's benchmark cycle slips. The senior operations role is therefore a production-engineering role disguised as operations: the hire designs the capacity plan, recruits the annotation and verification workforce, negotiates vendor contracts for GPU and storage burst capacity, and owns the runbook when a schema migration breaks the enrichment pipeline at 2 a.m.
Founder Thomas Aubry, the CTO, has framed the infrastructure gap bluntly: agents need fundamentally different data than chatbots — multi-step interaction trajectories, tool-use demonstrations, and multi-modal workflow data. The pipeline that produces that data is the product. Scaling it is the only way Alexandria becomes the "world's largest library of real-world business workflow datasets" the company claims it is building, and the only way the early-2027 data research lab launches with a live corpus instead of a slide deck.
Who's Buying? The Open Secret
Ooak Data's Y Combinator profile states it plainly: "Our traction speaks for itself: we've signed 7 figures deals with them, alongside recent partnerships with several smaller labs, all while being 100% bootstrapped." Industry context makes the identification straightforward. OpenAI and Anthropic have dominated frontier model development since 2023. Both have moved aggressively to secure specialized data pipelines: OpenAI through its Data Partnerships program launched in 2023, Anthropic via a five-year Claude 3 integration with Databricks valued at $100 million, and both through reported talks with biotech firms for domain-specific corpora. When a five-person, bootstrapped Paris startup announces seven-figure deals with them, the counterparties are effectively an open secret.
The contracts create immediate, concrete demand for Ooak Data's core product: turning raw corporate workflows into reinforcement-learning environments. The company's site describes the pipeline three ways: "We turn real company data into RL environments where AI agents learn to work in the real world," "We source real enterprise data, anonymize it into digital twins, and generate reinforcement learning environments with expert-level tasks," and "Frontier AI Labs: RL environments grounded in real enterprise data that push your models beyond synthetic benchmarks." Each phrasing points to the same operational burden. A single "full data ecosystem" — the unit Ooak Data aims to acquire and anonymize 300-plus times in the next six months, requires ingestion of multi-tool logs, access-control mapping, PII stripping, task-sequence reconstruction, and reward-function design before it becomes an RL environment an agent can train in. Seven-figure contracts imply delivery schedules. Delivery schedules require throughput. Throughput requires the operations layer the company is now hiring to build.
The smaller-lab partnerships compound the load. "Several smaller labs" means additional evaluation suites, custom task distributions, and likely faster iteration cycles than the frontier labs run. Ooak Data's own messaging frames this as a feature: "Synthetic benchmarks test what models can do in theory. Our data tests what they do in practice. Most evaluation frameworks test single-turn Q&A. We build multi-step, multi-tool environments that test what matters: can your agent actually complete a workflow?" That pitch — multi-step, multi-tool, real-world completion, is exactly what the largest labs need to close the gap between benchmark scores and production reliability. It is also exactly what demands a scalable, repeatable data-operations engine rather than a series of bespoke consulting engagements.
Bootstrapped revenue changes the calculus. Without venture capital to subsidize headcount ahead of revenue, each operations hire must pay for itself in processed ecosystems per month. The 300-ecosystem target over six months implies 50 per month, roughly two per business day, end to end, from acquisition through anonymized, task-annotated RL environment. The senior operations managers the company is recruiting will own that cadence. Their mandate is not "keep things running." It is to turn a founder-led, artisanal process into a production line that can honor seven-figure SLAs while the founder team focuses on the 2027 research-lab launch. The contracts are the forcing function. The hiring is the response. The pipeline is the product.
Why Agents Still Fail at Real Work
The performance gap between current AI agents and human workers remains stark. The best GPT-4-based agent achieved only 14.41% end-to-end task success compared to 78.24% for humans, a fivefold differential on tasks most knowledge workers perform routinely. Even the strongest frontier models fail approximately 40% of tasks in realistic workplace environments. GPT-5.2, the top performer in a December 2025 evaluation across 150 workplace tasks, reached 61% success. Humans hit 72.36%. The primary failure modes clustered around GUI grounding and operational knowledge, capabilities that demand visual, multi-modal training data that barely exists in current datasets.
| Metric | Best Agent | Human |
|---|---|---|
| GPT-4 agent success (WebArena) | 14.4% | 78.2% |
| GPT-5.2 success (Dec 2025) | 61% | 72.4% |
| Frontier model failure rate (realistic envs) | ~40% | — |
| Production agents needing human help within 10 steps | 68% | — |
| Customer-service task success (SOTA) | <50% | — |
| Consistency on repeated tasks (8×) | ~25% | — |
This is not a model architecture problem. Anthropic's 2024 guide to building effective agents identified the practical bottleneck as "the quality of tool definitions, context engineering, and evaluation data." The data and evaluation layer, not the model layer, is the binding constraint. Wing Venture Capital's 2025 analysis frames it in market terms: RL environments are playing for AI agents the same role EDA played for silicon design. Anthropic alone is estimated to spend tens of millions annually on RL environments, with three-to-fivefold growth expected into 2026.
They require a different kind of data. A chatbot takes a prompt and returns a response; its training data is input-output pairs. An agent takes a goal and executes a sequence of actions across multiple tools and systems. Its training data must capture trajectories: chains of reasoning, tool invocations, observations, corrections, and completions unfolding over many steps. This data is vanishingly rare in the wild. Synthetic trajectories inherit the limitations of the generating model; they do not contain patterns the model has never seen. The messy, ambiguous, context-dependent workflows of real companies are precisely what no model has been trained on.
Chen et al. (2023) showed that fine-tuning Llama2-7B with just 500 agent trajectories generated by GPT-4 produced a 77% performance increase on HotpotQA. Zeng et al. (2023) demonstrated that the quality and diversity of agent trajectory data is the bottleneck, not model architecture. Chen et al. (2024) found that current agent training corpora entangle format-following with agent reasoning, causing distribution shift, and naive fine-tuning introduces hallucinations as a side effect.
The hierarchy of agentic capabilities, derived from systematic failure analysis across frontier models, reveals where expanded real-world data moves the needle. Level 1 (tool use) and Level 2 (planning): weaker models fail here. Level 3 (adaptability) and Level 4 (groundedness): stronger models stumble. Level 5 (common-sense reasoning): the current frontier challenge, requiring world knowledge, contextual inference, and ambiguity resolution. Real company workflows, anonymized into digital twins while preserving structural fidelity, supply the multi-modal coherence that synthetic benchmarks lack: documents with real formatting inconsistencies, conversations with organizational context, project management tools with actual task dependencies.
Production reality confirms the gap. Pan et al. (2025) surveyed 306 practitioners and conducted 20 case studies across 26 domains. Their findings: 68% of production agents execute at most 10 steps before requiring human intervention. 70% rely on prompting off-the-shelf models instead of weight tuning. 74% depend primarily on human evaluation. Even state-of-the-art agents succeed on fewer than 50% of customer service tasks, and consistency drops to approximately 25% when the same task is repeated eight times. For enterprise deployment, a system that works half the time is not half as useful as one that works every time; it is essentially unusable where errors have consequences.
The shift from static datasets to RL environments is as fundamental as the shift from chatbots to agents. Agents learn through interaction: taking actions, observing outcomes, adjusting strategies. Ooak Data's Alexandria library, scaled to 300-plus full data ecosystems from real companies, converts raw corporate workflows into environments where agents train against the complexity they will actually face: multi-step, multi-tool, calibrated to challenge the latest frontier models. The same infrastructure supports evaluation, enabling controlled comparisons across domains within a single coherent world model. Tasks designed by domain experts ensure realistic edge cases. The Model Context Protocol provides structured tool access. Telemetry captures full trajectories for failure analysis.
This is the infrastructure that connects model capability to real-world performance. The model layer has had its revolution. The data layer is next.
Revenue First, Accelerator Second
Ooak Data reached the Summer 2026 Y Combinator batch without taking a dollar of outside capital. The company was founded in 2024 by Thomas Aubry, Grégoire Lamy, and Pierre-Louis Vouteau in Paris, and for its first two years the three founders plus one machine-learning engineer built the Alexandria pipeline on revenue from early contracts. A recent LinkedIn post from the team described the operation as "3 co-founders + 1 ML engineer, already working with leading AI labs, bootstrapped."
That self-funded phase matters because it forced the company to prove the data pipeline works before hiring at scale. The founders sold seven-figure contracts to them while the headcount sat at five. Those deals — signed before the YC batch began, created immediate demand for scaled data processing that the founding team could not meet alone. The hiring surge for seven such roles is the direct response to that demand.
Y Combinator accepted Ooak Data into the S26 cohort with Diana Hu as the primary partner. The accelerator slot does not change the company's capital structure — it remains bootstrapped, but it adds a signal that the pipeline and the lab contracts are credible enough to warrant the program's backing. The batch timing also aligns with the six-month window the company has set to acquire and anonymize 300-plus full data ecosystems from real-world companies. That window closes just as the S26 demo day approaches, giving the new hires a concrete deadline.
Aubry's background shapes the technical bar. Before Ooak Data he led the data and AI team at PayLead, worked as an applied ML scientist at Samsung AI, and co-founded Macro Vision on autonomous vehicles. The trio's decision to stay in Paris rather than relocate to San Francisco reflects a deliberate bet: the enterprise workflows they need to capture — Slack, Gmail, Notion, Jira, SharePoint, Teams, are dense in European mid-market companies that are easier to access from a Paris base.
The bootstrapped history also explains why the senior operations manager role is framed as "build the machine" rather than "keep things running." The first operations hire will inherit a pipeline that has already moved that data into RL environments for frontier labs. The job is to replicate that process at volume, not to design it from scratch. That distinction only exists because the founders proved the loop works with a team of four.
YC's network will help recruit the remaining six roles, but the company's leverage still comes from the contracts in hand. The seven-figure deals mean the next hires are funded by customer revenue, not dilution. That financial discipline is rare in the current AI data layer, where most competitors raised seed rounds before landing a single lab agreement. Ooak Data's path — revenue first, accelerator second, hiring third, is the constraint that makes the 2027 research lab target plausible rather than aspirational.
The January 2027 Inflection Point
The early-2027 target for a dedicated data research lab is not a horizon goal; it is a milestone with a workback plan that starts this summer. Ooak Data's Y Combinator S26 posting states the sequence plainly: over the next six months the company will achieve this, cement its lead in business-workflow data, and position itself to launch the lab. The lab, in turn, will "enable new services that help companies around the world automate their workflows with AI."
That six-month sprint is the critical path. Each acquired ecosystem — documents, communications, project-management tools, org charts, feeds the Alexandria library, which the company describes as that library. The anonymization pipeline strips such identifiers while preserving structure, relationships, and complexity. The output is not static corpora but such environments calibrated on the latest frontier models. Those environments are what they are already paying seven-figure contracts to access.
Senior operations managers are the lever that makes the sprint feasible. The role posting emphasizes that this is not a "keep things running" ops job; the hire will "scale the engine behind our data and RL projects." In practice that means turning a five-person, Paris-based team into a machine that can ingest, anonymize, and ship hundreds of digital twins on a predictable cadence. The founding GTM hire in the U.S. carries a target of ten signed datasets per month, a pace that implies the operations layer must handle parallel procurement, legal review, and engineering handoff without bottleneck.
The lab's research agenda will extend the work already published on Ooak's site: evaluation methodology, dataset design, and the gap between benchmark performance and real-world capability. The company's own research notes that 80 percent of AI projects fail in production because evaluation conditions do not match deployment conditions. Such benchmarks test theoretical capabilities; Alexandria tests practical performance. The lab will formalize that distinction into a service layer: evaluation infrastructure that lets enterprises test agents against realistic company environments before going to production.
Bootstrapped discipline shapes the timeline. With no outside capital beyond YC, the team has funded the first 300 ecosystems and the two anchor lab contracts from revenue. That constraint forces a hiring plan that is senior, small, and fast: exactly the profile the current operations search targets. The summer "elite squad" buildout is meant to lock in the data lead before competitors can replicate the procurement trust model that gets CEOs to share their most sensitive asset.
If the six-month acquisition rate holds, Alexandria will cross the 300-ecosystem threshold around January 2027. That is the inflection point: a library large enough to support continuous RL environment generation, diverse enough to stress-test any agent architecture, and proven enough to underwrite a research lab that sells evaluation as a service. The job posting that went live recently was the first brick. The machine is now being built.
Working in frontier tech? Zero G Talent tracks the openings: see every open World Labs role, browse frontier tech jobs, the companies hiring, and the people building the field.