Halluminate's Edge: RL Environments for Finance
A five-person startup builds the sandboxes, datasets, and evaluation layers that let model developers train and test computer-use agents on actual financial workflows. Halluminate entered Y Combinator's Summer 2025 batch with a thesis: the hard part of computer-use AI isn't the model — it's the environment.
Founded in 2024 by Wyatt Marshall and Jerry Wu, friends since their first week studying computer science at Cornell who had already spent seven years building together, Halluminate targets investment banking, private equity, and consulting workflows that are economically valuable, highly structured, and stubbornly resistant to automation.
Today's browser and computer-use agents face three compounding bottlenecks, per the company's Y Combinator launch post. Real-world testing is unsafe because agent actions trigger real trades, real emails, and real compliance flags. It's slow because you can't parallelize a live Salesforce instance. It's noisy, with captchas, auth walls, and layout shifts polluting every signal. The second bottleneck is data: high-quality benchmarks for due-diligence workflows don't exist on the open web. The third is evaluation: in most domains, only humans can reliably judge whether an agent's output is correct. As Brendan Foody of Mercor said in a Sequoia Capital interview, "You need humans almost definitionally to measure what is beyond the frontier of the model capabilities."
Halluminate's product stack attacks all three. Its sandbox layer spins up fully managed, parallelizable replicas of the tools knowledge workers actually use — Salesforce, Slack, ticketing systems, generic web environments — so agents can be trained and tested at scale without side effects. Its dataset layer produces proprietary benchmarks built from anonymized real transactions. Its evaluation layer routes agent outputs to expert annotators who score them against weighted rubrics. Paying customers already include the two largest browser-agent companies and several leading foundation-model labs.
Halluminate's Westworld Finance Diligence Bench, released in August 2026, contains 88 problems spanning the full arc of a PE deal, written and reviewed by practicing deal professionals. When seven frontier models were tested on the full suite, the best configuration scored 51 percent, Halluminate's LinkedIn post found. The gap between simulated and real performance is stark: a model trained only on simulation solved six-tenths of one percent of real tasks, Halluminate's LinkedIn data shows; a mixed-data model reached 23.4 percent, Halluminate's LinkedIn figures put it at.
That gap drives Halluminate's expert network. The company recruits professionals from Wharton, Harvard Business School, Stanford GSB, and alumni of Goldman Sachs, KKR, Bridgewater, General Atlantic, Centerview, McKinsey, and BCG to build the problems, rubrics, and benchmark deliverables that teach agents what a correct LBO model looks like. The work is fully remote, asynchronous, and requires a minimum of 20 hours per week for six weeks. Training takes roughly two weeks (15–20 hours) and pays a $3,000 milestone upon completion, Halluminate's website shows. Some contributors have earned over $200,000 while still in school, Halluminate's experts page reports.
Eight Posted Roles, Two Engines
Halluminate's job boards show the company building two distinct engines in parallel: the RL platform itself (environments, evals, sandboxes, post-training pipelines) and the finance domain layer that makes those environments economically meaningful. Postings across Y Combinator and StartupHub list eight roles split across that divide.
| Role | Location | Base Salary | Equity | Experience | Source |
|---|---|---|---|---|---|
| Strategic Project Lead | San Francisco | $150K – $225K | 0.15% – 0.30% | 5+ years | Y Combinator |
| Member of Technical Staff — Platform Engineering | San Francisco | $200K – $275K | 0.15% – 0.30% | 8+ years | Y Combinator |
| Member of Technical Staff — Research / Post-Training | San Francisco | $200K – $275K | 0.15% – 0.30% | 8+ years | Y Combinator |
| Chief of Staff | San Francisco | $160K – $210K | 0.15% – 0.30% | 8+ years | Y Combinator |
| Member of Technical Staff — Finance Researcher | San Francisco | $180K – $225K | 0.15% – 0.30% | 8+ years | Y Combinator |
| Head of Finance | San Francisco | $200K – $250K | — | — | StartupHub (Aug 2026) |
| People Platform Lead | San Francisco | $180K – $225K | — | — | StartupHub (Aug 2026) |
| Founding Member of Technical Staff — Finance Researcher | San Francisco | $180K – $225K | — | — | StartupHub (Aug 2026) |
The two finance-researcher listings — "Member of Technical Staff" on the YC board and "Founding Member of Technical Staff" on StartupHub — carry identical base ranges and almost certainly describe the same slot, titled differently for different audiences. That duplication aside, the seven distinct roles form three clusters.
Platform & Research Engineering
The Platform Engineering MTS role owns the sandbox infrastructure: fully managed, reproducible environments where browser- and computer-use agents can be trained and evaluated at scale. The Research / Post-Training MTS role sits downstream, designing the RL loops that turn those environments into better policies, including reward modeling, curriculum design, and offline-to-online transfer. Both require eight-plus years of engineering depth; the research role additionally expects publication-grade familiarity with post-training techniques for foundation models. Halluminate's own benchmark work, the Westworld bench's 88 tasks drawn from anonymized PE deals, gives a concrete sense of the evaluation surface these hires will extend.
Finance Domain Layer
The Finance Researcher (whether "Member" or "Founding Member") is the bridge role: a practitioner who has built or reviewed deal models, memoranda, and diligence workstreams in investment banking, private equity, or top-tier consulting, and who can translate that workflow into task specifications, reward functions, and expert annotations. The Head of Finance sits adjacent but distinct as a traditional finance operator who will own the company's own financial planning, fundraising support, and cap-table hygiene as Halluminate scales past its five-person core. The Strategic Project Lead, listed at five-plus years rather than eight, looks like the glue role: someone who can scope a customer engagement with a foundation-model lab, define the evaluation protocol, and coordinate the platform and research teams to deliver it.
Operations & People
Chief of Staff and People Platform Lead round out the seven. The Chief of Staff role, at eight-plus years, signals that founders Marshall and Wu want a partner who has already operated at the intersection of technical product and go-to-market in a high-velocity startup. The People Platform Lead, a title that leans platform rather than pure HR, suggests Halluminate intends to build internal tooling for talent density measurement, interview calibration, and perhaps the same kind of expert-network management it sells to model labs.
Across every technical role on the YC board, the equity band is identical: 0.15% – 0.30%. The base bands are also tight: $50K spreads for engineering, $50K for finance research, $50K for chief of staff.
The Documented Screen: Expert Contractor Onboarding
Halluminate's publicly documented screening and onboarding process applies to its expert contractor network, not to the full-time technical roles above. The company hires contractors on a rolling basis and reviews every application within three business days, per its public FAQ.
The first gate is a screener, a short video and quiz that explains Halluminate and the work in detail. It takes under 30 minutes, requires no preparation, and is mandatory before onboarding begins. Screener links expire after a set period; expired links are the single most common issue applicants hit.
Once the screener is complete, a background check runs through Checkr. Domestic checks typically take a few business days; international checks can take longer. No one starts delivering problems until the check clears, though candidates can work through onboarding paperwork in the meantime. Slack introductions and problem access wait for clearance.
The next step is signing the Rippling offer letter. Training access unlocks only after that signature. Training itself is milestone-based: contributors earn a $3,000 payment upon successful completion, defined as finishing the required modules and delivering two successful problems. Training time is not logged as work hours. Onboarding calls are recorded and shared by email and in Slack; attending live is optional.
After training, approved contributors become Core Contributors and move to hourly pay. Base rates are tied to background category: consulting starts at $100 per hour, investment banking and private equity at $150 per hour. Rates increase based on output metrics: problems delivered (roughly six within a set timeframe for consultants), total hours committed, responsiveness, and communication quality. Most strong contributors see a rate increase within their first month; senior writers who deliver consistently over time reach $225–$250 per hour.
The entire pipeline is built for a 1099 independent contractor model. Halluminate does not offer W-2 employment, H-1B transfers, or visa sponsorship. Candidates must be US-based with a US bank account for payroll compliance. A minimum commitment of the required weekly hours is required; the company says this threshold is the minimum needed to get up to speed and add value, so lower commitments (5–10 hours per week) are not accommodated. Hours can scale up to 60 per week depending on availability, with no fixed daily schedule.
Notably absent from this contractor process: coding challenges, RL algorithm whiteboarding, or model architecture discussions. The core requirement is finance domain expertise — IB, PE, or consulting experience — not engineering or coding skill. The work consists of creating complex financial deliverables (Excel models, PowerPoint presentations, financial analyses) and evaluating AI-generated outputs against those benchmarks. Candidates with adjacent backgrounds (FP&A, capital markets, asset management, investment management, corporate strategy) are considered, but most projects demand the deeper transactional experience.
Two Tracks, Different Profiles
Halluminate operates on two distinct tracks. The full-time technical roles, including Member of Technical Staff in Platform Engineering, Research/Post-Training, and Finance Researcher, plus Strategic Project Lead and Chief of Staff, recruit from frontier-AI pools: PhD researchers, systems engineers, and RL specialists. The company has not publicly documented the interview loop for these roles.
The volume side of the business is the expert-contractor network. That network hires domain practitioners who produce the gold-standard financial deliverables, such as Excel models, investment memos, due-diligence workbooks, and slide decks, that Halluminate uses to evaluate frontier models. The expert page lists target backgrounds explicitly: alumni of those firms, and MBA programs at Wharton, Harvard Business School, and Stanford GSB. Private-equity and investment-banking veterans command $150–$250 per hour; consultants, accountants, and FP&A professionals start at $100–$200. Rates climb to $225–$250 for contributors who consistently deliver high-quality evaluations.
In August 2026, Halluminate announced its "Department of Research" led by Alina Hyk and Victoria Knapp Perez. Perez authored the Westworld Finance Diligence Bench study, which consists of 88 problems drawn from anonymized real private-equity transactions and vetted by experts, and published the write-up, Weights & Biases report, dataset, and model weights publicly. The benchmark reveals what "passing" looks like for models, with Opus 5 leading on absolute score and GPT 5.6 Sol on token efficiency, and by extension what the human annotators grading those models must understand: the entire deal lifecycle, from data-room triage to IC memo.
The contractor training gate, which takes two weeks and 15–20 hours of work and pays a $3,000 milestone only after delivering two successful problems, filters for persistence and domain fluency. Candidates who cannot produce a defensible LBO model or diligence checklist in the sandbox don't collect the milestone and don't advance. The company's own LinkedIn posts show the output standard: "We asked seven frontier AI models to run a complete company acquisition due-diligence process. The best configuration scored 51%." The human experts grading those runs are the benchmark.
Preparing for the Expert Contractor Bar
Halluminate's contractor screen filters for fluency in the actual deliverables of investment banking, private equity, or consulting, and the ability to translate that fluency into structured tasks an RL agent can learn from. The company's public materials and contributor outcomes make the bar explicit.
Master the deliverable types Halluminate benchmarks. The platform's core work product is ground-truth financial artifacts: three-statement models, LBO waterfalls, comparable company analyses, precedent transaction decks, due-diligence memos, and IC investment memos. If you cannot build these from scratch in Excel and PowerPoint without templates, you will not pass the training period. Review the benchmark summary the team published on LinkedIn (August 2026); it describes the exact task taxonomy you will be asked to reproduce and evaluate.
Demonstrate recent, hands-on deal experience. Halluminate's expert network page lists compensation bands by background: Private Equity and Investment Banking at $150–250/hour, Consulting at $100–200/hour. The pricing signal is deliberate: the highest rates go to practitioners who have run live processes recently. If you left banking for corporate development three years ago, refresh by building a current LBO model for a public target using live filings. Contributors who reach $225–250/hour do so by consistently producing rubrics that the research team accepts without revision.
Prepare for the two-week paid training as a working audition. The onboarding is evaluation. You will be asked to draft realistic IB/PE/consulting tasks with ground-truth deliverables, then evaluate AI-generated outputs against your own rubrics. The day-to-day workflow mixes problem creation, output evaluation, and async review cycles through Slack and the Horizon platform. Before applying, complete at least one end-to-end cycle on your own: pick a public M&A announcement, build the full diligence workstream (quality of earnings, working capital, synergy model), write a rubric that scores each deliverable on a 0–1 scale, then have a peer grade your work against it. That loop — create, rubric, evaluate — is the job.
Signal availability that matches the commitment threshold. Halluminate requires that same commitment once you start. The rolling application lets you mark future availability windows, but the screen favors candidates who can commit to a near-term block. Contributors quoted on the site mention fitting the work around an MBA schedule; the asynchronous, self-scheduled structure supports that, but only if the weekly hour floor is real.
Show you understand the RL environment layer, not just the finance. Halluminate's product is sandbox environments for computer-use agents — simulated Excel, PowerPoint, browser, and data-room interfaces where models execute tasks. The Department of Research (launched August 2026 with the duo) published evidence that mixed simulated-and-real training outperforms real-only by 53% on transfer tasks (23.4% vs 15.3% success on real sites, according to Halluminate's LinkedIn post). You do not need to code, but you should be able to explain why a simulated data room with deterministic reset behavior produces better training signal than a live one. Read the Y Combinator company description: the bottleneck is "high-quality datasets for benchmarking and evaluations" and "realistic sandbox environments for safe and accurate testing/training." Your rubrics feed the first; your task design informs the second.
Leverage the referral channel if you have a connection. The platform attributes referrals automatically; several current contributors joined this way. The founders (Wyatt and Jerry, Cornell CS, 7+ years working together) also publish direct contact emails ([email protected], [email protected]) for introductions to researchers and RL/post-training practitioners. A cold email that references a specific benchmark result, for example, "GPT 5.6 Sol's harness-invariant scoring on your Westworld bench suggests your rubric design is model-agnostic," signals deeper engagement than a generic inquiry.
Treat other AI training platforms as complementary, not competitive. The FAQ states working on other platforms "is generally not an issue, but flag anything you think could conflict." If you have contributed to similar data-labeling services, disclose it proactively. Halluminate's differentiation is finance-specific RL environments; generic RLHF experience is a plus, but domain specificity is the filter.
The screen is narrow by design. The company's paying customers, those companies and leading foundation-model labs, need datasets that only practitioners who have lived the workflow can produce. When the next frontier model attempts a leveraged buyout in Halluminate's sandbox, the covenants will hold because the humans who built the test have signed them in real life.
Working in frontier tech? Zero G Talent tracks the openings: see every open ASML role, browse frontier tech jobs, openings at Stripe, and the people building the field.