How Work Gets Done
A small team in San Francisco and a distributed network of finance professionals produce benchmarks that frontier model makers cite in their own papers. The output looks like a research lab's. The operating rhythm feels like something else — a culture shaped by technical focus and lean structure, where hiring priorities reveal what the company values most in its early stage.
Halluminate, founded in 2024, calls itself a data research lab building benchmarks and reinforcement-learning environments for AI in knowledge work. Its first major release, the Westworld Finance Diligence Bench, comprises 88 problems that evaluate agents on a full company-acquisition due-diligence process. Each problem draws from anonymized private-transaction data, runs inside dynamic desktop environments, and is scored by a mixed pipeline of agentic, discrete, and binary verifiers written and reviewed by practicing deal professionals. A second benchmark, DealTrace, forces models to trace every claim in an investment recommendation back to the step that produced it. Both exist because the core team decided the field needed harder, more transparent tests than existing leaderboards offered. A third, WebBench, built with Skyvern, contains 2,454 public tasks across 452 live websites and separates information retrieval from actions that create, update, or delete data — a design choice that distinguishes model errors from infrastructure failures like CAPTCHAs, blocked proxies, and anti-bot defenses.
The salaried core (nine roles posted on Zero G Talent spanning platform engineering, research/post-training, data platform, forward-deployed engineering, people operations, and finance research) clusters in a San Francisco office that opens for marathon sessions around conference deadlines. When NeurIPS workshop submissions came due August 29, the office ran 10 a.m. to midnight. The same team presented DealTrace as an oral at EMNLP 2026 in Budapest. Conference cycles set the tempo: long stretches of environment construction and data curation punctuated by submission crunches where the entire technical staff converges on a single deadline.
Parallel to that core runs the expert network. Halluminate recruits domain experts with investment-banking, private-equity, and consulting backgrounds to author the financial deliverables (Excel models, PowerPoint decks, written analyses) that become ground truth for benchmark tasks. The work is fully remote and asynchronous. Experts commit a minimum of 20 hours per week for at least six consecutive weeks. Training pays a $1,500–$4,000 milestone upon completion; ongoing work runs $100–$250 per hour, with top contributors reaching $225–$250 per hour. Several have earned over $200,000 while still in school. The expert pool is not a labeling crowd; it is a curated cohort whose professional judgment defines what "correct" looks like on tasks that take hundreds of steps.
Benchmark development moves in a loop the team has made visible in its public write-ups. Task design starts with practicing professionals writing and reviewing problems. Environments are built to simulate the desktop tools those professionals use. Models run under different scaffolding levels (from no guidance to structured critique passes) and every run produces a trajectory, not just a score. The results get published with model-level breakdowns: which model explains 64 percent of score variance, which harness explains 6 percent, and where the best configuration still tops out at 51 percent mean score across 88 tasks. That transparency is deliberate. The lab treats its own output as a dataset for the field.
Decision-making traces back to that loop. Hiring prioritizes technical depth in platform engineering and post-training research, roles that build the infrastructure for running thousands of agent trajectories, while a People Platform Lead role signals early investment in the human side of a hybrid workforce. The finance researcher role sits at the intersection: translating domain expertise into benchmark specification. There is no product team, no sales function, no marketing hire. The lab publishes; the field adopts. Whether that rhythm holds as the team scales is the open question the next hiring cycle will answer.
The Values Driving the Work
Halluminate's operating principles read like a corrective to the hype cycle. The company presents itself not as an agent builder but as an evaluation infrastructure lab — a distinction that shapes every technical and hiring decision. "Evaluating agents means evaluating tasks, judges, and infrastructure," Marshall told the AI Engineer World's Fair audience. "Choose the task category before choosing an agent." That framing (tasks first, models second) recurs across their public work. The benchmarks publish mean scores and cost-per-run for every model-harness pair across multiple configurations, a level of methodological transparency rare in frontier AI.
The same rigor applies to browser-agent evaluation. Marshall's presentation distinguished model errors from infrastructure failures and argued that "the observe-decide-act loop makes latency a product constraint." BrowserBench exists to measure execution reliability across providers. The company also develops managed software sandboxes where agents practice realistic tasks "without the unpredictability and unintended consequences of operating directly on production websites." Agents, Marshall noted, "already produce surprising real-world side effects."
These technical commitments map to a stated philosophy: "The intelligence explosion is created by great teams solving hard problems at the frontier." The Series A blog post acknowledges that verifiable domains (math, code) are being mastered rapidly, but argues that "these same systems still struggle to build realistic financial models, deliver client-ready PowerPoint presentations, or grasp the nuance of well-balanced strategic analyses." Halluminate's bet is that generally intelligent AI coworkers for knowledge work — finance, consulting, insurance, operations, require evaluation infrastructure that mirrors the complexity of the work itself. "We have the recipe to build them," the post concludes.
The team composition reflects that recipe. The company describes an "interdisciplinary team of former founders, researchers, particle physicists, experts, and engineers from Meta, Scale AI, Capital One Labs, McKinsey, Goldman Sachs and more." Co-founder Wyatt Marshall studied computer science and philosophy at Cornell's Milstein Program before software and data engineering roles at early-stage New York startups; at MediaWallah he maintained production data infrastructure, built Python tooling for ETL workflows and AWS automation, and published the company's Python API package. Co-founder Jerry Wu's background is less public but the pairing (CS/philosophy plus deep technical execution) signals a value on conceptual clarity alongside shipping capability.
Profitability before scale appears deliberate. The Series A announcement notes the company grew from zero to a mid-eight-figure revenue run rate while remaining strongly profitable before raising $30 million, Fortune reported, led by Oak HC/FT (total capital: $38.5 million, Fortune's data shows). That trajectory — revenue-first, profitable, then capitalized for expansion, suggests a culture that treats funding as accelerant. The client list reinforces it: four of the top five closed-source U.S. AI labs are paying customers.
Hiring signals reinforce the same values. The board lists six open roles as of August 2026: Member of Technical Staff, the platform engineering, research/post-training, and data platform roles from the core team, plus Forward Deployed Engineer, the People Platform Lead, and the Finance Researcher. Salary bands cluster at $200k–$275k, Zero G Talent's figures put the top at $275k, for technical roles, $180k–$225k for the People Platform Lead and Finance Researcher. The existence of a People Platform Lead at roughly nine-person scale (nine salaried roles on the board) is unusual; it indicates an early, intentional investment in people operations as infrastructure rather than afterthought. The Finance Researcher role (a domain expert embedded in a technical team) mirrors the Westworld benchmark's reliance on practicing finance deal professionals as authors and reviewers. Forward Deployed Engineer suggests a customer-proximate engineering model, consistent with building evaluation environments for specific vertical workflows.
YouTube discussions from 2024–2025 elaborate the evaluation philosophy: supervision agents that evaluate other agents, reflection as an evaluation primitive, decomposing complex evals into simpler concrete steps, judge-versus-jury LLM ensembles, fine-tuning judge LLMs, and the necessity of human involvement in both building initial eval harnesses and improving them over time. "All of these techniques add up to SOTA evals in aggregate," the 2025 discussion concludes. The emphasis on building confidence in AI system fundamentals so that complexity can grow without causing a collapse reads as both a technical principle and an organizational one.
What the Hiring Bar Selects For
Halluminate's six open roles, all based in San Francisco and all posted after the company's Series A closed in October 2026, read like a map of the problems the team actually faces. Four carry the "Member of Technical Staff" title (Platform Engineering, Research/Post-Training, Forward Deployed Engineer, and Finance Researcher) with Platform Engineering, Research/Post-Training, and Forward Deployed Engineer in a $200,000–$275,000 band, according to Zero G Talent, signaling equal status.
| Role | Salary Band |
|---|---|
| Member of Technical Staff — Platform Engineering | $200k–$275k |
| Member of Technical Staff — Research/Post-Training | $200k–$275k |
| Member of Technical Staff — Forward Deployed Engineer | $200k–$275k |
| Member of Technical Staff — Finance Researcher | $180k–$225k |
| Data Platform Lead | $200k–$275k |
| People Platform Lead | $180k–$225k |
The Research/Post-Training role is the clearest signal. Halluminate's pivot from evaluation benchmarks to building RL environments for frontier labs — documented in a Dealroom note dated October 2026, means the team now lives in the post-training loop: designing reward functions, curating verification data, and iterating environment dynamics until a model's behavior matches what a domain expert would accept. The job description doesn't ask for "AI safety" branding; it asks for engineers who have shipped RLHF or RLAIF pipelines at scale and can debug why a reward model is over-optimizing a proxy. The company's own writing makes the standard explicit: "Verification quality beats volume. A weak or misaligned verifier teaches models the wrong thing, so Halluminate invests in subject-matter-expert review and alignment of task and verifier with what a real expert would reward." Candidates who have only fine-tuned on static datasets without building the evaluation harness that sits beside them will not clear this bar.
Platform Engineering and Data Platform Lead roles point to the same infrastructure burden. Serving that client roster with a sub-10-person team means every data pipeline, every environment container, every evaluation run must be reproducible and low-friction. The Westworld Finance Diligence Bench is not a static dataset; it is a living environment suite that must version cleanly, reset deterministically, and expose hooks for lab researchers to swap reward models. The hiring signal here is for engineers who have built data platforms that researchers actually trust, not just ingest pipelines that dump parquet files into a warehouse.
The Finance Researcher role reveals the domain-knowledge requirement that separates Halluminate from pure infra plays. The company chose finance as its entry vertical because it is "the world's largest knowledge-work service industry by headcount, relatively verifiable, and rich in sub-domains like investment banking and private equity." The researcher they want is not a quant who builds alpha models; it is someone who has lived inside an investment-banking diligence process, knows what a senior analyst checks before signing off on a working-capital adjustment, and can translate that procedural knowledge into verifier logic that catches an agent jumping straight to an answer without showing its work. The company's own note on "reward hacking comes in soft forms" — agents taking shortcuts a real analyst would never be trusted for, makes clear that the verifier must check process, not just output. That translation layer is the hiring filter.
Forward Deployed Engineer is the customer-facing counterpart. With fewer than 10 people generating mid-eight-figures revenue run rate, every lab integration is high-touch. The role selects for engineers who can sit across from a frontier-lab researcher, diagnose why the environment is not yielding the desired capability gain, and ship a fix that same week. It is a consulting skill set wrapped in an engineering title, and the compensation band matches the Platform Engineering role, signaling equal status.
The People Platform Lead, the only non-technical title in the set, is the outlier that proves the rule. A sub-10-person company hiring a dedicated People lead at $180,000–$225,000 is unusual. It reflects a deliberate bet that the cultural and operational infrastructure must scale before the headcount does. The hire will design the review cycles, the onboarding that teaches new engineers the verification philosophy, and the hiring process itself, which means the bar for every subsequent role will be set by someone who understands the technical standards above.
Taken together, the hiring mix selects for three intersecting traits: deep RL/post-training engineering fluency, domain expertise that can be codified into verifiers, and the autonomy to operate in a team where every person carries a critical path. The company does not hire for potential; it hires for verified capability in the exact loop it runs. The full-time roles all sit in San Francisco, five days a week, with equity grants of 0.10–0.20% for individual contributors. That compensation tier, combined with a team of roughly nine people, signals an environment where every hire carries disproportionate weight. Engineers who have shipped production ML infrastructure at scale — not just notebook prototypes, and who can own a vertical from data pipeline through to evaluation harness will thrive. The same goes for researchers who treat benchmark design as an engineering discipline: the Westworld Finance Diligence Bench and the DealTrace benchmark accepted at EMNLP 2026 are not academic exercises; they are the product.
The Expert Network contractor track selects for a different but overlapping profile. Pay runs $100–$250 an hour, but the first milestone ($1,500–$4,000) only pays out after self-paced training (15–20 hours) plus two accepted problem deliveries. The FAQ language on weekly commitment is inconsistent: one answer requires 20 hours for six consecutive weeks; another calls it a preference and invites applicants "regardless of availability." In practice, contractors who treat the role as a structured part-time engagement, can meet the 20-hour floor reliably, and possess verifiable deal-room or consulting experience will convert the hourly rate into meaningful income. Those who need guaranteed hours, W-2 benefits, or visa sponsorship will not find them here. The company explicitly states it cannot sponsor visas for contractors and operates on a 1099 basis only, even as the onboarding flow routes paperwork through Rippling and uses the word "payroll."
The on-site culture reinforces the self-selection. A LinkedIn post inviting the community to the office for a NeurIPS workshop deadline — doors open 10 a.m. to midnight, "food and good vibes provided", reads less like a perk and more like a filter. People who draw energy from that intensity, who want to debug a harness alongside the research team at 11 p.m., and who view the office as the primary collaboration tool will fit. People who need protected weekends, remote flexibility, or a clear separation between work and social space will chafe.
The Employee Record
Public employee-review platforms show effectively nothing. As of October 2026, Glassdoor returns no reviews for Halluminate. Trustpilot shows no page for halluminate.ai. The Better Business Bureau and usacomplaints.com each return no results. CourtListener's RECAP docket search finds zero filings. The California Attorney General's breach list shows no reported incidents. The research notes that employer-review sites and social-media forums (Reddit was inaccessible at query time) were not fully checked; but the major consumer-facing platforms that typically capture early-stage startup sentiment are blank.
That absence is itself a signal. Halluminate lists 2–10 employees on LinkedIn; Y Combinator's directory shows a team of five as of its Summer 2025 batch, while Fortune's October 2026 piece describes a nine-person San Francisco startup. At that scale, formal employee reviews are rare — most people who join either stay or leave before a review culture forms. The company's own LinkedIn page is the richest public window into how the team operates, and it reads like a research lab that happens to be a startup.
The page announces hires by name and specialty. An August 2026 post introduces Alina Hyk and Victoria Knapp Perez as the first members of the "Halluminate Department of Research," framing the addition as a research milestone rather than a headcount metric. A September 2026 post invites the external community into the SF office for a NeurIPS workshop submission day, "10am until midnight," with food, "good vibes," and direct feedback from "both our research and engineering teams." The same week, the company publishes benchmark results (Westworld Finance Diligence Bench, DealTrace) with model weights, datasets, and Weights & Biases reports all public. The tone is academic: "Go check it out and let us know what you think."
Those posts suggest a culture that values open research, conference participation, and community building: traits that align with the hiring mix (four technical research/engineering roles, one finance researcher, one people-platform lead on the current Ashby board). They also imply a pace tied to conference deadlines and public releases, not sprint cycles.
The only structured "worker voice" data comes from the Expert Network FAQ, which governs 1099 contractors, not W-2 employees. That FAQ acknowledges friction points: training is unpaid until a milestone is hit ("finishing the required modules plus delivering two successful problems"), takes roughly two weeks (typically 15–20 hours of work), and pays a one-time $1,500–$4,000 milestone. The FAQ gives conflicting weekly-hour expectations: reiterating the 20-hour/6-week requirement and the "regardless of availability" preference, while a third notes the schedule is "ad hoc given the contractor structure." It also mixes classification language, referring to "employment paperwork in Rippling" and "payroll compliance" for a role it simultaneously calls a 1099 independent-contractor position with no visa sponsorship. The FAQ directs access and account issues to [email protected] with a "resolved within 24 hours" claim.
No named current or former full-time employee is quoted in any public source reviewed. The founders — Jerry Wu (CEO) and Wyatt Marshall (CTO), are the only individuals who speak on the record, and they speak as principals, not employees. Wu told Fortune that the client roster includes those same five labs and that the firm had replicated that profitable revenue trajectory in the last 10 months; Fortune attributes those figures to Wu and does not present them as audited.
In short: the employee-experience record is thin because the company is tiny, young, and research-oriented. The public artifacts that exist — conference talks, open benchmarks, a late-night NeurIPS write-in, describe a team that works like a lab, publishes like a lab, and hires like a lab. Whether that translates to sustainable workload, career growth, or management quality for the next 20 hires is an open question the data cannot answer.
Who Thrives, Who Burns Out
The profile that succeeds at Halluminate is technically deep, domain-literate, comfortable with ambiguity, and willing to bet on a nine-person team that just closed a $30 million Series A. The profile that burns out is the one that mistakes the contractor role for a stable part-time job, or the full-time role for a managed career ladder. The small team in San Francisco that produces benchmarks cited by frontier labs is still the core; the question now is whether the operating rhythm that got them here survives the hiring the capital enables.
Working in AI? Zero G Talent tracks the openings: see every open Halluminate role, browse AI jobs, the companies hiring, and the people building the field.