The 1 Job at Centaur.ai Filters for Values, Not Just Code
One Role, By Design
Centaur.ai, a Y Combinator W19 company with 45 employees, has exactly one full-time position open as of October 2026. Not a hiring freeze. Not a stealth round. One role: Technical Product Manager, Labeling Tools, remote in the U.S., listed at $145,000 to $160,000 with a three-year experience floor, according to the Centaur.ai careers page. That is the entire aperture.
Centaur operates in a sector where headcount usually balloons. Data-labeling platforms, model-evaluation startups, and human-feedback marketplaces have spent three years raising rounds and posting dozens of requisitions at a time. Centaur took a different path. Its core product, Diagnosus, is a mobile app that turns medical annotation into a competitive game. Doctors, nurses, and students examine real X-rays on their phones; the same image routes to multiple professionals, and the system only accepts labels where consensus emerges. The AI trains on that overlap. The company calls it "collective intelligence, where humans and machines work together to outperform either one alone." That model — experts competing, AI learning from agreement — is the entire business. It does not require an army of labelers. It requires a platform that makes the competition work.
The single opening sits on top of that platform. The Technical Product Manager owns the labeling tools that power Diagnosus and the internal workflows that route cases, measure agreement, and surface the data the models consume. It is a product role, but the domain is the human-AI boundary: how experts interact with the interface, how disagreement gets resolved, how the signal feeds back into training. The job description asks for three-plus years of product experience. Centaur's careers page states: "We're looking for motivated, mission-driven, data nerds who share these values and want to bring them to life with us." The values are not decorative. They are the hiring rubric. In effect, Centaur's single opening creates a highly selective process where candidates must demonstrate alignment with the company's human-AI collective intelligence mission, not just technical competence.
Why only one role? Three factors coexist. First, the team is small (45 people after eight years), and a consultant network absorbs variable labor. Medical experts join as contractors, not employees, scaling the annotation supply without scaling headcount. The careers page invites board-certified physicians and recent graduates alike to "monetize your expertise" on flexible terms. Second, the product is mature enough that the labeling-tool surface area fits a single product manager; there is no second product line yet. Third, the company treats each addition as a cultural commitment. The W19 batch status and Boston base suggest a team that has survived multiple hype cycles without chasing growth for its own sake. The Guardian's January 2026 analysis of the AI bubble noted that "most of the companies would fail" and the survivors would be the ones with "open-source models that run on commodity hardware" and real utility. Centaur's consensus-driven medical data fits that description.
The consultant network listing on the same careers page reinforces the pattern. That is not a hiring pipeline; it is the supply side of the marketplace. The single full-time role is the demand side, the person who shapes the tools that make the marketplace efficient. Everything else about Centaur's process (the screens, the values filter, the human-AI competency test) flows from this constraint. One slot means the bar is not "qualified." The bar is "necessary."
The First Screen: What Kills Your Application Before a Human Sees It
At a frontier AI company running a single open requisition, the applicant pile doesn't shrink by accident. It shrinks by design, and the design starts before a recruiter opens a single file. Companies have been using AI to narrow candidate pools by crawling resumes for keywords, and the screening can eliminate you before you ever speak to a human. By 2026, AI screening software automates the most time-intensive parts of hiring, from resume parsing to initial assessments, letting recruiters spend reclaimed hours interviewing instead of skimming PDFs.
The mechanism is straightforward. An AI tool ingests every application, parses structure and semantics, then scores each candidate against a weighted rubric (required skills, domain experience, project complexity, publication record, open-source contributions). The rubric itself is the first filter: it encodes what the hiring team decided matters. At a company building human-AI collective intelligence, that rubric almost certainly weights evidence of human-AI collaboration higher than raw model-training throughput. A candidate who has shipped a product where a language model augments a human workflow, not just a model that benchmarks well, surfaces. A candidate whose resume lists "prompt engineering" as a buzzword without a shipped artifact sinks. The decision is made without a human in the loop, so a borderline score ends the process before anyone reads your resume a second time.
This isn't unique to Centaur. At Accenture, which ended fiscal 2025 with 779,000 employees across roughly 9,000 clients, scale forces automation. At Deloitte, with a workforce above 470,000 people across more than 150 countries, recruiters filter an enormous applicant pool before a human ever reads your file. Centaur operates at a different scale (one role, not thousands), but the logic inverts rather than disappears. Volume hiring uses automation to reduce a flood to a trickle. Selective hiring uses automation to verify that the trickle contains only candidates who meet a non-negotiable threshold. The three things that eliminate strong candidates at Deloitte are automated assessments, group cases, and shallow behavioral answers. At a one-role company, the automated assessment is the gate; the group case and behavioral depth come later, if you pass.
What the automated layer evaluates has expanded beyond keywords. The AI can evaluate communication skills, professionalism, and attitude, parsing video introductions, asynchronous responses, even code submissions for style and clarity. This effectively creates AI-powered work sample tests at scale, bringing more objectivity into evaluating soft skills and decision-making under pressure. But it also introduces new failure modes. It could penalize people with speech impediments or accents. It can flag candidates who use AI assistance during the screening itself, a growing cat-and-mouse game. Accenture permits AI for research and practice but bans it inside assessments and live interviews, stating that these uses violate its responsible AI principles and can result in disqualification; it uses proctoring tools and AI detection to keep the process fair. Evy, an assessment platform, detects when candidates use AI assistance during live assessments, protecting the reliability of technical and behavioral evaluation data. A company whose mission is human-AI collective intelligence faces a sharper version of this paradox: it wants candidates who use AI fluently, but it needs to verify the human contribution underneath.
The first screen also filters for signal that the candidate understands the assignment. Top talent stays available for only 10 to 20 days. Candidates who receive clarity on the process are more engaged, more prepared, and more likely to accept an offer. Candidates left guessing often withdraw quietly. At a one-role company, the cost of a false negative — filtering out the right person — is existential. The cost of a false positive — advancing the wrong person — is wasted interview cycles. The rubric therefore tilts conservative: require demonstrated human-AI workflow design, not just familiarity. Require public artifacts, such as a GitHub repo where an LLM routes tasks to a human reviewer, a blog post dissecting a failure mode, or a conference talk on evaluation methodology. The automated screen looks for those artifacts. If they're absent, the score drops. No human override.
Centaur's specific filter configuration (the exact keyword weights, the assessment platform, the go/no-go thresholds) is not publicly documented. The research covers the category, not the company. But the category logic is clear: at a frontier AI firm with one opening, the first screen is not a filter for competence. It is a filter for alignment. The resume that passes is the one that proves the candidate has already been working at the human-AI boundary and can show the receipts.
The Values Filter: Four Lines That Function as Code
Centaur's careers page states it plainly: "We've built a culture based on a set of core values for incredible people to do their best work. It repeats its earlier appeal." The four values — Act With Integrity, Create Change, Keep Learning, Every Voice Counts — operate as the primary filter for the single open role.
Start with Act With Integrity: "We do what's right, even when no one is watching and even when it's difficult. Trust is our foundation, and we earn it every day through our actions and openness." In a hiring context, this translates to a screen for candor over polish. Candidates who optimize answers for what they think interviewers want to hear, rather than describing what they actually did, including failures, fail this value. The integrity filter is structural: the simulation environment Centaur uses for evaluation rewards consistent decision-making under pressure, not rehearsed narratives.
Create Change: "We don't wait to be asked to make things better. We seek out problems worth solving and take ownership of driving solutions forward. This maps to a bias-for-action assessment. The company's consultant network pitch reinforces this: "Whether you have decades of experience or are recently board certified, AI teams need your leadership." Centaur hires people who have already demonstrated initiative in ambiguous environments. A resume that lists only assigned tasks, without evidence of self-directed problem identification, signals misalignment.
Keep Learning: "We're curious, humble, and hungry to grow. We stay adaptable by continuously learning from our wins, our losses, and each other. This is the hardest to fake. The simulation-based evaluation Centaur employs captures learning velocity in real time. Candidates who defend initial answers against contrary evidence score poorly. Those who adjust, cite new information, and explain the pivot score highly. This is not hypothetical; Andela, a remote-talent platform, already uses "real-world work simulations to predict on-the-job performance and to provide evidence of a technologist's problem-solving and decision-making abilities." Centaur applies the same logic to its own hiring.
Every Voice Counts: "Great ideas can come from anywhere. We value diverse perspectives and believe our best decisions come from listening to multiple viewpoints. Importantly, in the same way annotations are weighted, input is weighted for different decisions, revealing a specific operational nuance. The parenthetical about annotation weighting is a direct reference to Centaur's product: human-AI collective intelligence where human judgments are aggregated with learned weights. The hiring filter tests whether a candidate can operate in that system by contributing dissent when warranted, deferring when expertise lies elsewhere, and recognizing that not all input carries equal weight in every context. Glassdoor data shows 58% of Centaur Digital employees would recommend the company, with a 3.6 out of 5 rating for culture and values, suggesting the filter selects for people who actually experience the culture as described.
The counterpoint is necessary. The same analysis warns: "caution is needed – transparency and ethical guidelines must govern these AI evaluations to avoid simply baking in historical biases." Centaur's design originates from "a human-centered AI institute focused on interpretability and ethics," which the analysis calls "encouraging." But the values filter only works if the simulation itself reflects the values and if the weighting mechanism for "Every Voice Counts" doesn't silently replicate the very biases the company claims to mitigate. For the single candidate who clears the process, the test isn't whether they can recite the values. It's whether their decision trace in the simulation proves they already live them.
The Human-AI Competency: What Centaur Actually Tests For
Centaur.ai describes itself as a technology company specializing in high-quality data annotation and AI model evaluation, combining human expertise with advanced workflows to produce accurate, scalable training and evaluation data. That mission — human expertise channeled through structured workflows — maps directly onto the "centaur" paradigm the research literature defines: not full fusion (the cyborg model) and not abdication (the self-automator model), but directed co-creation. In the MIT Sloan study of consultants using generative AI, centaurs accounted for 14% of participants. They knew which questions they wanted answered and asked more-specific questions to get them. Unlike cyborgs, who engaged in conversational back-and-forth, centaurs maintained structured and controlled interactions with AI, harnessing it as a tool for targeted efficiency. Crucially, centaurs had the highest accuracy in their business recommendations, outperforming both cyborgs and self-automators.
Four competencies produce that outcome: deep domain knowledge in the relevant field; advanced capability in AI collaboration and prompt engineering; understanding of AI systems' strengths and limitations; and ability to orchestrate multiple AI agents effectively. A 2023 Stanford Institute for Human-Centered AI study found that AI-assisted professionals consistently outperformed both unaided humans and standalone AI in complex tasks like legal analysis and code debugging, but only when the human brought judgment to the interaction. The execution bottleneck is gone; the new bottleneck is judgment. AI can generate a product requirements document quickly. It cannot judge whether the content is reasonable, whether it fits the business, or whether it is worth shipping.
That judgment shows up in concrete behaviors the literature treats as markers of centaur competence. Practitioners who operate this way keep decision logs: recording assumptions, AI recommendations, final choices, outcomes, and deviation causes. They red-team AI output by asking about data sources, timeliness, missing actors, counterexamples, and opportunity cost. They avoid over-reliance — blindly trusting hallucinated data — and underutilization, ignoring data-driven insights. They recognize that poor interface design frustrates human users and that ethical gaps emerge if AI suggestions bypass scrutiny. The MIT researchers observed that centaurs, with their narrower use of AI, relied on their own domain expertise and increased that expertise through targeted questions like "What's the formula for compound annual growth rate?" Their approach was, in part, an effort to avoid becoming overly reliant on AI.
Self-automators saw no skills gains, while centaurs improved their own expertise through the interaction. Organizations that want centaur outcomes now build onboarding phases so workers can learn where the AI works well and where it doesn't, with performance feedback at each step. They add interfaces that visualize uncertainty in AI responses or default questions that force a pause before the human dives into conversation with the model.
For a company building annotation and evaluation workflows, the screening signal is not whether a candidate can write a prompt that works once. It is whether they can design a workflow where a human and an AI system repeatedly produce reliable output together, where the human knows when to trust the model, when to probe, and when to override. Centaur.ai's single opening does not reveal the exact rubric it uses to measure these competencies in a live candidate; the public record does not document it. The research describes the pattern; the company's interview pipeline is where the pattern becomes a test.
The Interview Pipeline: Stages, Duration, and Where Candidates Drop Off
Centaur has not published its interview loop for the single open role, and no candidate reports or company disclosures in the research describe its specific stages, timeline, or attrition points. What follows is the industry baseline for frontier AI engineering hiring (drawn from disclosed processes at 51 companies, structured hiring guides, and large-scale consulting benchmarks), so you can calibrate expectations. Where Centaur deviates from these patterns, the deviation is the signal.
The Standard Loop: Four to Five Rounds
Across 51 companies with public interview maps, a standard AI engineering process runs four to five rounds, each targeting a distinct competency. The first round is typically an asynchronous screen (a take-home assignment or recorded assessment) used by a third of those companies (17 of 51), with another five adding paid work trials. Analysis of 140-plus GitHub repos with verified candidate reports as of September 2026 shows the take-home is "the most revealing part of AI engineering interviews, testing full lifecycle development, evaluation, and production hygiene." Centaur's emphasis on human-AI collective intelligence suggests its take-home, if it uses one, would weigh prompt architecture, evaluation design, and workflow integration more heavily than model-tuning tricks.
Round two is usually a live technical session: system design, debugging a provided codebase, or walking through the take-home. Rounds three and four split between deep technical specialization (RL infrastructure, distributed inference, multi-agent evaluation) and a values or culture panel, often with a "bar raiser" style independent reviewer who holds veto power, a model Amazon popularized and structured hiring guides now recommend to limit groupthink. A final round with the hiring manager or a founder closes the loop.
Timeline: 18 Days at the Frontier, Weeks at Scale
Glassdoor's 2026 data across nearly 39,000 reported interviews puts the average time-to-hire at 29 days. Consulting candidates average 33 days; graduate programs stretch to 43 days because they run on fixed campus calendars rather than business need. Experienced hires into specialized practices can exceed two months when senior interviewers are hard to schedule. The best processes complete five rounds in 18 days by running asynchronous assessments in parallel with scheduling for the next live round. The pipeline never stalls because no single step waits on another to finish. Processes stretching past six weeks consistently lose preferred candidates to faster competitors.
Where Candidates Exit
Three choke points dominate drop-off. First, the initial screen: a former Bain interviewer said, "In my experience at Bain, the screening call was where a large share of applicants removed themselves. They had no clear reason for wanting the firm, and it showed inside ninety seconds." Automated scoring gates compound this, as Accenture's online assessment is scored against predefined criteria using the same automated gatekeeping described earlier. Second, the take-home: candidates who cannot allocate 8 to 12 hours over a weekend, or who treat it as a coding exercise rather than a product-and-evaluation exercise, self-select out. Third, scheduling silence: "Silence at Accenture is usually a scheduling problem rather than a rejection, because a single recruiter is filling several roles at once." They often withdraw quietly. Structured hiring guides advise communicating the full process upfront (number of rounds, format, expected timeline) because candidates who know what to expect stay engaged.
What This Means for Centaur's Single Slot
With one opening, Centaur can afford a tighter loop than a 779,000-person consultancy. It can run asynchronous technical screens in parallel with values conversations, assign a dedicated coordinator so silence doesn't read as rejection, and apply go/no-go criteria at every stage so no candidate advances past a failed core requirement. The research shows that rigor and speed are not opposites. They are a design choice. If Centaur's process takes 40 days and loses candidates at the take-home, that is a choice, not a constraint. The single role makes the cost of a mis-hire existential; the pipeline should reflect that.
What This Story Does Not Cover
Working in AI? Zero G Talent tracks the openings: see every open Databricks role, browse AI jobs, openings at Anthropic, and the people building the field.