Skip to main content
frontier

Braintrust Data Posts 28 AI Roles After $80M Series B, Salaries Up to $180K

By Marcus Bennett

The Hiring Surge

Braintrust Data lists 28 open roles today (13 in engineering, nine in sales and go-to-market, one executive, and one specialist) per Talent Giants tracking. The wave followed an $80 million Series B closed in February 2026. The company's own careers board shows a Developer Support Engineer in San Francisco at $157,000–$180,000, a counterpart in Singapore at SGD 140,000–166,000, a Data Engineer in San Francisco, a Strategic Account Executive for the West Region, and a Software Engineer focused on Developer Experience. The salaried band across roles runs $109,000–$170,000 with a median of $129,000.

Braintrust Data's AIR product, an AI recruiter that conducts structured interviews and skills-based assessments, advertises 80 percent faster screening, thousands of interviews per day across 16 languages, and 100 percent applicant coverage. Its Nexus automation product reports 84 percent of operations automated. Separately, Braintrust (braintrust.dev), the AI observability platform, lists customers including Notion, Stripe, Zapier, Vercel, Ramp, Dropbox, and Coursera. Coursera reports 45 times more feedback with AI grading; Notion reports under 24 hours to deploy a new frontier model. The observability platform's roadmap (tracing, Topics classification, Loop, Brainstore, native CI/CD enforcement via GitHub Actions) demands engineers who can design evals that run on every pull request and block merges that degrade agent quality. The AIR rubric selects for structured thinking, on-topic accuracy, and practical framing.

The open roles map to that product surface: engineering openings dominate a stack built on native SDKs, Brainstore, and the MCP server. The Developer Support roles signal customer-facing technical depth; the Strategic Account Executive hire signals enterprise motion; the Developer Experience role points to the SDK and tooling layer that makes the platform framework-agnostic in practice.

Inside the Screen: AIR's Published Rubric

AIR's evaluation rubric, as published on the product page, comprises eight scored dimensions:

  • Resume Match Score: AI ranking of resume relevance using skills extraction, keyword density, and experience mapping.
  • AI Interview Score: Composite across all dimensions.
  • Behavioral Analysis: Signals for problem-solving, collaboration, and adaptability mapped against 200-plus competency indicators.
  • Communication Rating: Clarity, structure, articulation; evaluates sentence coherence, vocabulary range, pacing, and filler-word frequency.
  • Technical Assessment: Role-specific knowledge and applied skill evaluation; domain question accuracy and practical framing.
  • Intent & Motivation: Signals of genuine interest and role-fit alignment; evaluates research depth, enthusiasm indicators, and career context.
  • Language Proficiency: Fluency assessment in 16-plus languages; evaluates grammar accuracy, vocabulary richness, and native-speaker proximity.
  • Fraud & Identity Check: Multi-signal verification: IP consistency, device fingerprinting, session continuity.

Public figures for the platform show 85 percent faster screening, a 45-to-12-day time-to-fill, 35 percent less first-year attrition across 30,000-plus candidate surveys, 90 percent of candidates saying they would do an AI interview again, and 87 percent feeling they had a fair shot to showcase skills.

Two Engineering Archetypes

Braintrust's 28-role wave clusters around two distinct engineering profiles visible in public postings.

AI Evaluation Engineer: the experiment designer

The evaluation-engineer listings (from the company careers page, third-party boards, and Zero G Talent's live board) center on designing and running evaluations of new AI capabilities, comparing frontier models, agent systems, and tool workflows, and turning emerging ideas into measurable benchmarks. Candidates must define datasets, tasks, and scoring logic; design realistic workloads reflecting production environments; create tests exposing failure modes and edge cases; build evaluation harnesses on Braintrust's platform; run comparisons across models, prompts, and agent approaches; analyze traces, outputs, and failure patterns; and publish results: technical posts, reproducible datasets, open-source reference implementations, and evaluation playbooks for agents, RAG systems, and LLM apps.

The "you are" bullets signal a specific mindset: an engineer who likes testing systems more than building features, enjoys breaking things and understanding why they fail, can design experiments isolating meaningful differences, understands LLMs, agents, and RAG systems, writes clearly for technical audiences, ships experiments quickly and iterates often, cares about methodology and reproducibility, and is curious, creative, and opinionated about AI evaluation. Required experience includes building or contributing to evaluation systems for LLM or agent applications, designing experiments comparing models, prompts, or AI architectures, writing Python code to run tests across models or APIs, building datasets or scoring logic for AI quality measurement, investigating model failures or unexpected behaviors, and publishing technical blog posts, research notes, or engineering write-ups.

This role sits at the intersection of engineering, experimentation, and technical storytelling. The customer roster (Notion, Stripe, Zapier, Vercel, Ramp, Box, OpenAI, Cloudflare) means evaluation patterns developed here become reference points for teams shipping production AI.

Developer Tools Engineer / Developer Experience Engineer: the platform builder

The developer-experience track appears as "Software Engineer, Developer Experience" (San Francisco) and that role in San Francisco and Singapore. The mandate is explicit: "We're looking for an Engineer who loves building tools for other engineers."

Where the evaluation engineer publishes benchmarks, the developer-experience engineer builds the SDK, CLI, dashboard widgets, and integration guides that let a Stripe or Notion engineer drop Braintrust into a CI/CD pipeline. The skill cluster skews toward API design, TypeScript/Python/Go client libraries, OpenTelemetry instrumentation, and documentation that reads like a tutorial rather than a spec sheet. The same "ship quickly, iterate often" velocity applies; the feedback loop is developer adoption (onboarding time, API error rates, feature adoption), not model accuracy deltas.

Side-by-side skill clusters
Dimension AI Evaluation Engineer Developer Tools / DX Engineer
Primary artifact Evaluation harnesses, benchmarks, datasets, scoring logic SDKs, CLIs, dashboard components, integration guides
Core languages Python (experiment runners, model APIs) TypeScript, Python, Go (client libraries, backend services)
Must-have domain knowledge LLM/agent/RAG internals, eval methodology, statistical rigor Developer-tool UX, API ergonomics, observability pipelines
Writing expectation Technical blog posts, reproducible experiment reports Tutorials, API reference docs, migration guides
Feedback loop Model/agent performance on held-out tasks Developer onboarding time, API error rates, feature adoption
Experience signals Published eval research, open-source eval contributions, failure-analysis postmortems Open-source dev-tool contributions, CLI/SDK design, developer-community engagement

Both profiles demand "curious, creative, and opinionated" practitioners who prototype fast. Braintrust's 51–200 headcount (role.com reports) means these two tracks will scale in parallel.

How the Bar Compares

The AI hiring landscape has shifted since 2018. Back then, Palantir stood nearly alone in hiring forward-deployed engineers. By late 2025, roughly one-third of hot AI companies tracked in industry surveys were recruiting for forward-deployed or deployment-engineering roles, and a hybrid "design engineer" — part designer, part engineer — had emerged as a distinct category (per Exponent interview data). Braintrust's current wave sits inside this transformation.

Candidates report walking into interviews at "hottest companies in the world for AI" only to discover late in the process that the role expects them to manage a team of agents, a requirement absent from the posting. Others describe being grilled on RAG internals, LLM-generated code quality, and low-level model behavior despite job descriptions that never mentioned LLMs (per that data). Interview formats are diverging: some companies still lean on LeetCode-style algorithmic puzzles; others have dropped them in favor of practical evaluations. Candidates say the latter signals a "really solid process" (per that data).

Compensation offers a clearer benchmark. Zero G Talent's board data shows Braintrust's salaried roles carrying a typical band of $109,000–$170,000 (median $129,000) across three recent listings; the company's own careers board shows the newly posted role in San Francisco at $157,000–$180,000 and Singapore equivalents at SGD 140,000–166,000. Those figures sit within the upper range for early-stage AI tooling companies in the Bay Area.

Friction shows in candidate behavior: a rise in "noping out" — candidates withdrawing mid-process after a call reveals unexpected agent-management expectations, opaque working-hour norms, or a mismatch between the advertised stack and the actual interview grilling (that data shows). Braintrust's public criteria let applicants self-select based on demonstrated fit.

Where Braintrust differs from the median is scope. Many peers hire one or two forward-deployed engineers for a single product line. Braintrust's 28-role push spans evaluation, developer experience, data engineering, and strategic sales — breadth signaling it is building an entire go-to-market and product engine around the same observability-and-prototyping competency cluster.

Candidate Tactics Inferred from Public Rubric

AIR's published mechanics suggest how applicants might prepare. The platform "presents candidates beyond just a resume" and guarantees qualified applicants an interview. Job seekers can focus less on polishing bullet points and more on the specific evaluation dimensions AIR measures: technical depth in the stated stack, communication clarity under time pressure, and the ability to walk through a problem live.

For Developer Support Engineer roles, that means rehearsing debugging walkthroughs and API troubleshooting narratives mapping to the scorecard's "Technical depth & problem-solving" and "Stakeholder communication" dimensions. The board data shows Braintrust hiring for Developer Experience and Data Engineer roles alongside support; candidates targeting those slots can build proof-of-concept integrations against Braintrust's public SDKs and log latency, error-rate, and cost traces.

AIR runs "back-to-back 15-minute phone screens" at scale, scoring each answer against predefined criteria. Candidates can use timed practice sessions: answer a technical prompt, self-score against a rubric reverse-engineered from the job description, then iterate. Consistency matters — AIR evaluates every applicant on the same rubric.

Salary transparency reshapes negotiation prep. The board's median band of $129,000 (range $109,000–$170,000) and the posted San Francisco figures give applicants a hard floor. Candidates enter final human-review stages with a narrow, data-backed ask.

Speed matters: AIR delivers scorecards within seconds of interview completion. Applicants who follow up within hours (referencing a specific scorecard criterion) signal responsiveness the Developer Support and Developer Experience roles demand. The hiring team can pull the raw transcript and see the follow-up in context.

The pattern: the screen is a structured, recorded, rubric-driven evaluation to prepare for like a practical exam. Candidates clearing it built the artifact, rehearsed the timed response, and treated the scorecard as the spec.

Ripple Effects on the Talent Pool

Braintrust's hiring wave arrives in a market stretched for engineers who can ship evaluation loops and agent observability tooling. First-party board data shows a such role in San Francisco at the previously noted range, a Software Engineer, Developer Experience role in the same city, a Strategic Account Executive (West Region) in San Francisco, and two such listings in Singapore at the previously noted SGD range.

The platform claims (2 million vetted professionals across 100-plus countries, the same screening speed and volume) describe a hiring engine that can compress a typical six-week funnel into days (AIR product page reports). If those numbers hold at scale, every candidate Braintrust moves through its pipeline at that velocity is a candidate removed from the market for other AI-infrastructure teams, evaluation-platform startups, and LLM-app builders recruiting from the same pool.

The "active observability" framing — agents fail differently than normal software, requiring instrumented evaluation (braintrust.dev reports) — has become a hiring signal. Engineers who build public eval harnesses, contribute to Braintrust's open-source SDKs, or publish agent-tracing write-ups are getting inbound from multiple companies.

Salary pressure is visible in the board data's spread. The $157,000–$180,000 band for a Developer Support Engineer (a role blending support escalation, SDK debugging, and eval-authoring) exceeds what many Series B AI startups budget for a senior backend engineer. The Singapore listings, converted at current rates, land near $103,000–$122,000, a premium over local median for developer-experience roles. Braintrust's talent-matching service markets "competitive, transparent pricing" (usebraintrust.com states), but the salary bands it posts for its own team set a floor rippling into offer negotiations elsewhere.

Research on quantified competitor reactions is thin; no public filings or surveys track offer-match rates or attrition tied specifically to Braintrust's hiring wave. First-party data and platform metrics show a company hiring at velocity for a narrow skill set, operating a matching engine that claims to compress a six-week funnel to 24 hours (AIR product page claims). In a talent market where the constraint is verified ability to ship evaluation infrastructure, that velocity compounds.

Can the Engine Sustain Itself?

Braintrust's current hiring wave (28 open roles spanning developer experience, developer support, data engineering, and strategic sales) sits on a platform the company built to accelerate recruiting for its clients. The firm's AI-driven screening tools, marketed to enterprises as cutting time-to-hire by that same margin, processing interviews at that volume across its client base (AIR product page's data shows). That creates a feedback loop: every hire improves the data training the matching models, which sharpens the next round of screening for clients.

The observability platform at braintrust.dev — "that platform for tracing production, running evals, and catching regressions" — signals where the next layer of demand will concentrate. Companies deploying agents at scale need engineers who can instrument, evaluate, and monitor those systems in production. Braintrust's product roadmap doubles as a hiring forecast: as the platform adds tracing, eval frameworks, and regression detection, the firm will need more evaluation engineers, developer tools engineers, and data engineers to build and maintain those capabilities.

Separately, Braintrust Group (braintrustgroup.com), a distinct entity offering live Agile and AI classes, coaching, private training, and digital transformation services, reports over 25,000 professionals have advanced through its courses. That alumni network operates independently from Braintrust Data's vetted global pool of 2 million professionals across 100-plus countries.

The strategic implication: Braintrust Data is simultaneously a consumer of elite AI talent and a supplier of the tooling defining how that talent gets discovered, assessed, and deployed for its clients. As those clients adopt similar stacks, the hiring bar Braintrust Data sets for its own roles becomes a reference point for the ecosystem. Competitors watching the 28-role push should note: the company building the screening infrastructure also calibrates it for its client base.

What happens after the current wave depends on how fast the observability platform's adoption curve steepens. If agent deployments move from pilot to production at the pace Braintrust's marketing suggests ("ship quality agents at scale"), the demand for eval and observability specialists will outstrip the training pipeline's output, forcing the firm to recruit aggressively from adjacent domains (ML ops, site reliability, compiler tooling). The next 12 months will reveal whether Braintrust's hiring engine keeps pace with its product's success — or whether the rubric filtering for agent-observability fluency today becomes the baseline the next wave of AI infrastructure companies copies tomorrow.


Working in frontier tech? Zero G Talent tracks the openings: see every open Braintrust Data role, browse frontier tech jobs, the companies hiring, and the people building the field.

Ready to Start Your Space Career?

Browse frontier jobs and find your next opportunity.

View frontier Jobs