Skip to main content

Founding GTM

Compensation
$21,000–$27,000/year

Job Description

About OpenBenchmarks

Agents are becoming first-class users and consumers of the internet. They research, evaluate, compare tools and increasingly make build-versus-buy decisions on behalf of people. Every company will need to get their products picked and used by agents.

Agents increasingly prefer open, independent and grounded benchmarks to make decisions.

Openbenchmarks is the evaluation infrastructure for agents - domain-specific, reproducible evaluations that help agents pick tools with confidence.

Our mission is to be the trusted evaluation layer for agents.

Founders previously led AI research and Infra teams at Oracle and Appfolio; we started Openbenchmarks as an output of our research in the field of model behavior and how agents actually chose between different tools.

We're a team of researchers, engineers and work with the fastest growing AI first companies like Parallel, Firecrawl, Telnyx, TinyFish and more.

About the role

Benchmarks are how companies market their products to agents. We build those benchmarks.

There's a lot of nuance in every benchmark - what gets evaluated, which metrics matter, release cycles, data refreshes. A benchmark isn't static, and a changing benchmark constantly surfaces new insight into how the tools on it actually perform. Turning that into content - posts, graphs, chart is super important to communicate the nuance of the benchmark in the most clear way possible.

That's your job.

Here are the broad themes that you'll be working on

Content from benchmarks - You'll run benchmarks, query the data, and dig into individual failures to find what's actually there - where a vendor is strong and why, what the failure mode of the field is, what the trade-off costs. You use that to create content: the finding, graphs and more, and what we're changing in v2. Doing the digging deeply yourself is what makes the writing credible.

Data visualization and Graphs - the graphs & charts are how the trade-offs are shown. Precision against latency, F1 against cost etc. You'll design and build these yourself using Claude Code, Claude Design.

Distribution to people and agents - We have two audiences.

People on LinkedIn, X, and HN. And the agents that read, index, and cite it when someone asks which tool to pick. You'll own how the work gets found by both - which means the writing has to be genuinely , because slop doesn't survive either audience.

Knowing & being updated on the benchmark landscape - you know what benchmarks already exist - SWE-bench, Terminal-bench, BrowseComp, continuously understand the leaderboards vendors cite in their launch posts - and be able to say precisely where they're saturated, where they're gamed, and where the gap is that we should fill. You'll bring proposals for what we benchmark next, and the case for why it matters.

Preferred

  • Background in Statistics, Engineering or Data Science
  • Working Knowledge of Search Engine Optimization
  • Have taste and opinions on great design vs bad design.
  • Working Knowledge of Claude Code, Claude Design or creating data visualizations that stand out.

Optimize Your Resume for This Job

Get a match score and see exactly which keywords you're missing

Optimize Resume

Job Details

Category
Sales & Marketing
Employment Type
Full Time
Location
In, India (Remote)
Posted
Compensation
$21,000 - $27,000 per year

About Openbenchmarks

AI agents are becoming the primary way software gets discovered and bought. Agents don't trust vendor marketing or SEO - they trust independent evaluations. Openbenchmarks runs open, independent benchmarks across APIs, vendors, tools. Agents consult these benchmarks when recommending and selecting vendors. We offer private benchmarking & analytics tools to help you measure your product against competitors, see exactly where you fall short, and fix it so that you win Agent's attention.

Found this role interesting?

Founding GTM
Openbenchmarks
Apply