Skip to main content
frontier

Arena Intelligence hits $100M ARR in 8 months, rattling AI benchmark trust

By Elena Petrova

The Scoreboard Starts Billing

Arena Intelligence reached $100 million in annualized run-rate revenue eight months after launching its commercial evaluation service in September 2025, Arena's blog reported — a pace that has triggered a crisis of confidence in traditional AI benchmarks and forced model labs to reassess how they validate performance and safety.

For two years prior, the platform now called Arena ran a free, crowdsourced leaderboard where anyone could pit two anonymous models against each other and vote on which answer felt better. Data accumulated: millions of comparisons, hundreds of models, a live ranking that frontier labs treated as the de facto standard. Then Arena started selling that signal back to the labs being ranked.

The company disclosed the figure in June 2026. The product, AI Evaluations, packages community preference data into structured analytics that model makers use to diagnose where their models lose to competitors and to guide post-training decisions. CEO Anastasios Angelopoulos clarified that the figure represents annualized run-rate based on consumption billing, not recurring subscriptions. Revenue moves with usage. In January 2026, when Arena announced its Series A at a post-money valuation, according to TechCrunch, led by Felicis and Andreessen Horowitz, that run-rate stood at the earlier level. Total funding now sits at that level from investors including Kleiner Perkins, Lightspeed, Laude Ventures, and UC Investments.

The company traces to a 2023 research project at UC Berkeley's SkyLab, where postdoctoral students Angelopoulos and Wei-Lin Chiang built Chatbot Arena under the LMSYS organization, advised by professor Ion Stoica, co-founder of Databricks. They incorporated as Arena Intelligence in April 2025. The free leaderboard remains the moat: 10 million monthly visitors, 700 million conversations, and 82 million human votes form the largest human-preference dataset for AI evaluation in the world. Agent Mode, launched a month before the revenue announcement, logged 5 million turns in its first month and is growing 10 percent week-over-week, extending the platform into agentic workflows.

Arena's only direct crowdsourced competitor, Yupp, shut down in March 2026, leaving the head-to-head comparison format effectively uncontested. The company now competes for the same dollar with human-labeling operations Mercor, Surge, and Scale AI, all of which assist model makers during post-training. Mercor's annualized revenue grew substantially earlier this year, up from half a billion last September; Handshake's AI-training revenue nearly doubled from the mid-hundreds to almost a billion in the same window. The combined spend across that category is a multi-billion-dollar line item that did not meaningfully exist three years ago.

Zero G Talent's board data shows Arena hiring across nine salaried roles: Data Scientist, Senior Software Engineer (Data Infrastructure), Staff Software Engineer (Product), Senior Software Engineer (ML Infrastructure), Founding Product Security Engineer, Software Engineer (Anti-Abuse), with salary bands in the typical range for senior technical roles. Two roles were added in the past week alone.

The business model inverts the traditional benchmark dynamic. Static benchmarks like MMLU or GSM8K are free, fixed, and increasingly gamed. Arena's leaderboard is live, human-driven, and now monetized through the very labs it ranks.

Category Metric Value Period / Note
Arena Funding & Valuation Series A $150M Jan 2026
Post-money valuation $1.7B Jan 2026
Annualized run-rate $30M → $100M Jan 2026 → Jun 2026
Total funding raised $250M As of Jun 2026
Competitor Revenue (Post-training labeling) Mercor annualized $500M → $1B Sep 2025 → 2026
Handshake AI-training $550M → ~$1B Sep 2025 → 2026
Arena Salary Bands (9 roles) Range $150K–$350K Per Zero G Talent
Research Grants Per project Up to $50K Q1 2026 deadline
Model Pricing (Leaderboard) Open model $0.50/M tokens 128k context example
Closed model $5/M tokens 200k context example

Why Enterprises Pay Seven Figures for Human Preference

Static benchmarks have a contamination problem. MMLU, GSM8K, and SWE-bench test against fixed datasets that model developers can access before evaluation, creating what Arena's research calls "benchmark contamination." Developers train on overlapping data, inflating scores without improving real-world performance. The result: a model aces the test but fails in production.

Arena's crowdsourced approach sidesteps this. Human evaluators generate novel prompts in real time. You cannot pre-train on preferences that haven't been expressed yet. The evaluation surface expands organically, covering whatever users actually ask about, not what benchmark authors anticipated.

Three structural flaws plague controlled benchmarks. First, contamination as described. Second, narrow scope: SWE-bench tests Python bug fixes in GitHub issues, a slice of coding work that doesn't generalize to what software engineers actually do daily. Third, proprietary evaluation: when AI labs run internal benchmarks, they grade their own homework. Enterprise buyers see the conflict.

Human preference evaluation dissolves all three. Anonymous public evaluators replace lab employees. The prompt distribution matches real usage. The methodology — blind side-by-side comparison with Elo and Bradley-Terry scoring, prevents gaming because neither the model nor the evaluator knows which system generated which response.

When Arena launched AI Evaluations in September 2025, it hit $30 million annualized run rate by December — less than four months.

"We cannot deploy AI responsibly without knowing how it delivers value to humans," said Anastasios Angelopoulos, co-founder and CEO. "To measure the real utility of AI, we need to put it in the hands of real users."

Enterprises pay for speed. Arena compresses the cycle between model release, user exposure, and comparative assessment. Deployment decisions now move faster than formal evaluation programs can keep up. A procurement team choosing between Claude, GPT-4o, and Gemini for legal document review needs task-specific signal — not a general MMLU score.

The platform serves economically valuable verticals: software engineering, law, medicine, scientific research. OpenAI, Google, and xAI all use Arena evaluations to benchmark models for production use cases. Since March 2024, Arena has tested proprietary and open-source models from major labs and small teams, including pre-release variants where community feedback directly shapes final releases.

"Progress in AI can't be measured in labs by benchmarks alone," said Peter Deng, general partner at Felicis. "It needs to take into account how real people want to use these systems and what they prefer."

Regulatory pressure compounds the demand. Documented AI performance assessments are becoming a compliance requirement. Model developers need third-party validation they can cite to enterprise customers. Enterprise buyers need evaluation they can trust for multi-million-dollar procurement decisions. The entire AI agent economy depends on comparative data telling teams which model to choose for which task.

Arena's dataset — 50 million votes across text, vision, web development, search, video, and image modalities; 400-plus model evaluations; 145,000 open-source battle data points across expert and occupational categories, provides that comparative foundation. The free public leaderboard generates the evaluator pool that subsidizes the enterprise business. Five million monthly users across 150 countries produce 60 million conversations monthly. That scale underwrites the enterprise business.

The Conflict-of-Interest Firestorm

Arena's business model creates a structural tension that the rating-agency analogy makes unavoidable. The platform sells evaluation services to the same model labs — OpenAI, Google, Anthropic, Meta, whose products it ranks for free on its public leaderboard. That "issuer pays" dynamic, which took decades to surface at Standard & Poor's and Moody's, appeared in AI evaluation within months of commercial launch. Critics flagged the conflict as the central risk to Arena's credibility: paying customers gain insight into evaluation methodology that free users don't see, and the labs being ranked are also the platform's revenue base.

The strongest scrutiny arrived in April 2025. Researchers from Cohere, Stanford, MIT, and the Allen Institute for AI published evidence that Arena had allowed a small group of "preferred providers" to test multiple model variants privately before public release, then surface only the best-performing version, according to the research.

Meta reportedly tested 27 variants of what became Llama 4 between January and March before publicizing a single score that ranked near the top. Google tested 10 variants before Gemma 3. By contrast, the startup Reka saw one private variant. The disparity was not marginal — it was systematic. The study showed that the more private variants a provider tests, the higher its expected Arena score for the model it eventually releases. A family of models that is on average weaker can rank higher than a stronger family simply by cherry-picking from a hidden pool, the study showed.

That private-testing advantage compounds with data-access asymmetry. Providers hosting their own models via first-party APIs capture 100 percent of the prompts submitted to their model. Models hosted through third-party platforms receive roughly 20 percent. During the study period (January 2024 through April 2025), just four providers (OpenAI, Google, Meta, Anthropic) collected an estimated 62.8 percent of all Chatbot Arena data, roughly 68 times more than the top academic and nonprofit labs combined, according to the research.

OpenAI and Google models sometimes hit a maximum daily sampling rate of 34 percent; Allen AI maxed out around 3.4 percent. A Cohere experiment confirmed the leverage: testing three variants instead of one lifted its prompt-collection share from 6 percent to over 19 percent, the study confirmed.

The leaderboard's integrity also depends on what disappears. Researchers identified 205 models that appeared effectively removed or sidelined without meeting Arena's published deprecation criteria. Only 47 models were officially marked deprecated. Silent deprecation breaks the Bradley-Terry model's assumption of a sufficiently interconnected comparison graph, because if models vanish before competing broadly, rankings become unstable and less meaningful. Simulations in the study demonstrated that frequent, opaque deprecation directly harms ranking accuracy.

Training on the test set deepens the loop. Win rates on Arena Hard jumped from roughly one in four to nearly one in two when researchers increased the proportion of Arena data in the training mix from zero to 70 percent. The platform's prompt distribution itself shifts: English usage declined while Chinese, Russian, and Korean prompts rose; monthly duplication rates averaged over 20 percent, peaking above 26 percent in March 2025. A 12,000-character limit truncates complex inputs. The user base skews technical. Optimizing for Arena success risks overfitting to these idiosyncrasies at the expense of broader capability.

Arena disputes the characterization of the April episode and says it tightened policies afterward. It argues that model submissions are processed blindly and that its crowdsourced voting pool is too large and distributed to manipulate easily. But the governance gap remains: no public disclosure of how many private variants each provider tests, no transparent limits on pre-release testing, no fair-sampling mechanism that prioritizes under-evaluated models, and no requirement to publish scores of all private variants. The researchers' five recommendations (prohibit score retraction, cap private variants, implement active sampling, enforce transparent deprecation rules, and disclose provider identities) describe the governance infrastructure Arena would need to separate its commercial business from its public authority. Until that wall exists, the platform that became the industry's scoreboard by being neutral now operates as a commercial entity whose customers are also its subjects.

Open Data Reshapes Startup Competition

Arena published its full leaderboard history on Hugging Face in early April 2026. The release spans three years, ten arenas, and 14 dataset subsets: text, vision, search, document, webdev, text-to-image, image-edit, text-to-video, image-to-video, and video-edit, each with style-controlled variants where applicable. Every subset ships in two splits: "latest" for the current snapshot and "full" for the complete historical record. The schema includes category, publish date, license type, organization, vote counts, price per million tokens, and max context window.

For a startup building on open-weight models, this changes the economics of evaluation. Before, a team had to run its own benchmarks against closed APIs or rely on static leaderboards that froze months ago. Now they can pull the full Text Arena trajectory (140,000 conversations, 145,000 battle data points) and see exactly how Llama, Qwen, or Mistral models have moved relative to GPT and Claude since May 2023. The mean score of the top five models has climbed from roughly 1,000 to nearly 1,500. That curve is public. So is the license column, which shows more than half the Text Arena entries carry open or non-proprietary licenses, while video and coding arenas remain dominated by closed models.

The cost and context columns matter as much as the scores. Arena's March update added price per million tokens and max context directly to the leaderboard view. A founder choosing between an open model and a closed model at different price points can compare them in the same table, filtered by license, ranked by human preference. No sales call required.

Big Tech's evaluation moats always relied on three things: proprietary test sets, compute for massive internal runs, and the prestige of "our benchmark says we win." Arena's dataset undercuts all three. The test set is now 82 million human votes across 700 million conversations — larger than any internal red-team corpus. The compute is distributed across 10 million monthly users. The prestige has shifted: investors and enterprise buyers cite Arena rank first, then ask for the internal numbers.

Independent researchers get a boost too. Arena committed significant funding per project for evaluation research, with a Q1 deadline that passed in March. The dataset's longitudinal structure (full splits with publish dates) lets academics track fine-tuning maturity across modalities. They can measure how fast open-source image editing catches up to proprietary video generation, or whether style-controlled rankings change the pecking order in coding tasks.

The counter-move from incumbents is predictable: private benchmarks, stricter NDAs, "trust us" marketing. But the data is already public. A two-person team in Berlin can now run the same longitudinal analysis that used to require a Google Brain allocation. That is the structural shift. The moat didn't just shrink — it opened.


Working in frontier tech? Zero G Talent tracks the openings: see every open Arena role, browse frontier tech jobs, the companies hiring, and the people building the field.

Ready to Start Your Space Career?

Browse frontier jobs and find your next opportunity.

View frontier Jobs