Article
AI Platforms Tech Giants

AI evaluation platform Arena nearly doubles valuation to $3.1 billion

The company, which started as a UC Berkeley research project, is turning its crowdsourced model rankings into an enterprise business.

by Ian Lyall
The image features a compass placed on top of a stock market trading sheet filled with numerical data. The compass symbolizes direction and decision-making in the financial world. — Credit: Photo by AbsolutVision on Unsplash c Photo by AbsolutVision on Unsplash

Arena, the artificial intelligence evaluation platform best known for its public model leaderboard, has raised $200 million in a Series B funding round that has nearly doubled its valuation in ten months.

The round was co-led by Lightspeed Venture Partners and Khosla Ventures, with Salesforce Ventures, Dell Technologies Capital, 01 Advisors, Endeavor Catalyst, a16z and Felicis also participating.

The new valuation of $3.1 billion compares with $1.7 billion when the company raised a $150 million Series A in January, at which point its annualised revenue was $30 million.

Fast-growing revenues

Revenue has since tripled. Arena said in June that its annualised run rate had reached $100 million, with the company evolving from a research tool into a commercial platform selling AI evaluation services to enterprise buyers.

Arena originated in 2023 as a research project at the University of California, Berkeley, where it crowdsourced head-to-head rankings of AI models from millions of users.

"AI is advancing faster than our ability to evaluate it," the company said in its funding announcement, "and static benchmarks break down once models recognise they're being tested."

Alignment Index

Alongside the fundraise, Arena released its Alignment Index, which scores 27 models across 90,000 real-world agent sessions on three categories of behavioural failure.

The first is unauthorised actions, where a model exceeds its instructions, for example by deleting files when asked only to edit a presentation.

The second is deceptive completions, where an agent falsely claims to have finished a task it did not perform, a pattern the company attributes to models trained to tell users what they want to hear rather than what is accurate.

The third is false attributions, where a model misinterprets what a user wants and performs a different action to the one requested.

OpenAI models currently head the preliminary alignment leaderboard, with Claude Opus 5.5 in sixth place and Claude Fable in ninth.

Neutral by design

Arena positions itself as a neutral third party with no commercial stake in the performance of any individual model. Its public leaderboard is free and non-commercial, used by journalists, AI developers and government policymakers. Enterprise customers pay for deeper benchmarking, model selection analytics and alignment monitoring.

The funding will be used to expand safety signals, agentic capabilities and support for additional modalities, the company said.


by Ian Lyall