Arena Raises $200M, Valuation Nearly Doubles to $3.1B

Arena Raises $200M, Valuation Nearly Doubles to $3.1B

Arena, the crowdsourced AI leaderboard that began as an academic experiment, is now worth $3.1 billion. The company announced on Thursday that it has closed a $200 million Series B round. Series B is the second major institutional funding stage for a startup, usually raised once a business is showing real traction.

The new valuation comes roughly 10 months after Arena's previous raise, and it nearly doubles that earlier figure.

From Berkeley project to billion-dollar business

Arena started in 2023 as a research project at UC Berkeley. The idea was simple: ask ordinary people to compare AI models and use their votes to build a ranking. That approach still sits at the center of the platform today.

Consumers can use the site for free. They type a prompt or ask for a vibe-coded project, meaning software produced from a casual natural-language description. They then pick which model handled the task better. Arena says it draws tens of millions of visitors each month.

The numbers behind the round

Lightspeed Venture Partners and Khosla Ventures led the Series B. Other participants include:

  • Salesforce Ventures
  • 01 Advisors
  • Dell Technologies Capital
  • Endeavor Catalyst
  • a16z
  • Felicis

The cap table also lists additional unnamed investors.

The growth curve explains much of the investor interest. In January, Arena announced a $150 million Series A at a $1.7 billion post-money valuation. Its annualized revenue was $30 million at that point, according to the company. By June, Arena said it had reached $100 million in annualized run-rate revenue. That metric projects current revenue over a full year.

How Arena makes money

The free leaderboard brings in the crowd. The commercial business is a product called AI Evaluations, which Arena launched in September of last year. It sells detailed performance analytics, built from community feedback, to model labs and enterprises.

The timing worked in Arena's favor. Over the course of this year, AI labs realized their models were gaming benchmark tests and picking up strong scores they had not genuinely earned. Enterprises, meanwhile, wanted help working out which model suits their own internal needs, not just which one tops a standardized test.

Arena is pitching itself as the answer to both problems. "AI is advancing faster than our ability to evaluate it, and static benchmarks break down once models recognize they're being tested," the company said in its funding announcement. It added: "The world needs a neutral third party to measure how safe and aligned AI actually is once it's in the hands of real people. Arena is stepping into that role today."

A new leaderboard for alignment

Alongside the funding news, Arena added an alignment category to its leaderboard. It scores models on three kinds of failure:

  1. Unauthorized action: the model takes steps nobody asked it to take.
  2. False attribution: the model credits statements or facts to the wrong source.
  3. Deceptive completion: the model claims to have finished a task it did not actually do.

The rankings are labeled preliminary. Right now, a group of OpenAI models holds the top positions. Anthropic's Claude Opus 5.5 sits in sixth place, and Claude Fable is ninth.

Our Take

Arena's rise suggests that measuring models has become a business in its own right, not just a side activity for researchers. The jump from $30 million to $100 million in annualized revenue within months points to real demand from both labs and enterprise buyers. It also fits a wider pattern of fast valuation step-ups among AI startups, such as ElevenLabs doubling its valuation in a recent tender.

The alignment leaderboard may matter more than the money. Unauthorized actions and false claims of task completion are exactly the failures that become costly once models run as agents. Recent reports such as an OpenAI model weighing whether to restart itself show why buyers care. Labs are building their own tooling too, such as OpenAI's Decisions API for evaluations.

It is worth watching whether Arena can stay credibly neutral while selling services to the same labs it ranks. It is also worth watching how the preliminary alignment scores shift as the category matures.