跳到主要內容
AI News HubLIVE
站內改寫4 分鐘閱讀

待翻譯:The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:From Berkeley’s Chatbot Arena to Agent Arena: preference rankings, cost-per-task frontiers, and the hard problems in measuring real-world AI utility.

來源TheSequence作者: Jesus Rodriguez
待翻譯:The Sequence Chat - Issue 930: Arena’s Anastasios Angelopoulos on Chatbot Arena, Evaluation, and What Models Actually Measure
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

We are back with our interview series and a very special guest today! Anastasios Angelopoulos has been at the center of how the industry measures model quality in the wild. We discuss the origins of Chatbot Arena, what an Arena score actually captures, how preference and factuality should (and shouldn’t) be combined, Agent Arena’s performance–cost frontier, AutoEval, and what evaluation looks like once harnesses and tools enter the ranking. Background 1. Could you introduce yourself and tell us how your work on reliable AI and uncertainty quantification led to Chatbot Arena — and eventually to leading Arena as a company? I’m Anastasios, co-founder and CEO of Arena. My path in AI has been pretty nonlinear. I started working in AI + medicine, then came to believe that one of the most important problems deploying AI in the medical sector is reliability. That led me to do a PhD at Berkeley focused on the theoretical foundations of AI reliability. In the process, I worked on many projects, one of which was Chatbot Arena, which evolved into Arena. Now, we are building an AI evaluation company that measures the performance and reliability of AI in real-world workflows across all economically valuable sectors. Reliable, responsible AI deployment has been the through-line in my career. The Main Sequence 2. When Chatbot Arena started at Berkeley, what was the original research bet? What does an Arena score actually measure today: model quality, human preference on Arena's traffic, expected usefulness, or something else? Chatbot Arena started as a side-project for folks in the Berkeley SkyLab. It was originally meant to gather evidence that the Berkeley OSS LLM, Vicuna, was better than competitor models, including one from Stanford (go Bears).Although these models performed similarly on benchmarks, when you chatted with them, it became clear that Vicuna was superior. This is how Battle Mode, and the idea of pairwise human preference grading, arose as a paradigm. Today, the platform has moved far beyond human preferences, ranking models’ task completion rates, hallucination rates, and much more. What makes Arena unique is the ability to rank models by their utility to real people during real workflows. 3. Arena-Rank made the ranking pipeline more transparent. Which statistical assumptions matter most, and how should readers interpret small score gaps or overlapping confidence intervals? The scores on Arena represent the utility of models to our userbase, which consists of tens of millions of monthly visitors. Under standard statistical assumptions, when two confidence intervals overlap, that tells us that two models are indistinguishable based on the available data. 4. Human preference can reward style, confidence, or verbosity even when an answer is wrong. Your factuality leaderboard adds a second signal. How should preference and truth be combined without hiding a value judgment in the weighting? There is no universal method for ranking models. Arena’s methodology takes a human stance, trying to reward models that benefit people.Human preference is one important aspect of benefiting folks, but we control for factors like style and verbosity, and implement factuality as yet another signal that can enhance models’ utility to humans. 5. The latest Agent Arena release introduces task categories, cost per task, and a performance-cost frontier. Is cost per successful task becoming a more meaningful unit than price per token — and how do you make those comparisons fair across tasks of very different difficulty? Because models differ in their token consumption per-task, assessing cost-per-task tends to be a more informative measure than price-per-token. Our randomized methodology ensures that the leaderboard remains fair across task difficulties, since every model gets tasks of the same difficulty on average. 6. Once model labs optimize for Arena, Goodhart's law becomes hard to avoid. How do you distinguish a genuinely better model from one tuned to Arena's prompts, voting behavior, or preferred style? The purpose of Arena is to measure the utility of AI to real people, and in doing so, incentivize the labs to develop models that benefit humanity.We are happy when model labs try to climb our leaderboard because it represents the real utility that millions of users experience on Arena. Our philosophy is that, if by climbing our leaderboard, labs make models that are more useful to our diverse, global community on Arena, it is good for the world. 7. AutoEval gives new models a Day-1 score before enough human votes arrive. How do you keep its reward model from simply encoding yesterday's preferences, and when should AutoEval abstain? We continually recalibrate the AutoEval model on a weekly basis to ensure it always reflects the most recent preferences. 8. As models become more agentic, multimodal, and dynamically routed, does the idea of a single "best model" start to break down? Are we ultimately evaluating a policy that chooses the right model, tools, and compute for each task — and what role does Arena play in that world? We’re soon to go far beyond ranking models! Harnesses and tools should also be part of the ranking conversation, and you’ll see more from us on these topics soon. 9. Arena has moved from a Berkeley research project to a company serving model labs and enterprises. What changed in that transition, and what must remain open, neutral, and auditable for Arena to preserve its scientific credibility? The main change has been that through the company, we have substantially more resources to help accomplish our mission. There has been no change to the level of openness and neutrality on Arena; we still have a strong academic research contingent, and the science is deep within our DNA. 10. What do you see as the hardest unsolved problems at the frontier of AI evaluation? Are the main bottlenecks richer environments, more reliable judges, measuring long-horizon agents, or separating the model from its tools and scaffolding — and which challenge is the field underestimating most? There are so many challenges in the field of AI evaluation, from social science questions – like how to properly define utility – to technical questions, like how to build personalized evaluations that give granular estimates of model performance down to the individual level. It is a rich research area with many ways to contribute. Miscellaneous 11. Give us one prediction about AI evaluation in 2027 that 90% of our readers would disagree with. The recent improvements in the performance of the Chinese model labs have made me feel that they may end up outperforming all American labs within the year. It is a controversial prediction, but I’d give it >50% odds. 12. Who is your favorite mathematician or computer scientist, and which idea from their work has most influenced how you think? Vladimir Vovk has been the most influential mathematician to me. I have been following his work on conformal prediction since I was a first-year PhD student and he has influenced me greatly. I essentially learned statistics by following his work. Another would be Emmanuel Candes, whose work on compressed sensing inspired me since my undergraduate days.I am lucky to have interacted with both of them, and it is really a huge privilege to be able to meet my heroes in real life. Thank you for joining The Sequence Chat.

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • From Berkeley’s Chatbot Arena to Agent Arena: preference rankings, cost-per-task frontiers, and the hard problems in measuring real-world AI utility.

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。