初创公司尽调
尽调报告 AI infrastructure / model evaluation Series A 2026-06-25

LMArena

这个评测品类由它定义,增长动能真实,但当前估值已经计入了超常增长和治理修复

LMArena 看起来已是 AI 评测类别龙头;但在治理、中立性和客户集中度问题解决前,$1.7B Series A 轮价格几乎不给安全边际。

封面要素

成立时间 01
2023 [CO001]
最新轮次 02
150 USD M [CO011]
已融资总额 04
250 USD M [CO013]
月活用户 06
5 M+ [CO017]
月度对话量 07
60 M+ [CO018]
已评测模型数 08
400+ [CO024]

公司概况

LMArena 源自 UC Berkeley LMSYS / Chatbot Arena 在 2023 年启动的研究项目,并在 2025 年 4 月商业化为 Arena Intelligence Inc.。它的核心产品是一个盲测、成对的人类偏好评测平台:用户比较匿名 AI 模型输出,由此生成公开排行榜和一条专有的真实世界偏好数据流。公司通过向模型实验室、企业和开发者提供付费 AI 评测服务变现,其中包括 2025 年 9 月推出的 AI Evaluations 产品,提供可审计性、代表性样本报告和服务级别协议。LMArena 的增长很少见:公开材料显示,月活用户 500 万、月度对话 6000 万、种子轮后累计投票 5000 万次,并在 2025 年 12 月达到 $30M 年化消费运行率。同一门生意也有结构性张力:最重要的客户和投资人,与平台所评测模型背后的实验室高度重叠。

官网
lmarena.ai
成立时间
2023-05-03
创始人
Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica
创立地点
UC Berkeley / San Francisco Bay Area, California
总部
San Francisco Bay Area, California, USA
产品
众包 AI 模型评测平台,覆盖文本、代码、搜索、视觉、图像、视频和智能体基准;另有付费企业评测产品,把基于社区的测试、代表性对战样本、排行榜基础设施和 SLA 支持的分析打包交付。
客户
Frontier AI 实验室、企业 AI 采购方、产品团队和开发者;这些客户需要用可比较的模型评测来支持发布决策、采购和产品质量衡量。
商业模式
评测即服务:在公开 Arena 排行榜和开放评测数据集之上,叠加付费定制与私有模型评测、企业基准测试及相关平台服务。
阶段
Series A
融资情况
2025 年 5 月以 $600M 投后估值完成 $100M 种子轮,随后在 2026 年 1 月以 $1.7B 投后估值完成 $150M Series A,披露融资总额约 $250M。
[CO001, CO005, CO006, CO007, CO011, CO013, CO017, CO018]

执行摘要

主要优势

  • 公开 Arena 排行榜已成为前沿模型发布的事实参照点,让 LMArena 拥有少见的数据和分发护城河。
  • 社区规模是真实的:5M 月用户、60M 月对话和 50M 投票,拼出一个新进入者很难快速复刻的评测数据集。
  • 公司商业化推进很快,2025 年 9 月推出 AI Evaluations,2025 年 12 月达到 $30M 年化运行率。
  • 产品扩展到搜索、编程、多模态和智能体基准,把可想象 ARR 天花板从单纯文本模型对比往上打开。
  • 一线投资人支持和扎实学术根基,弥补了有限运营历史,也提升了 AI 实验室和企业买家的信任。

主要风险

  • 结构性利益冲突明显:主要付费实验室和投资人与 LMArena 公开评测的模型主体存在重叠。
  • Leaderboard Illusion 论文和 Meta Maverick 事件已经挑战基准完整性,真实信誉风险上升。
  • $1.7B 估值约等于已披露 $30M 年化运行率的 57x,几乎容不下执行失误或倍数压缩。
  • 收入质量仍不透明,因为公司披露的是运行率代理指标,而不是审计收入、毛利率或客户集中度数据。
  • 对一家扮演裁判角色的公司来说,治理披露偏薄:公开证据尚未显示董事会构成、独立性控制或详细股权结构表条款。

未决问题

  • 头部客户集中度和净留存率尚未公开披露。
  • 毛利率、算力成本结构和烧钱速度仍无法从公开来源取得。
  • 尚无独立统计审计公开解决 2025 年提出的中立性担忧。
  • 董事会构成、投资人权利和清算优先权细节仍未公开。

目录

Chapter 01

01公司概览

1.1 身份与产品

LMArena 现在以「Arena」品牌运营,是一个 AI 模型评测平台,让用户通过匿名一对一对战比较前沿 AI 模型。用户同时向两个未标注模型提交提示词,投票选出更好的回答,然后才看到被评测的是哪些模型。聚合投票通过 Bradley-Terry / Elo 排名算法驱动公开排行榜,持续生成大型语言模型和多模态 AI 系统在真实世界中的表现对比。公司实体 Arena Intelligence Inc. 于 2025 年 4 月 18 日注册成立,把 Chatbot Arena 研究项目转为商业公司。总部位于 San Francisco Bay Area(具体地址未披露),与 UC Berkeley 渊源一致。主要网站为 arena.ai 和 lmarena.ai。LMArena 的商业产品 AI Evaluations 于 2025 年 9 月推出,为企业、模型实验室和开发者提供基于社区的表现分析、代表性反馈样本和服务级别协议。已具名商业客户包括 OpenAI、Google 和 xAI。商业模式是评测即服务:客户付费,让公司系统评估模型在软件工程、法律、医疗等领域的表现。平台已经远远超出纯文本 LLM 比较,扩展到 Search Arena、WebDev Arena、Vision Arena、文生图、文生视频,以及 2026 年 6 月推出的 Agent Arena。多模态扩张扩大了 LMArena 可评测的表面,也扩大了可触达客户群。截至 2026 年 6 月,Arena 品牌已经成为前沿 AI 模型事实上的公开排行榜,其排名被所有主要 AI 实验室写入产品发布、投资人材料和学术论文。 [CO001, CO002, CO005, CO007, CO008, CO009]

FO002: LMArena 商业架构流

LMArena 的社区输入、平台基础设施、数据产品和商业输出如何连接。

[CO007, CO017, CO020, CO024, CO025, CO026]
FO003: LMArena KPI 快照

截至 2026 年 1 月,在规模、财务和社区维度上的关键绩效指标。

数值为公司截至 2026 年 1 月披露;私营公司没有独立审计。ARR 指年化消耗运行率,不是 GAAP 收入。

[CO007, CO011, CO013, CO015, CO017, CO018]

1.2 创始人与领导层

LMArena 由三位具有 UC Berkeley 背景的人士共同创立。CEO Anastasios Angelopoulos 曾是 UC Berkeley 统计学与机器学习博士后研究员,也是公司主要公开发言人。联合创始人 Wei-Lin Chiang 在 Berkeley 完成分布式系统博士学位,是搭建支撑 Chatbot Arena 的 FastChat 服务框架的关键成员。UC Berkeley 计算机科学教授、连续创业者 Ion Stoica(Databricks 和 Anyscale 联合创始人)以联合创始人身份加入,带来商业化经验和产业可信度。最初的 Chatbot Arena 研究还包括 Lianmin Zheng 等更多贡献者,以及 Michael Jordan、Joseph Gonzalez 等教职顾问。创始团队展现出强创始人—市场匹配:Angelopoulos 和 Chiang 合著了奠定评测方法的关键学术论文,Stoica 在 Databricks 的履历则带来学术 spinout 少见的运营深度。关键人依赖偏高:Angelopoulos 同时担任科学负责人和 CEO。公司尚未公开任命非创始人的 C-suite 高管。董事会组成或独立董事也未披露,这对早期私营公司并不罕见。自注册成立以来,公司未披露领导层离职或治理变更。不过,2025 年基准完整性争议暴露了 Stoica(公开反驳 Leaderboard Illusion 发现)与外部研究者之间的张力,也凸显创始团队可信度对平台合法性有多关键。 [CO003, CO004, CO005, CO006, CO021, CO030]

领导层与创始人表
姓名角色背景创始人-市场匹配关键人风险
Anastasios AngelopoulosCEO、联合创始人UC Berkeley 博士后,Statistics/ML;Chatbot Arena 和 Arena-Hard 论文共同作者(arXiv:2403.04132、2406.11939)极高——评估方法论技术架构师,也是公司主要公开发声者极高——同时担任科学权威和 CEO;离职会影响产品可信度和投资者信心
Wei-Lin Chiang联合创始人UC Berkeley PhD,分布式系统;构建了支撑 Chatbot Arena 的 FastChat serving 框架;核心方法论论文共同作者高——原始平台架构师;技术执行深高——核心工程专长;公开曝光有限,说明可能承担关键运营角色
Ion Stoica联合创始人UC Berkeley CS 教授;Databricks(约 $43B 估值)和 Anyscale(Ray framework)联合创始人;连续创业者高——商业化经验和 venture-building 可信度中等——可能是顾问 / 董事会层面角色;既往公司建设记录降低单点依赖
Michael Jordan学术顾问(LMSYS 论文)UC Berkeley ML 先驱;在概率 ML、贝叶斯方法和统计学习理论上有基础性工作声誉型——学术声望为方法论主张背书低——顾问贡献;无运营角色
Joseph E. Gonzalez学术顾问(LMSYS 论文)UC Berkeley Systems+ML 教师;共同开发 Ray 分布式计算框架;核心论文共同作者技术型——系统专长与基础设施扩展相关低——顾问;无运营角色

董事会组成、非创始人 C-suite(CTO、CFO、VP Sales/BD)以及投资人委派董事未公开披露。枚举仅基于公开记录来源,且并不完整。

[CO003, CO004, CO005, CO006, CO033]

1.3 融资历史与投资人

LMArena 商业化前的资金来自资助和捐赠,包括 Google 的 Kaggle 平台、Andreessen Horowitz 和 Together AI 的贡献,主要以算力资源和现金形式支持研究基础设施。2025 年 5 月,公司注册成立后,LMArena 完成 $100M 种子轮,由 Andreessen Horowitz 与 UC Investments(University of California 的捐赠基金)共同领投,投后估值 $600M。Lightspeed Venture Partners、Felicis 和 Kleiner Perkins 也参与投资。2026 年 1 月,公司以 $1.7B 投后估值完成 $150M Series A,估值接近种子轮的三倍,由 Felicis 和 UC Investments 共同领投,Andreessen Horowitz、The House Fund、LDVP、Kleiner Perkins、Lightspeed 和 Laude Ventures 参投。截至 2026 年 1 月,融资总额约 $250M。UC Investments 同时是 University of California 捐赠基金管理方,也是 UC Berkeley spinout 的机构支持者;Andreessen Horowitz 的投资组合又与 Arena 排行榜上的模型(包括 Mistral)重叠。这些都构成外部批评者指出的潜在利益冲突。公司尚未公开披露二级交易或债务融资安排。 [CO010, CO011, CO012, CO013, CO014, CO021]

利益相关方或投资人图谱
利益相关方角色轮次重要性尽调问题
Felicis VenturesSeries A 领投方Series A共同领投;GP Peter Deng 在官方新闻稿中被引用;可能拥有董事席位或观察员权利董事席位构成;pro-rata 权利;term sheet 中的 full-ratchet 条款
UC Investments (University of California) 投资机构种子轮 + Series A 领投方种子轮 + Series A双重角色支持方:既是捐赠基金投资人,也是 UC Berkeley spinout 的机构支持者;存在结构性利益冲突捐赠基金角色与学术 IP 来源之间的利益冲突协议;来自 UC Berkeley 关联的信息权;IP 授权条款
Andreessen Horowitz (a16z) 投资机构投资人种子轮 + Series ATier-1 VC;两轮均参与;a16z 组合与 Arena 排行榜模型提供商存在重叠(Mistral 投资)组合冲突披露;Mistral 与 Arena 评估的接近程度;获得预发布模型测试优先访问的可能性
Kleiner Perkins投资人种子轮 + Series A老牌 VC;连续两轮参与传递 conviction;可能拥有董事观察员权利观察员席位条款;治理权;反稀释条款
Lightspeed Venture Partners投资人种子轮 + Series A活跃早期科技投资人;连续两轮参与信息权;未来轮次 pro-rata;治理结构
The House Fund投资人Series AUC Berkeley 关联基金;与创始机构存在关联方关系关联方交易披露;相对 arms-length 投资人的财务条款;治理重叠
LDVP投资人Series A科技 VC;关于条款或治理角色的公开信息有限基金重点;既往 AI 组合公司;治理参与程度
Laude Ventures投资人Series A较小轮次参与方;无额外公开信息基金背景;治理参与;pro-rata 权利

各投资人的单独投资金额未公开披露。UC Investments 同时作为捐赠基金投资人和 UC Berkeley 机构支持者,存在需要尽调的结构性利益冲突。二级股东、可转债或 SAFEs 均未公开披露。

[CO010, CO011, CO012]

1.4 增长指标与财务信号

到 2026 年 1 月,LMArena 报告称其月活用户超过 500 万,覆盖 150 个国家,每月产生 6000 万次对话。社区在文本、视觉、网页开发、搜索、视频和图像等模态中累计超过 5000 万次投票,并参与评测 400 多个不同 AI 模型。LMArena 还发布了来自专家和职业评测类别的 14.5 万个开源对战数据点。财务侧,公司 ARR 代理指标——年化消费运行率——在 2025 年 12 月超过 $30M,距离商业化推出约四个月。这个数字被表述为年化运行率,而非已实现 GAAP 收入;独立审计数据不可得。四个月内从零冲到 $30M 年化极不寻常,但商业服务结构(带 SLA 的交付物加社区评分者)意味着毛利率可能受到劳动力成本约束,公司尚未披露。LMArena 没有披露员工数、烧钱速度或毛利率,仅靠公开信息很难完成财务尽调。去重过滤器会移除约 10% 的提交投票;截至 2025 年 7 月方法更新,身份泄露检测移除的投票少于全部投票的 4%,说明公司在主动管理数据质量。 [CO015, CO017, CO018, CO019, CO020, CO023]

快照 KPI 表
指标数值 / 状态日期置信度缺口 / 尽调路径
Series A 后估值$1.7 billion(投后)Jan 2026未披露投前估值;无可用二级市场定价
累计融资额~$250 millionJan 2026新闻稿和多家新闻来源确认
种子轮估值$600 million(投后)May 2025公司宣布;TechCrunch 和 Bloomberg 确认
ARR(消耗 run rate)$30 million+Dec 2025年化 run-rate 代理;未经审计的 GAAP 收入;无 2026 年中更新
月活跃用户5 million+Jan 2026公司报告;「active」定义未说明;无独立审计
月度对话60 million+Jan 2026公司报告;无第三方验证
累计总投票50 million+Dec 2025包含已弃用模型对战;拆分不可得
已评估模型400+Dec 2025包含已弃用和私测模型;确切活跃数量未知
覆盖国家150Jan 2026反映用户地域;并非在 150 个国家注册实体
员工数未披露Jun 2026未公开披露;Series A 资金指定用于技术团队扩张
毛利率 / burn rate未披露Jun 2026私营公司;无财务披露;人工成本结构未知

数值来自公司新闻稿、TechCrunch 报道和投资人公告。「Consumption run rate」是 LMArena 对年化 ARR 代理的称呼;不等同于按 GAAP 确认的收入。Null / Undisclosed 字段表示没有公开披露,不是零值。

[CO011, CO013, CO014, CO015, CO017, CO018]

1.5 里程碑与负面事件

LMArena 从研究演示走到独角兽,历时约三年。最初的 Chatbot Arena 于 2023 年 5 月上线。学术论文快速跟进,最终以 Chatbot Arena 论文(arXiv:2403.04132)确立 Bradley-Terry 方法。到 2024 年,平台已成为事实上的参考基准,被每家主要 AI 实验室引用。首个重大负面事件发生在 2025 年初:Meta 在 Chatbot Arena 上测试了至少 27 个私有 Llama 4 模型变体,提交了一个为 Arena 优化的版本,得分接近榜首,而公开发布版本排名第 32。LMArena 随后道歉并更新排行榜政策。2025 年 4 月,一篇题为《The Leaderboard Illusion》(Cohere、Stanford、MIT、Ai2)的同行评议论文正式记录了对 LMArena 私有测试做法存在系统性偏差的指控。LMArena 联合创始人 Ion Stoica 公开称这些发现存在「不准确」。作为回应,LMArena 引入新的采样算法并发布更新后的透明度政策。公司 2025 年 4 月注册成立,2025 年 5 月融资 $100M,2025 年 9 月推出商业产品,并在 2026 年 1 月完成 Series A,跻身独角兽。截至 2026 年 6 月,未公开披露监管调查、诉讼、制裁或执法行动。 [CO001, CO002, CO005, CO016, CO027, CO028]

里程碑表
日期事件类型金额 / 状态参与方含义
May 2023Chatbot Arena 作为公开研究 demo 上线创立志愿研究项目UC Berkeley LMSYS 团队(Angelopoulos、Chiang 等)首个公开众包 LLM 排行榜;在技术社区内快速自然采用
Jun 2023MT-Bench 和 Chatbot Arena NeurIPS 论文提交产品研究发表(arXiv:2306.05685)Zheng、Chiang、Angelopoulos 等LLM-as-Judge 概念确立;30K 次对话和 3K 次专家投票公开发布
Mar 2024Chatbot Arena 正式平台论文发表产品研究发表(arXiv:2403.04132)Chiang、Zheng 等;累计 240K+ 投票Bradley-Terry/Elo 方法论正式化;所有主要 AI 实验室在产品公告中引用
Jun 2024Arena-Hard-Auto 和 BenchBuilder pipeline 论文发表产品研究发表(arXiv:2406.11939)Li、Chiang 等自动化基准策划约 $20/run;证明与人类偏好有 98.6% 相关性
Jan–Mar 2025Meta 在 Chatbot Arena 私下测试 ≥27 个 Llama 4 模型变体反向未披露私测;只提交优化后变体Meta AI、LMArena首次重大完整性争议;公开发布的 Llama 4 Maverick 排名第 32,而 arena 优化版本排名第 2
Apr 18, 2025Arena Intelligence Inc. 成立创立公司组建Angelopoulos、Chiang、Stoica从学术项目正式转为商业实体;融资流程启动
Apr 29, 2025《The Leaderboard Illusion》论文发表反向同行评审论文(arXiv:2504.20879)Singh、Hooker 等(Cohere、Stanford、MIT、Ai2)学术界正式质疑基准完整性;数据访问不对称已有记录;构成重大可信度风险
May 2025以 $600M 估值完成 $100M 种子轮融资融资 $100M;投后估值 $600M投资人:a16z、UC Investments、Lightspeed、Felicis、Kleiner Perkins首轮大型风险融资;种子轮即接近独角兽;显示 VC 看好 AI 评测市场
Sep 2025AI Evaluations 商业产品上线产品商业化发布LMArena 团队;OpenAI、Google、xAI 为锚定客户收入开始产生;验证 B2B 评测即服务模型
Dec 2025ARR 消耗口径运行率突破 $30M规模年化运行率 $30MLMArena 商业团队从零到 $30M 年化不到四个月;释放企业快速采用信号
Jan 2026以 $1.7B 估值完成 $150M Series A融资融资 $150M;投后估值 $1.7BFelicis(领投)、UC Investments(共同领投)、a16z、LDVP、Kleiner Perkins、Lightspeed、The House Fund、Laude Ventures达到独角兽状态;累计融资约 $250M;从产品发布到 Series A 约 7 个月
Apr 2026更新透明度政策;开源 Arena-Rank治理政策文件发布于 arena.ai/blog/policy/Arena 团队回应持续的基准完整性批评;正式写入抽样规则和数据共享承诺
Jun 2026推出采用因果推断方法的 Agent Arena产品新产品垂直Arena 团队借助处理效应估计,把评测 TAM 扩展到 agentic AI;这是 Series A 后首个重要产品发布

事件日期来自新闻稿和发布时间戳;精确到日的程度不一。负面事件按尽调范围纳入。金额 / 状态反映一手来源数字;未披露金额按未披露标注。

[CO001, CO002, CO005, CO010, CO011, CO016]
FO001: LMArena 公司里程碑时间线

从 2023 年研究项目上线到 2026 年 6 月的关键节点,包括融资事件、产品发布和负面事件。

日期根据新闻稿和发布时间戳估算;不同来源的日级精度不一。

[CO001, CO005, CO010, CO011, CO015, CO016]

1.6 展示项

Chapter 02

02市场分析

2.1 市场边界与定义

LMArena 应被放在 AI 评测与基准测试软件层中分析,而不是泛模型基础设施或可观测性公司。纳入的市场包括人类偏好基准测试、自动化回归测试、面向监管或重专业知识场景的领域评测工作流,以及更新的智能体评测产品——后者评估多步任务成功率,而不只看单轮聊天质量。可服务市场不包括基础模型训练算力、通用 MLOps 编排、广义开发者工具,以及未作为托管工作流软件销售的开源评测脚本。这个边界很重要:有些发布方测算的是宽口径评测平台品类,有些则测算狭义基准工具细分,导致 TAM 相差数倍。LMArena 自身商业信息强调法律、医疗和工程评测,说明它可变现市场更接近高风险验证支出,而不是单靠公开排行榜流量定义。[CM006, CM007, CM008, CM009, CM016, CM018]

市场定义表
类别纳入支出排除支出主要买家对 LMArena 的意义
公开基准和排行榜运营人类偏好基准、并排模型比较、公共信任信号、基准赞助核心训练算力、原始推理支出、通用流量变现前沿实验室、模型 API 厂商、基准赞助方这是为 LMArena 打开市场能见度的声誉切口
企业模型评测工作流回归测试、评测数据集、领域评分、发布闸门、QA 仪表盘通用 BI、工单系统,或无关的开发者生产力工具AI 平台团队、模型质量负责人、领域产品负责人经常性软件支出最可能从公开排行榜使用延续下来的场景就在这里
Agent 和工作流评测多步骤任务成功评分、因果轨迹、工作流基准、工具使用可靠性不带测量的一般 agent 编排、通用 copilots把 agents 嵌入工作流的应用 AI 团队Agent Arena 把市场从聊天机器人排名扩展到执行可靠性
受监管和高风险垂直验证法律、医疗、工程和合规敏感评测项目没有业务关键决策路径的消费者娱乐聊天排名领域负责人、风险负责人、质量团队LMArena 明确把这些领域作为可变现滩头阵地来营销
捆绑式平台评测嵌入云、MLOps 套件和实验平台的评测功能N/AAWS、Google、Microsoft、Databricks、MLflow 用户生态里确实存在这类支出,但独立供应商只能触达其中一部分
相邻但排除的基础设施训练数据管线、基础模型托管、推理服务、通用可观测性这些均不属于本范围内的评测软件市场基础设施团队和 CTO 预算排除这些项目可避免夸大可服务市场

边界行刻意限定在可变现评测软件,而不是所有 AI 工具;捆绑式平台评测作为背景纳入,但 LMArena 只能触达其中一部分。

[CM006, CM007, CM008, CM009, CM016, CM018]

2.2 市场规模与有争议估计

在已验证来源中,最干净的 2026 年宽口径市场锚点是 AI 模型评测平台市场报告:2025 年 $1.86B,2026 年增至 $2.36B,CAGR 为 27.3%,2030 年达到 $6.24B。这个视角大概率包含面向模型实验室、应用开发者和强合规组织销售的企业评测软件。Precedence Research 的窄口径视角则指向明显更小的基准工具细分,2026 年约 $0.85B,因为它似乎排除了部分打包或相邻工作流,并强调评测加基准工具,而非完整平台层。Gartner 的 2026 年 AI 支出预测和 Presenc AI 的生产采用调查都支持评测预算会从当前基数快速扩张,但无法解决边界问题。务实结论是:LMArena 真正可变现市场可能远小于宽口径 TAM;但如果它能拿下可信、嵌入工作流的支出,仍足以支撑一家有意义的独立公司。[CM001, CM002, CM003, CM004, CM005, CM021]

TAM/SAM/SOM 或规模测算视角表
来源年份范围指标数值增长 / 前景解读关键限制
The Business Research Company2026全球 AI 模型评测平台2025-2026 市场规模$1.86B(2025)至 $2.36B(2026)27.3% CAGR;2030 年达 $6.24B本章中经过验证的最佳宽口径平台 TAM 锚点方法细节只做摘要,未完整披露
Yahoo Finance / Research & Markets 转发2026全球 AI 模型评测平台2026 市场规模2026 年 $2.36B27.3% CAGR独立转发大体印证宽口径市场数字是联合发布的新闻稿报道,不是原始模型工作簿
Research & Markets2026全球 AI 模型评测平台2030 年预测2030 年达 $6.24B27.3% CAGR如果宽口径定义成立,则确认未来增长强劲仍是宽泛类别,细分拆解不清
Precedence Research2026模型评测和基准工具窄口径 2026 视角2026 年约 $0.85B约 7.3% CAGR可作为更窄基准工具类别的有用下限不能与宽口径平台 TAM 直接比较;类别看起来更窄
Precedence Research2025-2034模型评测和基准工具长期预测约 2034 年达 $9.57B长周期增长市场显示该类别的重要性和战略价值,包括 M&A 背景预测期限和范围不同于 2026 年宽口径市场来源
Gartner2026全球 AI 经济AI 软件支出2026 年 AI 软件 $453B属于全球 AI 支出 $2.59T 的一部分,同比 +47%支撑评测采购的上游预算池非常大不是直接的评测市场指标
Gartner2026全球 AI 经济AI 模型支出2026 年 AI 模型 $32.6B随软件支出同步快速扩张可作为前沿实验室预算可得性的有用代理仍然间接;没有隔离第三方评测供应商
基于已验证来源的作者估算2026与 LMArena 相关的独立评测 SAM / SOM分析区间SAM 约 $0.3-0.8B;SOM 约 $0.03-0.15B由宽口径 TAM、窄口径视角和已报道 ARR 推导可作为尽调决策区间,但不是出版方背书的估算缺少公开定价、预算占比或客户数量细节,无法精确验证

本表混合了出版方数字和一个明确标注的分析估算;用户在把各行视为可直接相加或互相矛盾之前,应先比较定义。

[CM001, CM002, CM003, CM004, CM005, CM021]
FM001: 市场规模视角

从广义 AI 评测平台,到与 LMArena 相关的 SAM,再到近期 SOM 的三层视角。

SAM 和 SOM 是根据已验证的广义 TAM、狭义基准测试视角和披露 ARR 推导出的分析区间;不是出版方发布的数字。

[CM001, CM022, CM023, CM043]
FM002: 市场估算区间

低位、中位和高位市场视角显示,品类定义如何改变 LMArena 市场规模的表观大小。

所有行均以十亿美元计;高位行合并了 2030-2034 年方向性品类端点,而不是单一年份。

[CM001, CM003, CM046]

2.3 买方分层与采购

LMArena 的买方地图有两个不同重心。第一类是前沿实验室和模型 API 供应商,它们在重大版本发布前后需要可信第三方基准、发布验证和竞争信号。第二类是监管或专业知识密集型企业,它们在法律、医疗和工程工作流中需要领域评测,因为失败成本高,公开消费者基准不够用。这些细分大概率采购方式不同:实验室从中央模型、安全或研究平台预算中支付评测费用;企业则常由 AI 平台负责人、产品 owner,或风险与质量团队采购。竞争也很混合。Scale AI、Arize、Galileo、Patronus、MLflow,以及大型云或平台厂商都在攻打相邻栈位;CoreWeave 收购 Weights & Biases 也说明,实验、可观测性和评测工作流正在企业采购中收敛。[CM010, CM011, CM012, CM013, CM014, CM015]

细分 / 买家地图
细分买家用户付费方工作流预算负责人采用触发因素
前沿模型实验室OpenAI、Google、xAI、Anthropic 类实验室评测研究员、发布经理、安全团队中央模型开发预算在发布前后对新模型做基准测试;与竞争对手比较模型质量 VP / 研究平台负责人竞争性发布节奏和可信外部证明需求
模型 API 厂商和平台提供商托管模型平台和 AI 云平台 PM、信任团队、GTM 团队产品或平台预算用外部和内部评测支撑企业销售和发布主张平台 GM 或产品负责人拥挤 API 市场里需要区分模型质量
法律 AI 厂商和企业法律工作流团队、法律科技买家律师、审阅员、AI 产品经理业务单元或创新预算验证领域准确性、引用质量和工作流可靠性法律创新负责人或 AI 平台负责人法律工作流中幻觉成本高
医疗和健康相关 AI 买家临床 AI 团队、医疗文档厂商临床医生、质量团队、模型验证人员产品、合规或临床运营预算评估安全性、术语准确性和失败阈值首席医疗 AI 负责人或质量负责人患者安全和合规风险让验证支出更容易被证明合理
工程 copilots 和工业知识工作流工程软件团队和应用 AI 小组工程师、分析师、技术审阅员R&D 或产品工程预算测试任务完成度、工具使用和领域正确性应用 AI 或工程系统负责人大规模推出前需要可衡量的生产力提升
基准生态合作伙伴赞助方、评测方和相邻工具厂商面向市场的研究、开发者关系、信任团队市场、产品或生态预算用基准塑造叙事、合作动作或集成工作流产品营销或生态负责人需要锚定类别可信度和公开可比性

买家行反映经验证产品信息和媒体报道支撑的最高概率付费细分,而不是完整客户名单。

[CM010, CM012, CM013, CM014, CM015, CM016]
FM003: 买方 / 细分市场地图

展示主要买方细分在不同采购标准下如何评估 LMArena。

序数标签概括每个细分市场对各项标准的相对重视程度,不是调研分数。

[CM015, CM017, CM018, CM044]

2.4 增长驱动与采用约束

多股力量支撑可信评测供应商跑赢市场。2026 年全球 AI 软件和模型支出激增,企业 AI 在大型公司中已高比例进入生产,向智能体转移又提高了对工作流级评测的需求,而不是静态提示词测试。基准竞争本身也会刺激需求,因为模型提供商需要外部证明点,企业需要独立质量信号。但这个品类也有真实约束。基准刷榜指控、众包评分者偏差问题,以及对公开排行榜的怀疑,都会削弱客户为未明确绑定生产结果的分数付费的意愿。独立供应商还可能受到捆绑云工具和 MLOps 工具的价格压力;公开定价披露有限,也让人很难区分持久软件需求和新鲜感驱动的试验。[CM026, CM027, CM028, CM029, CM030, CM031]

增长驱动因素和约束表
因素类型时点影响尽调问题
企业 AI 生产级采用扩大驱动因素2026-now更多生产工作负载带来回归、治理和发布测试的经常性需求生产 AI 团队中,目前有多少比例采购第三方评测,而不是自建?
上游 AI 软件和模型支出爆发驱动因素2026-2030评测预算可作为更大 AI 技术栈中一个小但扩大的比例增长管理层能否证明,随着 AI 项目成熟,既有客户支出在扩张?
Agentic AI 和工作流自动化驱动因素2026-2030多步骤 agents 需要工作流和因果评测,市场从聊天机器人排名向外扩有多少收入来自 agent 评测,而非经典排行榜工作流?
前沿实验室之间的基准竞争驱动因素当前发布竞争提高了对可信第三方测量和叙事控制的需求哪类买家主要把 LMArena 用于外部信号,而不是内部 QA?
基准刷榜指控约束当前信任流失会降低公开分数的变现能力,除非绑定受控企业工作流LMArena 向付费客户提供哪些反刷榜控制和审计轨迹?
众包评分者偏差和代表性批评约束当前开放 arena 投票可能无法满足需要领域扎根评测的受监管买家企业评测中,有多少比例使用筛选后的专家评分者或私有数据集?
云和 MLOps 平台的捆绑压力约束2026-2028如果评测从一个类别变成一项功能,独立供应商的定价权可能下降面对捆绑的 MLflow、超大规模云或可观测性工作流,LMArena 赢在哪里?
公开定价和合同披露稀少约束当前外部投资者无法独立把 TAM 转换成可预测的收入获取管理层能否披露 ACV 区间、续约率,以及席位或用量扩张动态?

驱动因素和约束只表示方向,不加权;其中若干因素可能利好类别需求,同时也压低独立供应商经济性。

[CM026, CM027, CM028, CM029, CM030, CM031]
FM004: 采用漏斗或价值链地图

评测软件从试验到持续治理支出的五阶段漏斗。

数值经过指数化,用来展示从试验到持久复购支出的收窄过程;不是实测转化率。

[CM026, CM027, CM045]

2.5 证据缺口与矛盾

主要尽调问题不是来源稀缺,而是来源不匹配。公开市场报告对品类边界意见不一;公司和媒体来源披露估值与轶事式客户名称,却没有足够合同细节来建模份额;竞争对手收入数据碎片化;本轮已验证来源也没有披露前沿实验室或企业 AI 预算中到底有多大比例花在评测软件上。因此,本章可以支持可信的区间判断,却不能给出单一精确的 SAM 或市场份额结论。投资人应保留这一矛盾,而不是强行压成点估计:相对狭义基准工具细分,LMArena 可能已经很大;相对更宽的 AI 评测平台机会,它仍然很小。要填补这个缺口,需要客户、定价和预算强度证据,而这些信息并不在已验证公开来源中。[CM003, CM024, CM038, CM039, CM040, CM041]

2.6 展示项

Chapter 03

03竞争对手

3.1 竞争格局:直接基准平台、自动化排行榜、数据标注商,以及自建替代方案

LMArena 位于 AI 模型评测和基准测试领域,目前没有单一竞争对手能复刻它的全栈。可触达的竞争集合有四层。第一,人类偏好平台:LMArena 是众包成对模型比较的主导公开平台;截至 2026 年 1 月,没有任何免费替代品能接近其 500 万月活用户或 6000 万月度对话量。第二,自动化学术排行榜:HuggingFace Open LLM Leaderboard(由 EleutherAI lm-evaluation-harness 驱动)、Stanford HELM 和 BenchLM 聚合 MMLU、GPQA Diamond、SWE-Bench 等标准化基准分数,覆盖数百个模型,但不采用人类偏好投票。这些工具免费且开源,但衡量的是任务准确性,不是整体用户满意度。第三,企业评测供应商:Scale AI 的 GenAI Platform 为企业和政府客户提供定制评测、微调和数据标注流水线,每个项目 $93 K–$400 K+;它不运营公开排行榜。第四,现状替代方案:AI 实验室可以用 EleutherAI harness 或 OpenAI Evals 自建内部评测流水线并运行自己的测试集,完全规避第三方依赖。潜在进入者包括任何想通过基准界面区分其模型市场的大型云厂商,或由大型 AI 实验室孵化的内部中立评测机构。真正重要的竞争维度是人类偏好规模、企业收入潜力、第三方独立性和方法可信度。[CP001, CP002, CP003, CP011, CP012, CP014]

竞争对手画像表
竞争对手类别规模 / 融资目标细分差异化限制
LMArena人类偏好排行榜 + 企业评测融资 $250 M,估值 $1.7 B(2026 年 1 月);月用户 5 MAI 实验室(基准营销)、企业(模型选择)、研究人员最大公开人类偏好数据集;跨模态;实时 Elo 排名收入来自它评测的同一批实验室;用户群偏向技术专业人士
Scale AI GenAI Platform 平台企业 AI 评测、数据标注、微调估值 $13.8 B(2024);据报客户包括 DoD、Meta、Mayo Clinic寻求私有且有 SLA 支撑的评测管线的企业和政府最大 RLHF 数据标注业务;定制私有评测;GPU 集群规模没有公开排行榜;对特定模型并不中立;定价不透明
HuggingFace Open LLM Leaderboard 排行榜自动化学术基准聚合器(开源)HuggingFace 支持(估值 $4.5 B);免费公开工具ML 研究人员、开源开发者、模型发布团队可复现基准;聚焦开放权重;由 EleutherAI harness 驱动没有人类偏好;仅覆盖技术维度;没有企业评测服务
Stanford HELM多维学术评测框架Stanford CRFM(学术);免费工具;无商业产品需要多轴模型评估的 AI 研究人员和政策受众同时覆盖准确性、校准、鲁棒性、偏见和效率周期性而非连续;没有人类偏好;没有企业收入
EleutherAI lm-evaluation-harness 工具开源评测框架社区资助(EleutherAI 非营利);免费;GitHub stars 约 70 K+ML 研究人员、排行榜运营方、构建定制 evals 的企业数据团队60+ 标准化学术基准;可 fork;支撑 HF 排行榜没有人类偏好;没有商业服务;没有 UI / 排行榜产品
OpenAI EvalsLLM 评测框架(开源 + 仪表盘)OpenAI(内部支持;无单独融资);免费框架构建定制 eval 管线的企业和 OpenAI 客户直接集成 OpenAI 模型;为企业用例提供仪表盘模型提供商拥有;独立性有限;聚焦 OpenAI 系列
BenchLM自动化 LLM 排行榜聚合器融资未知;免费公开工具企业和开发者模型选择;2026 年基准跟踪261 个模型、249 个基准;区分已验证与临时排名;包含价格 / 速度没有人类偏好;平台较新,品牌认知低于 HF / Arena
ArtificialAnalysis独立 AI 模型和 API 性能分析融资未知;免费公开工具比较速度、吞吐、成本和智能水平的企业 API 买家与提供商无关的延迟、吞吐、成本和智能基准没有人类偏好;没有评测即服务;模型广度有限,弱于 HF
[CP001, CP002, CP011, CP012, CP013, CP014]
FP001: 竞争定位图

LMArena 在人类偏好规模和企业服务深度上都有差异化位置;免费学术工具聚集在开放 / 研究象限;Scale AI 位于纯企业区。

分数是有证据支撑的序数判断,来自公开产品界面、论文和定价页;不是精确测量指标。

[CP024, CP025, CP011, CP012, CP014, CP015]

3.2 功能与能力对比:人类偏好 vs. 自动化基准

AI 评测的核心能力分野,在于人类偏好排行榜(LMArena 及其 LMSYS 前身)和自动化基准流水线(HuggingFace Open LLM Leaderboard、Stanford HELM、EleutherAI lm-evaluation-harness、OpenAI Evals)。LMArena 的 Elo 排名系统在 2024 年 Chatbot Arena 论文中有描述,它聚合用户提交的任意任务成对投票,形成一个持续更新、反映真实使用模式而非策划测试集的信号。实践中,这意味着 LMArena 对用户在日常任务中感知到的风格、流畅度和有用性信号更敏感——包括写代码、写作、回答问题——但可复现性弱于自动化基准,因为两个用户完全可能合理地偏好相反输出。 自动化排行榜可复现且任务特定。Stanford HELM 在统一 harness 上评估准确性、校准、鲁棒性、偏见和效率等维度,更适合需要受控测量的研究者。HuggingFace Open LLM Leaderboard 主要跟踪开源和开放权重模型在标准化学术测试上的表现。EleutherAI 的 harness 支撑两者,并且可以免费 fork。BenchLM 截至 2026 年 6 月聚合了 261 个模型的 249 个基准,并单独跟踪价格和速度。ArtificialAnalysis 提供独立 API 吞吐、延迟和成本基准。Scale AI 的企业评测产品填补的是另一块空白:针对客户具体生产用例的定制、私有、SLA 支持评测,而不是公开排行榜。这种模式不与 LMArena 争夺品牌认知,但可能争夺企业预算。 LMArena 的关键差异化在于网络效应(用户越多 → 投票越多 → 排名统计稳定性越强)、跨模态覆盖(截至 2025–2026 年覆盖文本、视觉、图像生成、视频、网页开发),以及与前沿实验室的大量预发布合作。它的弱点是方法可复现性、用户群偏向技术熟练的早期采用者,以及向同一批被排名实验室销售评测服务带来的内在张力。没有竞争对手把公开排行榜界面和企业付费评测服务放在同一个品牌下,这既是 LMArena 的结构性护城河,也是它的利益冲突风险。[CP006, CP007, CP008, CP009, CP013, CP015]

功能 / 能力矩阵
采购标准LMArenaScale AIHuggingFace Open LLM Leaderboard 排行榜Stanford HELMEleutherAI HarnessBenchLM / ArtificialAnalysis
人类偏好 / Elo 排名强(规模上独特)缺失(企业定制,不是 Elo)缺失缺失缺失缺失
实时 / 持续更新强(实时对战,24/7)未知(受 SLA 限制,未公开)部分(批量发布)周期性(非持续)按需(研究人员运行)部分(定期更新)
跨模态评估(视觉、图像、视频)强(截至 2026 年覆盖文本、视觉、图像、视频、Web 开发)未知(定制;未公开)部分(多模态推进中)部分(视觉基准有限)部分(多模态原型)部分(图像理解类别)
开源 / 学术基准覆盖部分(Arena-Hard、MT-Bench 衍生)缺失(私有管线)强(MMLU、GPQA、ARC、SWE-Bench)强(多轴学术套件)强(60+ 个基准)强(249 个基准)
企业评估服务(付费,SLA 支撑)强(AI Evaluations 产品 2025 年 9 月上线)强(核心产品;多年期合同)缺失缺失缺失缺失
独立 / 非模型供应商所有中(从被评估实验室获得收入;学术根基)中(数据标注客户与评估客户重叠)强(HuggingFace 是中立平台)强(Stanford 学术机构)强(非营利组织)强(独立分析机构)
公开排行榜 / 透明方法论强(开放方法论文;公开 Elo 分数)缺失(仅企业;结果私有)强(开放基准,可复现)强(多轴公开结果)强(开源;可复现)强(公开排行榜,区分已验证 / 临时排名)
开发者信号(GitHub stars / 社区)强(FastChat 38 K+ stars;平台原生社区)低(企业优先品牌;开源存在感有限)强(HF Spaces 社区;广泛 ML 生态)中(学术采用;研究者引用)强(70 K+ GitHub stars;支撑 HF 排行榜)中(增长中;无开源仓库)

能力等级(强 / 中 / 部分 / 缺失 / 未知)是基于截至 2026 年 6 月公开可访问的产品界面、论文和基准文档作出的序位判断;标为未知的单元格对应私有企业产品,能力可能存在但尚未被公开验证。

[CP006, CP007, CP008, CP013, CP015, CP017]
定价 / 打包对比
平台定价模式标价 / 合同区间包含能力公开定价可得性对买方的含义
LMArena AI Evaluations企业合同(按用量消耗)未公开披露;约 100 家企业客户贡献 $30 M 年化消耗规模,隐含平均约 $300 K/客户(估算,非公司披露)定制评估面板、社区反馈数据、SLA 交付、分析无(联系销售)企业可为优先评估付费,但标价不透明
Scale AI GenAI Platform 平台企业定制合同分析师报告显示每次项目 $93 K–$400 K+数据标注、模型微调、Agent 部署、评估管线无(联系销售)价格更高,但结果私有且保密;不依赖公开排行榜
HuggingFace Open LLM Leaderboard 排行榜免费(HuggingFace 补贴)$0开放基准提交、公开分数、可复现测试集完全公开无成本,但结果公开;无法针对具体企业用例定制
Stanford HELM免费(Stanford CRFM 补贴)$0多轴自动评估、公开结果完全公开学术严谨、多轴覆盖,但没有人类偏好或商业 SLA
EleutherAI lm-evaluation-harness 工具免费(开源)$0(基础设施成本自托管)60+ 个基准任务;自托管或云端运行完全公开控制力最高,但需要工程团队运营;无 UI 或排行榜
OpenAI Evals免费框架;使用 GPT-4o 评判按 API 付费$0 框架费;评判模型运行需 API 成本(约 $5–30/M 输出 tokens)定制 eval 模板、仪表盘、社区 eval 注册表完全公开(框架);API 定价公开深度集成 OpenAI,但由模型供应商所有;用于非 OpenAI 模型对比时受限
BenchLM / ArtificialAnalysis免费(广告支持的分析网站)$0LLM 排行榜、基准聚合、速度 / 成本 / 智能对比完全公开适合模型选择研究;无企业服务或定制评估

LMArena 的单客户平均值($300 K)来自用 $30 M 年化消耗规模除以约 100 家客户;这不是公司披露口径。Scale AI 定价来自第三方分析师聚合,不是公开价目表。

[CP036, CP037, CP038]
FP002: 功能广度 / 能力地图

LMArena 独家领先于人类偏好规模和跨模态实时覆盖;Scale AI 领先于私有企业深度;免费工具领先于可复现性和开源信号。

序数能力等级(强 / 中 / 部分 / 缺失 / 未知)是基于公开审查的产品界面和学术论文做出的有证据判断;标为缺失的单元格表示公开审查的产品界面确认没有该能力。

[CP006, CP007, CP017, CP018, CP019, CP023]

3.3 护城河持久性、切换成本与负面竞争证据

LMArena 的竞争持久性建立在三个因素上:同类最大规模的众包偏好数据集;由学术根基强化的社区信任(UC Berkeley LMSYS Org、FastChat 开源基础设施);以及深度引用网络效应——前沿实验室在营销材料和投资人沟通中引用 Arena 排名,从而抬高退出成本。AI 实验室的切换成本真实存在:如果 Arena 不再是市场信号,它们既有的基准投入就会失去营销价值。这在 LMArena 与客户实验室之间形成互相依赖,带来收入稳定性,也带来结构性独立风险。 关于护城河的负面证据可信且重要。Singh 等人在 2025 年发表的 Leaderboard Illusion 论文(Cohere、Stanford、MIT、Ai2)量化了数据访问不对称:Google 和 OpenAI 估计分别获得 Arena 总数据的 19.2% 和 20.4%,而 83 个开放权重模型合计只获得 29.7%。Meta 在 Llama 4 发布前测试了 27 个私有模型变体,只选择最高分版本;TechCrunch 证实,普通 Maverick 重新提交后排名第 32。LMArena 否认偏差定性,但承诺修改算法,说明批评具有运营层面的分量。第二个负面动态是 Goodhart 定律:当 Arena 排名开始驱动采购决策,实验室会针对 Arena 特定模式优化,信号质量随之下降,也提高了资源充足的竞争者声称方法更优的概率。 支撑 Arena 的 Elo 算法开源,任何资金充足的团队都能复制。难以复制的是社区。Scale AI 的企业触达更广,数据标注能力更深,但尚未建立公开偏好排行榜。真正的替代威胁来自内部自建:如果某个 AI 实验室认为 Arena 的利益冲突问题无解,它可以资助一个竞争性的中立机构。免费学术工具(HELM、HuggingFace)已经在研究用例中承担这类可信度功能,也限制了 LMArena 能为方法本身收取多少溢价。[CP028, CP029, CP030, CP034, CP035, CP036]

护城河耐久性 / 竞争风险登记表
护城河主张威胁严重性缓解措施 / 尽调问题
人类偏好数据集是该领域最大且引用最多的数据集(6 M+ 票、60 M 对话)资金充足的实验室或联盟可能在 18–24 个月内用等效用户激励搭出竞争性人类偏好平台确认 LMArena 的数据集是专有资产还是社区所有;评估与实验室的数据共享协议
网络效应:顶级实验室在 Arena 发布模型,因为社区有这个需求认为机制被操纵或商业上吃亏的实验室可能退出,并资助一个竞争性的中立机构跟踪是否有主要实验室公开减少 Arena 参与,或发起竞争项目
学术可信度(源自 UC Berkeley/LMSYS;9 篇已发表论文)The Leaderboard Illusion 论文(Singh et al., 2025)削弱可信度;后续研究可能加速声誉受损审查 LMArena 对选择性披露批评的方法论回应;确认算法改革是否落地
企业 AI Evaluations 收入为实验室客户制造付费切换成本收入来自被评估的同一批实验室,会形成结构性利益冲突,削弱独立性护城河要求管理层提供结构性独立政策(防火墙、编辑独立性);审查头部客户是否能影响排名
开源 FastChat 基础设施和开放数据发布维持开发者好感免费自动化基准工具(EleutherAI、HF、HELM)为学术用户提供可信度功能,压住 LMArena 的学术护城河上限评估企业收入增长后,LMArena 是否还能维持学术合作(论文、引用)
平台已成事实标准,用于模型发布基准测试Goodhart 定律:Arena 一旦成为目标,实验室会过度优化 Arena 特定模式,拉低信号质量,也给方法论攻击打开口子监测未来模型发布是否在新闻稿中特别引用 Arena 分数,并跟踪任何公开的 Arena 特定调优证据

严重性(高 / 中 / 低)反映截至 2026 年 6 月,基于公开证据对威胁重要性的定性判断;不是数值风险分。

[CP028, CP029, CP030, CP033, CP034, CP035]
FP003: 护城河 / 就绪度 KPI

LMArena 的护城河锚定在数据和社区规模;主要脆弱点是方法完整性和商业利益冲突。

[CP001, CP002, CP004, CP023, CP029, CP033]

3.4 展示项

Chapter 04

04财务

4.1 收入模式、定价与商业化进展

LMArena 的收入模式是从免费到企业的漏斗。免费公开排行榜每月服务 150 个国家的 500 万用户,生成 6000 万次模型比较对话,承担社区和信任建设层功能。企业、模型实验室和 AI 开发者为 LMArena 的商业 AI Evaluations 产品付费(2025 年 9 月推出),该产品提供基于真实用户反馈的定制评测小组、代表性数据样本,以及按 SLA 承诺的交付周期。LMArena 未发布公开价格表;所有企业合同都直接协商。三类已确认企业客户包括 AI 实验室(2026 年 1 月 Series A 新闻稿明确提及 OpenAI、Google、xAI)、软件企业,以及受监管专业垂直领域(法律、医疗、科学研究)。公司年化消费运行率在 2025 年 12 月超过 $30M——距离产品上线不到四个月。 收入确认是重要 caveat:LMArena 和 TechCrunch 都把这个数字描述为「consumption run rate」,而不是已实现年度数字。TechCrunch 指出,公司用消费率来描述其「年度经常性收入(ARR)」,反映的是按使用量计费,而非预付合同式 SaaS ARR。按 GetLatka 汇总的约 100 个企业客户计算,隐含平均合同额约为每客户 $300,000——这是用运行率除以客户数得出的估计,并非公司披露数字。实际 ACV 分布、合同期限和续约率均未公开。[CI001, CI002, CI003, CI004, CI005, CI006]

收入流表
收入流机制单位当前值 / 状态质量尽调问题
AI Evaluations 企业服务用社区人类反馈做付费定制评估面板,并按 SLA 交付企业合同(按消耗计费)截至 2025 年 12 月年化消耗规模为 $30 M(公司披露);约 100 家客户(分析师估算)中(消耗规模 ≠ 合同 ARR;无流失、NRR 或留存数据)获取合同 ACV 分布、续约率、NRR,以及计费是限时还是按用量
数据与分析授权(潜在)向企业和研究人员销售聚合偏好数据或模型性能洞察按数据集或订阅(未确认)尚未确认是独立收入线;公司已发布免费数据集低(该收入流无已确认收入)确认企业评估产品之外是否存在任何商业数据授权协议
合作评估费(潜在)AI 实验室为预发布模型评估名额的提前 / 优先访问付费按次评估,或纳入企业合同(结构未确认)未单独披露;可能打包进 AI Evaluations 合同低(无单独公开披露;无法从企业线拆分)确认预发布评估名额是否单独定价或打包;厘清收入确认方式

AI Evaluations 的当前值是截至 2025 年 12 月的「年化消耗规模」;这不是已实现的 12 个月收入。客户数(约 100)来自 GetLatka 聚合,不是公司披露。

[CI001, CI002, CI003, CI004, CI007, CI008]
定价 / 商业化表
产品 / 层级定价模式标价 / 合同区间包含能力公开价格可得性来源
免费公开排行榜免费(社区补贴)$0模型两两对战、Elo 排名、开放数据发布、多模态覆盖完全公开arena.ai(官方)
AI Evaluations — 企业按消耗计费合同(联系销售)未公开披露;隐含平均约 $300 K(分析师估算,未经公司确认)定制评估面板、社区反馈数据、SLA 交付、分析仪表盘无(联系 evaluations@lmarena.ai)arena.ai/blog/ai-evaluations/(官方)
研究开放数据免费$01.5 M+ 条社区提示词、145 K+ 个对战数据点(开源)完全公开(HuggingFace)arena.ai/blog/two-year-celebration/(官方)
预发布模型评估(合作实验室)打包或协商费用(不清楚)未公开披露优先评估名额、预发布分数可见性None从 TC 和 PRNewswire 报道推断
学术 / 开源层按公司承诺灵活定价未披露(较企业费率有折扣)访问评估社区和基础分析无(公司承诺支持非营利组织)arena.ai/blog/ai-evaluations/(官方)
[CI006, CI007, CI008, CI035]
FI001: 收入模型桥

免费社区活动驱动偏好数据价值,再通过 AI Evaluations 产品转化为企业评测收入。

收入、利润率和再投入数字均为近似值;毛利率未获确认。运行率是年化消耗率(公司披露),不是合同 ARR。

[CI001, CI002, CI007, CI008, CI030]

4.2 成本结构、单位经济与资本充足性

从公开证据推断,LMArena 的成本结构有三大类。第一,算力和基础设施:每月跨前沿模型服务 6000 万次对话,需要大量云算力。公开排行榜使用合作实验室提供的模型 API(可能根据合作协议以折扣或免费方式提供),但企业评测产品在规模化后很可能产生真实推理成本。第二,工程和研究人员:截至 2026 年 1 月 Series A 公告,公司有 41 名员工。按 Silicon Valley 技术初创公司的典型全包薪酬,这意味着年人员成本约 $15–25M。第三,社区管理、数据标注质量保障和 go-to-market 成本,这些无法从公开信息中测算。毛利率未披露;如果是边际评测成本较低的软件式模型,毛利率可能达到 60–80%,但每次对话算力成本高也可能把毛利率压到 40–60%。没有坚实依据的第三方估计。 近期资本充足性看起来很强。公司已在 2025 年 5 月种子轮($100M,估值 $600M)和 2026 年 1 月 Series A($150M,估值 $1.7B)累计融资 $250M。团队精简至 41 人,首个商业收入到 2025 年 9 月才建立,因此截至 2026 年初,年现金消耗大概率远低于 $30M。粗略估算显示,不再融资也至少有 18–36 个月 runway,但公司未披露官方烧钱率或现金余额。LMArena 表示将把 Series A 资金用于扩充技术团队、强化研究能力并搭建新平台功能——这些都会提高未来 burn。Recall Capital-LMArena feeder fund(SEC Form D,2026 年 2 月提交)确认了资本形成活动,但未披露 LMArena 自身资产负债表。未见公开报道显示公司有债务或项目融资义务。[CI009, CI010, CI011, CI012, CI013, CI014]

单位经济表
指标值 / null置信度为什么重要尽调问题
平均合同价值(ACV)~$300 K(估算:$30 M 年化消耗规模 ÷ 约 100 家客户)低(推导估算;两个输入都未经公司确认)决定收入扩展性和销售效率假设确认实际 ACV 分布、企业合同条款,以及客户数是否准确
毛利率n/a(未披露)承销关键;软件型模式应超过 60%;重计算可能压到 40-60%要求管理账提供毛利率;确认合作伙伴 API 成本是否由对方补贴
净收入留存(NRR)n/a(未披露)对按消耗计费的 SaaS,NRR 能区分扩张账户和一次性评估要求队列级 NRR;厘清 $30 M 年化消耗规模来自扩张账户还是稳定账户
获客成本(CAC)n/a(未披露)反映单位经济可行性和 GTM 效率要求按细分市场拆分 CAC(AI 实验室 vs. 企业);确认免费排行榜是否是主要获客渠道
销售周期长度n/a(未披露)大型 AI 实验室的企业评估合同可能涉及数月采购要求从首次接触到签署合同的平均时间;若实验室预算周期影响时点,需要标记
回本周期n/a(未披露)取决于 ACV 和 CAC;两者都未确认,无法估算ACV 和 CAC 确认后再推导;若回本超过 18 个月,需要标记
月度烧钱null(按约 41 名员工、每人 $300 K 全包平均成本 + 基础设施估算,为 $1.5–2.5 M/月)低(粗略估算;非公司披露)决定最低现金需求和融资触发点要求实际月度 P&L;确认基础设施成本和任何一次性费用
现金续航(自 2026 年 1 月起)null(按估算烧钱对比 $250 M 融资,估计 24–36 个月)低(取决于未经确认的烧钱)判断下一轮融资是近期必需还是机会型获取 Series A 交割时现金余额和当前月度现金流出

所有 null 值都表示未公开披露的指标;估算值(ACV、烧钱、现金续航)只是用于尽调框架的粗略推导。低置信度数值在承销前必须确认。

[CI004, CI005, CI014, CI015, CI017, CI018]
资本充足性表
项目来源 / 置信度含义
种子轮(2025 年 5 月)$100 M,投后估值 $600 M高(TechCrunch、PRNewswire 确认)确立商业化现金续航;由 a16z 和 UC Investments 领投
Series A(2026 年 1 月)$150 M,投后估值 $1.7 B高(PRNewswire 官方新闻稿,TechCrunch 确认)近期主要资本基础;用途:团队扩张、研究、平台功能
累计融资$250 M高(两轮已确认融资之和;未公开披露过过桥轮或可转债)对 41 人精简团队而言资本充足;当前规模下现金续航可能为 24–36+ 个月
月度烧钱估算$1.5–2.5 M/月(粗略估算)低(未披露;基于人数和基础设施假设推导)若月度烧钱为 $2 M,$250 M 意味着不含收入也有 10+ 年现金续航——偏保守;实际烧钱可能更高
手头现金(2026 年 1 月)未披露n/a需要管理层确认;关系到人员扩张和产品建设计划
债务 / 项目融资义务公开未见报告中(未发现公开备案或公告)假设资本结构干净;尽调时验证
下一轮触发点未披露;未见关于下一轮融资时间表的公开表述n/a按当前烧钱和收入轨迹,Series B 可能在 2026 年 1 月后 18–30 个月
Series A 资金计划用途扩大技术团队、强化研究能力、开发新功能(公司披露)中(公司披露;无预算拆分)人员增加会推高烧钱;需要扩张后的更新预测

月度烧钱是粗略估算;实际烧钱取决于人员增长节奏、基础设施扩容和数据质量投入。Series B 时间表估算仅作示意,不是公司表述。

[CI009, CI010, CI011, CI012, CI013, CI018]
FI002: 单位经济桥

企业合同先流经面板组建和推理成本池,再生成毛利;所有利润率输入仍为私有信息。

ACV($300 K)为估算。毛利率未知;区间反映类软件上限与重计算下限的差异。所有单位经济都需要管理层披露来确认。

[CI005, CI015, CI016, CI017, CI028]
FI004: 资本强度 / 现金流地图

在估计年度 burn 为 $20–45 M 的情况下,已融资 $250 M 意味着近期资本充足度强,但成本假设尚未确认。

所有成本项均由员工数和基础设施基准推导;$225 M 净现金估算仅作说明,不是公司披露数字。收入是运行率(年化消耗),不是 2025 年已实现收入。

[CI009, CI010, CI011, CI014, CI018, CI019]

4.3 财务风险、利益冲突与尽调障碍

最重要的财务风险是结构性的:LMArena 从同一批 AI 实验室——OpenAI、Google、xAI——获得收入,同时又在公开排行榜上对它们排名。CTOL Digital 在 2026 年 1 月融资后明确记录了这个悖论:$1.7B 估值「相当于该运行率的 57 倍,计入的不只是增长,还有一个假设:这种内在冲突可以被无限期管理」。如果企业客户开始相信 LMArena 排名受到商业关系影响(正如 Leaderboard Illusion 论文对数据访问不对称的指控),评测服务的定价权会下降,免费排行榜的信任优势——核心营销资产——也会同步坍塌。这种双重暴露很少见:多数 SaaS 公司面对的是客户流失风险;LMArena 面对的是可信度流失风险,同一事件既会丢掉客户,也会伤害带来新客户的免费产品。 第二个风险是收入集中度:约 100 个企业客户,且运行率由三大 AI 实验室主导,OpenAI、Google 或 xAI 中任何一个单一客户流失,都可能实质性削减收入。第三个风险是「consumption rate」表述:如果按使用量计费,任何单月收入都可能不可预测,$30M 年化数字可能只是用峰值月份外推,而非已签约的远期义务。第四,按运行率 57 倍估值意味着资本市场同时计入了高收入增长(收入必须超过 $100M 才能支撑 17 倍这一常见后期基准倍数)和持久毛利率——两者都无法从公开数据验证。企业采用放缓或方法信任受损,都可能触发显著估值重置,进而让后续融资或退出选择变复杂。[CI021, CI022, CI023, CI024, CI025, CI032]

公开财务缺口表
缺失的私有指标对分析的影响精确尽调路径
2025 年实际收入(非年化口径)无法验证 $30 M 年化消耗规模会转化为 $8–10 M 的 2025 年实际收入(产品上线 4 个月),还是因爬坡形态而更高 / 更低要求提供从产品上线(2025 年 9 月)到 2025 年 12 月的月度实际收入
NRR / 账户扩张数据没有 NRR,就无法区分增长型账户(利好)和一次性试点评估(不利于耐久性)分别获取 AI 实验室客户和企业客户的队列级 NRR;确认是否已有客户流失
毛利率(P&L 层面)毛利率决定 $30 M 年化消耗规模指向的是可扩展的高毛利业务,还是杠杆有限的服务交付模式要求管理账提供毛利行;确认合作伙伴 API 成本和基础设施成本如何处理
ACV 和合同条款无法验证隐含 $300 K ACV,也无法确认定价是限时、按用量还是按事件获取一份脱敏样本合同,以及各客户层级的 ACV 分布
客户集中度(top-3 收入占比)如果三个实验室客户(OpenAI、Google、xAI)贡献 50%+ 收入,任何单一客户流失都是重大事件要求按客户层级拆分收入;标记任何单一客户收入占比 >20% 的情况
截至 2026 年 6 月的人数和烧钱公司在 Series A 时(2026 年 1 月)有 41 名员工;若扩张计划已启动,烧钱会明显上升要求提供最近一个季度的当前人数、月度薪酬和基础设施成本运行速率

每个缺口都是承销 $1.7 B 估值所隐含收入质量和财务可持续性主张时需要补齐的关键变量。

[CI003, CI004, CI017, CI019, CI023, CI025]
FI003: 财务估算区间

关键财务变量的不确定区间很宽;57x 估值倍数已确认,但验证该倍数所需的收入和成本输入大多仍是私有信息。

除估值倍数外,所有区间均由员工数、典型硅谷薪酬、基础设施基准和披露的运行率数字推导而来。估值倍数中点(57x)根据已确认的 $1.7 B 估值和 $30 M 运行率计算。区间反映输入不确定性,也反映运行率与已实现收入之间的差别。

[CI002, CI011, CI014, CI017, CI018, CI032]

4.4 展示项

Chapter 05

05产品与技术

5.1 产品组合与 Arena 模态

LMArena 的核心面向用户产品是 arena.ai 上的 Arena 平台,它针对给定提示词展示两个匿名 AI 模型之间的对战,并记录社区偏好投票。截至 2026 年 6 月,平台支持十个不同评测 Arena:Text(通用聊天)、Code(智能体式网页开发和编码)、Search(检索增强生成)、Agent(自主多步任务完成)、Vision(多模态理解)、Text-to-Image(图像生成)、Text-to-Video(视频生成)、Image Edit(图像编辑模型)、Document(长文档分析),以及一个独立 Direct Chat 模式,用户可在没有成对对战结构的情况下与所选模型交互。Agent Arena 是最新、技术复杂度最高的 Arena,于 2026 年 6 月 4 日推出,它根据每周 2M+ 工具调用推导出的因果处理效应为编排模型排名。Code Arena 由早期 WebDev Arena(2024 年 12 月推出)重建而来,是一个实时、隔离的编码环境,模型可通过结构化工具调用自主创建、修改和执行文件,输出持久化在 Cloudflare R2,并通过 CodeMirror 6 展示。自 2026 年 5 月起,LMArena 还加入「Battles in Direct」,把 10% 的 Direct Chat 会话路由到匿名成对对战,以增加每日投票量。商业产品 AI Evaluations(2025 年 9 月推出)是一项卖给 AI 实验室和企业的合同评测服务,提供带代表性反馈样本和 SLA 支持交付周期的深度评测。所有公开模型评测都可在公开排行榜上免费获得;商业产品则增加私有、保密、受 SLA 管理的评测运行。 [CE001, CE002, CE008, CE012, CE029, CE032]

产品模块与 Arena 资产矩阵
Arena / 模块用户 / 买方状态 / 成熟度评估方法关键差异化尽调缺口
文本 Arena(Chatbot Arena)普通用户、研究人员、模型提供商已上线 / 成熟(自 2023 年起)Bradley-Terry 成对投票2.5 亿+ 真实对话;控制风格影响后的排名BT 估计尚无独立可靠性审计
Code Arena开发者、企业、模型提供商已上线 / 增长中(2026 年重建)生成式 Web 应用的成对投票;智能体工具调用环境持久会话、CodeMirror 6 实时预览、Cloudflare R2 快照功能正确性与偏好分歧尚未量化
Agent Arena高阶用户、企业、模型提供商已上线 / 早期(2026 年 6 月发布)因果追踪:多信号 RCT 框架方法论新颖;每周 200 万+ 工具调用;首个因果智能体排行榜信号选择和因果模型假设尚未经过同行评审
Search Arena研究人员、模型提供商、企业已上线 / 增长中Bradley-Terry;引用式随机化ICLR 2026 论文;2.4 万+ 场对战;接地质量分析仅支持 3 家提供商;地理定位功能有限
WebDev Arena开发者、模型提供商已上线 / 稳定(自 2024 年 12 月起)基于 Web 应用偏好投票的 Bradley-Terry8 万+ 票;主题建模分析;CSS / JS / HTML 生成Code Arena 的前身;方法论刷新仍待完成
Vision Arena多模态用户、模型提供商已上线 / 稳定图像理解任务上的 Bradley-Terry 成对投票覆盖主流多模态模型覆盖面相对文本 Arena 有限;未纳入图像保真度指标
Text-to-Image Arena创意用户、模型提供商已上线 / 增长中生成图像的成对偏好投票覆盖 Flux、Midjourney、DALL-E 等审美偏好与提示词遵循度没有拆开
Text-to-Video Arena创意用户、模型提供商已上线 / 早期生成视频片段的成对偏好投票早期阶段;2026 年新增 wan2.7-t2v、gemini-omni-flash投票量有限;时间一致性未单独评分
Document Arena企业用户、模型提供商已上线 / 增长中文档任务上的 Bradley-Terry 成对投票长上下文评估;2026 年加入多个排行榜尚未发布具体文档类型拆分
AI Evaluations(商业)AI 实验室、企业2025 年 9 月起 GA带 SLA 交付的私有评估;基于社区反馈2025 年 12 月 ARR run-rate 为 $30M;OpenAI、Google、xAI 被列为客户SLA 条款、安全披露、DPA 未公开
Direct Chat所有用户已上线 / 稳定不做评估;提供免费模型访问无付费墙访问前沿模型不产生收入;推理成本中心
AI Evaluations(API / pipeline 管线)企业、开发者路线图 / 计划中程序化提交评估可释放大规模自助评估能力尚未公布时间表;上线前风险

状态和投票数来自截至 2026 年 6 月的 LMArena 官方博客、排行榜更新日志和新闻稿。尽调缺口为研究员推断。

[CE001, CE002, CE008, CE012, CE029, CE032]
FE001: LMArena 产品架构栈

LMArena 平台的分层视图,从数据收集,到排名引擎,再到评测产品。

层级边界根据博客文章和 GitHub 仓库推断;内部服务拆分未公开披露。

[CE001, CE002, CE003, CE009, CE013]

5.2 技术架构与评测方法

所有 Arena 排行榜都使用 Bradley-Terry(BT)成对比较模型,根据胜负结果推断每个模型的潜在能力系数。LMArena 的 Arena-Rank Python 包以 Apache 2.0 开源,并发布在 GitHub 和 PyPI 上,实现了 BT 拟合、闭式置信区间计算,相比历史 FastChat 实现提速 30 倍。该包也支持重新加权来修正非均匀采样,意味着对战较少的模型不会被惩罚。风格控制扩展把额外协变量(token 长度、markdown 标题数、markdown 加粗数、markdown 列表数)叠加入 BT 回归,用于把内容实质与风格式格式化效应分开;系数估计显示长度是主导风格因素。Agent Arena 不用成对投票排名,而是使用因果追踪:Arena 把每次组件选择视为多干预随机对照试验中的一个 treatment,聚合多个行为信号(确认成功、表扬 / 投诉、可引导性、bash 错误恢复、工具幻觉),为每个模型形成单一净改进估计。Arena-Hard 流水线(BenchBuilder,arXiv:2406.11939)通过七标准难度标注器,从实时 Arena 数据集中抽取 hard prompts,自动构建基准;Arena-Hard-Auto v0.1 以约 $20 成本达到与人类偏好排名 98.6% 的一致性。Search Arena 方法以数据集和论文形式发表,并被 ICLR 2026 接收(arXiv:2506.05334),将 BT 模型扩展到来自 Perplexity、Gemini 和 OpenAI 的 11+ 搜索增强 LLM 系统。在分享任何对话数据之前,LMArena 使用 GCP Sensitive Data Protection API 移除个人身份信息。自 2025 年 7 月起,方法更新会公开记录在 Leaderboard Changelog。 [CE003, CE009, CE013, CE014, CE015, CE016]

技术与运营架构
层 / 组件作用依赖 / 实现风险
Arena-Rank(排名引擎)为所有 Arena 计算 BT 系数、置信区间和风格控制后的分数开源 Python 包(Apache 2.0);GitHub lmarena/arena-rank;PyPI arena-rank方法论开放 = 可复现,但也可被竞争对手复刻
Bradley-Terry 成对模型核心统计模型;从胜负对战中推断潜在能力自研 Python 实现;比 FastChat 基线快 30 倍;闭式 CI模型假设(IID 对战、无时间漂移)在重度采样操纵下可能失效
风格控制扩展在 BT 回归中把内容实力与格式效果拆开加性协变量回归;长度、markdown 标题、加粗、列表数量仅有四个风格协变量;更丰富的语义风格尚未捕捉
因果追踪(Agent Arena)通过多干预 RCT 对智能体组件排名;衡量净处理效应自研统计框架,结合 5 个行为信号方法论新颖;因果识别假设尚未由第三方验证
BenchBuilder / Arena-Hard 管线用 LLM 标注器从实时 Arena 数据自动生成高难基准子集arXiv:2406.11939;基于 GPT-4 的难度评分器;7 项难度标准依赖 GPT-4 做标注;标注器偏差会渗入基准质量
数据清洗(GCP Sensitive Data Protection API)公开数据发布前移除 PIIGoogle Cloud Platform API;任何共享数据集发布前调用依赖单一云厂商;未发布独立 PII 审计
Web 平台(arena.ai)承载对战、投票 UI、排行榜、Direct Chat、Agent Mode、Code ArenaJavaScript 前端;服务器架构未披露社区信任的单点故障;正常运行时间 SLA 未公开
Cloudflare R2 + CodeMirror 6(Code Arena 基础设施)存储持久代码会话;渲染实时 Web 应用预览Cloudflare R2 存快照;CodeMirror 6 做代码展示;流式前端数据驻留和保留政策未披露
FastChat(遗留)历史上的服务和训练平台;原 Chatbot Arena 后端GitHub lm-sys/FastChat;截至 2025 年主要处于维护模式已被主动降级优先级;仍在 FastChat 上的外部用户存在迁移风险
p2l 模型(HuggingFace)prompt-to-leaderboard 偏好模型;用于内部评估研究HuggingFace 上的 lmarena-ai 组织;参数规模 0.1B–7B生产用途未公开记录;目的和部署方式不清楚

架构根据官方博客、GitHub 仓库、PyPI 包和学术论文推断。服务器基础设施细节未公开披露。

[CE003, CE009, CE013, CE014, CE015, CE016]
FE002: LMArena 文本和代码对战用户工作流

从用户提交提示词,到投票记录,再到排行榜更新的端到端流程。

流程根据官方博客文章和政策页面重建;内部系统架构未披露。

[CE002, CE016, CE020, CE021]
FE003: 关键依赖地图

LMArena 产品在平台、数据、监管和合作伙伴上的关键依赖。

依赖关系根据公开文档推断;相对关键性为研究者判断。

[CE013, CE014, CE019, CE020, CE032]

5.3 部署、数据基础设施与开发者生态

Arena-Rank 排名引擎作为可用 pip 安装的 Python 包发布(pip install arena-rank),归属于 lmarena GitHub 组织,该组织也托管 list 仓库。FastChat 仓库(lm-sys/FastChat,现在主要处于维护模式)最初支撑 Chatbot Arena,并继续作为 Vicuna 模型权重的参考。lmarena-ai HuggingFace 组织托管内部用于偏好建模的 p2l(prompt-to-leaderboard)模型变体,并发布公开对战数据集(例如 arena-human-preference-140k)以支持外部研究。截至 2026 年 6 月,changelog 记录了各排行榜每周新增数个模型的持续节奏。Code Arena 的持久会话建立在 Cloudflare R2 存储之上,用 CodeMirror 6 展示源码视图和生成网页应用的实时渲染。Agent Mode 会话平均约 16.5 次结构化工具调用,约 75.6% 的会话至少使用一个工具;最高使用会话在单周内运行很长调用链,覆盖编码、文件创建和网页合成任务。在一个被测的 7 天窗口中,Agent Mode 通过成功的 write_file 调用写入 4030 万行代码,约每个编码会话 1000 行。Battles in Direct 自 2026 年 5 月集成后,把 10% 的直接聊天会话转化为成对对战,并对 BT 拟合应用位置偏差和同组织指示器修正。公司未公开记录面向第三方消费 Arena 分数或原始投票数据的 API,但会定期在 HuggingFace 发布开放数据集。 [CE013, CE014, CE015, CE019, CE030, CE031]

关键用户流程与用例表
用户任务传统流程LMArena 方案可衡量收益已知限制
选择生产模型前比较 LLM 质量跑内部 A/B 测试;依赖静态基准(MMLU、HumanEval)Text / Code Arena 成对对战投票;大规模观察真实用户偏好经过统计校准的排名,含 95% CI;覆盖风格控制和高难提示子集排名反映普通用户偏好,不一定贴合企业特定用例
为真实工作负载评测智能体编码模型静态代码正确性基准;合成测试套件Code Arena:隔离智能体环境;实时生成 Web 应用;持久会话捕捉智能体式、多轮行为;运行结果可分享正确性与审美偏好未拆开;编程语言多样性有限
评估面向检索任务的搜索增强 LLMSimpleQA 静态基准;人工标注Search Arena:众包投票评判多轮搜索-RAG 输出2.4 万+ 成对交互;引用数量分析;ICLR 2026 验证当前仅 3 家提供商;学术 / 领域覆盖有限
衡量自治智能体在业务任务中的表现人工评分会话;特定工具基准Agent Arena:跨每周 200 万+ 工具调用做因果追踪;5 信号排行榜基于 RCT 的因果处理效应;逐信号拆解可解释方法论尚未经过同行评审;信号集合会继续演进
为上线前模型获得带 SLA 的保密评估专有红队测试或供应商专属评估团队AI Evaluations:私有评估,含 SLA、代表性反馈和数据样本扎根真实世界人类偏好;交付时间线可认证定价、SLA 条款和安全细节未公开披露

工作流综合自 LMArena 博客、学术论文和 TechCrunch 报道。收益主张除标注为研究员推断外,均为公司表述。

[CE002, CE009, CE012, CE029]
FE004: 产品成熟度与能力地图

截至 2026 年 6 月,各评估 Arena 在四个能力维度上的成熟度。

成熟度评级是研究者基于截至 2026 年 6 月可得公开证据作出的判断。投票量分层:高 = 250M+;中 = 24k-80k;低 = <24k。

[CE002, CE008, CE009, CE035, CE036]

5.4 差异化、IP 与研究优势

LMArena 的核心差异化,是把大规模真实世界、野生环境数据与方法透明度结合起来。不同于静态学术基准(MMLU、HumanEval)或自动化 LLM-as-judge 流水线,Arena 捕捉跨语言、主题和技能水平的真实多轮用户交互。已发表研究验证了这种方法:最初的 Chatbot Arena 论文(arXiv:2403.04132)显示其与专家评分者一致性超过 80%;MT-Bench(NeurIPS 2023)证明 LLM-as-a-judge 可以达到人类级别的评分者间一致性。Arena-Hard 基准(arXiv:2406.11939)以一小部分成本实现了比 MT-Bench 高 3 倍的模型区分度。Search Arena ICLR 2026 论文(arXiv:2506.05334)把方法扩展到检索增强系统,并展示了偏好与引用数量、被引用来源类型之间的相关性。LMSYS(前身非营利组织)发布了支撑当前公司差异化的基础评测方法论文;这些论文在 AI 行业被广泛引用。HuggingFace 组织发布偏好数据集(140k+ 标注对战对),让外部研究者能够独立审计排名。公司 9 篇已发表论文和 15+ 篇博客文章覆盖评测方法,构成一组被认可的 IP。不过,方法资产被公开分享,意味着竞争对手可以复制算法;公司留下来的持久护城河,是对战规模带来的数据网络优势和社区信任。 [CE003, CE016, CE035, CE036, CE037, CE039]

路线图、发布与发展阶段
日期 / 时期功能 / 里程碑状态含义来源
May 2023Chatbot Arena(Text Arena)公开发布已完成确立众包评估作为可行方法论来源:lmsys.org/blog/2023-05-03-arena/
Jun 2023MT-Bench 多轮问题集和 LLM-as-a-judge 论文(NeurIPS 2023)已完成成对评估方法论获得学术验证arXiv:2306.05685
Mar 2024Chatbot Arena 技术论文发表(arXiv:2403.04132)已完成24 万+ 票;建立被产业引用的可信度arXiv:2403.04132
Apr 2024Arena-Hard pipeline(BenchBuilder)发布已完成自动化基准创建;与人类排名一致率 98.6%arXiv:2406.11939
Dec 2024WebDev Arena 发布已完成;被 Code Arena 取代首个面向编码的真实世界评估;收集 8 万+ 票来源:arena.ai/blog/webdev-arena/
Sep 2025AI Evaluations 商业产品发布GA;2025 年 12 月 ARR 达 $30M收入引擎;受 SLA 约束的企业和实验室评估服务来源:arena.ai/blog/ai-evaluations/
Mar 2026Search Arena 论文被 ICLR 2026 接收(arXiv:2506.05334)已接收并展示搜索增强 LLM 评估方法论获得同行评审验证arXiv:2506.05334
Mar 2026 (est.)Battles in Direct 实验开始(10% direct 会话)实验阶段提高投票量;校正位置偏差和同组织偏差来源:arena.ai/blog/leaderboard-changelog/
May 2026Battles in Direct 投票纳入排行榜已上线收窄置信区间;提示分布转向更难查询来源:arena.ai/blog/leaderboard-changelog/
May 2026Code Arena 发布(从 WebDev Arena 重建)已上线智能体编码环境;持久会话;Cloudflare R2 快照来源:arena.ai/blog/code-arena/
Jun 4, 2026Agent Arena 以因果追踪方法论发布已上线 / 早期首个基于 RCT 的智能体排行榜;分析每周 200 万+ 工具调用来源:arena.ai/blog/agent-arena-methodology/
2026(计划)Agent Arena 增加行为信号;扩展更多 Arena / 模态已宣布意向;无具体日期智能体评估平台扩张;信号集合会继续演进来源:arena.ai/blog/agent-arena-methodology/

日期来自博客文章、学术论文时间戳和排行榜更新日志。计划项基于已表述意向,不是已确认发布日期。

[CE008, CE012, CE029, CE033, CE035, CE036]

5.5 信任、安全、质量控制与负面证据

LMArena 的公开排行榜政策(最后更新于 2026 年 4 月 30 日)明确了模型上榜资格、采样规则(要求 ≥20% 的对战发生在公开可用模型之间)、匿名模型的预发布测试协议,以及数据共享政策:发布前用 GCP 的 Sensitive Data Protection API 清洗对话。可是,截至 2026 年 6 月,LMArena 尚未发布 SOC 2 Type II 或 ISO 27001 认证,也未公开披露第三方安全审计;help.arena.ai 的隐私政策页也缺少企业数据处理协议(DPA)细节。平台经历过两次严重的可信度挑战。第一,2025 年 4 月,Meta 提交了一个「experimental, chat-optimized」版本的 Llama 4 Maverick,在 Arena 排名第 2,但它并非公开发布模型;随后 LMArena 更新政策,并重新给公开版本打分,后者跌至第 32 名。第二,Cohere、Stanford、MIT 与 AI2 在 2025 年 4 月发表论文,指称部分大实验室(Meta、OpenAI、Google、Amazon)获得了不成比例的高采样率,并能选择性压低表现较差的预发布变体,构成基准博弈;LMArena 否认这些不准确之处,并指向其已发布的采样透明度。这些事件带来持续的声誉风险,也促使研究社区呼吁更严格的预发布限制、独立审计和算法透明度。AI Evaluations 商业产品提供 SLA 承诺,但详细的正常运行时间 SLA 条款和安全披露并未公开。 [CE020, CE021, CE022, CE023, CE024, CE025]

信任、质量与合规控制
控制 / 认证状态范围缺口 / 风险
排行榜政策(公开)已上线;最近更新于 2026 年 4 月 30 日模型准入、采样政策、预发布测试协议、数据共享规则政策由平台自我执行;没有独立审计方或执行机制
数据清洗(GCP Sensitive Data Protection API)用于公开数据发布Conversation 数据在 HuggingFace / 公开发布前移除 PII依赖单一供应商;清洗效果没有公开审计
开源排名方法论(Apache 2.0)已上线;GitHub / PyPI 上的 arena-rank v1Bradley-Terry 引擎、重加权、CI;驱动当前所有排行榜算法开放,但平台数据是专有的;完整复现受限
方法论的学术同行评审3 篇已发表论文,1 篇 ICLR 2026 接收论文论文:Chatbot Arena(arXiv 2403.04132)、MT-Bench(NeurIPS 2023)、Arena-Hard(arXiv 2406.11939)、Search Arena(ICLR 2026)Agent Arena 因果追踪方法论尚未经过同行评审
预发布测试协议2024 年 3 月起公布政策;2026 年 4 月更新模型提供商可测试匿名预发布模型;结果私下共享优先访问争议(Cohere / Stanford 论文)触发政策更新,但没有独立审计
企业 SLA(AI Evaluations)2025 年 9 月起 GA;提供 SLA为商业评估任务承诺交付时间线SLA 正常运行时间条款、赔偿和安全审计范围未公开披露
隐私政策help.arena.ai 上已有社区平台交互中的用户数据处理没有企业 DPA;GDPR / CCPA 合规细节未公开记录
SOC 2 Type II / ISO 27001未公开披露N/A —— 适用于企业云服务未发布认证是企业买家的尽调缺口

状态反映截至 2026 年 6 月 25 日的公开资料。「未公开披露」不等于确认没有;LMArena 可能持有未发布的认证。

[CE020, CE021, CE022, CE023, CE024, CE025]

5.6 附录

Chapter 06

06客户

6.1 客户基础分层

LMArena 的客户基础分成两个本质不同的群体,必须仔细区分。第一类是免费社区:来自 150+ 个国家、每月 500 万以上的全球用户,通过 arena.ai 免费使用平台。这些用户提交 prompt、为模型输出投票,支撑排行榜背后的集体偏好信号。社区覆盖开发者、研究人员、学生、知识工作者和爱好者;按两周年博客(2025 年 4 月)披露,约 41% 的对战涉及开源模型,说明用户基础技术含量较高。公司不从这一群体直接变现,但从中获得核心 IP(偏好数据)和品牌可信度。第二类是商业客户,即为 2025 年 9 月推出的 AI Evaluations 服务付费的 AI 实验室和企业。2026 年 1 月 Series A 新闻稿确认的具名付费客户包括 OpenAI、Google 和 xAI。Anthropic 与 Meta 是排行榜上的主要模型提供方,但公开披露并未单独确认它们是否为 AI Evaluations 付费订阅客户;TechCrunch 2026 年文章称 LMArena 与 OpenAI、Google 和 Anthropic 在模型提交上「partnered with」,这一类别同时包含商业和非商业关系。企业(非实验室公司用 AI Evaluations 为自身应用基准测试模型)是公司声明的第三个目标客群;目前没有公开披露任何具名的非实验室企业客户。 [CU001, CU002, CU003, CU004, CU005, CU006]

客户分层表
分层买方 / 用户 / 付款方用例规模收入 / 战略价值关键缺口
社区用户(免费)终端用户(研究人员、开发者、爱好者)免费模型对战;排行榜访问;Direct Chat500 万+ 月活用户;150+ 个国家无直接收入;核心 IP 和品牌护城河;PR / 增长引擎用户画像未按职业或行业公开拆分
AI 实验室模型提供商(非付费 / 免费)AI 实验室:Anthropic、Meta、Mistral、DeepSeek 等免费提交模型参与公开评估,获得排行榜排名和市场可信度评估 400+ 个模型;300+ 次预发布测试无直接收入;为平台价值提供模型供给在所有沟通中并未清楚区分非付费客户与付费客户
AI 实验室商业客户(付费)AI 实验室:OpenAI、Google、xAI(已点名)带 SLA 的私有 AI Evaluations;为模型改进提供全面反馈3 家已点名;其他客户未知产生收入;2025 年 12 月该分层 ARR run-rate 为 $30M客户数量、合同规模和单个客户贡献未披露
企业客户(付费)希望为生产用途评测 AI 模型的企业用 AI Evaluations 衡量特定领域(法律、医疗、软件)的模型表现规模未知;没有点名企业客户公司表述的目标客群;尚无公开证据显示已有非实验室企业客户没有点名企业客户;pipeline 和转化率未知
学术 / 研究用户(免费)大学研究人员、智库、开源贡献者用公开数据集做研究;把排行榜作为参照;分析对战数据数百篇研究论文引用 Arena无直接收入;建立学术可信度并验证方法论尚未宣布正式学术合作伙伴计划

分层基于官方新闻稿、博客和新闻报道。收入归因根据唯一披露的 ARR 数字估算;分层拆分未公开。

[CU001, CU002, CU003, CU004, CU005, CU006]
FU001: LMArena 客户旅程地图

从首次接触到商业扩张,梳理客户细分、采用入口和关键互动触点。

旅程阶段根据官方沟通和媒体报道推断;公开资料中没有客户研究或 NPS 数据。

[CU001, CU002, CU005, CU006]

6.2 采用轨迹与使用指标

LMArena 的社区采用增长很快,且证据链较完整。平台 2023 年 5 月作为 UC Berkeley 研究项目上线;到 2025 年 5 月,公司披露以 $600M 估值完成 $100M 种子轮,彼时已有显著牵引力。到 2026 年 1 月 Series A 公告时,LMArena 报告 5M+ 月度用户(高于 2025 年 9 月 AI Evaluations 博客中隐含的约 3M)、60M+ 月度对话、250M+ 累计对话和 2M+ 月度投票。Series A 博客把 2025 年 5 月种子轮以来的社区增长量化为「25x」。模型评估方面,LMArena 已评估 400+ 个公开模型和 300+ 个预发布模型变体。从 2025 年 9 月商业化上线到 2025 年 12 月(不到四个月),按 CEO Anastasios Angelopoulos 披露、并由 TechCrunch 2026 年 1 月文章和官方新闻稿确认,年化收入消耗 run-rate 达到 $30M。Series A 文章还披露,平台已收集 50M+ 社区投票、发布 145k+ 开源对战数据点,并完成 400+ 次跨模态模型评估。值得注意的是,平台增长看起来主要由拉动驱动:TechCrunch 播客提到,LMArena 排行榜「became something of an obsession among model makers」,意味着社区客群的获客成本较低。 [CU001, CU002, CU003, CU004, CU007, CU008]

客户增长与采用轨迹
指标数值日期来源置信度含义
月活用户3M+Sep 2025来源:arena.ai/blog/ai-evaluations/中(公司披露)商业化上线时,社区规模已经明显放大
月活用户5M+Jan 2026来源:PRNewswire / arena.ai/blog/series-a/高(多源)不到 2 年 MAU 增长 5 倍;自然拉力强
月度对话60M+Jan 2026PRNewswire / TechCrunch Jan 2026高(多源)平均每名用户每月产生约 12 次对话
累计对话250M+Jan 2026来源:PRNewswire / arena.ai/blog/ai-evaluations/中(公司披露)历史数据长尾丰富,增强偏好模型质量
月度投票2M+Sep 2025 估计来源:arena.ai/blog/ai-evaluations/中(公司披露)相对用户数的参与率较高;约 0.7 票 / 用户 / 月
社区累计投票50M+Jan 2026来源:arena.ai/blog/series-a/中(公司披露)累计投票语料支撑 BT 模型校准
已评估模型(公开)400+Apr 2025 / Jan 2026来源:arena.ai/blog/two-year-celebration/, arena.ai/blog/series-a/中(公司披露)覆盖面越宽,平台中立性的说法越站得住
发布前模型测试300+Apr 2025来源:arena.ai/blog/two-year-celebration/中(公司披露)实验室参与度强;也是操纵争议风险的来源
地理覆盖150+ 个国家Jan 2026PRNewswire中(公司披露)全球多样性增强偏好数据代表性
ARR 运行率年化 $30MDec 2025TechCrunch Jan 2026 / PRNewswire高(多源)从 Sep 2025 的零收入快速爬坡,验证商业模式
社区增长率同比 25 倍May 2025–Jan 2026来源:arena.ai/blog/series-a/低-中(公司口径)单个数据点;指标分母不清楚

除标注估计外,数值均为公司披露。「高」置信度需要两个独立来源。公开披露没有给出环比轨迹。

[CU001, CU002, CU003, CU007, CU008, CU009]
FU002: 采用与部署漏斗

从发现到社区投票、模型提交,再到商业评估:LMArena 的互动漏斗。

月访客是粗略估计;投票转化率由披露的 MAU 和月投票数推导。商业客户数是下限(仅计入具名客户)。

[CU001, CU002, CU007, CU008, CU009]

6.3 具名客户证明与生产证据

AI Evaluations 商业产品的具名付费客户证据有限,但指向明确。2026 年 1 月 PR Newswire 新闻稿称:「LMArena works with leading AI labs and enterprises, including OpenAI, Google and xAI, all drawing on LMArena's evaluations to improve their models for production use cases.」这是公司直接点名客户,也是目前最强的公开客户证明。另一个证据来自 Felicis 合伙人 Peter Deng,他在同一新闻稿中表示:「We're leading this round because LMArena has built the most trusted, reliable, real-world signal of AI performance. They have become essential infrastructure for every lab and enterprise.」公司的 Agent Arena 方法论博客记录了真实用户会话——一个实时体育电视赛程网站、一个自托管电影待看片单、一个自主水下航行器自动驾驶——作为生产式部署示例,不过这些是社区用户,不是付费企业客户。商业化之前,Google 的 Kaggle、a16z 和 Together AI 曾向 LMSYS(前身非营利组织)捐赠算力和 grant。lmsys.org 的历史上线记录显示,最初 2023 年的合作结构包括 UC Berkeley。尚未确认付费的模型提供方包括 Anthropic 和 Meta:二者经常因旗舰模型上榜而被提及,但 TechCrunch 2025 和 2026 年文章一直只把 OpenAI、Google 与 Anthropic 称为「partners」,而新闻稿中具名商业客户只有 OpenAI、Google 和 xAI。AI Evaluations 服务没有公开可得的政府采购记录、G2/Capterra 评价或独立第三方案例研究。 [CU005, CU006, CU012, CU013, CU014, CU015]

具名客户验证表
客户 / 用户分群部署 / 用例生产 vs. 试点成果证据证据限制
OpenAIAI 实验室(付费商业客户)AI Evaluations——评估 GPT 模型变体的人类偏好;旗舰模型登上公开排行榜生产(新闻稿具名)Jan 2026 Series A PRNewswire 新闻稿将其列为客户,并称其把评估用于生产没有 OpenAI 直接引述;仅为公司披露
GoogleAI 实验室(付费商业客户)AI Evaluations——评估 Gemini 模型变体;旗舰模型登上公开排行榜生产(新闻稿具名)Jan 2026 PRNewswire 新闻稿具名;Gemini 模型持续进入文本和视觉排行榜没有 Google 直接声明;公司披露
xAIAI 实验室(付费商业客户)AI Evaluations——评估 Grok 模型变体的人类偏好生产(新闻稿具名)Jan 2026 PRNewswire 新闻稿与 OpenAI、Google 一同具名三个具名客户中最新的一家;未披露产品使用细节
AnthropicAI 实验室(模型提供方;付费状态未确认)Claude 模型登上公开排行榜;据 TechCrunch 2026,模型提交为「合作」关系活跃生产(排行榜)Claude 被描述为赢得法律 / 医疗用例专家排行榜(TechCrunch 播客);合作伙伴关系已确认没有任何来源单独确认其为 AI Evaluations 付费客户
MetaAI 实验室(模型提供方;付费状态未确认;有负面历史)Llama 模型登上公开排行榜;Jan–Mar 2025 测试了 27 个发布前 Llama 4 变体活跃生产(排行榜)Cohere/Stanford 论文确认了大规模发布前测试;Llama 模型定期在排行榜更新操纵事件:实验性 Maverick 提交后排名第 #2;公开版本排名约第 32;关系受到负面影响
PerplexityAI 实验室(模型提供方;搜索竞技场合作伙伴)Sonar 模型登上 Search Arena 排行榜;引用风格协作活跃(排行榜)Search Arena 评估了 11 个 Perplexity 模型变体;风格随机化与提供方协作完成没有 AI Evaluations 付费订阅证据

具名付费客户(OpenAI、Google、xAI)来自 January 2026 PRNewswire 新闻稿。模型提供方参与(Anthropic、Meta、Perplexity)来自新闻文章和博客文章,但不能确认 AI Evaluations 商业订阅。公开渠道没有独立客户引述或证言。

[CU005, CU006, CU012, CU013, CU014, CU015]
FU003: 客户证据质量矩阵

按维度评估每个具名客户或模型提供商的证据质量。

证据质量评级反映截至 2026 年 6 月可得的最佳公开证据;“未确认”不等于“否”。

[CU005, CU006, CU012, CU013, CU014, CU021]

6.4 留存、耐久性与重复使用

LMArena 尚未公开披露 AI Evaluations 产品的正式留存指标,例如净收入留存(NRR)、总收入留存(GRR)或客户流失率。商业化时间还太短(2025 年 9 月上线),没有多年 cohort 数据。社区端的主要留存信号是平台活跃度:按 Series A 博客,2025 年 5 月至 2026 年 1 月社区同比增长 25x;250M+ 累计对话相对 60M/月活跃生成量,说明累计使用的大部分来自持续扩大的活跃社区,而非一次性访问。AI Evaluations 采用基于消耗的模型(「annualized consumption rate」而非订阅 ARR),因此留存取决于实验室和企业是否持续提交评估请求。不到四个月从 $0 爬到 $30M 年化,显示初始实验室需求强劲;但实验室发展内部评估能力后,这种需求能否持续,是关键风险。排行榜 changelog 显示模型持续新增(2026 年 6 月每周数个),可作为模型提供方持续参与的间接代理。对于一家估值 $1.7B、ARR run-rate 为 $30M 的公司而言,没有披露客户 logo 墙、案例研究站点或 testimonial 页面(除 Felicis 与 UC Investments 的投资人引述外),这一点值得注意。 [CU007, CU009, CU010, CU011, CU016, CU017]

留存、重复使用和满意度指标
指标数值 / 状态分群置信度尽调要求
Net Revenue Retention (NRR)未披露AI Evaluations 商业客户Unknown向 LMArena 索取实验室客户购买后第 1–6 个月的 NRR 或 cohort 分析
Gross Revenue Retention (GRR)未披露AI Evaluations 商业客户Unknown确认 Sep 2025 上线以来是否有商业客户流失
月度用户留存(社区)未披露;60M 月度对话和持续增长暗示较高社区(免费用户)低(推断)索取月活用户留存或会话频次数据
客户流失(商业)未披露;上线未满 4 个月,流失数据极少AI Evaluations 商业客户Unknown跟踪 OpenAI/Google/xAI 是否在下个周期续约
重复提交模型(模型提供方)300+ 次发布前测试和 400+ 次公开评估,意味着约 20+ 家实验室持续重复参与模型提供方(免费 + 付费)低-中(由平台数据推断)获取提交模型超过一次的独立提供方数量
间接留存信号:排行榜活动根据 changelog,June 2026 每周新增多个模型所有模型提供方中(观察)验证活跃模型新增与提供方满意度相关,而不是营销压力
用户满意度(社区)未正式披露;没有发布 NPS 或 CSAT 分数社区用户Unknown如有发布,查找 NPS 调查或用户评价数据

标注「未披露」的数值截至 June 25, 2026 未出现在公开来源中;这不代表数值为零。置信度评级反映可得证据。

[CU016, CU017, CU018]

6.5 反向证据:冲突、博弈与集中度风险

LMArena 的客户和模型提供方关系带来有充分记录的结构性风险。第一是利益冲突:付费使用 AI Evaluations、充当商业客户的同一批 AI 实验室(OpenAI、Google、Anthropic 等),也是其模型在公开排行榜上被排名的主体。这天然激励实验室优化 Arena 表现,而它们可以通过预发布测试协议做到这一点。TechCrunch 播客直接提出这个问题:「how a team like theirs can build a neutral benchmark when the companies they're ranking are also their backers.」第二是 Meta Llama 4 Maverick 事件(2025 年 4 月):Meta 向 Arena 提交了一个「chat-optimized experimental」版本的 Maverick,拿到第 2 名,但公开发布模型大约排在第 32。LMArena 更新政策并表示「Meta's interpretation of our policy did not match what we expect from model providers.」该事件说明,依赖实验室提交模型的商业合作结构,会给平台带来实际压力,去容纳损害排行榜完整性的行为。第三,Cohere/Stanford/MIT/AI2 研究(2025 年 4 月)指称,Meta 可在 2025 年 1 月至 3 月间私下测试 27 个模型变体,再选择表现最好者提交;顶级实验室获得不成比例的采样率;研究初步发现分享给 LMArena 后并未被其否认。LMArena 否认具体指控,但宣布了新的采样算法。这些指控尚未经过独立裁定。第四,客户集中:公开具名的付费客户只有三家(OpenAI、Google、xAI),LMArena 的商业收入高度集中;若其中任何一家实验室自建或签约替代评估供应商,对 ARR 的影响可能显著。 [CU005, CU006, CU012, CU019, CU020, CU021]

扩张、集中度风险和负面动态
扩张驱动 / 集中度风险影响证据尽调路径
客户集中度:3 家具名付费客户高——如果 OpenAI、Google 或 xAI 减少或取消 AI Evaluations 支出,单一客户收入冲击可能很重新闻稿只具名 3 家客户;没有其他具名商业客户确定每家具名客户收入占比;确认未具名商业客户数量
利益冲突:评分对象 = 付费方高——付费购买评估的实验室也是模型被排名的主体,带来操纵激励和可信度风险Meta 操纵事件(Apr 2025);Cohere/Stanford 论文指称存在优先访问结构性防护审计:索取采样算法和发布前测试限制细节
登陆扩张:一次性评估 → 重复合同中——如果实验室认可评估价值,按量消费模式支持扩张;没有多年合同证据CEO 提到「consumption rate」;$30M ARR 快速爬坡暗示重复使用向 LMArena 确认合同结构(现货 vs. 订阅 vs. 年度)
非实验室企业扩张中——潜在市场很大,但目前没有具名的非实验室企业客户Series A 新闻稿提到「enterprises」;TechCrunch 文章提到「software engineering, law, medicine」等目标领域识别任何正在推进的非实验室企业试用或试点
竞争替代风险:实验室内部搭建评估能力中——如果头部实验室(Google、OpenAI)搭出能满足自身需求的内部评估流水线,LMArena 商业需求可能见顶Scale AI 的 SEAL Showdown(Mashable 文章)作为竞争性评估产品上线跟踪实验室关系的规模 / 体量随时间增长还是趋稳
未来操纵事件带来的负面声誉风险高——如果再次出现高关注度操纵事件,社区信任和商业可信度可能同时快速流失Meta Llama 4 事件被 The Verge、TechCrunch 等广泛报道;LMArena 更新了政策,但结构性冲突仍在监测新的操纵指控;在下个评估周期复核更新后的采样政策

集中度和利益冲突评估基于截至 June 2026 的公开来源证据。

[CU019, CU020, CU021, CU022, CU023, CU024]
FU004: 社区采用增长:关键里程碑

离散采用里程碑显示,LMArena 社区从 2023 年到 2026 年 1 月的增长轨迹。

数值混合了不同指标(累计投票、数据集提示、MAU)来展示轨迹;彼此不能直接比较。MAU 是所列日期的月活跃用户;投票语料为累计值。240k 是 2024 年 3 月论文发布时的累计投票;截至 2026 年 1 月,累计投票报告为 50M+。

[CU001, CU007, CU008, CU009, CU010]

6.6 附录

Chapter 07

07风险

7.1 利益冲突与基准完整性风险

LMArena 的核心风险,是独立裁判身份与收入模型之间的结构性冲突。公司向同一批 AI 实验室——OpenAI、Google、xAI——收取付费评估服务费用,而这些实验室的模型又出现在其公开排行榜上。这种双重角色会制造真实或被感知的偏袒激励。来自 Cohere、Stanford、MIT 和 AI2 的学术研究人员在《The Leaderboard Illusion》(arXiv 2504.20879)中记录了这些动态:少数大型提供商可以私下测试多个模型变体,只选择披露最佳分数,并获得不成比例的更多评估对战。Meta 的 Llama 4 Maverick 事件进一步说明,即使没有明确违规,政策缝隙也足以让基准博弈发生:Meta 向 LMArena 提交了一个针对对话性优化、从未公开发布的模型,拿到前二排名,直到差异被公开曝光。LMArena 已在 2026 年 4 月更新政策,但商业依赖与中立感知之间的根本张力仍然存在。包括独立研究员 Gwern 在内的评论者称该排行榜是「a cancer」,并质疑其信号是否仍有科学价值。ucstrategies.com 分析发现,专门为 Arena 偏好调优的模型最多可把分数抬高 100 Elo;SurgeAI 分析则发现,评估者有 52% 的时间不同意 LMArena 投票。一旦可信度叙事永久裂开——例如出现高关注度偏见调查、付费实验室丑闻,或监管机构调查 AI 基准准确性声明——LMArena 面向消费者的排行榜和企业评估收入两端的产品价值主张会同时坍塌。 [CR001, CR002, CR003, CR004, CR005, CR006]

监管 / 法律风险登记表
风险 / 规则 / 案件司法辖区状态可能性严重性缓释剩余暴露尽调路径
GDPR 数据共享充分性——在没有明确细粒度同意的情况下与 AI 提供方共享用户提示EU / EEA活跃合规义务;政策于 Sep 2025 更新关键已更新隐私政策;数据共享前使用 GCP Sensitive Data Protection API如果 EU DPA 审计数据共享实践,存在执法风险;训练后模型的删除权存在缺口索取 DPA 往来函件;复核与 AI 提供方的数据处理协议
EU AI Act——GPAI 模型透明度文档要求EUGPAI 义务自 Aug 2025 生效;高风险系统义务于 Aug 2026 生效学术论文和开源 Arena-Rank 仓库记录了方法论公开渠道未找到正式 GPAI 文档证明;不清楚 LMArena 是「systemic risk」模型提供方还是下游部署方索取 LMArena 的 EU AI Act 合规备案或自评文件
Digital Services Act——规模扩大后可能触发 VLOP 义务EU平台达到阈值(>45M EU 月活用户)时适用低-中根据公开数字,截至 June 2026 平台低于 VLOP 阈值用户规模逼近 VLOP 触发线后,系统性风险评估、独立审计和透明度报告将成为强制要求监测 MAU 披露;聘请 DSA 法律顾问
美国州隐私法——CPRA、VCDPA、CPA 等美国(多州)需要持续合规隐私政策承认州法权利章节未找到外部隐私审计或 CPRA 证明;California AG 在 2026 年执法活跃索取隐私审计;验证 CPRA 合规计划
知识产权——用户提示版权和训练数据来源全球法律未定;AI 版权诉讼在 2025–2026 年全行业活跃服务条款授予 LMArena 使用用户内容的许可;提示与提供方共享如果用户提交内容在没有充分许可的情况下用于模型训练,存在版权侵权暴露复核服务条款;审计到提供方训练流水线的数据使用流
基准准确性主张——FTC 广告真实性风险美国截至运行日期没有已知执法行动方法论已开源;学术发表披露了统计局限任何声称 Arena 排名等同于「real-world performance」的商业表述,都可能受到 FTC Section 5 审查监测 FTC AI 指引;复核营销主张措辞

所有可能性 / 严重性评级都是作者基于公开证据和适用监管框架的评估;LMArena 没有公开披露官方监管往来函件。各行按严重性排序(关键 → 中)。截至 2026-06-25 的诉讼检索未发现将 Arena Intelligence, Inc. 列为当事方的活跃案件。

[CR007, CR008, CR009, CR010]
FR001: 风险热力图——按影响与可能性拆解 LMArena 风险清单

这张按严重度和可能性交叉的热力图,将八项主要风险映射到五档影响和四档可能性,显示利益冲突与 Goodhart 定律风险占据右上象限。

可能性和影响评级是作者基于公开证据和监管先例作出的评估;LMArena 未披露内部风险台账。

[CR001, CR007, CR012, CR017]

7.2 监管、法律与数据隐私风险

LMArena 面对复杂且快速变化的监管环境。法律实体 Arena Intelligence, Inc. d/b/a LMArena 处理来自 150 个国家、500 万月度用户的个人数据,因此触发欧盟《通用数据保护条例》(GDPR)、欧盟 AI Act(通用 AI 模型义务自 2025 年 8 月起生效)、Digital Services Act(规模达到时可能触发 VLOP 门槛),以及包括 California CPRA 在内的一系列美国州隐私法义务。LMArena 自身隐私政策(2025 年 9 月生效)明确提醒用户,prompt 和生成回复可能被公开共享,这带来 GDPR 第 7 条下用户同意充分性的风险。已训练评估模型的删除权挑战(GDPR 第 17 条)是整个 AI 行业已知合规缺口,目前没有明确的执法解释。欧盟 AI Act 关于通用 AI 文档和透明度的要求,也会增加运营负担。LMArena 将用户 prompt 分享给 AI 提供商用于模型改进,这种做法带来数据流义务,必须在各司法辖区用数据处理协议覆盖。欧盟 AI Act 对「高风险」AI 系统的定义尚未明确点名 AI 基准测试平台,但法律、医疗、就业等受监管领域的评估可能引来行业特定审查。截至运行日期,公开记录未发现直接涉及 LMArena 或 Arena Intelligence, Inc. 的诉讼;不过,公司尚未披露其合规认证状态或任何监管问询。用户提交 prompt 被用于训练带来的 IP 风险,在多个司法辖区仍是未定法律问题。 [CR007, CR008, CR009, CR010, CR011]

运营和安全风险登记表
失效模式可能性严重性缓释成熟度剩余暴露未解决缺口
AI 提供方 API 撤回——实验室因竞争争议或评分抗议撤销访问关键低(未披露与提供方的 SLA)被撤模型在完整排行榜中留下缺口;若为付费客户则带来收入风险未披露公开 API SLA 或冗余计划
国家行为体或竞争对手发起协同女巫 / 投票操纵活动中等(匿名投票;部分 bot 检测;重加权算法)分数污染很难实时发现;事后需要修订方法论没有公开异常检测披露;Battles in Direct 的位置偏差校正显示其具备被动响应能力
重大数据泄露——社区提示数据集或企业评估结果泄露低-中低-中(GCP Sensitive Data Protection;标准企业安全)用户信任受损;可能触发 GDPR 泄露通知义务;评估 IP 暴露给竞争对手截至运行日期未披露 SOC 2 或 ISO 27001 认证
云基础设施故障——battle 匹配或 API 路由持续停机中等(云提供商 SLA;假设多区域部署但未确认)实时评估连续性中断;可能违反企业 SLA未找到公开状态页 SLA 或正常运行时间承诺
方法论过拟合——排名与真实世界模型质量的偏离加速中等(开源方法论;学术论文发表;政策更新)排行榜失去可信信号;开发者和企业流失Leaderboard Illusion 论文后,没有发布评分方法论外部统计审计
模型蒸馏攻击——对手系统性挖掘 Arena 评估分布低(数据按设计部分公开;未披露反爬控制)竞争信息泄露;Arena 数据优势被侵蚀AI 安全文献记录了蒸馏风险;LMArena 未披露应对措施

缓释成熟度评级仅基于公开可得信息。截至 2026-06-25,LMArena 未发布内部安全或基础设施审计结果。各行按严重性排序。

[CR012, CR013, CR014]
FR002: 风险传导图——LMArena 风险如何流向收入与可信度

这张有向无环图展示,根因风险如何传导到收入、用户参与、监管暴露和估值等下游影响。

[CR001, CR003, CR007, CR017]

7.3 运营、技术与安全风险

LMArena 的运营风险集中在平台可靠性、数据质量完整性,以及社区生成评估数据集的安全。每月 6000 万次对话流经其基础设施,一旦持续宕机或发生数据泄露,会直接伤害 LMArena 区别于静态基准的实时、连续评估主张。平台高度依赖云基础设施(同时运行多个实时 AI 模型 API 的算力成本),其价格和可用性由第三方控制,其中包括 LMArena 正在评估的同一批 AI 实验室。如果某个提供商撤回 API 访问——当实验室质疑评分或竞争格局变化时,这并非不可行——排行榜会立刻出现缺口。统计完整性是另一项运营风险:LMArena 的 Elo/Bradley-Terry 评分依赖足够均匀且无偏的对战分布;任何系统性 prompt injection、协同投票活动或针对平台用户基础的女巫攻击,都会污染排名信号,而且未必能被实时发现。2026 年 5 月转向「Battles in Direct」后,10% 的 direct-chat 会话生成对战投票,需要事后校正位置偏差和同组织指标偏差,这说明方法论更新会在事后改变模型排名。评估数据集是具有商业价值的资产,如何防范模型蒸馏攻击或未经授权抓取,仍是持续担忧;尤其是领先 AI 实验室有财务动机研究评估分布。约 41 人的员工规模,对于这一量级的平台形成了关键人集中风险。 [CR012, CR013, CR014, CR015, CR016]

合作伙伴和依赖风险登记表
依赖交易对手角色集中度失效场景严重性缓释剩余暴露
AI 实验室收入——前三大付费实验室(OpenAI、Google、xAI)OpenAI / Google / xAI 客户主要商业客户高(估计占 $30M ARR 的 >50%)客户退出评估合同;排行榜数据冲突关键多实验室客户基础;开源社区中立性承诺如果一家 Tier-1 实验室流失,收入可能断崖式下滑
AI 提供方 API 访问——battle 中所有模型都需要实时 API所有主要实验室模型能力提供方高(每家实验室控制自己的 API)实验室撤回 API 访问权限;模型从排行榜下线API 访问历史上较稳定;政策承诺纳入公开模型没有合同锁定保障;任何实验室都可短期通知后撤回
云计算基础设施GCP / 主要云服务商计算、存储、数据保护服务高(明确提及 GCP Sensitive Data Protection)云服务中断、涨价或条款变化标准云企业协议;假设已有冗余若依赖单一云,整个平台可能停摆
UC Berkeley / 学术研究管线UC Berkeley SkyLab人才来源;研究背书;过往资助记录核心研究人员转向行业;学术合作降温团队已独立注册公司;研究发表仍在继续若学术切分演变为对立,可信度会受损
投资人关系 — Felicis(领投)与 UC InvestmentsFelicis / UC Investments资本提供方;战略锚点中(Felicis 的 Peter Deng 曾任 OpenAI 员工)若被评测公司与 Arena 排名发生争议,投资人可能出现利益冲突治理结构未公开披露;Felicis GP 的 OpenAI 背景带来观感风险若投资人冲突公开浮出,独立性认知会受损

收入集中度估算基于约 100 个客户和已披露 $30M ARR,并假设早期 B2B SaaS 常见的帕累托分布。按 2026-06-25 的公开信息,LMArena 与 AI 提供商之间尚未披露合同条款或正式 SLA。

[CR017, CR018, CR019]
FR003: 依赖地图——关键平台依赖

展示 LMArena 对云服务商、AI 实验室 API、学术人才管线和投资方的关键基础设施与商业依赖。

[CR013, CR018]

7.4 合作伙伴集中、财务与执行风险

LMArena 的商业模式依赖少数大型 AI 实验室和企业,贡献其 $30M 年化收入的大部分。2026 年初约有 100 家付费客户,头部客户为 OpenAI、Google 和 xAI,收入集中度具有实质性:失去一两家 Tier-1 实验室关系,就可能带来不成比例的收入下滑。同一批实验室既是 LMArena 最好的客户,也是最有动机博弈其排行榜的潜在对手,这一张力没有干净解法。$1.7B 估值对应 $30M ARR 约 57x 收入倍数,定价中包含很高的增长预期:公司必须守住信任、扩展新领域,并向现有客户 upsell。41 名员工的执行产能,相对 LMArena 已承诺的产品面(Agent Arena、WebDev Arena、Vision、Video、Search、Code、Document 排行榜,加商业评估)偏薄。公司起源于 UC Berkeley 研究项目,创始团队的学术关系和发表义务可能分散管理层对商业执行的注意力。现金消耗动态未披露;以 $250M 融资和 41 人团队看,runway 似乎较长,但每月 60M 次对话的基础设施成本很高。未来融资风险中等但并非为零:如果 AI 评估市场未能增长到 2030 年预计的 $3.8B,后续轮次可能会反映失望定价。$382,500 的 Recall Capital-LMArena SEC Form D 备案(2026 年 2 月)反映了二级市场投资者需求,但不能证明公司自身财务健康。 [CR017, CR018, CR019, CR020, CR021]

人员与执行风险登记表
角色 / 职能依赖或缺口可能性严重性缓解措施尽调路径
CEO — Anastasios Angelopoulos创始 CEO;公开发言人;学术可信度锚点;所有战略关系都经由他推进低-中关键联合创始人 Wei-Lin Chiang 任 CTO;Ion Stoica 任顾问接班计划和关键人条款未公开披露;需索取董事会文件
统计方法团队核心 Elo / Bradley-Terry 评分能力集中在约 41 人的小型学术团队开源 Arena-Rank 仓库;学术论文提供外部校验确认方法团队规模和梯队深度;要求披露员工人数
企业销售与客户成功尚未公开任命企业销售负责人;商业产品于 2025 年 9 月推出ARR 快速增至 $30M,说明早期 GTM 已有牵引索取组织架构图;确认 VP Sales 招聘进度
安全与合规负责人41 人团队中未看到 SOC 2 / GDPR DPO / EU AI Act 合规岗位的公开证据已有隐私政策;提及 GCP 数据保护工具核实 DPO 指定情况;索取合规组织结构
研究科学家留存人才市场竞争激烈;前沿 AI 实验室薪酬为创业公司的 2–3 倍股权;使命驱动文化;UC 关系审查期权池规模;留存悬崖日期

41 人员工数来自 Latka(2026 年 1 月)。组织架构图未公开。接班和关键人安排未公开披露。各行按严重性排序。

[CR020, CR021]
缓解措施与否决标准
风险可监测触发项阈值 / 事件行动含义
利益冲突导致可信度坍塌Tier-1 AI 媒体报道情绪;批评 Arena 方法论的学术论文数量单季度出现两篇或以上同行评审论文,记录 LMArena 未能反驳的系统性偏差投资逻辑破裂:退出或暂停评估;只有方法论经独立审计后才重新开启
顶级实验室客户流失与 OpenAI、Google、xAI 的商业评测协议续约情况前三大实验室中任一家公开终止或公开质疑评测结果监控合同续约日期;若续约存疑则升级处理
监管执法行动披露 EU DPA 询问、FTC 信息要求或州 AG 调查针对 Arena Intelligence, Inc. 的任何正式监管调查启动投资逻辑破裂:冻结新增资本部署,等待法律解决
开发者社区放弃基准可信度GitHub stars、HuggingFace 排行榜引用,以及开发者论坛对 Arena 方法论的提及模型发布公告中的 Arena 引用率同比下降 >30%投资逻辑预警:调查根因并评估缓解空间
收入集中度断崖董事会材料披露前三大客户 ARR 占比单一客户超过 ARR 的 40%,并释放不满信号投资逻辑预警:下一轮融资前必须实现多元化
关键人物离职Angelopoulos 或 Chiang 的公开公告或 LinkedIn 更新任一创始人在 Series A 交割后 24 个月内离职投资逻辑破裂:立即暂停,等待接班安排清晰

触发阈值是作者基于同类早期 AI 基础设施公司尽调实践设定的启发式标准。所有触发项都需结合语境判断——单一负面事件不一定单独击穿投资逻辑。

[CR003, CR006, CR017]

7.5 附录

Chapter 08

08估值

8.1 融资与估值背景

Arena Intelligence, Inc.(d/b/a LMArena)在 2026 年 1 月宣布完成 $150M Series A,投后估值 $1.7B。此前公司于 2025 年 5 月以 $600M 估值完成 $100M 种子轮,约七个月内累计融资达到 $250M。Series A 由 Felicis 和 UC Investments(University of California)领投,Andreessen Horowitz、The House Fund、LDVP、Kleiner Perkins、Lightspeed Venture Partners 和 Laude Ventures 参投。公司年化「consumption run rate」——公司称其等同于 ARR——在 2025 年 12 月超过 $30M,距离 2025 年 9 月推出首个商业产品(AI Evaluations)不到四个月。按 $30M run rate 计算,投后估值 / ARR 倍数约为 57x,而晚期 Series A SaaS 公司通常在 8-15x。不过,四个月内从 $0 ARR(2025 年 5 月)到 $30M ARR(2025 年 12 月)的轨迹极其罕见,投资人定价的是未来增长,而不是当前收入。SEC EDGAR Form D(accession 0002113470-26-000001,2026 年 2 月提交)中「Recall Capital-LMArena」二级基金规模为 $382,500,显示二级市场需求,其隐含估值与 Series A 价格一致。股权结构细节——从种子轮到 Series A 的稀释(种子轮约出售 17%,Series A 约出售 9%)——基于披露的投后估值和轮次规模,隐含 Series A 前企业价值为 $1.55B。未发现可转债、优先清算权或既往 down-round 风险的公开证据;公司八个月内估值 step-up 为 2.83x。 [CV001, CV002, CV003, CV004, CV005, CV041]

推荐摘要表
维度评估备注
推荐跟踪转为买入取决于方法论审计 + 入场价 <30x NTM ARR
置信度证据足以形成判断;核心风险是结构性的,且仍未解决
风险评级利益冲突 + 极高倍数 + 未审计治理
估值立场偏高57x ARR 倍数位于最高十分位;$1.7B 估值下安全边际不足
决策含义不按当前价格领投或共同领投;继续跟踪治理改善若 ARR 达到 $57M+,或入场价降至 $1.0–1.2B,则重新评估

评估截至 2026-06-25,依据公开披露的融资数据(Series A 投后估值 $1.7B、截至 2025 年 12 月 ARR $30M)和公开风险证据。未审阅内部财务模型、董事会材料或数据室。

[CV001, CV015]
正反投资逻辑表
维度正向投资逻辑反向投资逻辑什么会改变判断
市场结构前沿模型快速增多,中立评测基础设施变成必需品;目前没有可信的独立替代者能做到同等规模AI 实验室会逐步把评测内化,或组建联盟以绕开第三方费用实验室自建评测达到可比覆盖(5M+ 用户);或 LMArena 失去前三大付费实验室中的 2 家
网络效应2025 年 5 月以来,5M 月活用户 / 60M 次对话 / 400+ 个模型评测形成数据护城河社区投票可被操纵;Elo 分数会变成优化代理,而不再代表质量2026 年批评 Arena 方法论的学术论文数量翻倍;开发者引用率下滑
利益冲突透明度、开源方法论和更新后的政策缓解了风险;丑闻之后,社区信任仍然守住同行评审研究已有记录;只要收入依赖被评测实验室,根本张力就无法消除独立审计确认方法论完整性;或来自非实验室客户的收入多元化超过 ARR 的 60%
收入轨迹从 $0 到 $30M ARR 只用 4 个月,表现异常强;说明 AI 评测有强产品市场匹配收入集中在约 100 个客户;前三大实验室很可能占多数;任何流失都会被放大客户 cohort 数据显示前三大集中度 <20%,且 NRR >110%
估值支撑可比 AI 基础设施公司的加权平均倍数可支撑高增长先发者 30-40x ARR57x ARR 且未披露盈利路径,相比企业 SaaS 中位数偏贵若到 2026 年底 ARR 增至 $75M+,当前价格隐含倍数会压缩至 23x

论点综合了公开来源中的证实与反向视角。未审阅公司内部材料。五行对应五个关键投资维度;每个维度都有独立证据支撑。

[CV011, CV012, CV013, CV014]

8.2 估值分析——多重视角

收入倍数分析:在 $30M ARR 和 $1.7B 估值下,ARR 倍数为 57x,把 LMArena 放在后期 AI 基础设施融资轮次的最高十分位。AI 评估基础设施的可比公司很少:Weights and Biases 于 2025 年 3 月被 CoreWeave 以 $1.4B 收购(训练 / 评估平台,工具链更宽);Scale AI 的估值已达数十亿美元区间,收入规模也大一个数量级。早期纯 AI 基准测试公司没有直接公开可比。最相关的类比是 AI「卖铲人」基础设施:LMArena 作为中立数据层的定位,类似 Bloomberg 或 S&P Global 服务金融市场——一个具备高切换成本的可信数据特许权。按 The Business Research Company,AI 模型评估市场 2025 年为 $1.86B,预计 2026 年达到 $2.36B,2030 年以 27.3% CAGR 增至 $6.24B。如果 LMArena 到 2028 年拿下 $3B 市场的 5-10%,收入为 $150-300M,按 15-20x 倍数,对应 $2.25-6B 估值——这是公开可比退出的 base-to-bull 区间。增长可持续性:四个月达到 $30M,隐含每月新增 $7.5M 的增长速度。即便未来 12 个月只维持该速度的 50%,也会新增 $45M ARR,到 2026 年底达到 $75M ARR。若倍数压缩到 25x(符合成熟企业 SaaS 特许经营),估值为 $1.875B——大致持平当前。若按 20x 倍数(高增长公开 SaaS 中位数),$75M ARR 对应 $1.5B。因此,当前 $1.7B 估值取决于:(a)维持快速 ARR 增长,(b)守住平台溢价倍数,或(c)两者兼具。下方情景分析更明确地展示结果分布。 [CV006, CV007, CV008, CV009, CV010]

牛 / 基准 / 熊情景表
情景2026 年底 ARR入场收入倍数隐含估值(2028 年退出)关键风险概率信号
$120M+基于 $120M 为 14x$3–6B(按平台 25-50x ARR)可信度保持;前三大实验室续约;方法论审计通过需要持续每月新增 $7.5M ARR,且零可信度事件
基准$60–75M基于 $70M 为 23–28x$1.4–2.1B(按 20-30x ARR)轻微信任摩擦;吸收 1–2 轮方法论批评周期需要当前 ARR 增速的 50%;考虑到 $250M 资本和 41 人团队,具备可行性
$15–25M(流失)基于 $20M 为 68–113x — 不可持续$200–500M(平台陷入困境)可信度坍塌;顶级实验室客户退出;竞争性评测平台夺走市场由一次重大公开调查或监管执法触发

ARR 预测从 2025 年 12 月 $30M 运行率外推,不反映任何内部预测或指引。估值区间使用公开可比 AI 基础设施倍数(The Business Research Company;Latka;二级市场报道)。概率信号列描述把某个情景成立所需的证据,并非正式概率估计。

[CV008, CV009, CV016]
可比估值表
可比公司类型收入 / ARR估值倍数与 LMArena 的相关性局限
Weights & Biases(CoreWeave 收购,2025 年 3 月)私营 → 被收购~$150M ARR(估计)$1.4B 收购价~9x ARRAI 开发者平台,具备模型评测和实验跟踪;产品上最直接的参照工具覆盖比 LMArena 更广;收购价不是独立估值;收入未披露
Scale AI私营,后期$500M+ 收入(2025 年报道)报道区间 $14–29B28-58x 收入为企业和政府提供数据标注、RLHF、模型评测;相邻市场Scale AI 做直接标注劳务;利润率结构不同;收入基础更宽
Hugging Face私营~$70M ARR(2024 年估计)$4.5B(2023 年融资)~64x ARRAI 模型枢纽,带社区评测功能;作为「中立」基础设施,品牌相关不是以基准评测为核心的业务;社区托管才是主产品;估值时间更早
Bloomberg LP私营~$6.5B 收入估值 $80–100B~13-15x 收入拥有可信中立评分的数据特许经营(如 Bloomberg Intelligence);长期类比规模、资产类别和市场结构差异很大;30+ 年特许经营
S&P Global(评级分部)上市公司(SPGI)评级分部 ~$4.5B市值 $130B+~29x 分部收入受监管认可的中立裁判,具备结构性护城河;数据特许经营的长期类比S&P 拥有监管授权和市场锁定;LMArena 处在未监管的基准评测市场
Arize AI(AI 监控)私营未披露$148M 融资(2024 年)N/A面向生产环境的 AI 模型监控平台;相邻评测市场产品焦点不同(监控 vs 基准评测);规模较小
Aporia Technologies私营未披露~$50M 融资(2024 年)N/A负责任 AI 监控与偏差检测;相邻治理市场聚焦合规;没有社区评测组件

所有可比财务数据均来自公开报道(Latka、TechCrunch、分析师来源)或可获得的 SEC 文件。私营公司收入为估算或第三方报道近似值。倍数按报道 / 估算数据计算,应视为方向性而非精确值。

[CV007, CV008]
FV002: 估值敏感性——ARR 增长 vs 退出倍数

展示在不同 ARR 增长情景($45M、$75M、$120M)和退出倍数情景(15x、25x、40x)下,2028 年隐含退出估值如何变化,覆盖从压力情景到牛市情景的区间。

ARR 预测由 2025 年 12 月 $30M 基线外推;退出倍数基于公开 AI 基础设施 SaaS 可比公司。数值单位为 USD 百万美元。不是财务模型输出,仅作示意。

[CV009, CV010, CV016]
FV003: 估值 / 回报区间——以 $1.7B 进入 LMArena

低 / 基准 / 高三档 2028 年退出估值区间及假设,展示以 $1.7B 入场与示意摊薄入场价对应的回报倍数。

所有数字均为 USD 百万美元(2028 年退出价值)。熊市情景假设可信度崩塌、ARR 流失至 <$25M;基准情景假设持续增长且倍数适度压缩;牛市情景假设突破式增长并维持平台溢价倍数。

[CV016]

8.3 投资论点与反论点

投资论点建立在三项耐久的结构性优势上:(1)社区网络效应——5M 月度用户生成 60M 次对话,形成数据护城河,即便投入大量资本,也极难从零复制;(2)关键判断角色的在位优势——排行榜排名会影响 AI 实验室营销、开发者采用和企业采购,为 LMArena 的付费客户创造切换成本;(3)市场时点——随着竞争性前沿模型增多、企业买家面对真实选择复杂度,AI 评估中的真实世界人类偏好数据会越来越有价值,而不是更少。反论点集中在四个担忧:(1)评估收入与被评估主体之间的结构性利益冲突(Leaderboard Illusion 论文已有记录)可能触发信任坍塌,同时摧毁公开排行榜和企业收入;(2)Goodhart 定律动态——当 Arena 成为事实标准,实验室会专门为它优化,削弱信号的真实世界相关性;(3)收入倍数极端,定价中包含既要持续执行、又要保持可信度的增长轨迹;(4)平台依赖风险存在:付费客户本身控制着平台运转所需的 API 访问,形成结构性杠杆失衡。总体看,在合适进入价格下,论点强于反论点;但当前 57x ARR 倍数在冲突未解决时留下的安全边际不足。 [CV011, CV012, CV013, CV014]

投资逻辑击穿与否决触发表
触发项阈值 / 事件对投资逻辑的传导行动含义
方法论独立性破裂同行评审论文或监管询问记录显示,商业收入影响了 LMArena 公开排行榜排名可信度根基坍塌;企业客户退出;估值跌至困境水平立即暂停 / 退出;若无补救,投资逻辑破裂
顶级实验室客户流失OpenAI、Google 或 xAI 任一家公开终止评测合同或移除模型收入断崖($30M ARR 中多数面临风险);社区覆盖出现缺口投资逻辑预警;升级前先评估集中度和替代管线
ARR 增长停滞截至 2025 年 12 月的 $30M ARR 到 2026 年 Q3 未能达到 $50M+意味着产品市场匹配比估值假设更窄;57x 静态倍数无法自洽重新评估投资逻辑;入场价应反映增长停滞倍数(20-25x)
开发者引用率坍塌AI 实验室公告中的 Arena 排行榜引用同比下降 >30%失去事实标准地位会削弱定价权和获客能力按季度监控;若趋势持续 2 个季度,投资逻辑实质转弱
监管执法行动针对 Arena Intelligence, Inc. 的任何 EU DPA、FTC 或州 AG 正式调查法律成本、整改负担和声誉损害;可能约束运营立即暂停;重新评级前评估严重性和司法辖区暴露
关键人物离职Anastasios Angelopoulos 或 Wei-Lin Chiang 在 Series A 交割后 18 个月内离职学术可信度锚点和客户关系流失;关键增长拐点面临接班风险投资逻辑破裂;暂停,等待接班和治理清晰

触发阈值是作者基于同类 AI 基础设施投资尽调实践设定的监控启发式标准。它们不反映公司指引。

[CV015, CV017, CV018]
FV001: 推荐逻辑流——从证据到确信度

从市场规模和社区验证出发,经过风险评估和估值证据,最终落到“跟踪,满足条件后买入”的建议。

[CV011, CV012, CV015]

8.4 建议、情景与最终尽调要求

总体建议为「track」,并在满足以下条件时具备转向「buy」路径:(a)收到回应 arXiv 2504.20879 发现的独立方法论审计;(b)以低于未来十二个月 ARR 30x 的估值进入(意味着若按当前 $1.7B 价格承诺投资,ARR 至少需达到 $57M);(c)商业合同条款确认,付费实验室客户相对于公开排行榜方法论没有获得偏好评估定价。鉴于利益冲突和倍数压缩风险,风险评级为「high」。对该建议的置信度为「medium」——证据足以形成观点,但关键风险因素具有结构性且未解决。估值立场是「stretched」。公司的轨迹、社区参与度和先发优势真实存在;但为这些优势付出的价格并不便宜。牛市情景要求到 2026 年底 ARR 达到 $75M+ 且可信度不破;熊市情景是可信度坍塌,收入和估值一起压缩至 $200-400M(类似其他行业的基准声誉危机)。持有期较长且有强治理影响力的投资人,如果以 30x forward ARR 或更低价格进入,在 base case 下有可信的风险 / 回报情景。 [CV015, CV016, CV017, CV018, CV019]

最终尽调问题表
主题缺失证据为什么重要负责人 / 尽调路径
独立方法论审计对 Leaderboard Illusion 论文发现和 LMArena 公开回应做外部统计复现核心可信度资产未经审计;自我证明不足以支撑投资承诺聘请独立学者或咨询机构;把交付结果作为投资条件
客户集中度与 NRR按客户层级拆分 ARR、前三大客户占比、净收入留存率和 cohort 流失数据汇总 $30M ARR 掩盖了极高估值下的集中度风险向数据室索取;若无法提供,考虑延长 90 天尽调
股权结构与清算优先权完整资本结构表,包括期权池、优先权条款、反稀释条款和 pro-rata 权利57x ARR 意味着持有期较长;清算顺位会显著影响回报情景标准 Series A 尽调;通常可在数据室获得
EU AI Act 合规状态EU AI Act(2025 年 8 月)要求的 GPAI 模型文档、数据来源记录和版权政策不合规会带来 EU 市场风险,并引发受监管行业企业客户顾虑索取法律意见;承诺投资前请 EU 律师审阅
基础设施与 SLA 协议云服务商协议(GCP 合同)、AI 提供商 API 访问协议和企业评测 SLA平台可用性和收入连续性依赖未披露的第三方协议在数据室索取;特别关注提供商终止通知期
员工数与招聘计划组织架构图,列出现有缺口、支撑每名员工 $10M+ ARR 的招聘计划,以及创始人关键人协议41 人支撑 $30M ARR 已很精简;扩至 $100M ARR 需要显著补组织在启动会议中向管理层索取

尽调问题按接近当前 $1.7B 估值入场的投资人优先级排序。第 1 和第 2 项被视为买入决策的阻断项。第 3–6 项对投资结构设计和交割后监控有实质影响。

[CV019]
FV004: 投资 KPI——IC 可用评分卡

围绕七个评估维度给出 IC 可用评分,反映截至运行日期支撑各维度的证据质量和完整度。

分数是作者仅基于可得公开证据按 1–10 分作出的判断,不是算法输出。治理分数因利益冲突未解决、缺少独立审计而受罚;估值分数因 57x ARR 倍数而受罚。

[CV015, CV012, CV013]

8.5 附录

免责声明

本报告基于截至 2026-06-25 的公开且可访问来源,不构成投资、法律或会计建议;在作出投资决策前,应补充管理层尽调、客户访谈和一手财务文件。

证据索引

结论
编号陈述可信度来源
CO001 Chatbot Arena was launched in May 2023 as a public demo by the LMSYS research group at UC Berkeley's Sky Computing Lab. SO010, SO011
CO002 The original Chatbot Arena platform was developed under UC Berkeley's LMSYS (Large Model Systems) research group, operating within the Sky Computing Lab. SO010, SO019
CO003 Anastasios Angelopoulos and Wei-Lin Chiang are co-founders of LMArena (Arena Intelligence Inc.). SO001, SO003
CO004 Ion Stoica, a UC Berkeley CS professor and co-founder of Databricks and Anyscale, is a co-founder of Arena Intelligence Inc. SO019, SO004
CO005 Arena Intelligence Inc. was formally incorporated on April 18, 2025. SO020, SO016
CO006 Anastasios Angelopoulos serves as CEO of Arena Intelligence Inc. SO003, SO001
CO007 LMArena's platform enables users to submit prompts to two anonymous AI models simultaneously, vote on the preferred response, and then see which models they compared — feeding a public leaderboard. SO001, SO010, SO011
CO008 LMArena's headquarters is located in the San Francisco Bay Area; the specific office address is not publicly disclosed. SO001, SO019
CO009 LMArena now operates under the brand name 'Arena'; it was previously branded 'LMArena' and before that 'Chatbot Arena.' SO001, SO018
CO010 LMArena raised a $100 million seed round in May 2025, co-led by Andreessen Horowitz and UC Investments, at a $600 million post-money valuation. SO005, SO004, SO006
CO011 LMArena raised $150 million in Series A financing in January 2026, co-led by Felicis and UC Investments, at a post-money valuation of $1.7 billion. SO003, SO004, SO002
CO012 LMArena's Series A investors include Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed Venture Partners, and Laude Ventures alongside lead investors Felicis and UC Investments. SO002, SO003
CO013 LMArena's total capital raised as of January 2026 is approximately $250 million across seed and Series A rounds. SO004, SO003
CO014 LMArena's $1.7 billion Series A post-money valuation represents nearly triple its $600 million seed-round valuation achieved approximately nine months earlier. SO003, SO004
CO015 LMArena's annualized consumption run rate surpassed $30 million in December 2025, approximately four months after commercial product launch. SO003, SO004
CO016 LMArena launched its commercial AI Evaluations product in September 2025. SO007, SO004
CO017 As of January 2026, LMArena reports more than 5 million monthly users across 150 countries. SO003, SO004
CO018 LMArena processes more than 60 million conversations per month across its evaluation platform. SO003, SO004
CO019 LMArena's community accumulated over 50 million votes across text, vision, web development, search, video, and image modalities by December 2025. SO002, SO003
CO020 LMArena has evaluated more than 400 distinct AI models including both open-weight and proprietary systems since founding. SO002, SO003
CO021 LMArena's pre-commercial research was supported by grants and donations from Google's Kaggle platform, Andreessen Horowitz, and Together AI, primarily in compute resources and cash. SO005, SO015
CO022 LMArena's leaderboard uses a Bradley-Terry model / Elo-based rating system with pairwise crowdsourced comparisons to rank AI models. SO011, SO012, SO010
CO023 LMArena released 145,000 open-source battle data points from expert and occupational evaluation categories in late 2025. SO002
CO024 LMArena's AI Evaluations commercial product includes SLAs with committed delivery timelines, representative feedback samples, and community-grounded performance analytics for model labs and enterprises. SO007, SO003
CO025 LMArena's named commercial customers include OpenAI, Google, and xAI, which use its evaluation services to improve their production models. SO003, SO004
CO026 LMArena has expanded beyond text evaluation to include Search Arena, WebDev Arena, Vision Arena, text-to-image, video generation, and Agent Arena as of June 2026. SO024, SO025, SO026, SO009
CO027 In April 2025, researchers from Cohere, Stanford, MIT, and Ai2 published 'The Leaderboard Illusion,' alleging LMArena systematically allowed certain AI companies to privately test multiple model variants and selectively disclose only top-performing scores. SO014, SO016
CO028 The Leaderboard Illusion paper identified that Meta privately tested at least 27 Llama 4 model variants on Chatbot Arena between January and March 2025, ahead of the public release. SO014, SO016
CO029 The Leaderboard Illusion paper estimated that Google and OpenAI each received approximately 19–20% of all Chatbot Arena battle data, while 83 combined open-weight models received only approximately 29.7% of total data. SO014
CO030 LMArena co-founder Ion Stoica publicly characterized The Leaderboard Illusion paper as containing 'inaccuracies' and 'questionable analysis,' and LMArena invited all model providers to submit more models for testing. SO016
CO031 Critics including researchers at the Allen Institute for AI and King's College London argued that LMArena's user base skews toward technical programmers and AI enthusiasts, making its benchmark unrepresentative of general end-user preferences. SO015, SO020
CO032 LMArena published updated transparency and leaderboard policies as of April 30, 2026, committing to open-sourcing evaluation pipelines and releasing portions of data to support auditing. SO008
CO033 Ion Stoica previously co-founded Databricks (valued at approximately $43 billion) and Anyscale (the Ray computing framework), providing the founding team with deep commercialization experience. SO019, SO016
CO034 LMArena reached a $1.7 billion valuation approximately three years after founding and within seven months of launching its first commercial product. SO004, SO003
CO035 LMArena's commercial business model charges AI labs and enterprises for evaluation services including community-grounded model performance analysis across domains such as software engineering, law, and medicine. SO003, SO007
CO036 LMArena launched Agent Arena in June 2026 with a causal inference methodology using treatment effect estimation to evaluate multi-component AI agents. SO024
CO037 In early April 2025, Meta submitted a specially arena-optimized, unreleased variant of Llama 4 Maverick to Chatbot Arena that ranked 2nd overall; when the standard public release was subsequently scored, it ranked 32nd. SO027, SO028
CO038 LMArena's deduplication system filters approximately 10% of submitted votes and its identity-leak detection pipeline removes fewer than 4% of votes, per July 2025 methodology updates. SO009
CO039 LMArena has not publicly disclosed its total headcount; the company's Series A press release stated funds would be used to expand the technical team. SO003, SO002
CO040 No debt facilities, convertible notes, secondary transactions, or SAFEs have been publicly disclosed for LMArena as of June 2026; financing has been entirely via equity rounds. SO003, SO004
CO041 The Leaderboard Illusion paper estimated that even limited additional Chatbot Arena battle data can yield relative performance gains of up to 112% on Arena-specific benchmarks, demonstrating the value of asymmetric data access. SO014
CO042 LMArena's Chatbot Arena paper (arXiv:2403.04132) reported crowdsourced human votes achieve over 80% agreement with expert rater judgments, establishing methodological validity for the platform's ranking approach. SO011, SO012
CO043 Arena Intelligence Inc. was spun out from UC Berkeley's LMSYS research group with UC Investments (University of California) serving as both lead investor and institutional backer of the spinout; the formal IP licensing terms governing transfer of the Chatbot Arena methodology from the university to the commercial entity are not publicly disclosed. SO005, SO020
CM001 The broad AI model evaluation platform market is sized at $1.86B in 2025 and $2.36B in 2026, implying 27.3% growth into 2026. SM001, SM002, SM003
CM002 The same broad market lens projects the AI model evaluation platform category to reach about $6.24B by 2030 at a 27.3% CAGR. SM001, SM003
CM003 A narrower model evaluation and benchmarking tools lens implies roughly a $0.85B market in 2026 at around 7.3% CAGR, materially below the broad platform estimate. SM004
CM004 Gartner forecasts worldwide AI spending to rise 47% to about $2.59T in 2026, creating a large upstream budget pool for evaluation software. SM005, SM013
CM005 Presenc AI reported that 78% of Global 2000 companies had at least one AI workload in production in Q1 2026, consistent with fast-rising demand for recurring evaluation and governance workflows. SM005, SM006
CM006 LMArena competes in the evaluation software layer that includes public benchmarking, release testing, domain scoring, and workflow-level quality measurement rather than core model training or hosting. SM014, SM016, SM017
CM007 Chatbot Arena introduced LMArena's core pairwise human-preference comparison model, making public leaderboard trust a foundational but not sufficient part of the company's market. SM014, SM015
CM008 Agent Arena broadens LMArena from single-turn chatbot ranking into agent and workflow evaluation, increasing the relevance of multi-step enterprise use cases. SM016, SM017
CM009 The sharp difference between broad and narrow market estimates is best explained by scope: broad reports appear to include wider enterprise evaluation-platform workflows, while narrow reports emphasize benchmarking tools only. SM001, SM003, SM004
CM010 Scale AI launched SEAL Showdown in 2026 with respondents across 100+ countries, 70+ languages, and 200+ professional domains, positioning it as a directly competitive benchmark product. SM008, SM009, SM025
CM011 Scale says GPT-5 tops all SEAL categories while Gemini 2.5 Pro leads most LMArena categories, showing that benchmark outcomes are sensitive to task mix and rater composition. SM008, SM009
CM012 Future AGI's competitor roundup places Galileo, Arize, Patronus, and MLflow in the same evaluation-tool conversation as LMArena, implying a fragmented specialist field. SM004, SM007
CM013 Precedence Research lists major platform and infrastructure vendors such as AWS, Google, Microsoft, IBM, Databricks, Hugging Face, and Arize AI among category participants, implying competition from bundled as well as standalone products. SM004, SM007
CM014 CoreWeave's roughly $1.4B acquisition of Weights & Biases in March 2025 shows strategic M&A appetite around evaluation-adjacent workflows such as experimentation, observability, and model quality. SM004, SM021
CM015 TechCrunch reported that LMArena reached a $1.7B valuation, about $30M ARR, and commercial customers including OpenAI, Google, and xAI four months after launching its product. SM012, SM024
CM016 LMArena says its commercial product targets law, medicine, and engineering workflows, signaling that high-stakes domain evaluation is central to its monetization strategy. SM012, SM016
CM017 Likely budget owners for evaluation software are AI platform leads, model quality owners, or domain product teams rather than generalized IT procurement alone. SM007, SM016, SM017
CM018 Frontier labs typically buy evaluation for model release, trust, and competitive benchmarking, while enterprises buy it for domain reliability, governance, and workflow quality. SM016, SM017, SM012
CM019 The buyer universe is concentrated at the frontier-lab end but much broader across regulated and expertise-heavy enterprise verticals, creating two different procurement motions. SM012, SM016, SM007
CM020 Law, medicine, and engineering are promising beachheads because domain failures there are expensive enough to justify paid evaluation rather than relying on public leaderboard performance alone. SM012, SM016
CM021 A practical LMArena-relevant 2026 SAM is a subset of the broad $2.36B TAM, likely concentrated in frontier labs plus evaluation-heavy enterprise verticals rather than the full surrounding AI-tooling economy. SM001, SM004, SM012, SM016
CM022 A plausible 2026 SAM range for standalone trusted evaluation software relevant to LMArena is about $0.3-0.8B after excluding most bundled hyperscaler, generic MLOps, and non-paid benchmarking activity. SM001, SM004, SM016
CM023 A plausible near-term SOM range for LMArena is about $30-150M because the company has reported roughly $30M ARR and could deepen spend within existing frontier-lab and high-stakes enterprise customers. SM012, SM013, SM016
CM024 The market should be treated as contested because a narrow benchmarking-tools lens can be roughly one-third or less of the broad platform estimate depending on what is included. SM001, SM003, SM004
CM025 Gartner's 2026 forecast includes about $453B of AI software spend and $32.6B of AI model spend, indicating that evaluation budgets need only capture a small fraction of upstream AI activity to support category growth. SM005, SM013
CM026 As AI workloads move into production at large enterprises, evaluation shifts from one-off benchmarking toward recurring regression testing, governance, and release gating. SM005, SM006, SM016
CM027 The rise of agentic AI increases evaluation demand because buyers need to measure multi-step task completion and causal workflow reliability rather than only single-response quality. SM016, SM017
CM028 Benchmark rivalry among frontier labs is itself a demand driver because model providers want external proof points to market releases and defend product claims. SM008, SM014, SM015
CM029 New benchmark launches such as SEAL Showdown confirm that benchmark design is now a contested product category rather than a settled research utility. SM008, SM009, SM018
CM030 Commercializing evaluation in law, medicine, and engineering makes the category more monetizable because quality signals are tied to high-cost business workflows instead of casual consumer usage. SM012, SM016
CM031 Even without exact public compliance budgets, regulated and high-stakes deployments are likely to sustain third-party evaluation demand because internal trust and audit requirements are higher than for generic chat use cases. SM016, SM017, SM005
CM032 The Leaderboard Illusion paper argues that benchmark contamination, hidden sampling choices, and over-interpretation of small score gaps can distort arena-style rankings. SM010, SM018, SM019
CM033 TechCrunch reported in 2024 that Chatbot Arena's user base and interaction style may bias results, weakening its usefulness as a universal benchmark. SM010, SM011
CM034 TechCrunch reported in 2025 that LM Arena faced accusations of helping top labs game its benchmark, creating a direct commercial trust risk for public-score-based products. SM018, SM020
CM035 The Verge's coverage of benchmark gaming around Meta's Llama 4 Maverick suggests that optimization against public benchmarks is ecosystem-wide rather than unique to LMArena. SM018, SM020
CM036 Critics of crowdsourced AI benchmarks argue that rater representativeness and benchmark ethics are material issues, implying that some enterprise buyers may prefer curated private evaluations over open-arena votes. SM019, SM011, SM010
CM037 Standalone vendors like LMArena likely face pricing pressure wherever evaluation is bundled into broader cloud, experimentation, or observability stacks. SM004, SM007, SM016
CM038 Public pricing and average contract values for LMArena and most direct competitors are undisclosed, preventing reliable bottom-up revenue-to-market-share validation from open sources. SM007, SM012, SM016
CM039 No verified public source in this run discloses what share of frontier-lab or enterprise AI budgets is actually allocated to evaluation tooling. SM012, SM016
CM040 Competitor revenue disclosure is sparse for Galileo, Patronus, Arize, Future AGI, and other specialist vendors, making independent startup-share mapping incomplete. SM004, SM007
CM041 LMArena's reported roughly $30M ARR suggests it may already represent a meaningful share of a narrow benchmarking-tools niche while remaining only a small share of the broad AI evaluation platform market. SM001, SM004, SM012
CM042 Because broad and narrow market definitions differ so much, investors should validate how much of LMArena's revenue comes from public benchmarks, agent evaluation, and enterprise domain workflows before relying on TAM-based share math. SM001, SM004, SM016, SM017
CM043 A three-tier lens of $2.36B TAM, about $0.3-0.8B SAM, and about $0.03-0.15B SOM is directionally consistent with the verified evidence but remains an analytical construct rather than a published market model. SM001, SM004, SM012, SM016
CM044 LMArena's buyer-segment matrix is driven more by use-case needs such as benchmark credibility, workflow depth, and domain rigor than by simple company-size segmentation. SM007, SM016, SM017, SM008
CM045 Evaluation spend narrows from open experimentation to recurring governance spend as AI systems move into production, which is why adoption-funnel economics depend on deployment depth rather than benchmark traffic alone. SM006, SM016, SM017
CM046 A low-mid-high range of roughly $0.85B, $2.36B, and $9.57B shows how the category can look small, medium, or strategic depending on whether the lens is narrow 2026 benchmarking, broad 2026 platforms, or longer-horizon strategic infrastructure. SM001, SM004
CP001 LMArena's platform serves more than 5 million monthly users across 150 countries as of January 2026. SP007, SP008
CP002 LMArena generates more than 60 million model-comparison conversations per month as of January 2026. SP007, SP008
CP003 LMArena users span 150 countries according to the company's January 2026 Series A announcement. SP008
CP004 LMArena has evaluated more than 400 public models and conducted more than 300 pre-release tests across multiple modalities as of the two-year anniversary in 2025. SP007
CP005 LMArena has released more than 1.5 million community-contributed prompts as open data for research use, supporting its open-access mission. SP007
CP006 LMArena uses a pairwise comparison approach and Elo-style statistical ranking across blind model battles, as described in the 2024 Chatbot Arena paper by Chiang, Zheng et al. SP001, SP023
CP007 LMArena publicly launched its commercial AI Evaluations enterprise product in September 2025, marking its entry into the paid evaluation-as-a-service market. SP008
CP008 LMArena has established partnerships with OpenAI, Google, Anthropic, Meta, and xAI to make their flagship models available for community evaluation on the platform. SP007, SP008
CP009 FastChat, the open-source platform backing LMArena, has powered over 10 million chat requests for 70+ LLMs, according to the GitHub repository. SP011
CP010 The LMSYS Org at UC Berkeley, which originated LMArena, reports 15+ projects, 79K+ GitHub stars, and 1,000+ contributors across its open-source research portfolio. SP023
CP011 Scale AI operates the Scale GenAI Platform, offering enterprise evaluation, data labeling, and agent deployment services to clients including Meta, Mayo Clinic, and defense agencies. SP016, SP017
CP012 Scale AI's enterprise evaluation product is private and SLA-backed, with no public leaderboard, targeting a different buyer than LMArena's public benchmark community. SP016, SP017
CP013 Scale AI's enterprise engagement pricing reportedly starts at approximately $93,000 per year, with complex projects reaching $400,000 or more, per analyst aggregations. SP016
CP014 The HuggingFace Open LLM Leaderboard uses automated benchmark pipelines to evaluate open-source models on tasks including MMLU, GPQA, ARC, and SWE-Bench, without human preference voting. SP014, SP012
CP015 The HuggingFace Open LLM Leaderboard is freely accessible to model submitters and public readers, with no commercial evaluation service attached. SP014
CP016 Stanford HELM evaluates language models holistically across multiple automated dimensions including accuracy, robustness, calibration, bias, efficiency, and toxicity. SP018, SP019
CP017 Stanford HELM is an academic, non-commercial tool with no enterprise evaluation service; it is freely accessible to any researcher or practitioner. SP018
CP018 EleutherAI's lm-evaluation-harness provides over 60 standardized academic benchmarks and powers the HuggingFace Open LLM Leaderboard. SP012
CP019 EleutherAI's lm-evaluation-harness has no commercial evaluation product and is freely available as open-source software, with the harness running on any infrastructure the user controls. SP012
CP020 OpenAI Evals is an open-source framework for evaluating LLMs that can now be run directly in the OpenAI Dashboard, with a community benchmark registry and enterprise custom eval support. SP013
CP021 BenchLM tracked 261 models across 249 benchmarks as of June 2026, distinguishing verified from provisional rankings, with pricing and speed data included. SP020
CP022 ArtificialAnalysis provides independent, provider-agnostic AI model performance analytics including intelligence, output speed, latency, and cost benchmarking, without a commercial evaluation service. SP021
CP023 None of the major free leaderboard competitors — HuggingFace Open LLM Leaderboard, Stanford HELM, EleutherAI harness, BenchLM, or ArtificialAnalysis — offer an enterprise evaluation-as-a-service product with SLA commitments. SP012, SP014, SP018, SP020, SP021
CP024 LMArena is the only platform that combines large-scale human-preference ranking (5 M+ monthly users) with a commercial enterprise evaluation service in a single brand and infrastructure, as of June 2026. SP001, SP007, SP008, SP016
CP025 LMArena benefits from network effects: a larger and more active user community generates more preference votes, which strengthens the statistical stability of Elo rankings, making the platform more attractive to labs seeking reliable signal. SP001, SP007
CP026 The Elo ranking algorithm used by LMArena is open-source and publicly documented, meaning the mathematical mechanism of ranking can be reproduced by any sufficiently resourced team. SP001, SP003
CP027 OpenAI Evals is built primarily for evaluating OpenAI models, which limits its independence as a neutral tool for cross-provider model comparison. SP013
CP028 Singh et al. (2025) found that Meta tested 27 private LLM variants on Chatbot Arena in the lead-up to the Llama 4 public release, selecting only the best-performing score for public disclosure. SP003, SP005
CP029 Singh et al. (2025) estimated that Google and OpenAI received 19.2% and 20.4% of all Chatbot Arena battle data respectively, while 83 combined open-weight models received only 29.7%. SP003, SP005
CP030 After the gaming incident, Meta's vanilla Maverick model (unoptimized) was ranked approximately 32nd on the LMArena leaderboard, below GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro. SP006
CP031 TechCrunch reported that Chatbot Arena's user base is skewed toward tech and AI professionals; the top questions in the LMSYS-Chat-1M dataset pertain to programming, software bugs, and app design rather than general consumer use. SP004
CP032 Researchers including Yuchen Lin (Allen Institute for AI) and Mike Cook (King's College London) raised concerns that LMArena's evaluation lacks construct validity, meaning it is unclear whether user votes reliably measure real model quality. SP004
CP033 LMArena earns revenue by selling AI evaluation services to the same AI labs (OpenAI, Google, xAI) it ranks on its public leaderboard, creating a structural conflict of interest. SP005, SP008
CP034 LMArena stated in response to the Singh et al. study that it has published information on pre-release testing since March 2024 and that it does not favor any model provider over another. SP005
CP035 LMArena committed to create a new sampling algorithm to address concerns about unequal model battle frequency, acknowledging operational merit in some of the Leaderboard Illusion critiques. SP005
CP036 LMArena has not published a public price list for its AI Evaluations enterprise product; enterprise buyers must contact the sales team for pricing. SP008, SP009
CP037 Scale AI enterprise evaluation engagements reportedly start at approximately $93,000 per year according to third-party analyst aggregations, with complex projects reaching $400,000 or more. SP016
CP038 HuggingFace Open LLM Leaderboard, Stanford HELM, EleutherAI lm-evaluation-harness, BenchLM, and ArtificialAnalysis are all available free of charge with no enterprise evaluation contract required. SP012, SP014, SP018, SP020, SP021
CP039 LMArena's cumulative preference vote dataset (6 M+ votes across 60 M conversations) is the largest publicly known human-preference benchmark dataset for LLMs, with no comparable free alternative. SP007, SP008, SP001
CP040 AI labs cite their Arena Elo scores in press releases and marketing materials, creating a dependency on LMArena's ranking signal for product positioning that raises the switching cost of defecting to an alternative leaderboard. SP004, SP005
CP041 No competitor has replicated LMArena's combination of real-time human-preference data at population scale with a commercial enterprise evaluation product as of June 2026, as inferred from publicly reviewed product surfaces. SP007, SP016, SP020, SP021
CI001 LMArena publicly launched its commercial AI Evaluations enterprise product in September 2025, marking the company's first commercial revenue-generating product. SI001, SI002
CI002 The disclosed $30 million annualized consumption run rate implies roughly $2.5 million of December 2025 monthly revenue at the point LMArena reported the metric. SI001, SI002
CI003 TechCrunch noted that LMArena calls this figure a "consumption rate" — as it "describes its annual recurring revenue (ARR)" — making clear that this is a usage-based annualized run rate, not a contracted ARR with forward booking guarantees. SI001, SI002
CI004 GetLatka estimates approximately 100 enterprise clients for LMArena as of early 2026; this is an analyst aggregation, not a company-confirmed figure. SI005
CI005 Dividing the $30 million annualized consumption run rate by ~100 enterprise clients implies an average contract value of approximately $300,000 per client; this is a derived estimate, not a company-stated figure. SI001, SI005
CI006 OpenAI, Google, and xAI are confirmed as enterprise clients of LMArena's AI Evaluations service, per the January 2026 Series A press release from LMArena. SI001, SI004
CI007 LMArena earns revenue by providing paid AI evaluation services to AI labs and enterprises; the free public leaderboard is not the direct monetized product. SI001, SI011
CI008 LMArena's AI Evaluations product includes comprehensive in-depth evaluations based on community feedback, auditability through representative data samples, and SLA-committed delivery timelines. SI011
CI009 LMArena raised $100 million in a seed round in May 2025 at a post-money valuation of $600 million, led by Andreessen Horowitz and UC Investments, with Lightspeed, Felicis, and Kleiner Perkins also participating. SI003, SI004
CI010 LMArena raised $150 million in a Series A in January 2026 at a post-money valuation of $1.7 billion, led by Felicis and UC Investments with participation from a16z, Kleiner Perkins, Lightspeed, House Fund, LDVP, and Laude Ventures. SI001, SI004
CI011 LMArena has raised $250 million in total across its seed round and Series A, in approximately seven months from May 2025 to January 2026. SI002, SI008
CI012 The LMArena Series A investor group includes UC Investments (managing University of California public funds), which the company's CIO cited as validation of LMArena's role as critical AI evaluation infrastructure. SI001
CI013 An SEC Form D filing (accession 0002113470-26-000001, filed 2026-02-26) documents the formation of Recall Capital-LMArena, a Delaware LLC venture capital feeder fund established to invest in LMArena. SI007, SI024
CI014 LMArena had approximately 41 employees as of the January 2026 Series A announcement, according to TechCrunch's reporting. SI002
CI015 LMArena's primary cost categories are inferred to be AI inference compute for serving 60 million monthly conversations, engineering headcount, and community management; gross margin is not publicly disclosed. SI001, SI014
CI016 LMArena operates a pure software and services model with no physical manufacturing, hardware capital expenditure, or project-finance obligations reported publicly. SI001, SI011
CI017 LMArena's gross margins are not publicly disclosed; analyst estimates suggest a range of 40–70%, depending on whether AI inference costs for partner model hosting are subsidized or borne by LMArena directly. SI011, SI006
CI018 With $250 million raised and a 41-person team as of January 2026, LMArena's estimated cash runway is 24–36 months from the Series A close, though no official burn rate has been confirmed. SI002, SI005
CI019 LMArena has not publicly disclosed monthly burn rate, current cash on hand, or debt obligations; these metrics are fully private. SI002
CI020 LMArena stated it will use the Series A funds to operate its platform, expand its technical team, and strengthen its research capabilities. SI001
CI021 LMArena earns revenue from AI labs (OpenAI, Google, xAI) it evaluates and publicly ranks, creating a structural conflict of interest between its commercial incentive and its neutrality claim. SI006, SI010
CI022 CTOL Digital calculated that LMArena's $1.7 billion valuation is approximately 57 times the $30 million annualized consumption run rate, meaning the valuation embeds an expectation that the commercial conflict can be managed indefinitely. SI006
CI023 LMArena's revenue was concentrated in its first four months of commercial operation; customer concentration risk is high given that three confirmed clients (OpenAI, Google, xAI) are the same entities ranked on the public leaderboard. SI001, SI006
CI024 If enterprise clients reduce engagement with LMArena's evaluation service in response to perceived bias in rankings — as documented in the Singh et al. Leaderboard Illusion paper — the revenue model and the free leaderboard's trust premium would be impaired simultaneously. SI009, SI010, SI006
CI025 No independent audit of LMArena's evaluation methodology, conflict-of-interest management policy, or financial controls has been published as of June 2026. SI006, SI009
CI026 LMArena's GTM motion is a freemium funnel: the free public leaderboard attracts AI labs and enterprise technical teams, which then convert to the paid AI Evaluations service. SI001, SI011
CI027 LMArena has stated that all publicly released models will be evaluated under the same methodology regardless of commercial interest, maintaining the free public leaderboard as a trust anchor for the GTM funnel. SI011
CI028 LMArena has not published information about sales cycle length, CAC, payback period, or enterprise win rates for the AI Evaluations product. SI002
CI029 Both LMArena and TechCrunch used the phrase "consumption rate" rather than a standard SaaS "ARR" metric for the $30 million figure, signaling that the revenue structure may be transaction-based rather than subscription-based. SI001, SI002
CI030 The $30 million annualized consumption run rate is an annualized pace of revenue observed in December 2025, not the full-year 2025 realized revenue figure, which would be substantially lower given product launch in September 2025. SI001, SI002
CI031 If LMArena's billing is usage-based (consumption), enterprise clients may reduce evaluation volume in slow periods without formal churn, making the run rate a less reliable indicator of future revenue than contracted ARR would be. SI002, SI006
CI032 LMArena's $1.7 billion post-money Series A valuation implies a revenue multiple of approximately 57x the $30 million annualized consumption run rate, consistent with early-stage high-growth software but requiring significant revenue growth to justify at later stages. SI006, SI002
CI033 UC Investments, which manages investment assets for the University of California system, led both the seed and Series A rounds, providing institutional and academic credibility to the investment thesis. SI003, SI004
CI034 Andreessen Horowitz (a16z) participated in both the May 2025 seed round and the January 2026 Series A, indicating high-conviction early backing from a top venture firm. SI003, SI004
CI035 LMArena has not published a public price list for its AI Evaluations enterprise product; enterprise buyers must contact the team at evaluations@lmarena.ai. SI011
CI036 The Leaderboard Illusion paper's finding of data-access asymmetry creates a reputational risk that could reduce enterprise clients' willingness to pay LMArena for evaluation services if the methodology is perceived as commercially compromised. SI009, SI010
CI037 Bloomberg reported in April 2025 — one month before the formal seed announcement — that Chatbot Arena was becoming a "real company," providing early public evidence of the commercialization timeline. SI003
CI038 LMArena has released 1.5 million+ community prompts and 145K+ battle data points as open data, which supports its academic credibility narrative but generates no direct revenue. SI004, SI014
CE001 LMArena operates a multi-arena AI evaluation platform at arena.ai covering text, code, search, agent, vision, image generation, video generation, image editing, and document modalities. SE001, SE002, SE005
CE002 The core platform mechanic is a pairwise battle where two anonymous models respond to the same prompt and users vote on the preferred output. SE001, SE014
CE003 All Arena text, code, search, and vision leaderboards use the Bradley-Terry statistical model to infer latent skill coefficients from pairwise win/loss battle outcomes. SE009, SE014, SE015, SE018
CE004 LMArena had 5 million or more monthly users as of January 2026. SE003, SE005
CE005 LMArena has logged more than 250 million real conversations across all arenas since its inception. SE003
CE006 LMArena generates more than 2 million preference votes per month from community users. SE003
CE007 LMArena's community spans users in more than 150 countries. SE003
CE008 Agent Arena was launched on June 4, 2026 as LMArena's newest evaluation arena, ranking orchestrator models for autonomous multi-step task completion. SE005, SE006
CE009 Agent Arena ranks models using causal tracing, treating each component selection as a treatment in a multi-intervention randomized controlled trial measuring five behavioral signals. SE006
CE010 WebDev Arena, launched December 2024, collected over 80,000 community votes on AI-generated web applications before being superseded by Code Arena in 2026. SE007
CE011 Search Arena supports 11 models from three providers (Perplexity, Gemini, and OpenAI) and has collected over 24,000 paired multi-turn user interactions. SE008, SE017
CE012 Code Arena was rebuilt from WebDev Arena with a new isolated agentic coding environment, persistent sessions stored in Cloudflare R2, and live preview rendering via CodeMirror 6. SE013, SE005
CE013 Arena-Rank is an open-source Python package (Apache 2.0) published on GitHub under the lmarena organization and installable from PyPI as 'arena-rank'. SE009, SE001, SE020, SE023
CE014 Arena-Rank implements Bradley-Terry ranking with closed-form confidence-interval calculation and a 30x speedup over the historical FastChat implementation. SE009, SE001, SE020
CE015 FastChat (GitHub: lm-sys/FastChat) was the original open-source platform powering Chatbot Arena but is now primarily in maintenance mode, with ranking code migrated to Arena-Rank. SE019, SE009
CE016 LMArena's style-control extension adds length, markdown header count, bold count, and list count as regression covariates in the BT model to isolate substance from formatting effects; response length is the dominant style factor. SE010, SE015
CE017 The Arena-Hard BenchBuilder pipeline (arXiv:2406.11939) automatically extracts hard prompts from live Arena data using a seven-criterion hardness labeler and generates benchmarks that achieve 98.6% agreement with human preference rankings. SE016, SE012
CE018 LMArena defines prompt hardness using seven criteria including domain knowledge, problem-solving complexity, and real-world applicability; approximately 20% of Arena prompts have a hardness score of 6 or higher. SE011
CE019 The lmarena-ai HuggingFace organization hosts multiple p2l (prompt-to-leaderboard) preference models ranging from 135M to 7B parameters and releases public battle datasets. SE022
CE020 LMArena uses GCP's Sensitive Data Protection API to remove personally identifiable information from conversation data before any public release or data sharing. SE004
CE021 LMArena's leaderboard policy was last updated April 30, 2026 and specifies model eligibility criteria, sampling policies, pre-release testing protocols, and data sharing rules. SE004, SE005
CE022 LMArena's sampling policy requires that at least 20% of all battles involve only publicly available models, with remaining capacity available for unreleased or experimental models. SE004
CE023 LMArena allows model providers to test unreleased models anonymously, shares results privately with the provider, and then removes the model before any public listing. SE004, SE025
CE024 In April 2025 Meta submitted an 'experimental, chat-optimized' version of Llama 4 Maverick (not the public release) to LMArena, achieving a #2 ranking; when the public version was scored, it fell to approximately 32nd place. SE025, SE024
CE025 Following the Meta Llama 4 incident, LMArena updated its leaderboard policies to require that pre-release model variants be explicitly labeled as customized and not represent the publicly released model. SE025, SE024
CE026 A paper by researchers from Cohere, Stanford, MIT, and AI2 (April 2025) alleged that Meta, OpenAI, Google, and Amazon received disproportionately high sampling rates and could suppress low-scoring pre-release variants, constituting benchmark gaming. SE024, SE025
CE027 LMArena denied the claims in the Cohere/Stanford study as containing 'inaccuracies and questionable analysis,' asserting that all model providers are allowed to submit more models for testing and that the leaderboard remains fair. SE024
CE028 Independent researchers have noted that LMArena's user base is skewed toward technical and developer prompts, making it less representative of general-population or enterprise preferences. SE024
CE029 Agent Arena analyzes five behavioral signals: confirmed success, praise vs. complaint, steerability, bash-error recovery, and tool hallucination rate. SE006
CE030 In a recent 7-day window, Agent Mode recorded 160,480 agent tasks on LMArena's platform, with code writing as the largest category at 17.5%. SE006
CE031 Agent Mode issued approximately 2 million structured tool calls in one 7-day period, including 936,000 bash calls and 550,000 file-write operations. SE006
CE032 Code Arena uses Cloudflare R2 for persistent session storage and CodeMirror 6 for source-code display and live preview rendering of generated web applications. SE013
CE033 Battles in Direct, launched May 2026, converts 10% of Direct Chat sessions into anonymous pairwise battles and applies position-bias and same-org-indicator corrections in the BT model. SE005
CE034 Arena-Rank's open-source implementation achieves a 30x speedup over the historical FastChat-based BT implementation and uses reweighting to correct for non-uniform model sampling. SE009, SE001, SE020
CE035 Arena-Hard-Auto v0.1 achieves 98.6% agreement with human preference rankings from Chatbot Arena and provides 3x higher model separation than MT-Bench. SE016, SE018
CE036 The Search Arena paper was accepted at ICLR 2026, providing peer-reviewed validation of LMArena's search-augmented LLM evaluation methodology. SE017, SE005
CE037 LMArena has not publicly disclosed SOC 2 Type II, ISO 27001, or any independent third-party security audit for its AI Evaluations commercial product or the arena.ai platform. SE004
CE038 LMArena's help.arena.ai privacy policy exists but lacks enterprise data processing agreement (DPA) terms, and GDPR/CCPA compliance details are not publicly documented. SE004
CE039 LMArena has evaluated over 400 public models and over 300 pre-release model variants across all its arenas since the platform launched in 2023. SE005
CE040 LMArena has released 1.5 million prompts and over 145,000 battle data points publicly for open research as of early 2026. SE005
CE041 Agent Arena sessions average approximately 16.5 structured tool calls, with about 75.6% of sessions using at least one tool in a measured 7-day window. SE006
CE042 Agent Mode wrote 40.3 million lines of code through successful write_file calls in one 7-day window, approximately 1,000 lines per coding session. SE006
CU001 LMArena had 5 million or more monthly active users as of January 2026, spanning more than 150 countries. SU001, SU009
CU002 LMArena generates more than 60 million conversations per month and has accumulated more than 250 million conversations in total as of January 2026. SU001, SU009
CU003 LMArena's community generates more than 2 million preference votes per month. SU013
CU004 LMArena serves two distinct customer populations: a free community of millions of monthly users providing preference votes, and a small paying commercial segment of AI labs and enterprises purchasing AI Evaluations. SU001, SU013
CU005 OpenAI, Google, and xAI are named paying customers of LMArena's AI Evaluations commercial product, drawing on evaluations to improve their models for production use cases. SU001, SU009
CU006 LMArena partnered with OpenAI, Google, and Anthropic to make their flagship models available for community evaluation; Anthropic's paying status as an AI Evaluations subscriber is not separately confirmed. SU001, SU009
CU007 LMArena's community grew approximately 25 times between May 2025 (seed round) and January 2026 (Series A) as reported by the company. SU005
CU008 LMArena had approximately 3 million or more monthly users at the time of its commercial AI Evaluations product launch in September 2025. SU013
CU009 LMArena has evaluated more than 400 public models and more than 300 pre-release model variants across its arenas as of early 2026. SU005, SU006
CU010 The LMArena community released more than 50 million votes and 1.5 million open prompts by January 2026. SU005
CU011 LMArena's AI Evaluations commercial product achieved an annualized consumption run-rate of $30 million in December 2025, less than four months after its September 2025 launch. SU001, SU009
CU012 Meta submitted 27 Llama 4 model variants to LMArena for private pre-release testing between January and March 2025, then publicly disclosed only the score of the highest-performing experimental variant. SU016, SU017
CU013 The publicly released version of Meta Llama 4 Maverick ranked approximately 32nd on the LMArena leaderboard, versus the experimental version which had ranked #2. SU007, SU017
CU014 Anthropic's Claude models were described as currently winning the expert leaderboard for legal and medical use cases in January 2026. SU003
CU015 Perplexity's Sonar models and Google Gemini were the top-ranked models in the Search Arena as of the ICLR 2026 paper, with Perplexity-Sonar-Reasoning-Pro and Gemini-2.5-Pro at the top. SU019, SU024
CU016 LMArena has not publicly disclosed net revenue retention, gross revenue retention, or customer churn rates for its AI Evaluations commercial product. SU011
CU017 The leaderboard changelog shows multiple model additions per week across all arenas in June 2026, serving as an indirect proxy for ongoing model-provider engagement. SU015
CU018 LMArena describes its revenue as 'annualized consumption rate' rather than annual recurring revenue (ARR), implying a consumption-based billing model rather than committed subscriptions. SU001
CU019 The structural conflict of interest in which the same AI labs that pay for AI Evaluations also have their models ranked on the public leaderboard was identified by independent journalists as a key credibility risk. SU002, SU016
CU020 A paper from Cohere, Stanford, MIT, and AI2 alleged in April 2025 that Meta, OpenAI, Google, and Amazon received disproportionately high sampling rates, enabling benchmark gaming; LMArena denied specific claims and announced a new sampling algorithm. SU016, SU003
CU021 LMArena updated its leaderboard policy after the Meta Llama 4 incident and stated that 'Meta's interpretation of our policy did not match what we expect from model providers.' SU017, SU025
CU022 As of June 2026, only three companies (OpenAI, Google, xAI) are publicly named as paying customers of LMArena's AI Evaluations service; no enterprise non-lab customers have been publicly named. SU001, SU009
CU023 LMArena's revenue is highly concentrated in three named AI lab customers; departure of any one of these would represent a material revenue risk given the $30M ARR run-rate and unknown diversification. SU001
CU024 LMArena's community growth trajectory—25x from May 2025 to January 2026—implies organic conversion of community interest into lab evaluation demand, though the precise pathway from free user to commercial customer is not documented. SU023, SU009
CU025 Prior to commercialization, Google's Kaggle, Andreessen Horowitz, and Together AI had donated compute, cloud credits, and cash to LMSYS as corporate sponsors, creating an indirect financial relationship with leaderboard participants. SU016, SU010
CU026 Felicis General Partner Peter Deng stated in the LMArena Series A press release that LMArena has 'become essential infrastructure for every lab and enterprise' and that Felicis led the round because of LMArena's trustworthy real-world performance signal. SU001, SU009
CU027 Scale AI launched a competing benchmarking product, SEAL Showdown, as a direct rival to LMArena's evaluation platform, representing competitive pressure in the AI evaluation market. SU004, SU016
CU028 OpenTools.AI independently reported that LMArena was 'under fire' for benchmark bias in 2025, reflecting industry-wide concern beyond the single Cohere/Stanford paper. SU003
CU029 PitchBook tracks Arena Intelligence (LMArena) with a confirmed valuation history from the $600M seed in May 2025 to the $1.7B Series A in January 2026. SU006, SU009
CU030 LMArena's original domain lmarena.ai now redirects to arena.ai, reflecting the company's rebranding as it expanded beyond language model evaluation to a multi-arena platform. SU007
CU031 EDGAR full-text search confirms LMArena (entity 'Recall Capital-LMArena') has an active Form D filing for the Series A under CIK 0002113470, providing independent corroboration of the funding event. SU008, SU009
CU032 Anthropic's Claude is described by TechCrunch podcast as currently winning LMArena's expert leaderboard for legal and medical professional use cases as of January 2026, indicating provider engagement depth. SU002, SU009
CU033 LMArena's two-year anniversary blog (April 2025) confirms approximately 41% of battles involve open-source models, indicating a mixed commercial and research user base that constrains pure commercial curation. SU024, SU023
CU034 No independent third-party reviews on G2, Capterra, or Gartner Peer Insights for LMArena's AI Evaluations commercial product are publicly available as of June 2026. SU011
CU035 LMArena's AI Evaluations commercial product targets software engineering, law, medicine, and scientific research as economically valuable industries for paid evaluation services. SU001, SU009
CR001 LMArena earns revenue by selling paid AI evaluation services to AI labs (OpenAI, Google, xAI) whose models simultaneously appear on LMArena's public leaderboard, creating a structural conflict of interest between commercial and evaluation roles. SR015, SR011, SR014
CR002 The Leaderboard Illusion paper (arXiv 2504.20879) found that a small number of providers could privately test multiple model variants and selectively disclose only their best scores, resulting in biased Arena rankings. SR002, SR001
CR003 Meta privately tested 27 Llama-4 model variants on Chatbot Arena in the lead-up to its Llama 4 release, selecting the best-performing variant for its public score. SR002, SR001, SR003
CR004 Meta submitted an "experimental chat version" of Llama 4 Maverick "optimized for conversationality" to LMArena that achieved a top-two leaderboard ranking, but this model was not the same version released publicly. SR003, SR004
CR005 SurgeAI's analysis of 500 LMArena votes found that evaluators disagreed with LMArena's outcomes 52% of the time, with "confidence beats accuracy and formatting beats facts." SR006
CR006 Independent researcher Gwern described LMArena as "a cancer" and questioned whether it is worth running, reflecting significant reputational erosion among technical users. SR006
CR007 Arena Intelligence, Inc. (d/b/a LMArena) processes personal data from users in EU member states and is subject to GDPR compliance obligations including data subject rights, lawful basis for processing, and international data transfer safeguards. SR009, SR010
CR008 LMArena's privacy policy (effective September 2025) explicitly warns users that prompts, votes/ratings, and other user content may be shared publicly and with AI providers as part of the evaluation process. SR009, SR013
CR009 The EU AI Act's general-purpose AI (GPAI) model obligations, which came into force in August 2025, impose transparency documentation, training data summaries, and copyright policy requirements on AI platform operators. SR012, SR010
CR010 No litigation directly naming Arena Intelligence, Inc. or LMArena as a plaintiff or defendant was found in public records as of 2026-06-25. SR009
CR011 LMArena uses GCP's Sensitive Data Protection API to remove personal and sensitive data before sharing conversation data with model providers or publishing it publicly, and GCP is the confirmed primary cloud infrastructure provider. SR015, SR009
CR012 LMArena's leaderboard methodology update in May 2026 (Battles in Direct) discovered and corrected two new biases: position bias favoring Model A, and an advantage for models sharing an organization with prior conversational context. SR029, SR008
CR013 LMArena processes 60 million monthly user conversations across 150 countries on cloud infrastructure, creating platform reliability, data security, and regulatory compliance obligations at scale. SR015, SR016
CR014 No public disclosure of SOC 2, ISO 27001, or equivalent security certification for Arena Intelligence, Inc. has been found as of the run date. SR009
CR015 Chatbot Arena's Elo/Bradley-Terry scoring methodology relies on sufficient uniformity of battle distribution; coordinated voting campaigns, prompt injection, or sybil attacks could corrupt the ranking signal without immediate detection. SR002, SR006
CR016 AI providers including OpenAI and Google have financial incentives to study and potentially overfit to the Arena evaluation distribution, as documented by the finding that access to Arena data yields up to 112% relative performance gains on Arena Hard. SR002, SR005
CR017 LMArena had approximately 100 paying customers as of early 2026, generating $30M ARR, implying material revenue concentration risk with the top AI lab customers. SR026, SR015
CR018 LMArena publicly confirmed that OpenAI, Google, and xAI draw on its evaluations to improve their models, making these companies simultaneously its best customers and its most motivated potential evaluators of its gaming policies. SR015, SR020
CR019 No contractual SLA or public agreement guaranteeing sustained AI provider API access to LMArena's platform has been publicly disclosed; any lab can withdraw its model. No known instance of API withdrawal has occurred as of the run date. SR007
CR020 LMArena employs approximately 41 people as of January 2026, representing a lean team relative to the platform's operational scope and regulatory obligations. SR026
CR021 The lead investor in LMArena's Series A (Peter Deng of Felicis) previously worked at OpenAI, creating an appearance-of-conflict risk between Felicis's portfolio interest and LMArena's independence claims regarding OpenAI model evaluations. SR015, SR025
CR022 Arena Hard performance — a synthetic benchmark LMArena maintains — can be improved by up to 112% relative gains with additional Arena battle data, according to the Leaderboard Illusion researchers' conservative estimates. SR002, SR001
CR023 The Leaderboard Illusion paper found that Google and OpenAI each received an estimated 19.2% and 20.4% of all Chatbot Arena data, while 83 open-weight models combined received only approximately 29.7% of total data. SR002, SR001
CR024 LMArena updated its sampling policy post-April 2026 to guarantee that at least 20% of all battles involve only publicly available models, and committed to reweighting scoring so that sampling probabilities do not bias Arena scores. SR007, SR008
CR025 The SEC EDGAR filing for "Recall Capital-LMArena a Series of CGF2021 LLC" (Form D, filed 2026-02-26, accession 0002113470-26-000001) is a secondary-market venture fund raising $382,500 specifically to invest in LMArena. SR024, SR014
CR026 LMArena's Agent Arena leaderboard launched on June 4, 2026, expanding into agentic evaluation with behavioral signals like file downloads, disapproval events, retries, and steerability rather than static preference votes alone. SR008, SR016
CR027 US state privacy laws including California's CPRA impose separate data-subject rights and business compliance obligations on LMArena that are distinct from GDPR requirements. SR009, SR012
CR028 The EU AI Act imposes penalties of up to 7% of global annual turnover for prohibited AI practices and up to 3% for other violations, applicable to AI platforms with EU operations. SR012, SR010
CR029 CTOL Digital Solutions noted that LMArena's $1.7 billion valuation implies roughly 57x the company's annualized revenue run rate of $30M, pricing in substantial growth expectations that depend on sustained trust in benchmark neutrality. SR011
CR030 LMArena's commercial evaluation product, AI Evaluations (launched September 2025), provides paid evaluation services to enterprises and AI labs, creating a revenue stream directly tied to the labs it ranks publicly. SR017, SR015
CR031 The AI model evaluation platform market is projected to grow from $1.86 billion in 2025 to $2.36 billion in 2026 and $6.24 billion by 2030, at a 27.3-27.5% CAGR. SR023
CR032 LMArena's open-source Arena-Rank repository and methodology academic papers provide external auditability of the leaderboard scoring approach, which is a partial mitigation against benchmark-integrity allegations. SR007, SR016
CR033 LMArena's Series A (January 2026) achieved a post-money valuation of $1.7 billion, nearly triple the $600 million seed valuation from May 2025, on $250M total capital raised. SR014, SR015, SR022
CR034 The Leaderboard Illusion researchers found that proprietary/closed models are sampled at higher battle rates and have fewer models removed from Arena compared to open-weight alternatives, creating a data access asymmetry. SR002, SR018
CR035 A TechCrunch analysis noted that LMArena's commercial relationships with OpenAI, Google, and Anthropic raise questions about whether the benchmark can be trusted to assess AI models without corporate influence clouding the process. SR001, SR011
CR036 LMArena was previously funded through grants and donations from Google's Kaggle, Andreessen Horowitz, and Together AI — organizations with direct stakes in the models being evaluated. SR005, SR022
CR037 LMArena stated that its commercial evaluation product provides the same methodology to all paying customers without preferential treatment, and that the public leaderboard will always be available freely. SR017
CR038 The LMArena policy (updated April 30, 2026) now requires that if a publicly released model differs from the pre-release version tested on Arena, Arena will remove the model from the leaderboard until it can be re-evaluated under the requirements of this policy. SR007, SR008
CR039 TechCrunch noted that a Llama 4 Maverick vanilla release ranked 32nd on the LMArena leaderboard after the experimental version that ranked second was withdrawn, demonstrating a 30-rank gap attributable to benchmark optimization. SR004, SR003
CR040 At 41 employees and $30M ARR, LMArena's revenue-per-employee ratio is approximately $730K/FTE, suggesting significant infrastructure leverage but thin organizational depth for regulatory compliance, security, and enterprise scale-up. SR026, SR015
CR041 LMArena's 5 million monthly users across 150 countries is substantially below the 45 million EU monthly active user threshold that would trigger VLOP designation under the Digital Services Act, making VLOP obligations unlikely in the near term. SR015, SR012
CR042 LMArena has not disclosed any data breach incidents involving the community evaluation dataset or user prompt data as of the run date, but uses GCP security tools rather than published third-party security certifications. SR009, SR015
CV001 LMArena raised $150 million in a Series A round in January 2026 at a post-money valuation of $1.7 billion, led by Felicis and UC Investments, with participation from a16z, Kleiner Perkins, Lightspeed, The House Fund, LDVP, and Laude Ventures. SV001, SV002
CV002 LMArena's annualized "consumption run rate" surpassed $30 million in December 2025, less than four months after launching its first commercial product (AI Evaluations) in September 2025. SV002, SV001
CV003 LMArena raised a $100 million seed round in May 2025 at a $600 million valuation, bringing total capital raised to $250 million by January 2026 across two rounds in approximately seven months. SV001, SV007
CV004 The Latka database records that LMArena sold approximately 17% at the seed round ($100M / $600M) and approximately 9% at the Series A ($150M / $1.7B), providing an implied pre-money Series A enterprise value of approximately $1.55 billion. SV003, SV001
CV005 At $30M ARR and a $1.7B valuation, LMArena's implied ARR multiple is approximately 57x — placing it in the top decile of AI infrastructure Series A rounds in 2025–2026. SV013, SV001
CV006 The $30M ARR figure represents December 2025 monthly revenue annualized (run rate), not a full-year booked revenue number; the company had fewer than four months of commercial operations as of the Series A announcement. SV002, SV008
CV007 The AI model evaluation platform market was valued at $1.86 billion in 2025 and is projected to reach $2.36 billion in 2026 and $6.24 billion by 2030 at a 27.3–27.5% CAGR, according to The Business Research Company's 2026 market report. SV010, SV006
CV008 Weights & Biases (an AI developer platform with model evaluation capabilities) was acquired by CoreWeave for approximately $1.4 billion in March 2025, providing a transaction comparable for AI evaluation infrastructure valuation. SV012, SV011
CV009 A bull case for LMArena requires $120M+ ARR by 2028 and sustained platform multiple above 25x, implying a $3–6B exit value; this requires sustaining the $7.5M/month new ARR velocity seen in the first four months of commercial operations. SV008, SV001
CV010 A base case for LMArena projects $60–75M ARR by end-2026 (at 50% of current ARR velocity), implying a $1.4–2.1B valuation at a 20–30x ARR multiple — roughly in line with the current $1.7B mark, providing limited upside from today's entry price. SV001, SV013
CV011 LMArena's core competitive advantage is a community data moat: 5 million monthly users across 150 countries generating over 60 million conversations monthly, creating a real-time human preference dataset that is extremely difficult to replicate. SV002, SV009
CV012 LMArena's leaderboard has become a standard reference in AI lab product launches, developer procurement decisions, and media coverage, creating network effects and switching costs that reinforce its incumbent position. SV022, SV026
CV013 The structural conflict of interest — where LMArena earns revenue from the same labs it evaluates — creates an existential risk to the investment thesis; a single high-profile investigation confirming revenue influenced rankings would collapse both leaderboard credibility and enterprise revenue simultaneously. SV013, SV016
CV014 Goodhart's Law dynamic — where labs optimize specifically for Arena rather than for genuine capability improvement — degrades the leaderboard's real-world signal value over time, as documented in the Leaderboard Illusion paper and corroborated by multiple independent analyses. SV014, SV016
CV015 The overall investment recommendation for LMArena is "track" with a conditional buy signal requiring an independent methodology audit and an entry price below 30x next-twelve-months ARR. Risk rating is "high"; valuation stance is "stretched." SV013, SV001
CV016 The bear case for LMArena implies a valuation range of $200–500M following a credibility collapse, driven by ARR churn to below $25M and multiple compression to 10–20x distressed ARR — representing a 70–88% loss from the $1.7B entry. SV013, SV014
CV017 The primary thesis-break trigger is discovery of documentary evidence that commercial revenue relationships influenced LMArena's public leaderboard outcomes; secondary triggers include churn of any top-3 lab customers or ARR failing to reach $50M+ by Q3 2026. SV013, SV016
CV018 Peter Deng, the Felicis general partner leading LMArena's Series A round, previously worked at OpenAI — one of LMArena's paying evaluation customers — creating an appearance-of-conflict risk between Felicis's portfolio interest and LMArena's independence claims regarding OpenAI evaluations. SV002, SV022
CV019 The minimum blocking diligence items before a buy decision are: (1) independent statistical audit of the Arena methodology; (2) customer cohort data showing top-3 customer share below 40% or NRR above 110%; (3) cap table with liquidation preference terms; (4) EU AI Act compliance self-assessment. SV013, SV016
CV020 The SEC Form D filing for "Recall Capital-LMArena a Series of CGF2021 LLC" (accession 0002113470-26-000001, filed February 2026) is a $382,500 secondary venture fund, demonstrating secondary-market demand for LMArena exposure at valuations consistent with the Series A. SV015, SV001
CV021 Long-run structural analogies for LMArena's potential platform value include Bloomberg LP and S&P Global's ratings segment, which operate as neutral data arbiters with structural moats and high customer switching costs — though these analogies require LMArena to first resolve its conflict-of-interest and achieve regulatory equivalence. SV006, SV013
CV022 LMArena's product expansion into Agent Arena (June 2026), WebDev Arena, Search Arena, Vision, Video, and coding leaderboards suggests the company is actively increasing its potential ARR ceiling by expanding beyond text model evaluation. SV020, SV009
CV023 A CTOL analysis noted that LMArena's 57x ARR multiple prices in the assumption that the conflict between being a revenue-generating evaluation service and an independent benchmark arbiter "can be managed indefinitely" — a structural assumption that has not been independently validated. SV013, SV001
CV024 LMArena employs approximately 41 people as of January 2026 with approximately $730K revenue per employee, indicating high capital efficiency but thin organizational depth for maintaining a $1.7B asset at scale. SV003, SV002
CV025 No down-round risk has been evidenced in LMArena's financing history; the company raised at 2.83x step-up from seed ($600M) to Series A ($1.7B) in eight months on genuine commercial traction. SV001, SV003
CV026 The revenue concentration risk at LMArena is heightened by its approximately 100 paying customers, where the top AI labs (OpenAI, Google, xAI) likely represent a disproportionate share of $30M ARR per standard B2B Pareto distributions. SV003, SV013
CV027 The AI model evaluation platform market's 27.3% CAGR projection implies that if LMArena captures a 5-10% market share in a $3B+ market by 2028, its revenue would be $150-300M — sufficient to justify the current valuation at 10-15x revenue multiples characteristic of scaled data platforms. SV010, SV013
CV028 LMArena's Arena Intelligence, Inc. d/b/a structure was incorporated in 2025 in Delaware; the company is headquartered in San Francisco, California, and is a private company with no SEC reporting obligations beyond Form D filings. SV021, SV015
CV029 No public evidence of preferred stock liquidation preferences, anti-dilution ratchets, or convertible notes has been disclosed by LMArena or its investors as of the run date, though these terms are standard in Series A financings and their absence from public disclosure does not indicate they do not exist.
CV030 LMArena's investor base (Felicis, UC Investments, a16z, Kleiner Perkins, Lightspeed) includes firms with direct investments in AI labs that are LMArena's paying customers, creating a potential for governance conflicts that could compromise the independence of the evaluation platform. SV025, SV007
CV031 LMArena's prior pre-commercial funding through grants from Google's Kaggle and donations from Andreessen Horowitz and Together AI — organizations with evaluated models on the platform — established a pattern of commercial entanglement that the new corporate structure has not fully resolved. SV025, SV007
CV032 Scale AI, the closest large-scale comparable in AI data labeling and evaluation services, was valued in the $14–29B range on reported revenue substantially larger than LMArena's current $30M ARR, suggesting LMArena's 57x multiple is significantly richer than Scale AI's implied multiple on comparable revenue. SV010, SV013
CV033 LMArena generated 50 million votes, 400+ model evaluations, and added 145,000 open-source battle data points to the community between the seed round (May 2025) and the Series A (January 2026), demonstrating substantial community engagement growth. SV009, SV002
CV034 LMArena's AI Evaluations commercial product provides paid evaluation services including comprehensive in-depth evaluations based on community feedback, auditability through representative data samples, and service-level agreements with committed delivery timelines. SV019, SV002
CV035 No institutional investor has publicly disclosed a reduction in LMArena position, secondary sale, or hedge of LMArena exposure since the Series A close as of the run date; the Recall Capital secondary fund implies continued secondary-market demand. SV015, SV001
CV036 LMArena's Agent Arena leaderboard (launched June 4, 2026) represents a new evaluation domain targeting agentic AI systems, expanding the platform's addressable commercial market beyond text model benchmarking. SV020, SV026
CV037 The minimum entry valuation offering adequate risk/return profile is estimated below $1.0–1.2B (approximately 30x $40M NTM ARR), providing sufficient discount to the current $1.7B mark to compensate for conflict-of-interest, concentration, and multiple-compression risks. SV013, SV001
CV038 LMArena's exit readiness is limited in the near term: the company is 18 months old commercially, has no disclosed profitability path, and an IPO would require 2–3 years of operating history and revenue scale above $150M+ to command strong public market reception; strategic acquisition by a hyperscaler is the most plausible nearer-term exit scenario. SV001, SV022
CV039 LMArena's product expansion into evaluation of models across text, code, vision, video, search, documents, and agents (seven distinct modalities as of June 2026) expands the potential ARR ceiling and reduces concentration risk in the core text model benchmarking segment. SV020, SV009
CV040 The OfficeChai analysis notes that LMArena's AI testing market is estimated at $800–900M in 2025, projected to grow to $3.8 billion by 2032, providing a TAM growth context for the $1.7B valuation. SV006, SV010
CV041 Reuters syndication through U.S. News reported that LMArena's valuation tripled to $1.7 billion in about eight months, underscoring how quickly private-market pricing expanded between the $600M seed round and the January 2026 Series A. SV033, SV007
来源
编号出版方标题引文
SO001 Arena Intelligence Inc. About Arena | Crowdsourced AI Model Evaluation Platform Created by researchers from UC Berkeley, Arena (formerly LMArena) is a community-powered platform for understanding AI performance in the real world.
SO002 Arena Intelligence Inc. Fueling the World's Most Trusted AI Evaluation Platform (Series A Blog) We've raised $150M of Series A funding led by Felicis and UC Investments (University of California).
SO003 PR Newswire / LMArena LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product.
SO004 TechCrunch LMArena lands $1.7B valuation four months after launching its product LMArena raised a $150 million Series A at a post-money valuation of $1.7 billion.
SO005 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M LM Arena... has raised $100 million in a seed funding round that values the organization at $600 million.
SO006 Bloomberg Popular AI Ranking Website Chatbot Arena Is Becoming a Real Company
SO007 Arena Intelligence Inc. New Product: AI Evaluations This service offers enterprises, model labs, and developers comprehensive evaluation services grounded in real-world human feedback.
SO008 Arena Intelligence Inc. Arena Leaderboard Policy
SO009 Arena Intelligence Inc. Arena Leaderboard Changelog
SO010 LMSYS / UC Berkeley Chatbot Arena: New Leaderboard & More Models We are releasing Chatbot Arena, an open-source evaluation platform for LLMs.
SO011 arXiv / UC Berkeley LMSYS Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies.
SO012 arXiv / UC Berkeley Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
SO013 arXiv / UC Berkeley From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
SO014 arXiv / Cohere, Stanford, MIT, Ai2 The Leaderboard Illusion We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired.
SO015 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark The evaluation is not reproducible, and the limited data released by LMSYS makes it challenging to study the limitations of models in depth.
SO016 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification.
SO017 TechCrunch Here's why most AI benchmarks tell us so little
SO018 TechCrunch The leaderboard 'you can't game,' funded by the companies it ranks
SO019 Founded.com How two Berkeley roommates built a $1.7B startup that helps you decide which AI to use The founders behind AI evaluation platform Arena, formerly known as LMArena, have turned that confusion into a business now worth $1.7 billion.
SO020 Winbuzzer Experts Challenge Validity and Ethics of Crowdsourced AI Benchmarks Like LMArena Chatbot Arena hasn't shown that voting for one output over another actually correlates with preferences, however they may be defined.
SO021 Hugging Face lmarena-ai Organization on Hugging Face
SO022 Andreessen Horowitz (a16z) Announcing Our Latest Open Source AI Grants
SO023 Mashable SEA LMArena has some competition: Scale AI launches Seal Showdown, a new benchmarking tool Critics say that LMArena's system favors frontier models from big AI companies like Google, xAI, and OpenAI.
SO024 Arena Intelligence Inc. Agent Arena: Causal Evaluation of Agents in the Real World
SO025 Arena Intelligence Inc. Introducing the Search Arena: Evaluating Search-Enabled AI
SO026 Arena Intelligence Inc. WebDev Arena: A Live LLM Leaderboard for Web App Development
SO027 The Verge Meta got caught gaming AI benchmarks Meta's interpretation of our policy did not match what we expect from model providers.
SO028 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SO029 arXiv / Search Arena (ICLR 2026) Search Arena: A Crowd-Sourced, Human-Preference Dataset for Search-Augmented LLMs
SM001 The Business Research Company Artificial Intelligence (AI) Model Evaluation Platform Market Report The AI model evaluation platform market grows from $1.86 billion in 2025 to $2.36 billion in 2026 at a CAGR of 27.3%.
SM002 Yahoo Finance AI Model Evaluation Platform Market article
SM003 Research & Markets AI Model Evaluation Platform Market Report
SM004 Precedence Research Model Evaluation and Benchmarking Tools Market
SM005 Gartner Gartner forecasts worldwide AI spending to grow 47 percent in 2026
SM006 Presenc AI Enterprise AI adoption statistics 2026 78% of Global 2000 companies have at least one AI workload in production in Q1 2026.
SM007 Future AGI Top 5 LLM Evaluation Tools 2025
SM008 Scale AI SEAL Showdown SEAL Showdown includes users in 100+ countries, 70+ languages, and 200+ professional domains.
SM009 Mashable SEA LMArena has some competition: Scale AI launches SEAL Showdown
SM010 arXiv The Leaderboard Illusion
SM011 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark
SM012 TechCrunch LMArena lands $1.7B valuation four months after launching its product
SM013 PR Newswire LMArena raises $150 million to build the world's most trusted AI evaluation platform
SM014 arXiv Chatbot Arena paper
SM015 LMSYS Chatbot Arena launch post
SM016 Arena AI AI Evaluations
SM017 Arena AI Agent Arena methodology
SM018 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark
SM019 Winbuzzer Experts challenge validity and ethics of crowdsourced AI benchmarks like LMArena
SM020 The Verge Meta Llama 4 Maverick benchmarks gaming
SM021 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SM022 arXiv Arena-Hard paper
SM023 arXiv MT-Bench paper
SM024 Founded LMArena founders profile
SM025 Scale AI SEAL Showdown
SP001 arXiv (Wei-Lin Chiang, Lianmin Zheng, et al. — UC Berkeley / LMSYS) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies.
SP002 arXiv (Tianle Li, Wei-Lin Chiang, et al. — UC Berkeley / LMSYS) From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20.
SP003 arXiv (Singh et al. — Cohere, Stanford, MIT, Ai2) The Leaderboard Illusion Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data.
SP004 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark The distribution of testing data may not accurately reflect the target market's real human users. Moreover, the platform's evaluation process is largely uncontrollable.
SP005 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark LM Arena allowed some industry-leading AI companies like Meta, OpenAI, Google, and Amazon to privately test several variants of AI models, then not publish the scores of the lowest performers.
SP006 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark The unmodified Maverick, 'Llama-4-Maverick-17B-128E-Instruct,' was ranked below models including OpenAI's GPT-4o, Anthropic's Claude 3.5 Sonnet, and Google's Gemini 1.5 Pro as of Friday.
SP007 LMArena (Arena Intelligence) Celebrating Community Impact at LMArena (Two-Year Celebration) 400+ models have been evaluated across various modalities (Text, Vision, Text-to-Image, WebDev and more!). 300+ evaluations have been pre-release.
SP008 LMArena (Arena Intelligence) New Product: AI Evaluations This service offers enterprises, model labs, and developers comprehensive evaluation services grounded in real-world human feedback.
SP009 LMArena (Arena Intelligence) Arena AI: The Official AI Ranking & LLM Leaderboard
SP010 LMArena (Arena Intelligence) Arena Leaderboard
SP011 LMArena / LMSYS Org (GitHub) FastChat: An open platform for training, serving, and evaluating large language models FastChat powers Chatbot Arena (lmarena.ai), serving over 10 million chat requests for 70+ LLMs.
SP012 EleutherAI (GitHub) lm-evaluation-harness: A framework for few-shot evaluation of language models Over 60 standard academic benchmarks for LLMs, with hundreds of subtasks and variants implemented.
SP013 OpenAI (GitHub) openai/evals: Evals is a framework for evaluating LLMs and LLM systems You can now configure and run Evals directly in the OpenAI Dashboard.
SP014 Hugging Face Open LLM Leaderboard
SP015 LMArena (Hugging Face Space) Arena Leaderboard (HuggingFace)
SP016 Scale AI Scale AI — Reliable AI Systems Benchmarking the frontier of AI capability with expert-level evaluations.
SP017 Scale AI Scale GenAI Platform Every agent is built and tested against your specific enterprise standards — your workflows, your rules, your definition of good — before it ever touches production.
SP018 Stanford CRFM Holistic Evaluation of Language Models (HELM) — Latest
SP019 Stanford CRFM Holistic Evaluation of Language Models (HELM) — Classic
SP020 BenchLM LLM Leaderboard 2026 — Compare 261 AI Models Across 249 Benchmarks 261 models · 249 benchmarks. The most comprehensive LLM comparison tool — 249 benchmarks, real pricing, and runtime data in one place.
SP021 Artificial Analysis AI Model & API Providers Analysis
SP022 MetaTech.dev LMArena AI Benchmarking Crisis: Why Model Rankings Are Broken The problem becomes even more concerning when we look at critical applications... This disconnect between benchmark success and practical reliability is exactly what's wrong with current AI benchmarking approaches.
SP023 LMSYS Org (UC Berkeley) LMSYS Org — Large Model Systems Organization The Large Model Systems Organization develops large models and systems that are open, accessible, and scalable.
SP024 U.S. Securities and Exchange Commission (EDGAR) EDGAR Search Results — Recall Capital-LMArena a Series of CGF2021 LLC
SP025 U.S. Securities and Exchange Commission (EDGAR) Form D — Recall Capital-LMArena a Series of CGF2021 LLC (search index entry)
SI001 LMArena (PR Newswire) LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product.
SI002 TechCrunch LMArena lands $1.7B valuation four months after launching its product LMArena's annualized "consumption rate" — as the company describes its annual recurring revenue (ARR) — of $30 million as of December, less than four months after launch.
SI003 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M LM Arena, a crowdsourced benchmarking project that major AI labs rely on to test and market their AI models, has raised $100 million in a seed funding round that values the organization at $600 million.
SI004 LMArena (Arena Intelligence) Fueling the World's Most Trusted AI Evaluation Platform (Series A Blog) This year we saw our community grow by over 25x alongside rapid adoption by AI labs who trust this platform as a gold-standard for evaluating real-world model performance.
SI005 GetLatka LMArena Revenue 2026: $30M ARR, $1.7B Valuation In 2026, LMArena's revenue reached $30M. LMArena reached a $1.7B valuation in 2026, set during its Series A round.
SI006 CTOL Digital LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges The company claims over $30 million in annualized revenue from selling evaluation services to these labs, launching its commercial product only in September 2025. At 57 times that run rate, the valuation prices in not just growth, but the assumption that this inherent conflict can be managed indefinitely.
SI007 U.S. Securities and Exchange Commission Form D — Recall Capital-LMArena a Series of CGF2021 LLC
SI008 Wired LMArena Raises $150 Million to Evaluate the World's Most Powerful AI Models
SI009 arXiv (Singh et al. — Cohere, Stanford, MIT, Ai2) The Leaderboard Illusion
SI010 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark
SI011 LMArena (Arena Intelligence) New Product: AI Evaluations LMArena earns revenue by providing paid AI evaluation services to AI labs and enterprises.
SI012 LMArena (Arena Intelligence) Arena AI: The Official AI Ranking & LLM Leaderboard
SI013 LMArena (Arena Intelligence) Arena Leaderboard
SI014 LMArena (Arena Intelligence) Celebrating Community Impact at LMArena (Two-Year Celebration)
SI015 MetaTech.dev LMArena AI Benchmarking Crisis: Why Model Rankings Are Broken
SI016 LMSYS Org (UC Berkeley) LMSYS Org — Large Model Systems Organization
SI017 BenchLM LLM Leaderboard 2026 — Compare 261 AI Models Across 249 Benchmarks
SI018 LMArena / LMSYS Org (GitHub) FastChat: An open platform for training, serving, and evaluating large language models
SI019 Scale AI Scale AI — Reliable AI Systems
SI020 Scale AI Scale GenAI Platform
SI021 arXiv (Wei-Lin Chiang, Lianmin Zheng et al.) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
SI022 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark
SI023 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SI024 U.S. Securities and Exchange Commission (EDGAR) EDGAR Search Results — Recall Capital-LMArena a Series of CGF2021 LLC
SI025 U.S. Securities and Exchange Commission (EDGAR search index) SEC EDGAR Full-Text Search — Form D filings mentioning LMArena
SI026 VentureBeat LMArena raises $150M at $1.7B valuation to build AI evaluation infrastructure
SI027 Bloomberg UC Berkeley AI Benchmarking Group Chatbot Arena Raises $100 Million
SI028 BusinessWire LMArena Raises $150 Million, Achieves $1.7B Valuation
SI029 StartupWired AI startup LMArena triples valuation to $1.7B in 2026
SE001 Arena Intelligence Inc. Arena AI: The Official AI Ranking & LLM Leaderboard
SE002 Arena Intelligence Inc. About Arena — Crowdsourced AI Model Evaluation Platform
SE003 Arena Intelligence Inc. New Product: AI Evaluations LMArena has already logged 250M+ real conversations, 2M+ monthly votes, and has 3M+ monthly users.
SE004 Arena Intelligence Inc. Arena Leaderboard Policy
SE005 Arena Intelligence Inc. Leaderboard Changelog
SE006 Arena Intelligence Inc. Agent Arena: Causal Evaluation of Agents in the Real World The methodology powering the Agent Arena Leaderboard is different from our previous arenas. Rather than pairwise votes, rankings are calculated using a methodology we call causal tracing.
SE007 Arena Intelligence Inc. WebDev Arena: A Live LLM Leaderboard for Web App Development
SE008 Arena Intelligence Inc. Introducing the Search Arena: Evaluating Search-Enabled AI
SE009 Arena Intelligence Inc. Arena-Rank: Open Sourcing the Leaderboard Methodology Arena-Rank, an open-source Python package for ranking that powers the LMArena leaderboard!
SE010 Arena Intelligence Inc. Does Style Matter in AI Evaluations?
SE011 Arena Intelligence Inc. Introducing Hard Prompts Category in Chatbot Arena
SE012 Arena Intelligence Inc. The Arena-Hard Pipeline
SE013 Arena Intelligence Inc. The Next Stage of AI Coding Evaluation Is Here Record: Every model action (file creation, edit, or execution) is logged and versioned. Snapshots are stored in Cloudflare R2.
SE014 arXiv (Chiang et al., UC Berkeley / LMSYS) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
SE015 arXiv / NeurIPS 2023 (Zheng et al., UC Berkeley / LMSYS) Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
SE016 arXiv (Li et al., UC Berkeley / LMSYS) From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
SE017 arXiv / ICLR 2026 (Miroyan et al., UC Berkeley) Search Arena: Analyzing Search-Augmented LLMs
SE018 NeurIPS 2023 Proceedings (Zheng et al.) Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023)
SE019 GitHub (lm-sys) FastChat: An Open Platform for Training, Serving, and Evaluating LLMs
SE020 GitHub (lmarena) arena-rank: Source Code of Arena Leaderboard Methodology
SE021 GitHub (lmarena org) Arena GitHub Organization
SE022 HuggingFace / lmarena-ai (via Wayback) lmarena-ai (Arena) Organization on HuggingFace
SE023 PyPI arena-rank PyPI package
SE024 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification.
SE025 The Verge Meta got caught gaming AI benchmarks Meta's interpretation of our policy did not match what we expect from model providers.
SE026 LMSYS (UC Berkeley SkyLab) Chatbot Arena: New features and a Elo Rating System (original launch blog)
SE027 U.S. Securities and Exchange Commission SEC Form D: Recall Capital-LMArena (Arena Intelligence Series A vehicle)
SU001 PRNewswire (Arena Intelligence Inc.) LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform LMArena works with leading AI labs and enterprises, including OpenAI, Google and xAI, all drawing on LMArena's evaluations to improve their models for production use cases.
SU002 TechCrunch Podcast The PhD students who became the judges of the AI industry how a team like theirs can build a neutral benchmark when the companies they're ranking are also their backers
SU003 OpenTools.AI LM Arena Under Fire: Allegations of Benchmark Bias Stir AI Industry
SU004 Scale AI SEAL Showdown: Scale's AI Evaluation Platform
SU005 AI Wiki LMArena.org — AI Wiki
SU006 PitchBook Arena Intelligence Company Profile
SU007 Arena Intelligence Inc. LMArena (original domain) — redirects to arena.ai
SU008 U.S. Securities and Exchange Commission (EDGAR) EDGAR Full-Text Search: LMArena Form D filings
SU009 TechCrunch LMArena lands $1.7B valuation four months after launching its product That trajectory, and the startup's popularity, were enough for VCs to pile in for the Series A.
SU010 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SU011 Arena Intelligence Inc. Arena AI: The Official AI Ranking & LLM Leaderboard
SU012 Arena Intelligence Inc. About Arena — Crowdsourced AI Model Evaluation Platform
SU013 Arena Intelligence Inc. New Product: AI Evaluations LMArena has already logged 250M+ real conversations, 2M+ monthly votes, and has 3M+ monthly users.
SU014 Arena Intelligence Inc. Arena Leaderboard Policy
SU015 Arena Intelligence Inc. Leaderboard Changelog
SU016 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others.
SU017 The Verge Meta got caught gaming AI benchmarks Meta's interpretation of our policy did not match what we expect from model providers.
SU018 arXiv (Chiang et al., UC Berkeley / LMSYS) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
SU019 arXiv / ICLR 2026 (Miroyan et al.) Search Arena: Analyzing Search-Augmented LLMs
SU020 GitHub (lm-sys) FastChat: An Open Platform for Training, Serving, and Evaluating LLMs
SU021 HuggingFace / lmarena-ai (via Wayback) lmarena-ai (Arena) Organization on HuggingFace
SU022 Arena Intelligence Inc. Agent Arena: Causal Evaluation of Agents in the Real World In a sample of the heaviest real sessions we saw: a live sports-TV schedule site, an autonomous-underwater-vehicle autopilot, a self-hosted movie-watchlist app.
SU023 Arena Intelligence Inc. Fueling the World's Most Trusted AI Evaluation Platform (Series A blog) This year we saw our community grow by over 25x alongside rapid adoption by AI labs who trust this platform as a gold-standard for evaluating real-world model performance.
SU024 Arena Intelligence Inc. Celebrating Community Impact at LMArena (Two-Year Anniversary)
SU025 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark The release version of Llama 4 has been added to LMArena after it was found out they cheated, but you probably didn't see it because you have to scroll down to 32nd place.
SU026 Discord (LMArena Community) Arena Discord Community Server
SU027 Tech in Asia a16z, Lightspeed back $150M Series A of AI model evaluator LMArena
SU028 U.S. Securities and Exchange Commission (EDGAR) EDGAR Full-Text Search: Arena Intelligence Form D 2025-2026
SR001 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark "Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification." — Sara Hooker, Cohere VP of AI Research
SR002 arXiv (Cohere, Stanford, MIT, AI2) The Leaderboard Illusion (arXiv:2504.20879) "We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired."
SR003 The Verge Meta got caught gaming AI benchmarks "Meta's interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that 'Llama-4-Maverick-03-26-Experimental' was a customized model to optimize for human preference."
SR004 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SR005 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark "Companies can continually optimize their models to better align with the LMSYS user distribution, possibly leading to unfair competition and a less meaningful evaluation."
SR006 UCStrategies AI Benchmarks Are a Game Now — And the Industry Is Cheating to Win "SurgeAI analyzed 500 LMArena votes and disagreed with 52%, finding that 'confidence beats accuracy and formatting beats facts.'"
SR007 LMArena Arena Leaderboard Policy (Last Updated April 30, 2026)
SR008 LMArena Leaderboard Changelog
SR009 Arena Intelligence, Inc. LMArena Privacy Policy (Previous Version, Effective 2025-09-05) "Arena Intelligence, Inc. d/b/a LMArena provides a platform for using, comparing, rating, testing, evaluating, and ranking third-party AI models."
SR010 European Commission (Your Europe) Data protection under GDPR — Your Europe
SR011 CTOL Digital Solutions LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges "LMArena has positioned itself as the independent arbiter of model performance, yet it derives revenue from the same AI labs it evaluates — OpenAI, Google, and xAI among them."
SR012 Didit AI Compliance in the LLM Era: Regulatory Guide 2026
SR013 LMArena Arena AI: The Official AI Ranking & LLM Leaderboard (Homepage)
SR014 TechCrunch LMArena lands $1.7B valuation four months after launching its product
SR015 PR Newswire LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform
SR016 LMArena Fueling the World's Most Trusted AI Evaluation Platform (Series A blog)
SR017 LMArena New Product: AI Evaluations
SR018 ByteIota LMArena Raises $150M at $1.7B Valuation in 4 Months
SR019 StartupWired AI Startup LMArena Triples Valuation to $1.7B in 2026
SR020 The AI Insider LMArena Secures $150M to Build the World's Most Trusted AI Evaluation Platform
SR021 OfficeChai AI Evaluation Platform LMArena Raises Series A At Valuation Of $1.7 Billion "LMArena, formally known as Arena Intelligence Inc., was founded in 2025 by Anastasios N. Angelopoulos (CEO), Wei-Lin Chiang (CTO), and Ion Stoica (Advisor)."
SR022 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SR023 The Business Research Company AI Model Evaluation Platform Market Size and Trends Report 2026
SR024 U.S. Securities and Exchange Commission Form D: Recall Capital-LMArena a Series of CGF2021 LLC (EDGAR filing 0002113470-26-000001)
SR025 TechCrunch The leaderboard 'you can't game,' funded by the companies it ranks (video)
SR026 Latka LMArena Revenue 2026: $30M ARR, $1.7B Valuation
SR027 CTOL Digital Solutions LMArena Raises $150 Million — conflict-of-interest analysis
SR028 ByteIota LMArena — open-source model bias and methodological flaws
SR029 LMArena Leaderboard Changelog — Battles in Direct Update (May 12, 2026) "We observed two new biases in the Battles in Direct voting data and corrected for them in the Bradley-Terry fit: the first is a position bias favoring Model A; the second is an advantage given to models that share an organization with the prior turns of context."
SR030 StartupWired LMArena Triples Valuation — Risks and Challenges Ahead section "As usage grows, so do demands for data security, fairness, and governance. Any misstep could damage credibility, which forms the core of LMArena's value proposition."
SV001 TechCrunch LMArena lands $1.7B valuation four months after launching its product "The startup bolted out of the gate as a commercial venture with a $100 million seed round in May at a $600 million valuation. This new round means it raised $250 million in about seven months."
SV002 PR Newswire LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform "LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product."
SV003 Latka LMArena Revenue 2026: $30M ARR, $1.7B Valuation
SV004 StartupWired AI Startup LMArena Triples Valuation to $1.7B in 2026
SV005 The AI Insider LMArena Secures $150M to Build the World's Most Trusted AI Evaluation Platform
SV006 OfficeChai AI Evaluation Platform LMArena Raises Series A At Valuation Of $1.7 Billion "The company is positioning itself in what analysts estimate is an $800-900 million AI testing market in 2025, projected to grow to $3.8 billion by 2032."
SV007 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SV008 ByteIota LMArena Raises $150M at $1.7B Valuation in 4 Months "Four months to $30 million: In May 2025, LMArena raised $100M at $600M; by September, launched AI Evaluations; three months later, hit $30M annualized run rate."
SV009 LMArena Fueling the World's Most Trusted AI Evaluation Platform (Series A blog) "Since announcing our $100M Seed round last year in May, LMArena has grown far faster than we imagined. In a matter of months, the community has contributed 50 million votes."
SV010 The Business Research Company AI Model Evaluation Platform Market Size and Trends Report 2026 "AI Model Evaluation Platform market size has reached $1.86 billion in 2025; expected to grow to $6.24 billion in 2030 at a CAGR of 27.5%."
SV011 Research and Markets AI Model Evaluation Platform Market Report 2026
SV012 The Business Research Company Human-In-The-Loop AI Market Size and Drivers Report 2026 "In March 2025, CoreWeave Inc. acquired Weights & Biases Inc. for approximately $1.4 billion."
SV013 CTOL Digital Solutions LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges "At 57 times that run rate, the valuation prices in not just growth, but the assumption that this inherent conflict can be managed indefinitely."
SV014 UCStrategies AI Benchmarks Are a Game Now — And the Industry Is Cheating to Win
SV015 U.S. Securities and Exchange Commission Form D: Recall Capital-LMArena a Series of CGF2021 LLC (EDGAR accession 0002113470-26-000001) "Recall Capital-LMArena a Series of CGF2021 LLC — Pooled Investment Fund / Venture Capital Fund — Amount Sold: $382,500"
SV016 arXiv (Cohere, Stanford, MIT, AI2) The Leaderboard Illusion (arXiv:2504.20879)
SV017 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark
SV018 LMArena Arena Leaderboard Policy (Last Updated April 30, 2026)
SV019 LMArena New Product: AI Evaluations
SV020 LMArena Leaderboard Changelog
SV021 Arena Intelligence, Inc. LMArena Privacy Policy (Previous Version, Effective 2025-09-05) "Arena Intelligence, Inc. d/b/a LMArena provides a platform for using, comparing, rating, testing, evaluating, and ranking third-party AI models."
SV022 TechCrunch The leaderboard 'you can't game,' funded by the companies it ranks (video)
SV023 The Verge Meta got caught gaming AI benchmarks
SV024 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SV025 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark
SV026 Grokipedia Arena (LMArena) — Grokipedia overview "As of March 5, 2026, the top models on the text leaderboard are claude-opus-4-6 (1504 Elo), gemini-3.1-pro-preview (1500 Elo); total 5,430,034 votes collected across the platform."
SV027 TLDL AI Company Rankings 2026: Revenue, Funding & Valuation Data for 2,000+ Companies "Private funding for AI startups topped $150 billion over the trailing twelve months; foundation model companies raising $80 billion in 2025."
SV028 Wellows 85 Hottest AI Startups to Watch in 2026 [By Valuation, Funding, & Growth] "Anysphere (Cursor): AI coding assistant, $29.3B valuation, $1B ARR; Harvey: legal AI; LMArena reached a $1.7 billion valuation in under four months."
SV029 AgentMarketCap LMArena's $1.7B Valuation in 4 Months: Why AI Evaluation Is the New Data Labeling "Scale AI was valued at $7B in 2021. In June 2025, Meta acquired a 49% stake for $14.3 billion — the largest VC transaction of 2025 — implicitly valuing Scale above $29 billion."
SV030 Axis Intelligence Research AI Copyright Lawsuits 2026: Status Tracker — Updated Monthly "$50 billion+: Cumulative legal exposure across all active AI copyright and related IP cases; Anthropic's confirmed settlement in Bartz v. Anthropic: $1.5 billion covering ~482,000 works."
SV031 U.S. Securities and Exchange Commission (EDGAR) EDGAR Company Search: Recall Capital-LMArena a Series of CGF2021 LLC (CIK 0002113470)
SV032 Copyright Alliance AI Copyright Lawsuit Developments in 2025: A Year in Review "Anthropic's $1.5B settlement in Bartz v. Anthropic required payment of approximately $3,000 for each of the 482,460 books downloaded from pirate libraries — the first publicly confirmed pricing benchmark for AI training on pirated content."
SV033 U.S. News & World Report / Reuters AI Startup LMArena Triples Its Valuation to $1.7 Billion in Latest Fundraise "LMArena said on Tuesday its valuation had tripled to $1.7 billion in about eight months, following a new funding round where it raised $150 million."