LMArena
这个评测品类由它定义,增长动能真实,但当前估值已经计入了超常增长和治理修复
LMArena 看起来已是 AI 评测类别龙头;但在治理、中立性和客户集中度问题解决前,$1.7B Series A 轮价格几乎不给安全边际。
封面要素
公司概况
LMArena 源自 UC Berkeley LMSYS / Chatbot Arena 在 2023 年启动的研究项目,并在 2025 年 4 月商业化为 Arena Intelligence Inc.。它的核心产品是一个盲测、成对的人类偏好评测平台:用户比较匿名 AI 模型输出,由此生成公开排行榜和一条专有的真实世界偏好数据流。公司通过向模型实验室、企业和开发者提供付费 AI 评测服务变现,其中包括 2025 年 9 月推出的 AI Evaluations 产品,提供可审计性、代表性样本报告和服务级别协议。LMArena 的增长很少见:公开材料显示,月活用户 500 万、月度对话 6000 万、种子轮后累计投票 5000 万次,并在 2025 年 12 月达到 $30M 年化消费运行率。同一门生意也有结构性张力:最重要的客户和投资人,与平台所评测模型背后的实验室高度重叠。
- 成立时间
- 2023-05-03
- 创始人
- Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica
- 创立地点
- UC Berkeley / San Francisco Bay Area, California
- 总部
- San Francisco Bay Area, California, USA
- 产品
- 众包 AI 模型评测平台,覆盖文本、代码、搜索、视觉、图像、视频和智能体基准;另有付费企业评测产品,把基于社区的测试、代表性对战样本、排行榜基础设施和 SLA 支持的分析打包交付。
- 客户
- Frontier AI 实验室、企业 AI 采购方、产品团队和开发者;这些客户需要用可比较的模型评测来支持发布决策、采购和产品质量衡量。
- 商业模式
- 评测即服务:在公开 Arena 排行榜和开放评测数据集之上,叠加付费定制与私有模型评测、企业基准测试及相关平台服务。
- 阶段
- Series A
- 融资情况
- 2025 年 5 月以 $600M 投后估值完成 $100M 种子轮,随后在 2026 年 1 月以 $1.7B 投后估值完成 $150M Series A,披露融资总额约 $250M。
执行摘要
主要优势
- 公开 Arena 排行榜已成为前沿模型发布的事实参照点,让 LMArena 拥有少见的数据和分发护城河。
- 社区规模是真实的:5M 月用户、60M 月对话和 50M 投票,拼出一个新进入者很难快速复刻的评测数据集。
- 公司商业化推进很快,2025 年 9 月推出 AI Evaluations,2025 年 12 月达到 $30M 年化运行率。
- 产品扩展到搜索、编程、多模态和智能体基准,把可想象 ARR 天花板从单纯文本模型对比往上打开。
- 一线投资人支持和扎实学术根基,弥补了有限运营历史,也提升了 AI 实验室和企业买家的信任。
主要风险
- 结构性利益冲突明显:主要付费实验室和投资人与 LMArena 公开评测的模型主体存在重叠。
- Leaderboard Illusion 论文和 Meta Maverick 事件已经挑战基准完整性,真实信誉风险上升。
- $1.7B 估值约等于已披露 $30M 年化运行率的 57x,几乎容不下执行失误或倍数压缩。
- 收入质量仍不透明,因为公司披露的是运行率代理指标,而不是审计收入、毛利率或客户集中度数据。
- 对一家扮演裁判角色的公司来说,治理披露偏薄:公开证据尚未显示董事会构成、独立性控制或详细股权结构表条款。
未决问题
- 头部客户集中度和净留存率尚未公开披露。
- 毛利率、算力成本结构和烧钱速度仍无法从公开来源取得。
- 尚无独立统计审计公开解决 2025 年提出的中立性担忧。
- 董事会构成、投资人权利和清算优先权细节仍未公开。
目录
01公司概览
1.1 身份与产品
LMArena 现在以「Arena」品牌运营,是一个 AI 模型评测平台,让用户通过匿名一对一对战比较前沿 AI 模型。用户同时向两个未标注模型提交提示词,投票选出更好的回答,然后才看到被评测的是哪些模型。聚合投票通过 Bradley-Terry / Elo 排名算法驱动公开排行榜,持续生成大型语言模型和多模态 AI 系统在真实世界中的表现对比。公司实体 Arena Intelligence Inc. 于 2025 年 4 月 18 日注册成立,把 Chatbot Arena 研究项目转为商业公司。总部位于 San Francisco Bay Area(具体地址未披露),与 UC Berkeley 渊源一致。主要网站为 arena.ai 和 lmarena.ai。LMArena 的商业产品 AI Evaluations 于 2025 年 9 月推出,为企业、模型实验室和开发者提供基于社区的表现分析、代表性反馈样本和服务级别协议。已具名商业客户包括 OpenAI、Google 和 xAI。商业模式是评测即服务:客户付费,让公司系统评估模型在软件工程、法律、医疗等领域的表现。平台已经远远超出纯文本 LLM 比较,扩展到 Search Arena、WebDev Arena、Vision Arena、文生图、文生视频,以及 2026 年 6 月推出的 Agent Arena。多模态扩张扩大了 LMArena 可评测的表面,也扩大了可触达客户群。截至 2026 年 6 月,Arena 品牌已经成为前沿 AI 模型事实上的公开排行榜,其排名被所有主要 AI 实验室写入产品发布、投资人材料和学术论文。 [CO001, CO002, CO005, CO007, CO008, CO009]
LMArena 的社区输入、平台基础设施、数据产品和商业输出如何连接。
[CO007, CO017, CO020, CO024, CO025, CO026]截至 2026 年 1 月,在规模、财务和社区维度上的关键绩效指标。
数值为公司截至 2026 年 1 月披露;私营公司没有独立审计。ARR 指年化消耗运行率,不是 GAAP 收入。
[CO007, CO011, CO013, CO015, CO017, CO018]1.2 创始人与领导层
LMArena 由三位具有 UC Berkeley 背景的人士共同创立。CEO Anastasios Angelopoulos 曾是 UC Berkeley 统计学与机器学习博士后研究员,也是公司主要公开发言人。联合创始人 Wei-Lin Chiang 在 Berkeley 完成分布式系统博士学位,是搭建支撑 Chatbot Arena 的 FastChat 服务框架的关键成员。UC Berkeley 计算机科学教授、连续创业者 Ion Stoica(Databricks 和 Anyscale 联合创始人)以联合创始人身份加入,带来商业化经验和产业可信度。最初的 Chatbot Arena 研究还包括 Lianmin Zheng 等更多贡献者,以及 Michael Jordan、Joseph Gonzalez 等教职顾问。创始团队展现出强创始人—市场匹配:Angelopoulos 和 Chiang 合著了奠定评测方法的关键学术论文,Stoica 在 Databricks 的履历则带来学术 spinout 少见的运营深度。关键人依赖偏高:Angelopoulos 同时担任科学负责人和 CEO。公司尚未公开任命非创始人的 C-suite 高管。董事会组成或独立董事也未披露,这对早期私营公司并不罕见。自注册成立以来,公司未披露领导层离职或治理变更。不过,2025 年基准完整性争议暴露了 Stoica(公开反驳 Leaderboard Illusion 发现)与外部研究者之间的张力,也凸显创始团队可信度对平台合法性有多关键。 [CO003, CO004, CO005, CO006, CO021, CO030]
| 姓名 | 角色 | 背景 | 创始人-市场匹配 | 关键人风险 |
|---|---|---|---|---|
| Anastasios Angelopoulos | CEO、联合创始人 | UC Berkeley 博士后,Statistics/ML;Chatbot Arena 和 Arena-Hard 论文共同作者(arXiv:2403.04132、2406.11939) | 极高——评估方法论技术架构师,也是公司主要公开发声者 | 极高——同时担任科学权威和 CEO;离职会影响产品可信度和投资者信心 |
| Wei-Lin Chiang | 联合创始人 | UC Berkeley PhD,分布式系统;构建了支撑 Chatbot Arena 的 FastChat serving 框架;核心方法论论文共同作者 | 高——原始平台架构师;技术执行深 | 高——核心工程专长;公开曝光有限,说明可能承担关键运营角色 |
| Ion Stoica | 联合创始人 | UC Berkeley CS 教授;Databricks(约 $43B 估值)和 Anyscale(Ray framework)联合创始人;连续创业者 | 高——商业化经验和 venture-building 可信度 | 中等——可能是顾问 / 董事会层面角色;既往公司建设记录降低单点依赖 |
| Michael Jordan | 学术顾问(LMSYS 论文) | UC Berkeley ML 先驱;在概率 ML、贝叶斯方法和统计学习理论上有基础性工作 | 声誉型——学术声望为方法论主张背书 | 低——顾问贡献;无运营角色 |
| Joseph E. Gonzalez | 学术顾问(LMSYS 论文) | UC Berkeley Systems+ML 教师;共同开发 Ray 分布式计算框架;核心论文共同作者 | 技术型——系统专长与基础设施扩展相关 | 低——顾问;无运营角色 |
董事会组成、非创始人 C-suite(CTO、CFO、VP Sales/BD)以及投资人委派董事未公开披露。枚举仅基于公开记录来源,且并不完整。
[CO003, CO004, CO005, CO006, CO033]1.3 融资历史与投资人
LMArena 商业化前的资金来自资助和捐赠,包括 Google 的 Kaggle 平台、Andreessen Horowitz 和 Together AI 的贡献,主要以算力资源和现金形式支持研究基础设施。2025 年 5 月,公司注册成立后,LMArena 完成 $100M 种子轮,由 Andreessen Horowitz 与 UC Investments(University of California 的捐赠基金)共同领投,投后估值 $600M。Lightspeed Venture Partners、Felicis 和 Kleiner Perkins 也参与投资。2026 年 1 月,公司以 $1.7B 投后估值完成 $150M Series A,估值接近种子轮的三倍,由 Felicis 和 UC Investments 共同领投,Andreessen Horowitz、The House Fund、LDVP、Kleiner Perkins、Lightspeed 和 Laude Ventures 参投。截至 2026 年 1 月,融资总额约 $250M。UC Investments 同时是 University of California 捐赠基金管理方,也是 UC Berkeley spinout 的机构支持者;Andreessen Horowitz 的投资组合又与 Arena 排行榜上的模型(包括 Mistral)重叠。这些都构成外部批评者指出的潜在利益冲突。公司尚未公开披露二级交易或债务融资安排。 [CO010, CO011, CO012, CO013, CO014, CO021]
| 利益相关方 | 角色 | 轮次 | 重要性 | 尽调问题 |
|---|---|---|---|---|
| Felicis Ventures | Series A 领投方 | Series A | 共同领投;GP Peter Deng 在官方新闻稿中被引用;可能拥有董事席位或观察员权利 | 董事席位构成;pro-rata 权利;term sheet 中的 full-ratchet 条款 |
| UC Investments (University of California) 投资机构 | 种子轮 + Series A 领投方 | 种子轮 + Series A | 双重角色支持方:既是捐赠基金投资人,也是 UC Berkeley spinout 的机构支持者;存在结构性利益冲突 | 捐赠基金角色与学术 IP 来源之间的利益冲突协议;来自 UC Berkeley 关联的信息权;IP 授权条款 |
| Andreessen Horowitz (a16z) 投资机构 | 投资人 | 种子轮 + Series A | Tier-1 VC;两轮均参与;a16z 组合与 Arena 排行榜模型提供商存在重叠(Mistral 投资) | 组合冲突披露;Mistral 与 Arena 评估的接近程度;获得预发布模型测试优先访问的可能性 |
| Kleiner Perkins | 投资人 | 种子轮 + Series A | 老牌 VC;连续两轮参与传递 conviction;可能拥有董事观察员权利 | 观察员席位条款;治理权;反稀释条款 |
| Lightspeed Venture Partners | 投资人 | 种子轮 + Series A | 活跃早期科技投资人;连续两轮参与 | 信息权;未来轮次 pro-rata;治理结构 |
| The House Fund | 投资人 | Series A | UC Berkeley 关联基金;与创始机构存在关联方关系 | 关联方交易披露;相对 arms-length 投资人的财务条款;治理重叠 |
| LDVP | 投资人 | Series A | 科技 VC;关于条款或治理角色的公开信息有限 | 基金重点;既往 AI 组合公司;治理参与程度 |
| Laude Ventures | 投资人 | Series A | 较小轮次参与方;无额外公开信息 | 基金背景;治理参与;pro-rata 权利 |
各投资人的单独投资金额未公开披露。UC Investments 同时作为捐赠基金投资人和 UC Berkeley 机构支持者,存在需要尽调的结构性利益冲突。二级股东、可转债或 SAFEs 均未公开披露。
[CO010, CO011, CO012]1.4 增长指标与财务信号
到 2026 年 1 月,LMArena 报告称其月活用户超过 500 万,覆盖 150 个国家,每月产生 6000 万次对话。社区在文本、视觉、网页开发、搜索、视频和图像等模态中累计超过 5000 万次投票,并参与评测 400 多个不同 AI 模型。LMArena 还发布了来自专家和职业评测类别的 14.5 万个开源对战数据点。财务侧,公司 ARR 代理指标——年化消费运行率——在 2025 年 12 月超过 $30M,距离商业化推出约四个月。这个数字被表述为年化运行率,而非已实现 GAAP 收入;独立审计数据不可得。四个月内从零冲到 $30M 年化极不寻常,但商业服务结构(带 SLA 的交付物加社区评分者)意味着毛利率可能受到劳动力成本约束,公司尚未披露。LMArena 没有披露员工数、烧钱速度或毛利率,仅靠公开信息很难完成财务尽调。去重过滤器会移除约 10% 的提交投票;截至 2025 年 7 月方法更新,身份泄露检测移除的投票少于全部投票的 4%,说明公司在主动管理数据质量。 [CO015, CO017, CO018, CO019, CO020, CO023]
| 指标 | 数值 / 状态 | 日期 | 置信度 | 缺口 / 尽调路径 |
|---|---|---|---|---|
| Series A 后估值 | $1.7 billion(投后) | Jan 2026 | 高 | 未披露投前估值;无可用二级市场定价 |
| 累计融资额 | ~$250 million | Jan 2026 | 高 | 新闻稿和多家新闻来源确认 |
| 种子轮估值 | $600 million(投后) | May 2025 | 高 | 公司宣布;TechCrunch 和 Bloomberg 确认 |
| ARR(消耗 run rate) | $30 million+ | Dec 2025 | 中 | 年化 run-rate 代理;未经审计的 GAAP 收入;无 2026 年中更新 |
| 月活跃用户 | 5 million+ | Jan 2026 | 中 | 公司报告;「active」定义未说明;无独立审计 |
| 月度对话 | 60 million+ | Jan 2026 | 中 | 公司报告;无第三方验证 |
| 累计总投票 | 50 million+ | Dec 2025 | 中 | 包含已弃用模型对战;拆分不可得 |
| 已评估模型 | 400+ | Dec 2025 | 中 | 包含已弃用和私测模型;确切活跃数量未知 |
| 覆盖国家 | 150 | Jan 2026 | 中 | 反映用户地域;并非在 150 个国家注册实体 |
| 员工数 | 未披露 | Jun 2026 | 低 | 未公开披露;Series A 资金指定用于技术团队扩张 |
| 毛利率 / burn rate | 未披露 | Jun 2026 | 低 | 私营公司;无财务披露;人工成本结构未知 |
数值来自公司新闻稿、TechCrunch 报道和投资人公告。「Consumption run rate」是 LMArena 对年化 ARR 代理的称呼;不等同于按 GAAP 确认的收入。Null / Undisclosed 字段表示没有公开披露,不是零值。
[CO011, CO013, CO014, CO015, CO017, CO018]1.5 里程碑与负面事件
LMArena 从研究演示走到独角兽,历时约三年。最初的 Chatbot Arena 于 2023 年 5 月上线。学术论文快速跟进,最终以 Chatbot Arena 论文(arXiv:2403.04132)确立 Bradley-Terry 方法。到 2024 年,平台已成为事实上的参考基准,被每家主要 AI 实验室引用。首个重大负面事件发生在 2025 年初:Meta 在 Chatbot Arena 上测试了至少 27 个私有 Llama 4 模型变体,提交了一个为 Arena 优化的版本,得分接近榜首,而公开发布版本排名第 32。LMArena 随后道歉并更新排行榜政策。2025 年 4 月,一篇题为《The Leaderboard Illusion》(Cohere、Stanford、MIT、Ai2)的同行评议论文正式记录了对 LMArena 私有测试做法存在系统性偏差的指控。LMArena 联合创始人 Ion Stoica 公开称这些发现存在「不准确」。作为回应,LMArena 引入新的采样算法并发布更新后的透明度政策。公司 2025 年 4 月注册成立,2025 年 5 月融资 $100M,2025 年 9 月推出商业产品,并在 2026 年 1 月完成 Series A,跻身独角兽。截至 2026 年 6 月,未公开披露监管调查、诉讼、制裁或执法行动。 [CO001, CO002, CO005, CO016, CO027, CO028]
| 日期 | 事件 | 类型 | 金额 / 状态 | 参与方 | 含义 |
|---|---|---|---|---|---|
| May 2023 | Chatbot Arena 作为公开研究 demo 上线 | 创立 | 志愿研究项目 | UC Berkeley LMSYS 团队(Angelopoulos、Chiang 等) | 首个公开众包 LLM 排行榜;在技术社区内快速自然采用 |
| Jun 2023 | MT-Bench 和 Chatbot Arena NeurIPS 论文提交 | 产品 | 研究发表(arXiv:2306.05685) | Zheng、Chiang、Angelopoulos 等 | LLM-as-Judge 概念确立;30K 次对话和 3K 次专家投票公开发布 |
| Mar 2024 | Chatbot Arena 正式平台论文发表 | 产品 | 研究发表(arXiv:2403.04132) | Chiang、Zheng 等;累计 240K+ 投票 | Bradley-Terry/Elo 方法论正式化;所有主要 AI 实验室在产品公告中引用 |
| Jun 2024 | Arena-Hard-Auto 和 BenchBuilder pipeline 论文发表 | 产品 | 研究发表(arXiv:2406.11939) | Li、Chiang 等 | 自动化基准策划约 $20/run;证明与人类偏好有 98.6% 相关性 |
| Jan–Mar 2025 | Meta 在 Chatbot Arena 私下测试 ≥27 个 Llama 4 模型变体 | 反向 | 未披露私测;只提交优化后变体 | Meta AI、LMArena | 首次重大完整性争议;公开发布的 Llama 4 Maverick 排名第 32,而 arena 优化版本排名第 2 |
| Apr 18, 2025 | Arena Intelligence Inc. 成立 | 创立 | 公司组建 | Angelopoulos、Chiang、Stoica | 从学术项目正式转为商业实体;融资流程启动 |
| Apr 29, 2025 | 《The Leaderboard Illusion》论文发表 | 反向 | 同行评审论文(arXiv:2504.20879) | Singh、Hooker 等(Cohere、Stanford、MIT、Ai2) | 学术界正式质疑基准完整性;数据访问不对称已有记录;构成重大可信度风险 |
| May 2025 | 以 $600M 估值完成 $100M 种子轮 | 融资 | 融资 $100M;投后估值 $600M | 投资人:a16z、UC Investments、Lightspeed、Felicis、Kleiner Perkins | 首轮大型风险融资;种子轮即接近独角兽;显示 VC 看好 AI 评测市场 |
| Sep 2025 | AI Evaluations 商业产品上线 | 产品 | 商业化发布 | LMArena 团队;OpenAI、Google、xAI 为锚定客户 | 收入开始产生;验证 B2B 评测即服务模型 |
| Dec 2025 | ARR 消耗口径运行率突破 $30M | 规模 | 年化运行率 $30M | LMArena 商业团队 | 从零到 $30M 年化不到四个月;释放企业快速采用信号 |
| Jan 2026 | 以 $1.7B 估值完成 $150M Series A | 融资 | 融资 $150M;投后估值 $1.7B | Felicis(领投)、UC Investments(共同领投)、a16z、LDVP、Kleiner Perkins、Lightspeed、The House Fund、Laude Ventures | 达到独角兽状态;累计融资约 $250M;从产品发布到 Series A 约 7 个月 |
| Apr 2026 | 更新透明度政策;开源 Arena-Rank | 治理 | 政策文件发布于 arena.ai/blog/policy/ | Arena 团队 | 回应持续的基准完整性批评;正式写入抽样规则和数据共享承诺 |
| Jun 2026 | 推出采用因果推断方法的 Agent Arena | 产品 | 新产品垂直 | Arena 团队 | 借助处理效应估计,把评测 TAM 扩展到 agentic AI;这是 Series A 后首个重要产品发布 |
事件日期来自新闻稿和发布时间戳;精确到日的程度不一。负面事件按尽调范围纳入。金额 / 状态反映一手来源数字;未披露金额按未披露标注。
[CO001, CO002, CO005, CO010, CO011, CO016]从 2023 年研究项目上线到 2026 年 6 月的关键节点,包括融资事件、产品发布和负面事件。
日期根据新闻稿和发布时间戳估算;不同来源的日级精度不一。
[CO001, CO005, CO010, CO011, CO015, CO016]1.6 展示项
02市场分析
2.1 市场边界与定义
LMArena 应被放在 AI 评测与基准测试软件层中分析,而不是泛模型基础设施或可观测性公司。纳入的市场包括人类偏好基准测试、自动化回归测试、面向监管或重专业知识场景的领域评测工作流,以及更新的智能体评测产品——后者评估多步任务成功率,而不只看单轮聊天质量。可服务市场不包括基础模型训练算力、通用 MLOps 编排、广义开发者工具,以及未作为托管工作流软件销售的开源评测脚本。这个边界很重要:有些发布方测算的是宽口径评测平台品类,有些则测算狭义基准工具细分,导致 TAM 相差数倍。LMArena 自身商业信息强调法律、医疗和工程评测,说明它可变现市场更接近高风险验证支出,而不是单靠公开排行榜流量定义。[CM006, CM007, CM008, CM009, CM016, CM018]
| 类别 | 纳入支出 | 排除支出 | 主要买家 | 对 LMArena 的意义 |
|---|---|---|---|---|
| 公开基准和排行榜运营 | 人类偏好基准、并排模型比较、公共信任信号、基准赞助 | 核心训练算力、原始推理支出、通用流量变现 | 前沿实验室、模型 API 厂商、基准赞助方 | 这是为 LMArena 打开市场能见度的声誉切口 |
| 企业模型评测工作流 | 回归测试、评测数据集、领域评分、发布闸门、QA 仪表盘 | 通用 BI、工单系统,或无关的开发者生产力工具 | AI 平台团队、模型质量负责人、领域产品负责人 | 经常性软件支出最可能从公开排行榜使用延续下来的场景就在这里 |
| Agent 和工作流评测 | 多步骤任务成功评分、因果轨迹、工作流基准、工具使用可靠性 | 不带测量的一般 agent 编排、通用 copilots | 把 agents 嵌入工作流的应用 AI 团队 | Agent Arena 把市场从聊天机器人排名扩展到执行可靠性 |
| 受监管和高风险垂直验证 | 法律、医疗、工程和合规敏感评测项目 | 没有业务关键决策路径的消费者娱乐聊天排名 | 领域负责人、风险负责人、质量团队 | LMArena 明确把这些领域作为可变现滩头阵地来营销 |
| 捆绑式平台评测 | 嵌入云、MLOps 套件和实验平台的评测功能 | N/A | AWS、Google、Microsoft、Databricks、MLflow 用户 | 生态里确实存在这类支出,但独立供应商只能触达其中一部分 |
| 相邻但排除的基础设施 | 训练数据管线、基础模型托管、推理服务、通用可观测性 | 这些均不属于本范围内的评测软件市场 | 基础设施团队和 CTO 预算 | 排除这些项目可避免夸大可服务市场 |
边界行刻意限定在可变现评测软件,而不是所有 AI 工具;捆绑式平台评测作为背景纳入,但 LMArena 只能触达其中一部分。
[CM006, CM007, CM008, CM009, CM016, CM018]2.2 市场规模与有争议估计
在已验证来源中,最干净的 2026 年宽口径市场锚点是 AI 模型评测平台市场报告:2025 年 $1.86B,2026 年增至 $2.36B,CAGR 为 27.3%,2030 年达到 $6.24B。这个视角大概率包含面向模型实验室、应用开发者和强合规组织销售的企业评测软件。Precedence Research 的窄口径视角则指向明显更小的基准工具细分,2026 年约 $0.85B,因为它似乎排除了部分打包或相邻工作流,并强调评测加基准工具,而非完整平台层。Gartner 的 2026 年 AI 支出预测和 Presenc AI 的生产采用调查都支持评测预算会从当前基数快速扩张,但无法解决边界问题。务实结论是:LMArena 真正可变现市场可能远小于宽口径 TAM;但如果它能拿下可信、嵌入工作流的支出,仍足以支撑一家有意义的独立公司。[CM001, CM002, CM003, CM004, CM005, CM021]
| 来源 | 年份 | 范围 | 指标 | 数值 | 增长 / 前景 | 解读 | 关键限制 |
|---|---|---|---|---|---|---|---|
| The Business Research Company | 2026 | 全球 AI 模型评测平台 | 2025-2026 市场规模 | $1.86B(2025)至 $2.36B(2026) | 27.3% CAGR;2030 年达 $6.24B | 本章中经过验证的最佳宽口径平台 TAM 锚点 | 方法细节只做摘要,未完整披露 |
| Yahoo Finance / Research & Markets 转发 | 2026 | 全球 AI 模型评测平台 | 2026 市场规模 | 2026 年 $2.36B | 27.3% CAGR | 独立转发大体印证宽口径市场数字 | 是联合发布的新闻稿报道,不是原始模型工作簿 |
| Research & Markets | 2026 | 全球 AI 模型评测平台 | 2030 年预测 | 2030 年达 $6.24B | 27.3% CAGR | 如果宽口径定义成立,则确认未来增长强劲 | 仍是宽泛类别,细分拆解不清 |
| Precedence Research | 2026 | 模型评测和基准工具 | 窄口径 2026 视角 | 2026 年约 $0.85B | 约 7.3% CAGR | 可作为更窄基准工具类别的有用下限 | 不能与宽口径平台 TAM 直接比较;类别看起来更窄 |
| Precedence Research | 2025-2034 | 模型评测和基准工具 | 长期预测 | 约 2034 年达 $9.57B | 长周期增长市场 | 显示该类别的重要性和战略价值,包括 M&A 背景 | 预测期限和范围不同于 2026 年宽口径市场来源 |
| Gartner | 2026 | 全球 AI 经济 | AI 软件支出 | 2026 年 AI 软件 $453B | 属于全球 AI 支出 $2.59T 的一部分,同比 +47% | 支撑评测采购的上游预算池非常大 | 不是直接的评测市场指标 |
| Gartner | 2026 | 全球 AI 经济 | AI 模型支出 | 2026 年 AI 模型 $32.6B | 随软件支出同步快速扩张 | 可作为前沿实验室预算可得性的有用代理 | 仍然间接;没有隔离第三方评测供应商 |
| 基于已验证来源的作者估算 | 2026 | 与 LMArena 相关的独立评测 SAM / SOM | 分析区间 | SAM 约 $0.3-0.8B;SOM 约 $0.03-0.15B | 由宽口径 TAM、窄口径视角和已报道 ARR 推导 | 可作为尽调决策区间,但不是出版方背书的估算 | 缺少公开定价、预算占比或客户数量细节,无法精确验证 |
本表混合了出版方数字和一个明确标注的分析估算;用户在把各行视为可直接相加或互相矛盾之前,应先比较定义。
[CM001, CM002, CM003, CM004, CM005, CM021]从广义 AI 评测平台,到与 LMArena 相关的 SAM,再到近期 SOM 的三层视角。
SAM 和 SOM 是根据已验证的广义 TAM、狭义基准测试视角和披露 ARR 推导出的分析区间;不是出版方发布的数字。
[CM001, CM022, CM023, CM043]低位、中位和高位市场视角显示,品类定义如何改变 LMArena 市场规模的表观大小。
所有行均以十亿美元计;高位行合并了 2030-2034 年方向性品类端点,而不是单一年份。
[CM001, CM003, CM046]2.3 买方分层与采购
LMArena 的买方地图有两个不同重心。第一类是前沿实验室和模型 API 供应商,它们在重大版本发布前后需要可信第三方基准、发布验证和竞争信号。第二类是监管或专业知识密集型企业,它们在法律、医疗和工程工作流中需要领域评测,因为失败成本高,公开消费者基准不够用。这些细分大概率采购方式不同:实验室从中央模型、安全或研究平台预算中支付评测费用;企业则常由 AI 平台负责人、产品 owner,或风险与质量团队采购。竞争也很混合。Scale AI、Arize、Galileo、Patronus、MLflow,以及大型云或平台厂商都在攻打相邻栈位;CoreWeave 收购 Weights & Biases 也说明,实验、可观测性和评测工作流正在企业采购中收敛。[CM010, CM011, CM012, CM013, CM014, CM015]
| 细分 | 买家 | 用户 | 付费方 | 工作流 | 预算负责人 | 采用触发因素 |
|---|---|---|---|---|---|---|
| 前沿模型实验室 | OpenAI、Google、xAI、Anthropic 类实验室 | 评测研究员、发布经理、安全团队 | 中央模型开发预算 | 在发布前后对新模型做基准测试;与竞争对手比较 | 模型质量 VP / 研究平台负责人 | 竞争性发布节奏和可信外部证明需求 |
| 模型 API 厂商和平台提供商 | 托管模型平台和 AI 云 | 平台 PM、信任团队、GTM 团队 | 产品或平台预算 | 用外部和内部评测支撑企业销售和发布主张 | 平台 GM 或产品负责人 | 拥挤 API 市场里需要区分模型质量 |
| 法律 AI 厂商和企业 | 法律工作流团队、法律科技买家 | 律师、审阅员、AI 产品经理 | 业务单元或创新预算 | 验证领域准确性、引用质量和工作流可靠性 | 法律创新负责人或 AI 平台负责人 | 法律工作流中幻觉成本高 |
| 医疗和健康相关 AI 买家 | 临床 AI 团队、医疗文档厂商 | 临床医生、质量团队、模型验证人员 | 产品、合规或临床运营预算 | 评估安全性、术语准确性和失败阈值 | 首席医疗 AI 负责人或质量负责人 | 患者安全和合规风险让验证支出更容易被证明合理 |
| 工程 copilots 和工业知识工作流 | 工程软件团队和应用 AI 小组 | 工程师、分析师、技术审阅员 | R&D 或产品工程预算 | 测试任务完成度、工具使用和领域正确性 | 应用 AI 或工程系统负责人 | 大规模推出前需要可衡量的生产力提升 |
| 基准生态合作伙伴 | 赞助方、评测方和相邻工具厂商 | 面向市场的研究、开发者关系、信任团队 | 市场、产品或生态预算 | 用基准塑造叙事、合作动作或集成工作流 | 产品营销或生态负责人 | 需要锚定类别可信度和公开可比性 |
买家行反映经验证产品信息和媒体报道支撑的最高概率付费细分,而不是完整客户名单。
[CM010, CM012, CM013, CM014, CM015, CM016]展示主要买方细分在不同采购标准下如何评估 LMArena。
序数标签概括每个细分市场对各项标准的相对重视程度,不是调研分数。
[CM015, CM017, CM018, CM044]2.4 增长驱动与采用约束
多股力量支撑可信评测供应商跑赢市场。2026 年全球 AI 软件和模型支出激增,企业 AI 在大型公司中已高比例进入生产,向智能体转移又提高了对工作流级评测的需求,而不是静态提示词测试。基准竞争本身也会刺激需求,因为模型提供商需要外部证明点,企业需要独立质量信号。但这个品类也有真实约束。基准刷榜指控、众包评分者偏差问题,以及对公开排行榜的怀疑,都会削弱客户为未明确绑定生产结果的分数付费的意愿。独立供应商还可能受到捆绑云工具和 MLOps 工具的价格压力;公开定价披露有限,也让人很难区分持久软件需求和新鲜感驱动的试验。[CM026, CM027, CM028, CM029, CM030, CM031]
| 因素 | 类型 | 时点 | 影响 | 尽调问题 |
|---|---|---|---|---|
| 企业 AI 生产级采用扩大 | 驱动因素 | 2026-now | 更多生产工作负载带来回归、治理和发布测试的经常性需求 | 生产 AI 团队中,目前有多少比例采购第三方评测,而不是自建? |
| 上游 AI 软件和模型支出爆发 | 驱动因素 | 2026-2030 | 评测预算可作为更大 AI 技术栈中一个小但扩大的比例增长 | 管理层能否证明,随着 AI 项目成熟,既有客户支出在扩张? |
| Agentic AI 和工作流自动化 | 驱动因素 | 2026-2030 | 多步骤 agents 需要工作流和因果评测,市场从聊天机器人排名向外扩 | 有多少收入来自 agent 评测,而非经典排行榜工作流? |
| 前沿实验室之间的基准竞争 | 驱动因素 | 当前 | 发布竞争提高了对可信第三方测量和叙事控制的需求 | 哪类买家主要把 LMArena 用于外部信号,而不是内部 QA? |
| 基准刷榜指控 | 约束 | 当前 | 信任流失会降低公开分数的变现能力,除非绑定受控企业工作流 | LMArena 向付费客户提供哪些反刷榜控制和审计轨迹? |
| 众包评分者偏差和代表性批评 | 约束 | 当前 | 开放 arena 投票可能无法满足需要领域扎根评测的受监管买家 | 企业评测中,有多少比例使用筛选后的专家评分者或私有数据集? |
| 云和 MLOps 平台的捆绑压力 | 约束 | 2026-2028 | 如果评测从一个类别变成一项功能,独立供应商的定价权可能下降 | 面对捆绑的 MLflow、超大规模云或可观测性工作流,LMArena 赢在哪里? |
| 公开定价和合同披露稀少 | 约束 | 当前 | 外部投资者无法独立把 TAM 转换成可预测的收入获取 | 管理层能否披露 ACV 区间、续约率,以及席位或用量扩张动态? |
驱动因素和约束只表示方向,不加权;其中若干因素可能利好类别需求,同时也压低独立供应商经济性。
[CM026, CM027, CM028, CM029, CM030, CM031]评测软件从试验到持续治理支出的五阶段漏斗。
数值经过指数化,用来展示从试验到持久复购支出的收窄过程;不是实测转化率。
[CM026, CM027, CM045]2.5 证据缺口与矛盾
主要尽调问题不是来源稀缺,而是来源不匹配。公开市场报告对品类边界意见不一;公司和媒体来源披露估值与轶事式客户名称,却没有足够合同细节来建模份额;竞争对手收入数据碎片化;本轮已验证来源也没有披露前沿实验室或企业 AI 预算中到底有多大比例花在评测软件上。因此,本章可以支持可信的区间判断,却不能给出单一精确的 SAM 或市场份额结论。投资人应保留这一矛盾,而不是强行压成点估计:相对狭义基准工具细分,LMArena 可能已经很大;相对更宽的 AI 评测平台机会,它仍然很小。要填补这个缺口,需要客户、定价和预算强度证据,而这些信息并不在已验证公开来源中。[CM003, CM024, CM038, CM039, CM040, CM041]
2.6 展示项
03竞争对手
3.1 竞争格局:直接基准平台、自动化排行榜、数据标注商,以及自建替代方案
LMArena 位于 AI 模型评测和基准测试领域,目前没有单一竞争对手能复刻它的全栈。可触达的竞争集合有四层。第一,人类偏好平台:LMArena 是众包成对模型比较的主导公开平台;截至 2026 年 1 月,没有任何免费替代品能接近其 500 万月活用户或 6000 万月度对话量。第二,自动化学术排行榜:HuggingFace Open LLM Leaderboard(由 EleutherAI lm-evaluation-harness 驱动)、Stanford HELM 和 BenchLM 聚合 MMLU、GPQA Diamond、SWE-Bench 等标准化基准分数,覆盖数百个模型,但不采用人类偏好投票。这些工具免费且开源,但衡量的是任务准确性,不是整体用户满意度。第三,企业评测供应商:Scale AI 的 GenAI Platform 为企业和政府客户提供定制评测、微调和数据标注流水线,每个项目 $93 K–$400 K+;它不运营公开排行榜。第四,现状替代方案:AI 实验室可以用 EleutherAI harness 或 OpenAI Evals 自建内部评测流水线并运行自己的测试集,完全规避第三方依赖。潜在进入者包括任何想通过基准界面区分其模型市场的大型云厂商,或由大型 AI 实验室孵化的内部中立评测机构。真正重要的竞争维度是人类偏好规模、企业收入潜力、第三方独立性和方法可信度。[CP001, CP002, CP003, CP011, CP012, CP014]
| 竞争对手 | 类别 | 规模 / 融资 | 目标细分 | 差异化 | 限制 |
|---|---|---|---|---|---|
| LMArena | 人类偏好排行榜 + 企业评测 | 融资 $250 M,估值 $1.7 B(2026 年 1 月);月用户 5 M | AI 实验室(基准营销)、企业(模型选择)、研究人员 | 最大公开人类偏好数据集;跨模态;实时 Elo 排名 | 收入来自它评测的同一批实验室;用户群偏向技术专业人士 |
| Scale AI GenAI Platform 平台 | 企业 AI 评测、数据标注、微调 | 估值 $13.8 B(2024);据报客户包括 DoD、Meta、Mayo Clinic | 寻求私有且有 SLA 支撑的评测管线的企业和政府 | 最大 RLHF 数据标注业务;定制私有评测;GPU 集群规模 | 没有公开排行榜;对特定模型并不中立;定价不透明 |
| HuggingFace Open LLM Leaderboard 排行榜 | 自动化学术基准聚合器(开源) | HuggingFace 支持(估值 $4.5 B);免费公开工具 | ML 研究人员、开源开发者、模型发布团队 | 可复现基准;聚焦开放权重;由 EleutherAI harness 驱动 | 没有人类偏好;仅覆盖技术维度;没有企业评测服务 |
| Stanford HELM | 多维学术评测框架 | Stanford CRFM(学术);免费工具;无商业产品 | 需要多轴模型评估的 AI 研究人员和政策受众 | 同时覆盖准确性、校准、鲁棒性、偏见和效率 | 周期性而非连续;没有人类偏好;没有企业收入 |
| EleutherAI lm-evaluation-harness 工具 | 开源评测框架 | 社区资助(EleutherAI 非营利);免费;GitHub stars 约 70 K+ | ML 研究人员、排行榜运营方、构建定制 evals 的企业数据团队 | 60+ 标准化学术基准;可 fork;支撑 HF 排行榜 | 没有人类偏好;没有商业服务;没有 UI / 排行榜产品 |
| OpenAI Evals | LLM 评测框架(开源 + 仪表盘) | OpenAI(内部支持;无单独融资);免费框架 | 构建定制 eval 管线的企业和 OpenAI 客户 | 直接集成 OpenAI 模型;为企业用例提供仪表盘 | 模型提供商拥有;独立性有限;聚焦 OpenAI 系列 |
| BenchLM | 自动化 LLM 排行榜聚合器 | 融资未知;免费公开工具 | 企业和开发者模型选择;2026 年基准跟踪 | 261 个模型、249 个基准;区分已验证与临时排名;包含价格 / 速度 | 没有人类偏好;平台较新,品牌认知低于 HF / Arena |
| ArtificialAnalysis | 独立 AI 模型和 API 性能分析 | 融资未知;免费公开工具 | 比较速度、吞吐、成本和智能水平的企业 API 买家 | 与提供商无关的延迟、吞吐、成本和智能基准 | 没有人类偏好;没有评测即服务;模型广度有限,弱于 HF |
LMArena 在人类偏好规模和企业服务深度上都有差异化位置;免费学术工具聚集在开放 / 研究象限;Scale AI 位于纯企业区。
分数是有证据支撑的序数判断,来自公开产品界面、论文和定价页;不是精确测量指标。
[CP024, CP025, CP011, CP012, CP014, CP015]3.2 功能与能力对比:人类偏好 vs. 自动化基准
AI 评测的核心能力分野,在于人类偏好排行榜(LMArena 及其 LMSYS 前身)和自动化基准流水线(HuggingFace Open LLM Leaderboard、Stanford HELM、EleutherAI lm-evaluation-harness、OpenAI Evals)。LMArena 的 Elo 排名系统在 2024 年 Chatbot Arena 论文中有描述,它聚合用户提交的任意任务成对投票,形成一个持续更新、反映真实使用模式而非策划测试集的信号。实践中,这意味着 LMArena 对用户在日常任务中感知到的风格、流畅度和有用性信号更敏感——包括写代码、写作、回答问题——但可复现性弱于自动化基准,因为两个用户完全可能合理地偏好相反输出。 自动化排行榜可复现且任务特定。Stanford HELM 在统一 harness 上评估准确性、校准、鲁棒性、偏见和效率等维度,更适合需要受控测量的研究者。HuggingFace Open LLM Leaderboard 主要跟踪开源和开放权重模型在标准化学术测试上的表现。EleutherAI 的 harness 支撑两者,并且可以免费 fork。BenchLM 截至 2026 年 6 月聚合了 261 个模型的 249 个基准,并单独跟踪价格和速度。ArtificialAnalysis 提供独立 API 吞吐、延迟和成本基准。Scale AI 的企业评测产品填补的是另一块空白:针对客户具体生产用例的定制、私有、SLA 支持评测,而不是公开排行榜。这种模式不与 LMArena 争夺品牌认知,但可能争夺企业预算。 LMArena 的关键差异化在于网络效应(用户越多 → 投票越多 → 排名统计稳定性越强)、跨模态覆盖(截至 2025–2026 年覆盖文本、视觉、图像生成、视频、网页开发),以及与前沿实验室的大量预发布合作。它的弱点是方法可复现性、用户群偏向技术熟练的早期采用者,以及向同一批被排名实验室销售评测服务带来的内在张力。没有竞争对手把公开排行榜界面和企业付费评测服务放在同一个品牌下,这既是 LMArena 的结构性护城河,也是它的利益冲突风险。[CP006, CP007, CP008, CP009, CP013, CP015]
| 采购标准 | LMArena | Scale AI | HuggingFace Open LLM Leaderboard 排行榜 | Stanford HELM | EleutherAI Harness | BenchLM / ArtificialAnalysis |
|---|---|---|---|---|---|---|
| 人类偏好 / Elo 排名 | 强(规模上独特) | 缺失(企业定制,不是 Elo) | 缺失 | 缺失 | 缺失 | 缺失 |
| 实时 / 持续更新 | 强(实时对战,24/7) | 未知(受 SLA 限制,未公开) | 部分(批量发布) | 周期性(非持续) | 按需(研究人员运行) | 部分(定期更新) |
| 跨模态评估(视觉、图像、视频) | 强(截至 2026 年覆盖文本、视觉、图像、视频、Web 开发) | 未知(定制;未公开) | 部分(多模态推进中) | 部分(视觉基准有限) | 部分(多模态原型) | 部分(图像理解类别) |
| 开源 / 学术基准覆盖 | 部分(Arena-Hard、MT-Bench 衍生) | 缺失(私有管线) | 强(MMLU、GPQA、ARC、SWE-Bench) | 强(多轴学术套件) | 强(60+ 个基准) | 强(249 个基准) |
| 企业评估服务(付费,SLA 支撑) | 强(AI Evaluations 产品 2025 年 9 月上线) | 强(核心产品;多年期合同) | 缺失 | 缺失 | 缺失 | 缺失 |
| 独立 / 非模型供应商所有 | 中(从被评估实验室获得收入;学术根基) | 中(数据标注客户与评估客户重叠) | 强(HuggingFace 是中立平台) | 强(Stanford 学术机构) | 强(非营利组织) | 强(独立分析机构) |
| 公开排行榜 / 透明方法论 | 强(开放方法论文;公开 Elo 分数) | 缺失(仅企业;结果私有) | 强(开放基准,可复现) | 强(多轴公开结果) | 强(开源;可复现) | 强(公开排行榜,区分已验证 / 临时排名) |
| 开发者信号(GitHub stars / 社区) | 强(FastChat 38 K+ stars;平台原生社区) | 低(企业优先品牌;开源存在感有限) | 强(HF Spaces 社区;广泛 ML 生态) | 中(学术采用;研究者引用) | 强(70 K+ GitHub stars;支撑 HF 排行榜) | 中(增长中;无开源仓库) |
能力等级(强 / 中 / 部分 / 缺失 / 未知)是基于截至 2026 年 6 月公开可访问的产品界面、论文和基准文档作出的序位判断;标为未知的单元格对应私有企业产品,能力可能存在但尚未被公开验证。
[CP006, CP007, CP008, CP013, CP015, CP017]| 平台 | 定价模式 | 标价 / 合同区间 | 包含能力 | 公开定价可得性 | 对买方的含义 |
|---|---|---|---|---|---|
| LMArena AI Evaluations | 企业合同(按用量消耗) | 未公开披露;约 100 家企业客户贡献 $30 M 年化消耗规模,隐含平均约 $300 K/客户(估算,非公司披露) | 定制评估面板、社区反馈数据、SLA 交付、分析 | 无(联系销售) | 企业可为优先评估付费,但标价不透明 |
| Scale AI GenAI Platform 平台 | 企业定制合同 | 分析师报告显示每次项目 $93 K–$400 K+ | 数据标注、模型微调、Agent 部署、评估管线 | 无(联系销售) | 价格更高,但结果私有且保密;不依赖公开排行榜 |
| HuggingFace Open LLM Leaderboard 排行榜 | 免费(HuggingFace 补贴) | $0 | 开放基准提交、公开分数、可复现测试集 | 完全公开 | 无成本,但结果公开;无法针对具体企业用例定制 |
| Stanford HELM | 免费(Stanford CRFM 补贴) | $0 | 多轴自动评估、公开结果 | 完全公开 | 学术严谨、多轴覆盖,但没有人类偏好或商业 SLA |
| EleutherAI lm-evaluation-harness 工具 | 免费(开源) | $0(基础设施成本自托管) | 60+ 个基准任务;自托管或云端运行 | 完全公开 | 控制力最高,但需要工程团队运营;无 UI 或排行榜 |
| OpenAI Evals | 免费框架;使用 GPT-4o 评判按 API 付费 | $0 框架费;评判模型运行需 API 成本(约 $5–30/M 输出 tokens) | 定制 eval 模板、仪表盘、社区 eval 注册表 | 完全公开(框架);API 定价公开 | 深度集成 OpenAI,但由模型供应商所有;用于非 OpenAI 模型对比时受限 |
| BenchLM / ArtificialAnalysis | 免费(广告支持的分析网站) | $0 | LLM 排行榜、基准聚合、速度 / 成本 / 智能对比 | 完全公开 | 适合模型选择研究;无企业服务或定制评估 |
LMArena 的单客户平均值($300 K)来自用 $30 M 年化消耗规模除以约 100 家客户;这不是公司披露口径。Scale AI 定价来自第三方分析师聚合,不是公开价目表。
[CP036, CP037, CP038]LMArena 独家领先于人类偏好规模和跨模态实时覆盖;Scale AI 领先于私有企业深度;免费工具领先于可复现性和开源信号。
序数能力等级(强 / 中 / 部分 / 缺失 / 未知)是基于公开审查的产品界面和学术论文做出的有证据判断;标为缺失的单元格表示公开审查的产品界面确认没有该能力。
[CP006, CP007, CP017, CP018, CP019, CP023]3.3 护城河持久性、切换成本与负面竞争证据
LMArena 的竞争持久性建立在三个因素上:同类最大规模的众包偏好数据集;由学术根基强化的社区信任(UC Berkeley LMSYS Org、FastChat 开源基础设施);以及深度引用网络效应——前沿实验室在营销材料和投资人沟通中引用 Arena 排名,从而抬高退出成本。AI 实验室的切换成本真实存在:如果 Arena 不再是市场信号,它们既有的基准投入就会失去营销价值。这在 LMArena 与客户实验室之间形成互相依赖,带来收入稳定性,也带来结构性独立风险。 关于护城河的负面证据可信且重要。Singh 等人在 2025 年发表的 Leaderboard Illusion 论文(Cohere、Stanford、MIT、Ai2)量化了数据访问不对称:Google 和 OpenAI 估计分别获得 Arena 总数据的 19.2% 和 20.4%,而 83 个开放权重模型合计只获得 29.7%。Meta 在 Llama 4 发布前测试了 27 个私有模型变体,只选择最高分版本;TechCrunch 证实,普通 Maverick 重新提交后排名第 32。LMArena 否认偏差定性,但承诺修改算法,说明批评具有运营层面的分量。第二个负面动态是 Goodhart 定律:当 Arena 排名开始驱动采购决策,实验室会针对 Arena 特定模式优化,信号质量随之下降,也提高了资源充足的竞争者声称方法更优的概率。 支撑 Arena 的 Elo 算法开源,任何资金充足的团队都能复制。难以复制的是社区。Scale AI 的企业触达更广,数据标注能力更深,但尚未建立公开偏好排行榜。真正的替代威胁来自内部自建:如果某个 AI 实验室认为 Arena 的利益冲突问题无解,它可以资助一个竞争性的中立机构。免费学术工具(HELM、HuggingFace)已经在研究用例中承担这类可信度功能,也限制了 LMArena 能为方法本身收取多少溢价。[CP028, CP029, CP030, CP034, CP035, CP036]
| 护城河主张 | 威胁 | 严重性 | 缓解措施 / 尽调问题 |
|---|---|---|---|
| 人类偏好数据集是该领域最大且引用最多的数据集(6 M+ 票、60 M 对话) | 资金充足的实验室或联盟可能在 18–24 个月内用等效用户激励搭出竞争性人类偏好平台 | 高 | 确认 LMArena 的数据集是专有资产还是社区所有;评估与实验室的数据共享协议 |
| 网络效应:顶级实验室在 Arena 发布模型,因为社区有这个需求 | 认为机制被操纵或商业上吃亏的实验室可能退出,并资助一个竞争性的中立机构 | 中 | 跟踪是否有主要实验室公开减少 Arena 参与,或发起竞争项目 |
| 学术可信度(源自 UC Berkeley/LMSYS;9 篇已发表论文) | The Leaderboard Illusion 论文(Singh et al., 2025)削弱可信度;后续研究可能加速声誉受损 | 高 | 审查 LMArena 对选择性披露批评的方法论回应;确认算法改革是否落地 |
| 企业 AI Evaluations 收入为实验室客户制造付费切换成本 | 收入来自被评估的同一批实验室,会形成结构性利益冲突,削弱独立性护城河 | 高 | 要求管理层提供结构性独立政策(防火墙、编辑独立性);审查头部客户是否能影响排名 |
| 开源 FastChat 基础设施和开放数据发布维持开发者好感 | 免费自动化基准工具(EleutherAI、HF、HELM)为学术用户提供可信度功能,压住 LMArena 的学术护城河上限 | 低 | 评估企业收入增长后,LMArena 是否还能维持学术合作(论文、引用) |
| 平台已成事实标准,用于模型发布基准测试 | Goodhart 定律:Arena 一旦成为目标,实验室会过度优化 Arena 特定模式,拉低信号质量,也给方法论攻击打开口子 | 中 | 监测未来模型发布是否在新闻稿中特别引用 Arena 分数,并跟踪任何公开的 Arena 特定调优证据 |
严重性(高 / 中 / 低)反映截至 2026 年 6 月,基于公开证据对威胁重要性的定性判断;不是数值风险分。
[CP028, CP029, CP030, CP033, CP034, CP035]LMArena 的护城河锚定在数据和社区规模;主要脆弱点是方法完整性和商业利益冲突。
[CP001, CP002, CP004, CP023, CP029, CP033]3.4 展示项
04财务
4.1 收入模式、定价与商业化进展
LMArena 的收入模式是从免费到企业的漏斗。免费公开排行榜每月服务 150 个国家的 500 万用户,生成 6000 万次模型比较对话,承担社区和信任建设层功能。企业、模型实验室和 AI 开发者为 LMArena 的商业 AI Evaluations 产品付费(2025 年 9 月推出),该产品提供基于真实用户反馈的定制评测小组、代表性数据样本,以及按 SLA 承诺的交付周期。LMArena 未发布公开价格表;所有企业合同都直接协商。三类已确认企业客户包括 AI 实验室(2026 年 1 月 Series A 新闻稿明确提及 OpenAI、Google、xAI)、软件企业,以及受监管专业垂直领域(法律、医疗、科学研究)。公司年化消费运行率在 2025 年 12 月超过 $30M——距离产品上线不到四个月。 收入确认是重要 caveat:LMArena 和 TechCrunch 都把这个数字描述为「consumption run rate」,而不是已实现年度数字。TechCrunch 指出,公司用消费率来描述其「年度经常性收入(ARR)」,反映的是按使用量计费,而非预付合同式 SaaS ARR。按 GetLatka 汇总的约 100 个企业客户计算,隐含平均合同额约为每客户 $300,000——这是用运行率除以客户数得出的估计,并非公司披露数字。实际 ACV 分布、合同期限和续约率均未公开。[CI001, CI002, CI003, CI004, CI005, CI006]
| 收入流 | 机制 | 单位 | 当前值 / 状态 | 质量 | 尽调问题 |
|---|---|---|---|---|---|
| AI Evaluations 企业服务 | 用社区人类反馈做付费定制评估面板,并按 SLA 交付 | 企业合同(按消耗计费) | 截至 2025 年 12 月年化消耗规模为 $30 M(公司披露);约 100 家客户(分析师估算) | 中(消耗规模 ≠ 合同 ARR;无流失、NRR 或留存数据) | 获取合同 ACV 分布、续约率、NRR,以及计费是限时还是按用量 |
| 数据与分析授权(潜在) | 向企业和研究人员销售聚合偏好数据或模型性能洞察 | 按数据集或订阅(未确认) | 尚未确认是独立收入线;公司已发布免费数据集 | 低(该收入流无已确认收入) | 确认企业评估产品之外是否存在任何商业数据授权协议 |
| 合作评估费(潜在) | AI 实验室为预发布模型评估名额的提前 / 优先访问付费 | 按次评估,或纳入企业合同(结构未确认) | 未单独披露;可能打包进 AI Evaluations 合同 | 低(无单独公开披露;无法从企业线拆分) | 确认预发布评估名额是否单独定价或打包;厘清收入确认方式 |
AI Evaluations 的当前值是截至 2025 年 12 月的「年化消耗规模」;这不是已实现的 12 个月收入。客户数(约 100)来自 GetLatka 聚合,不是公司披露。
[CI001, CI002, CI003, CI004, CI007, CI008]| 产品 / 层级 | 定价模式 | 标价 / 合同区间 | 包含能力 | 公开价格可得性 | 来源 |
|---|---|---|---|---|---|
| 免费公开排行榜 | 免费(社区补贴) | $0 | 模型两两对战、Elo 排名、开放数据发布、多模态覆盖 | 完全公开 | arena.ai(官方) |
| AI Evaluations — 企业 | 按消耗计费合同(联系销售) | 未公开披露;隐含平均约 $300 K(分析师估算,未经公司确认) | 定制评估面板、社区反馈数据、SLA 交付、分析仪表盘 | 无(联系 evaluations@lmarena.ai) | arena.ai/blog/ai-evaluations/(官方) |
| 研究开放数据 | 免费 | $0 | 1.5 M+ 条社区提示词、145 K+ 个对战数据点(开源) | 完全公开(HuggingFace) | arena.ai/blog/two-year-celebration/(官方) |
| 预发布模型评估(合作实验室) | 打包或协商费用(不清楚) | 未公开披露 | 优先评估名额、预发布分数可见性 | None | 从 TC 和 PRNewswire 报道推断 |
| 学术 / 开源层 | 按公司承诺灵活定价 | 未披露(较企业费率有折扣) | 访问评估社区和基础分析 | 无(公司承诺支持非营利组织) | arena.ai/blog/ai-evaluations/(官方) |
免费社区活动驱动偏好数据价值,再通过 AI Evaluations 产品转化为企业评测收入。
收入、利润率和再投入数字均为近似值;毛利率未获确认。运行率是年化消耗率(公司披露),不是合同 ARR。
[CI001, CI002, CI007, CI008, CI030]4.2 成本结构、单位经济与资本充足性
从公开证据推断,LMArena 的成本结构有三大类。第一,算力和基础设施:每月跨前沿模型服务 6000 万次对话,需要大量云算力。公开排行榜使用合作实验室提供的模型 API(可能根据合作协议以折扣或免费方式提供),但企业评测产品在规模化后很可能产生真实推理成本。第二,工程和研究人员:截至 2026 年 1 月 Series A 公告,公司有 41 名员工。按 Silicon Valley 技术初创公司的典型全包薪酬,这意味着年人员成本约 $15–25M。第三,社区管理、数据标注质量保障和 go-to-market 成本,这些无法从公开信息中测算。毛利率未披露;如果是边际评测成本较低的软件式模型,毛利率可能达到 60–80%,但每次对话算力成本高也可能把毛利率压到 40–60%。没有坚实依据的第三方估计。 近期资本充足性看起来很强。公司已在 2025 年 5 月种子轮($100M,估值 $600M)和 2026 年 1 月 Series A($150M,估值 $1.7B)累计融资 $250M。团队精简至 41 人,首个商业收入到 2025 年 9 月才建立,因此截至 2026 年初,年现金消耗大概率远低于 $30M。粗略估算显示,不再融资也至少有 18–36 个月 runway,但公司未披露官方烧钱率或现金余额。LMArena 表示将把 Series A 资金用于扩充技术团队、强化研究能力并搭建新平台功能——这些都会提高未来 burn。Recall Capital-LMArena feeder fund(SEC Form D,2026 年 2 月提交)确认了资本形成活动,但未披露 LMArena 自身资产负债表。未见公开报道显示公司有债务或项目融资义务。[CI009, CI010, CI011, CI012, CI013, CI014]
| 指标 | 值 / null | 置信度 | 为什么重要 | 尽调问题 |
|---|---|---|---|---|
| 平均合同价值(ACV) | ~$300 K(估算:$30 M 年化消耗规模 ÷ 约 100 家客户) | 低(推导估算;两个输入都未经公司确认) | 决定收入扩展性和销售效率假设 | 确认实际 ACV 分布、企业合同条款,以及客户数是否准确 |
| 毛利率 | n/a(未披露) | 承销关键;软件型模式应超过 60%;重计算可能压到 40-60% | 要求管理账提供毛利率;确认合作伙伴 API 成本是否由对方补贴 | |
| 净收入留存(NRR) | n/a(未披露) | 对按消耗计费的 SaaS,NRR 能区分扩张账户和一次性评估 | 要求队列级 NRR;厘清 $30 M 年化消耗规模来自扩张账户还是稳定账户 | |
| 获客成本(CAC) | n/a(未披露) | 反映单位经济可行性和 GTM 效率 | 要求按细分市场拆分 CAC(AI 实验室 vs. 企业);确认免费排行榜是否是主要获客渠道 | |
| 销售周期长度 | n/a(未披露) | 大型 AI 实验室的企业评估合同可能涉及数月采购 | 要求从首次接触到签署合同的平均时间;若实验室预算周期影响时点,需要标记 | |
| 回本周期 | n/a(未披露) | 取决于 ACV 和 CAC;两者都未确认,无法估算 | ACV 和 CAC 确认后再推导;若回本超过 18 个月,需要标记 | |
| 月度烧钱 | null(按约 41 名员工、每人 $300 K 全包平均成本 + 基础设施估算,为 $1.5–2.5 M/月) | 低(粗略估算;非公司披露) | 决定最低现金需求和融资触发点 | 要求实际月度 P&L;确认基础设施成本和任何一次性费用 |
| 现金续航(自 2026 年 1 月起) | null(按估算烧钱对比 $250 M 融资,估计 24–36 个月) | 低(取决于未经确认的烧钱) | 判断下一轮融资是近期必需还是机会型 | 获取 Series A 交割时现金余额和当前月度现金流出 |
所有 null 值都表示未公开披露的指标;估算值(ACV、烧钱、现金续航)只是用于尽调框架的粗略推导。低置信度数值在承销前必须确认。
[CI004, CI005, CI014, CI015, CI017, CI018]| 项目 | 值 | 来源 / 置信度 | 含义 |
|---|---|---|---|
| 种子轮(2025 年 5 月) | $100 M,投后估值 $600 M | 高(TechCrunch、PRNewswire 确认) | 确立商业化现金续航;由 a16z 和 UC Investments 领投 |
| Series A(2026 年 1 月) | $150 M,投后估值 $1.7 B | 高(PRNewswire 官方新闻稿,TechCrunch 确认) | 近期主要资本基础;用途:团队扩张、研究、平台功能 |
| 累计融资 | $250 M | 高(两轮已确认融资之和;未公开披露过过桥轮或可转债) | 对 41 人精简团队而言资本充足;当前规模下现金续航可能为 24–36+ 个月 |
| 月度烧钱估算 | $1.5–2.5 M/月(粗略估算) | 低(未披露;基于人数和基础设施假设推导) | 若月度烧钱为 $2 M,$250 M 意味着不含收入也有 10+ 年现金续航——偏保守;实际烧钱可能更高 |
| 手头现金(2026 年 1 月) | 未披露 | n/a | 需要管理层确认;关系到人员扩张和产品建设计划 |
| 债务 / 项目融资义务 | 公开未见报告 | 中(未发现公开备案或公告) | 假设资本结构干净;尽调时验证 |
| 下一轮触发点 | 未披露;未见关于下一轮融资时间表的公开表述 | n/a | 按当前烧钱和收入轨迹,Series B 可能在 2026 年 1 月后 18–30 个月 |
| Series A 资金计划用途 | 扩大技术团队、强化研究能力、开发新功能(公司披露) | 中(公司披露;无预算拆分) | 人员增加会推高烧钱;需要扩张后的更新预测 |
月度烧钱是粗略估算;实际烧钱取决于人员增长节奏、基础设施扩容和数据质量投入。Series B 时间表估算仅作示意,不是公司表述。
[CI009, CI010, CI011, CI012, CI013, CI018]企业合同先流经面板组建和推理成本池,再生成毛利;所有利润率输入仍为私有信息。
ACV($300 K)为估算。毛利率未知;区间反映类软件上限与重计算下限的差异。所有单位经济都需要管理层披露来确认。
[CI005, CI015, CI016, CI017, CI028]在估计年度 burn 为 $20–45 M 的情况下,已融资 $250 M 意味着近期资本充足度强,但成本假设尚未确认。
所有成本项均由员工数和基础设施基准推导;$225 M 净现金估算仅作说明,不是公司披露数字。收入是运行率(年化消耗),不是 2025 年已实现收入。
[CI009, CI010, CI011, CI014, CI018, CI019]4.3 财务风险、利益冲突与尽调障碍
最重要的财务风险是结构性的:LMArena 从同一批 AI 实验室——OpenAI、Google、xAI——获得收入,同时又在公开排行榜上对它们排名。CTOL Digital 在 2026 年 1 月融资后明确记录了这个悖论:$1.7B 估值「相当于该运行率的 57 倍,计入的不只是增长,还有一个假设:这种内在冲突可以被无限期管理」。如果企业客户开始相信 LMArena 排名受到商业关系影响(正如 Leaderboard Illusion 论文对数据访问不对称的指控),评测服务的定价权会下降,免费排行榜的信任优势——核心营销资产——也会同步坍塌。这种双重暴露很少见:多数 SaaS 公司面对的是客户流失风险;LMArena 面对的是可信度流失风险,同一事件既会丢掉客户,也会伤害带来新客户的免费产品。 第二个风险是收入集中度:约 100 个企业客户,且运行率由三大 AI 实验室主导,OpenAI、Google 或 xAI 中任何一个单一客户流失,都可能实质性削减收入。第三个风险是「consumption rate」表述:如果按使用量计费,任何单月收入都可能不可预测,$30M 年化数字可能只是用峰值月份外推,而非已签约的远期义务。第四,按运行率 57 倍估值意味着资本市场同时计入了高收入增长(收入必须超过 $100M 才能支撑 17 倍这一常见后期基准倍数)和持久毛利率——两者都无法从公开数据验证。企业采用放缓或方法信任受损,都可能触发显著估值重置,进而让后续融资或退出选择变复杂。[CI021, CI022, CI023, CI024, CI025, CI032]
| 缺失的私有指标 | 对分析的影响 | 精确尽调路径 |
|---|---|---|
| 2025 年实际收入(非年化口径) | 无法验证 $30 M 年化消耗规模会转化为 $8–10 M 的 2025 年实际收入(产品上线 4 个月),还是因爬坡形态而更高 / 更低 | 要求提供从产品上线(2025 年 9 月)到 2025 年 12 月的月度实际收入 |
| NRR / 账户扩张数据 | 没有 NRR,就无法区分增长型账户(利好)和一次性试点评估(不利于耐久性) | 分别获取 AI 实验室客户和企业客户的队列级 NRR;确认是否已有客户流失 |
| 毛利率(P&L 层面) | 毛利率决定 $30 M 年化消耗规模指向的是可扩展的高毛利业务,还是杠杆有限的服务交付模式 | 要求管理账提供毛利行;确认合作伙伴 API 成本和基础设施成本如何处理 |
| ACV 和合同条款 | 无法验证隐含 $300 K ACV,也无法确认定价是限时、按用量还是按事件 | 获取一份脱敏样本合同,以及各客户层级的 ACV 分布 |
| 客户集中度(top-3 收入占比) | 如果三个实验室客户(OpenAI、Google、xAI)贡献 50%+ 收入,任何单一客户流失都是重大事件 | 要求按客户层级拆分收入;标记任何单一客户收入占比 >20% 的情况 |
| 截至 2026 年 6 月的人数和烧钱 | 公司在 Series A 时(2026 年 1 月)有 41 名员工;若扩张计划已启动,烧钱会明显上升 | 要求提供最近一个季度的当前人数、月度薪酬和基础设施成本运行速率 |
每个缺口都是承销 $1.7 B 估值所隐含收入质量和财务可持续性主张时需要补齐的关键变量。
[CI003, CI004, CI017, CI019, CI023, CI025]关键财务变量的不确定区间很宽;57x 估值倍数已确认,但验证该倍数所需的收入和成本输入大多仍是私有信息。
除估值倍数外,所有区间均由员工数、典型硅谷薪酬、基础设施基准和披露的运行率数字推导而来。估值倍数中点(57x)根据已确认的 $1.7 B 估值和 $30 M 运行率计算。区间反映输入不确定性,也反映运行率与已实现收入之间的差别。
[CI002, CI011, CI014, CI017, CI018, CI032]4.4 展示项
05产品与技术
5.1 产品组合与 Arena 模态
LMArena 的核心面向用户产品是 arena.ai 上的 Arena 平台,它针对给定提示词展示两个匿名 AI 模型之间的对战,并记录社区偏好投票。截至 2026 年 6 月,平台支持十个不同评测 Arena:Text(通用聊天)、Code(智能体式网页开发和编码)、Search(检索增强生成)、Agent(自主多步任务完成)、Vision(多模态理解)、Text-to-Image(图像生成)、Text-to-Video(视频生成)、Image Edit(图像编辑模型)、Document(长文档分析),以及一个独立 Direct Chat 模式,用户可在没有成对对战结构的情况下与所选模型交互。Agent Arena 是最新、技术复杂度最高的 Arena,于 2026 年 6 月 4 日推出,它根据每周 2M+ 工具调用推导出的因果处理效应为编排模型排名。Code Arena 由早期 WebDev Arena(2024 年 12 月推出)重建而来,是一个实时、隔离的编码环境,模型可通过结构化工具调用自主创建、修改和执行文件,输出持久化在 Cloudflare R2,并通过 CodeMirror 6 展示。自 2026 年 5 月起,LMArena 还加入「Battles in Direct」,把 10% 的 Direct Chat 会话路由到匿名成对对战,以增加每日投票量。商业产品 AI Evaluations(2025 年 9 月推出)是一项卖给 AI 实验室和企业的合同评测服务,提供带代表性反馈样本和 SLA 支持交付周期的深度评测。所有公开模型评测都可在公开排行榜上免费获得;商业产品则增加私有、保密、受 SLA 管理的评测运行。 [CE001, CE002, CE008, CE012, CE029, CE032]
| Arena / 模块 | 用户 / 买方 | 状态 / 成熟度 | 评估方法 | 关键差异化 | 尽调缺口 |
|---|---|---|---|---|---|
| 文本 Arena(Chatbot Arena) | 普通用户、研究人员、模型提供商 | 已上线 / 成熟(自 2023 年起) | Bradley-Terry 成对投票 | 2.5 亿+ 真实对话;控制风格影响后的排名 | BT 估计尚无独立可靠性审计 |
| Code Arena | 开发者、企业、模型提供商 | 已上线 / 增长中(2026 年重建) | 生成式 Web 应用的成对投票;智能体工具调用环境 | 持久会话、CodeMirror 6 实时预览、Cloudflare R2 快照 | 功能正确性与偏好分歧尚未量化 |
| Agent Arena | 高阶用户、企业、模型提供商 | 已上线 / 早期(2026 年 6 月发布) | 因果追踪:多信号 RCT 框架 | 方法论新颖;每周 200 万+ 工具调用;首个因果智能体排行榜 | 信号选择和因果模型假设尚未经过同行评审 |
| Search Arena | 研究人员、模型提供商、企业 | 已上线 / 增长中 | Bradley-Terry;引用式随机化 | ICLR 2026 论文;2.4 万+ 场对战;接地质量分析 | 仅支持 3 家提供商;地理定位功能有限 |
| WebDev Arena | 开发者、模型提供商 | 已上线 / 稳定(自 2024 年 12 月起) | 基于 Web 应用偏好投票的 Bradley-Terry | 8 万+ 票;主题建模分析;CSS / JS / HTML 生成 | Code Arena 的前身;方法论刷新仍待完成 |
| Vision Arena | 多模态用户、模型提供商 | 已上线 / 稳定 | 图像理解任务上的 Bradley-Terry 成对投票 | 覆盖主流多模态模型 | 覆盖面相对文本 Arena 有限;未纳入图像保真度指标 |
| Text-to-Image Arena | 创意用户、模型提供商 | 已上线 / 增长中 | 生成图像的成对偏好投票 | 覆盖 Flux、Midjourney、DALL-E 等 | 审美偏好与提示词遵循度没有拆开 |
| Text-to-Video Arena | 创意用户、模型提供商 | 已上线 / 早期 | 生成视频片段的成对偏好投票 | 早期阶段;2026 年新增 wan2.7-t2v、gemini-omni-flash | 投票量有限;时间一致性未单独评分 |
| Document Arena | 企业用户、模型提供商 | 已上线 / 增长中 | 文档任务上的 Bradley-Terry 成对投票 | 长上下文评估;2026 年加入多个排行榜 | 尚未发布具体文档类型拆分 |
| AI Evaluations(商业) | AI 实验室、企业 | 2025 年 9 月起 GA | 带 SLA 交付的私有评估;基于社区反馈 | 2025 年 12 月 ARR run-rate 为 $30M;OpenAI、Google、xAI 被列为客户 | SLA 条款、安全披露、DPA 未公开 |
| Direct Chat | 所有用户 | 已上线 / 稳定 | 不做评估;提供免费模型访问 | 无付费墙访问前沿模型 | 不产生收入;推理成本中心 |
| AI Evaluations(API / pipeline 管线) | 企业、开发者 | 路线图 / 计划中 | 程序化提交评估 | 可释放大规模自助评估能力 | 尚未公布时间表;上线前风险 |
状态和投票数来自截至 2026 年 6 月的 LMArena 官方博客、排行榜更新日志和新闻稿。尽调缺口为研究员推断。
[CE001, CE002, CE008, CE012, CE029, CE032]LMArena 平台的分层视图,从数据收集,到排名引擎,再到评测产品。
层级边界根据博客文章和 GitHub 仓库推断;内部服务拆分未公开披露。
[CE001, CE002, CE003, CE009, CE013]5.2 技术架构与评测方法
所有 Arena 排行榜都使用 Bradley-Terry(BT)成对比较模型,根据胜负结果推断每个模型的潜在能力系数。LMArena 的 Arena-Rank Python 包以 Apache 2.0 开源,并发布在 GitHub 和 PyPI 上,实现了 BT 拟合、闭式置信区间计算,相比历史 FastChat 实现提速 30 倍。该包也支持重新加权来修正非均匀采样,意味着对战较少的模型不会被惩罚。风格控制扩展把额外协变量(token 长度、markdown 标题数、markdown 加粗数、markdown 列表数)叠加入 BT 回归,用于把内容实质与风格式格式化效应分开;系数估计显示长度是主导风格因素。Agent Arena 不用成对投票排名,而是使用因果追踪:Arena 把每次组件选择视为多干预随机对照试验中的一个 treatment,聚合多个行为信号(确认成功、表扬 / 投诉、可引导性、bash 错误恢复、工具幻觉),为每个模型形成单一净改进估计。Arena-Hard 流水线(BenchBuilder,arXiv:2406.11939)通过七标准难度标注器,从实时 Arena 数据集中抽取 hard prompts,自动构建基准;Arena-Hard-Auto v0.1 以约 $20 成本达到与人类偏好排名 98.6% 的一致性。Search Arena 方法以数据集和论文形式发表,并被 ICLR 2026 接收(arXiv:2506.05334),将 BT 模型扩展到来自 Perplexity、Gemini 和 OpenAI 的 11+ 搜索增强 LLM 系统。在分享任何对话数据之前,LMArena 使用 GCP Sensitive Data Protection API 移除个人身份信息。自 2025 年 7 月起,方法更新会公开记录在 Leaderboard Changelog。 [CE003, CE009, CE013, CE014, CE015, CE016]
| 层 / 组件 | 作用 | 依赖 / 实现 | 风险 |
|---|---|---|---|
| Arena-Rank(排名引擎) | 为所有 Arena 计算 BT 系数、置信区间和风格控制后的分数 | 开源 Python 包(Apache 2.0);GitHub lmarena/arena-rank;PyPI arena-rank | 方法论开放 = 可复现,但也可被竞争对手复刻 |
| Bradley-Terry 成对模型 | 核心统计模型;从胜负对战中推断潜在能力 | 自研 Python 实现;比 FastChat 基线快 30 倍;闭式 CI | 模型假设(IID 对战、无时间漂移)在重度采样操纵下可能失效 |
| 风格控制扩展 | 在 BT 回归中把内容实力与格式效果拆开 | 加性协变量回归;长度、markdown 标题、加粗、列表数量 | 仅有四个风格协变量;更丰富的语义风格尚未捕捉 |
| 因果追踪(Agent Arena) | 通过多干预 RCT 对智能体组件排名;衡量净处理效应 | 自研统计框架,结合 5 个行为信号 | 方法论新颖;因果识别假设尚未由第三方验证 |
| BenchBuilder / Arena-Hard 管线 | 用 LLM 标注器从实时 Arena 数据自动生成高难基准子集 | arXiv:2406.11939;基于 GPT-4 的难度评分器;7 项难度标准 | 依赖 GPT-4 做标注;标注器偏差会渗入基准质量 |
| 数据清洗(GCP Sensitive Data Protection API) | 公开数据发布前移除 PII | Google Cloud Platform API;任何共享数据集发布前调用 | 依赖单一云厂商;未发布独立 PII 审计 |
| Web 平台(arena.ai) | 承载对战、投票 UI、排行榜、Direct Chat、Agent Mode、Code Arena | JavaScript 前端;服务器架构未披露 | 社区信任的单点故障;正常运行时间 SLA 未公开 |
| Cloudflare R2 + CodeMirror 6(Code Arena 基础设施) | 存储持久代码会话;渲染实时 Web 应用预览 | Cloudflare R2 存快照;CodeMirror 6 做代码展示;流式前端 | 数据驻留和保留政策未披露 |
| FastChat(遗留) | 历史上的服务和训练平台;原 Chatbot Arena 后端 | GitHub lm-sys/FastChat;截至 2025 年主要处于维护模式 | 已被主动降级优先级;仍在 FastChat 上的外部用户存在迁移风险 |
| p2l 模型(HuggingFace) | prompt-to-leaderboard 偏好模型;用于内部评估研究 | HuggingFace 上的 lmarena-ai 组织;参数规模 0.1B–7B | 生产用途未公开记录;目的和部署方式不清楚 |
架构根据官方博客、GitHub 仓库、PyPI 包和学术论文推断。服务器基础设施细节未公开披露。
[CE003, CE009, CE013, CE014, CE015, CE016]从用户提交提示词,到投票记录,再到排行榜更新的端到端流程。
流程根据官方博客文章和政策页面重建;内部系统架构未披露。
[CE002, CE016, CE020, CE021]LMArena 产品在平台、数据、监管和合作伙伴上的关键依赖。
依赖关系根据公开文档推断;相对关键性为研究者判断。
[CE013, CE014, CE019, CE020, CE032]5.3 部署、数据基础设施与开发者生态
Arena-Rank 排名引擎作为可用 pip 安装的 Python 包发布(pip install arena-rank),归属于 lmarena GitHub 组织,该组织也托管 list 仓库。FastChat 仓库(lm-sys/FastChat,现在主要处于维护模式)最初支撑 Chatbot Arena,并继续作为 Vicuna 模型权重的参考。lmarena-ai HuggingFace 组织托管内部用于偏好建模的 p2l(prompt-to-leaderboard)模型变体,并发布公开对战数据集(例如 arena-human-preference-140k)以支持外部研究。截至 2026 年 6 月,changelog 记录了各排行榜每周新增数个模型的持续节奏。Code Arena 的持久会话建立在 Cloudflare R2 存储之上,用 CodeMirror 6 展示源码视图和生成网页应用的实时渲染。Agent Mode 会话平均约 16.5 次结构化工具调用,约 75.6% 的会话至少使用一个工具;最高使用会话在单周内运行很长调用链,覆盖编码、文件创建和网页合成任务。在一个被测的 7 天窗口中,Agent Mode 通过成功的 write_file 调用写入 4030 万行代码,约每个编码会话 1000 行。Battles in Direct 自 2026 年 5 月集成后,把 10% 的直接聊天会话转化为成对对战,并对 BT 拟合应用位置偏差和同组织指示器修正。公司未公开记录面向第三方消费 Arena 分数或原始投票数据的 API,但会定期在 HuggingFace 发布开放数据集。 [CE013, CE014, CE015, CE019, CE030, CE031]
| 用户任务 | 传统流程 | LMArena 方案 | 可衡量收益 | 已知限制 |
|---|---|---|---|---|
| 选择生产模型前比较 LLM 质量 | 跑内部 A/B 测试;依赖静态基准(MMLU、HumanEval) | Text / Code Arena 成对对战投票;大规模观察真实用户偏好 | 经过统计校准的排名,含 95% CI;覆盖风格控制和高难提示子集 | 排名反映普通用户偏好,不一定贴合企业特定用例 |
| 为真实工作负载评测智能体编码模型 | 静态代码正确性基准;合成测试套件 | Code Arena:隔离智能体环境;实时生成 Web 应用;持久会话 | 捕捉智能体式、多轮行为;运行结果可分享 | 正确性与审美偏好未拆开;编程语言多样性有限 |
| 评估面向检索任务的搜索增强 LLM | SimpleQA 静态基准;人工标注 | Search Arena:众包投票评判多轮搜索-RAG 输出 | 2.4 万+ 成对交互;引用数量分析;ICLR 2026 验证 | 当前仅 3 家提供商;学术 / 领域覆盖有限 |
| 衡量自治智能体在业务任务中的表现 | 人工评分会话;特定工具基准 | Agent Arena:跨每周 200 万+ 工具调用做因果追踪;5 信号排行榜 | 基于 RCT 的因果处理效应;逐信号拆解可解释 | 方法论尚未经过同行评审;信号集合会继续演进 |
| 为上线前模型获得带 SLA 的保密评估 | 专有红队测试或供应商专属评估团队 | AI Evaluations:私有评估,含 SLA、代表性反馈和数据样本 | 扎根真实世界人类偏好;交付时间线可认证 | 定价、SLA 条款和安全细节未公开披露 |
工作流综合自 LMArena 博客、学术论文和 TechCrunch 报道。收益主张除标注为研究员推断外,均为公司表述。
[CE002, CE009, CE012, CE029]截至 2026 年 6 月,各评估 Arena 在四个能力维度上的成熟度。
成熟度评级是研究者基于截至 2026 年 6 月可得公开证据作出的判断。投票量分层:高 = 250M+;中 = 24k-80k;低 = <24k。
[CE002, CE008, CE009, CE035, CE036]5.4 差异化、IP 与研究优势
LMArena 的核心差异化,是把大规模真实世界、野生环境数据与方法透明度结合起来。不同于静态学术基准(MMLU、HumanEval)或自动化 LLM-as-judge 流水线,Arena 捕捉跨语言、主题和技能水平的真实多轮用户交互。已发表研究验证了这种方法:最初的 Chatbot Arena 论文(arXiv:2403.04132)显示其与专家评分者一致性超过 80%;MT-Bench(NeurIPS 2023)证明 LLM-as-a-judge 可以达到人类级别的评分者间一致性。Arena-Hard 基准(arXiv:2406.11939)以一小部分成本实现了比 MT-Bench 高 3 倍的模型区分度。Search Arena ICLR 2026 论文(arXiv:2506.05334)把方法扩展到检索增强系统,并展示了偏好与引用数量、被引用来源类型之间的相关性。LMSYS(前身非营利组织)发布了支撑当前公司差异化的基础评测方法论文;这些论文在 AI 行业被广泛引用。HuggingFace 组织发布偏好数据集(140k+ 标注对战对),让外部研究者能够独立审计排名。公司 9 篇已发表论文和 15+ 篇博客文章覆盖评测方法,构成一组被认可的 IP。不过,方法资产被公开分享,意味着竞争对手可以复制算法;公司留下来的持久护城河,是对战规模带来的数据网络优势和社区信任。 [CE003, CE016, CE035, CE036, CE037, CE039]
| 日期 / 时期 | 功能 / 里程碑 | 状态 | 含义 | 来源 |
|---|---|---|---|---|
| May 2023 | Chatbot Arena(Text Arena)公开发布 | 已完成 | 确立众包评估作为可行方法论 | 来源:lmsys.org/blog/2023-05-03-arena/ |
| Jun 2023 | MT-Bench 多轮问题集和 LLM-as-a-judge 论文(NeurIPS 2023) | 已完成 | 成对评估方法论获得学术验证 | arXiv:2306.05685 |
| Mar 2024 | Chatbot Arena 技术论文发表(arXiv:2403.04132) | 已完成 | 24 万+ 票;建立被产业引用的可信度 | arXiv:2403.04132 |
| Apr 2024 | Arena-Hard pipeline(BenchBuilder)发布 | 已完成 | 自动化基准创建;与人类排名一致率 98.6% | arXiv:2406.11939 |
| Dec 2024 | WebDev Arena 发布 | 已完成;被 Code Arena 取代 | 首个面向编码的真实世界评估;收集 8 万+ 票 | 来源:arena.ai/blog/webdev-arena/ |
| Sep 2025 | AI Evaluations 商业产品发布 | GA;2025 年 12 月 ARR 达 $30M | 收入引擎;受 SLA 约束的企业和实验室评估服务 | 来源:arena.ai/blog/ai-evaluations/ |
| Mar 2026 | Search Arena 论文被 ICLR 2026 接收(arXiv:2506.05334) | 已接收并展示 | 搜索增强 LLM 评估方法论获得同行评审验证 | arXiv:2506.05334 |
| Mar 2026 (est.) | Battles in Direct 实验开始(10% direct 会话) | 实验阶段 | 提高投票量;校正位置偏差和同组织偏差 | 来源:arena.ai/blog/leaderboard-changelog/ |
| May 2026 | Battles in Direct 投票纳入排行榜 | 已上线 | 收窄置信区间;提示分布转向更难查询 | 来源:arena.ai/blog/leaderboard-changelog/ |
| May 2026 | Code Arena 发布(从 WebDev Arena 重建) | 已上线 | 智能体编码环境;持久会话;Cloudflare R2 快照 | 来源:arena.ai/blog/code-arena/ |
| Jun 4, 2026 | Agent Arena 以因果追踪方法论发布 | 已上线 / 早期 | 首个基于 RCT 的智能体排行榜;分析每周 200 万+ 工具调用 | 来源:arena.ai/blog/agent-arena-methodology/ |
| 2026(计划) | Agent Arena 增加行为信号;扩展更多 Arena / 模态 | 已宣布意向;无具体日期 | 智能体评估平台扩张;信号集合会继续演进 | 来源:arena.ai/blog/agent-arena-methodology/ |
日期来自博客文章、学术论文时间戳和排行榜更新日志。计划项基于已表述意向,不是已确认发布日期。
[CE008, CE012, CE029, CE033, CE035, CE036]5.5 信任、安全、质量控制与负面证据
LMArena 的公开排行榜政策(最后更新于 2026 年 4 月 30 日)明确了模型上榜资格、采样规则(要求 ≥20% 的对战发生在公开可用模型之间)、匿名模型的预发布测试协议,以及数据共享政策:发布前用 GCP 的 Sensitive Data Protection API 清洗对话。可是,截至 2026 年 6 月,LMArena 尚未发布 SOC 2 Type II 或 ISO 27001 认证,也未公开披露第三方安全审计;help.arena.ai 的隐私政策页也缺少企业数据处理协议(DPA)细节。平台经历过两次严重的可信度挑战。第一,2025 年 4 月,Meta 提交了一个「experimental, chat-optimized」版本的 Llama 4 Maverick,在 Arena 排名第 2,但它并非公开发布模型;随后 LMArena 更新政策,并重新给公开版本打分,后者跌至第 32 名。第二,Cohere、Stanford、MIT 与 AI2 在 2025 年 4 月发表论文,指称部分大实验室(Meta、OpenAI、Google、Amazon)获得了不成比例的高采样率,并能选择性压低表现较差的预发布变体,构成基准博弈;LMArena 否认这些不准确之处,并指向其已发布的采样透明度。这些事件带来持续的声誉风险,也促使研究社区呼吁更严格的预发布限制、独立审计和算法透明度。AI Evaluations 商业产品提供 SLA 承诺,但详细的正常运行时间 SLA 条款和安全披露并未公开。 [CE020, CE021, CE022, CE023, CE024, CE025]
| 控制 / 认证 | 状态 | 范围 | 缺口 / 风险 |
|---|---|---|---|
| 排行榜政策(公开) | 已上线;最近更新于 2026 年 4 月 30 日 | 模型准入、采样政策、预发布测试协议、数据共享规则 | 政策由平台自我执行;没有独立审计方或执行机制 |
| 数据清洗(GCP Sensitive Data Protection API) | 用于公开数据发布 | Conversation 数据在 HuggingFace / 公开发布前移除 PII | 依赖单一供应商;清洗效果没有公开审计 |
| 开源排名方法论(Apache 2.0) | 已上线;GitHub / PyPI 上的 arena-rank v1 | Bradley-Terry 引擎、重加权、CI;驱动当前所有排行榜 | 算法开放,但平台数据是专有的;完整复现受限 |
| 方法论的学术同行评审 | 3 篇已发表论文,1 篇 ICLR 2026 接收论文 | 论文:Chatbot Arena(arXiv 2403.04132)、MT-Bench(NeurIPS 2023)、Arena-Hard(arXiv 2406.11939)、Search Arena(ICLR 2026) | Agent Arena 因果追踪方法论尚未经过同行评审 |
| 预发布测试协议 | 2024 年 3 月起公布政策;2026 年 4 月更新 | 模型提供商可测试匿名预发布模型;结果私下共享 | 优先访问争议(Cohere / Stanford 论文)触发政策更新,但没有独立审计 |
| 企业 SLA(AI Evaluations) | 2025 年 9 月起 GA;提供 SLA | 为商业评估任务承诺交付时间线 | SLA 正常运行时间条款、赔偿和安全审计范围未公开披露 |
| 隐私政策 | help.arena.ai 上已有 | 社区平台交互中的用户数据处理 | 没有企业 DPA;GDPR / CCPA 合规细节未公开记录 |
| SOC 2 Type II / ISO 27001 | 未公开披露 | N/A —— 适用于企业云服务 | 未发布认证是企业买家的尽调缺口 |
状态反映截至 2026 年 6 月 25 日的公开资料。「未公开披露」不等于确认没有;LMArena 可能持有未发布的认证。
[CE020, CE021, CE022, CE023, CE024, CE025]5.6 附录
06客户
6.1 客户基础分层
LMArena 的客户基础分成两个本质不同的群体,必须仔细区分。第一类是免费社区:来自 150+ 个国家、每月 500 万以上的全球用户,通过 arena.ai 免费使用平台。这些用户提交 prompt、为模型输出投票,支撑排行榜背后的集体偏好信号。社区覆盖开发者、研究人员、学生、知识工作者和爱好者;按两周年博客(2025 年 4 月)披露,约 41% 的对战涉及开源模型,说明用户基础技术含量较高。公司不从这一群体直接变现,但从中获得核心 IP(偏好数据)和品牌可信度。第二类是商业客户,即为 2025 年 9 月推出的 AI Evaluations 服务付费的 AI 实验室和企业。2026 年 1 月 Series A 新闻稿确认的具名付费客户包括 OpenAI、Google 和 xAI。Anthropic 与 Meta 是排行榜上的主要模型提供方,但公开披露并未单独确认它们是否为 AI Evaluations 付费订阅客户;TechCrunch 2026 年文章称 LMArena 与 OpenAI、Google 和 Anthropic 在模型提交上「partnered with」,这一类别同时包含商业和非商业关系。企业(非实验室公司用 AI Evaluations 为自身应用基准测试模型)是公司声明的第三个目标客群;目前没有公开披露任何具名的非实验室企业客户。 [CU001, CU002, CU003, CU004, CU005, CU006]
| 分层 | 买方 / 用户 / 付款方 | 用例 | 规模 | 收入 / 战略价值 | 关键缺口 |
|---|---|---|---|---|---|
| 社区用户(免费) | 终端用户(研究人员、开发者、爱好者) | 免费模型对战;排行榜访问;Direct Chat | 500 万+ 月活用户;150+ 个国家 | 无直接收入;核心 IP 和品牌护城河;PR / 增长引擎 | 用户画像未按职业或行业公开拆分 |
| AI 实验室模型提供商(非付费 / 免费) | AI 实验室:Anthropic、Meta、Mistral、DeepSeek 等 | 免费提交模型参与公开评估,获得排行榜排名和市场可信度 | 评估 400+ 个模型;300+ 次预发布测试 | 无直接收入;为平台价值提供模型供给 | 在所有沟通中并未清楚区分非付费客户与付费客户 |
| AI 实验室商业客户(付费) | AI 实验室:OpenAI、Google、xAI(已点名) | 带 SLA 的私有 AI Evaluations;为模型改进提供全面反馈 | 3 家已点名;其他客户未知 | 产生收入;2025 年 12 月该分层 ARR run-rate 为 $30M | 客户数量、合同规模和单个客户贡献未披露 |
| 企业客户(付费) | 希望为生产用途评测 AI 模型的企业 | 用 AI Evaluations 衡量特定领域(法律、医疗、软件)的模型表现 | 规模未知;没有点名企业客户 | 公司表述的目标客群;尚无公开证据显示已有非实验室企业客户 | 没有点名企业客户;pipeline 和转化率未知 |
| 学术 / 研究用户(免费) | 大学研究人员、智库、开源贡献者 | 用公开数据集做研究;把排行榜作为参照;分析对战数据 | 数百篇研究论文引用 Arena | 无直接收入;建立学术可信度并验证方法论 | 尚未宣布正式学术合作伙伴计划 |
分层基于官方新闻稿、博客和新闻报道。收入归因根据唯一披露的 ARR 数字估算;分层拆分未公开。
[CU001, CU002, CU003, CU004, CU005, CU006]从首次接触到商业扩张,梳理客户细分、采用入口和关键互动触点。
旅程阶段根据官方沟通和媒体报道推断;公开资料中没有客户研究或 NPS 数据。
[CU001, CU002, CU005, CU006]6.2 采用轨迹与使用指标
LMArena 的社区采用增长很快,且证据链较完整。平台 2023 年 5 月作为 UC Berkeley 研究项目上线;到 2025 年 5 月,公司披露以 $600M 估值完成 $100M 种子轮,彼时已有显著牵引力。到 2026 年 1 月 Series A 公告时,LMArena 报告 5M+ 月度用户(高于 2025 年 9 月 AI Evaluations 博客中隐含的约 3M)、60M+ 月度对话、250M+ 累计对话和 2M+ 月度投票。Series A 博客把 2025 年 5 月种子轮以来的社区增长量化为「25x」。模型评估方面,LMArena 已评估 400+ 个公开模型和 300+ 个预发布模型变体。从 2025 年 9 月商业化上线到 2025 年 12 月(不到四个月),按 CEO Anastasios Angelopoulos 披露、并由 TechCrunch 2026 年 1 月文章和官方新闻稿确认,年化收入消耗 run-rate 达到 $30M。Series A 文章还披露,平台已收集 50M+ 社区投票、发布 145k+ 开源对战数据点,并完成 400+ 次跨模态模型评估。值得注意的是,平台增长看起来主要由拉动驱动:TechCrunch 播客提到,LMArena 排行榜「became something of an obsession among model makers」,意味着社区客群的获客成本较低。 [CU001, CU002, CU003, CU004, CU007, CU008]
| 指标 | 数值 | 日期 | 来源 | 置信度 | 含义 |
|---|---|---|---|---|---|
| 月活用户 | 3M+ | Sep 2025 | 来源:arena.ai/blog/ai-evaluations/ | 中(公司披露) | 商业化上线时,社区规模已经明显放大 |
| 月活用户 | 5M+ | Jan 2026 | 来源:PRNewswire / arena.ai/blog/series-a/ | 高(多源) | 不到 2 年 MAU 增长 5 倍;自然拉力强 |
| 月度对话 | 60M+ | Jan 2026 | PRNewswire / TechCrunch Jan 2026 | 高(多源) | 平均每名用户每月产生约 12 次对话 |
| 累计对话 | 250M+ | Jan 2026 | 来源:PRNewswire / arena.ai/blog/ai-evaluations/ | 中(公司披露) | 历史数据长尾丰富,增强偏好模型质量 |
| 月度投票 | 2M+ | Sep 2025 估计 | 来源:arena.ai/blog/ai-evaluations/ | 中(公司披露) | 相对用户数的参与率较高;约 0.7 票 / 用户 / 月 |
| 社区累计投票 | 50M+ | Jan 2026 | 来源:arena.ai/blog/series-a/ | 中(公司披露) | 累计投票语料支撑 BT 模型校准 |
| 已评估模型(公开) | 400+ | Apr 2025 / Jan 2026 | 来源:arena.ai/blog/two-year-celebration/, arena.ai/blog/series-a/ | 中(公司披露) | 覆盖面越宽,平台中立性的说法越站得住 |
| 发布前模型测试 | 300+ | Apr 2025 | 来源:arena.ai/blog/two-year-celebration/ | 中(公司披露) | 实验室参与度强;也是操纵争议风险的来源 |
| 地理覆盖 | 150+ 个国家 | Jan 2026 | PRNewswire | 中(公司披露) | 全球多样性增强偏好数据代表性 |
| ARR 运行率 | 年化 $30M | Dec 2025 | TechCrunch Jan 2026 / PRNewswire | 高(多源) | 从 Sep 2025 的零收入快速爬坡,验证商业模式 |
| 社区增长率 | 同比 25 倍 | May 2025–Jan 2026 | 来源:arena.ai/blog/series-a/ | 低-中(公司口径) | 单个数据点;指标分母不清楚 |
除标注估计外,数值均为公司披露。「高」置信度需要两个独立来源。公开披露没有给出环比轨迹。
[CU001, CU002, CU003, CU007, CU008, CU009]从发现到社区投票、模型提交,再到商业评估:LMArena 的互动漏斗。
月访客是粗略估计;投票转化率由披露的 MAU 和月投票数推导。商业客户数是下限(仅计入具名客户)。
[CU001, CU002, CU007, CU008, CU009]6.3 具名客户证明与生产证据
AI Evaluations 商业产品的具名付费客户证据有限,但指向明确。2026 年 1 月 PR Newswire 新闻稿称:「LMArena works with leading AI labs and enterprises, including OpenAI, Google and xAI, all drawing on LMArena's evaluations to improve their models for production use cases.」这是公司直接点名客户,也是目前最强的公开客户证明。另一个证据来自 Felicis 合伙人 Peter Deng,他在同一新闻稿中表示:「We're leading this round because LMArena has built the most trusted, reliable, real-world signal of AI performance. They have become essential infrastructure for every lab and enterprise.」公司的 Agent Arena 方法论博客记录了真实用户会话——一个实时体育电视赛程网站、一个自托管电影待看片单、一个自主水下航行器自动驾驶——作为生产式部署示例,不过这些是社区用户,不是付费企业客户。商业化之前,Google 的 Kaggle、a16z 和 Together AI 曾向 LMSYS(前身非营利组织)捐赠算力和 grant。lmsys.org 的历史上线记录显示,最初 2023 年的合作结构包括 UC Berkeley。尚未确认付费的模型提供方包括 Anthropic 和 Meta:二者经常因旗舰模型上榜而被提及,但 TechCrunch 2025 和 2026 年文章一直只把 OpenAI、Google 与 Anthropic 称为「partners」,而新闻稿中具名商业客户只有 OpenAI、Google 和 xAI。AI Evaluations 服务没有公开可得的政府采购记录、G2/Capterra 评价或独立第三方案例研究。 [CU005, CU006, CU012, CU013, CU014, CU015]
| 客户 / 用户 | 分群 | 部署 / 用例 | 生产 vs. 试点 | 成果证据 | 证据限制 |
|---|---|---|---|---|---|
| OpenAI | AI 实验室(付费商业客户) | AI Evaluations——评估 GPT 模型变体的人类偏好;旗舰模型登上公开排行榜 | 生产(新闻稿具名) | Jan 2026 Series A PRNewswire 新闻稿将其列为客户,并称其把评估用于生产 | 没有 OpenAI 直接引述;仅为公司披露 |
| AI 实验室(付费商业客户) | AI Evaluations——评估 Gemini 模型变体;旗舰模型登上公开排行榜 | 生产(新闻稿具名) | Jan 2026 PRNewswire 新闻稿具名;Gemini 模型持续进入文本和视觉排行榜 | 没有 Google 直接声明;公司披露 | |
| xAI | AI 实验室(付费商业客户) | AI Evaluations——评估 Grok 模型变体的人类偏好 | 生产(新闻稿具名) | Jan 2026 PRNewswire 新闻稿与 OpenAI、Google 一同具名 | 三个具名客户中最新的一家;未披露产品使用细节 |
| Anthropic | AI 实验室(模型提供方;付费状态未确认) | Claude 模型登上公开排行榜;据 TechCrunch 2026,模型提交为「合作」关系 | 活跃生产(排行榜) | Claude 被描述为赢得法律 / 医疗用例专家排行榜(TechCrunch 播客);合作伙伴关系已确认 | 没有任何来源单独确认其为 AI Evaluations 付费客户 |
| Meta | AI 实验室(模型提供方;付费状态未确认;有负面历史) | Llama 模型登上公开排行榜;Jan–Mar 2025 测试了 27 个发布前 Llama 4 变体 | 活跃生产(排行榜) | Cohere/Stanford 论文确认了大规模发布前测试;Llama 模型定期在排行榜更新 | 操纵事件:实验性 Maverick 提交后排名第 #2;公开版本排名约第 32;关系受到负面影响 |
| Perplexity | AI 实验室(模型提供方;搜索竞技场合作伙伴) | Sonar 模型登上 Search Arena 排行榜;引用风格协作 | 活跃(排行榜) | Search Arena 评估了 11 个 Perplexity 模型变体;风格随机化与提供方协作完成 | 没有 AI Evaluations 付费订阅证据 |
具名付费客户(OpenAI、Google、xAI)来自 January 2026 PRNewswire 新闻稿。模型提供方参与(Anthropic、Meta、Perplexity)来自新闻文章和博客文章,但不能确认 AI Evaluations 商业订阅。公开渠道没有独立客户引述或证言。
[CU005, CU006, CU012, CU013, CU014, CU015]按维度评估每个具名客户或模型提供商的证据质量。
证据质量评级反映截至 2026 年 6 月可得的最佳公开证据;“未确认”不等于“否”。
[CU005, CU006, CU012, CU013, CU014, CU021]6.4 留存、耐久性与重复使用
LMArena 尚未公开披露 AI Evaluations 产品的正式留存指标,例如净收入留存(NRR)、总收入留存(GRR)或客户流失率。商业化时间还太短(2025 年 9 月上线),没有多年 cohort 数据。社区端的主要留存信号是平台活跃度:按 Series A 博客,2025 年 5 月至 2026 年 1 月社区同比增长 25x;250M+ 累计对话相对 60M/月活跃生成量,说明累计使用的大部分来自持续扩大的活跃社区,而非一次性访问。AI Evaluations 采用基于消耗的模型(「annualized consumption rate」而非订阅 ARR),因此留存取决于实验室和企业是否持续提交评估请求。不到四个月从 $0 爬到 $30M 年化,显示初始实验室需求强劲;但实验室发展内部评估能力后,这种需求能否持续,是关键风险。排行榜 changelog 显示模型持续新增(2026 年 6 月每周数个),可作为模型提供方持续参与的间接代理。对于一家估值 $1.7B、ARR run-rate 为 $30M 的公司而言,没有披露客户 logo 墙、案例研究站点或 testimonial 页面(除 Felicis 与 UC Investments 的投资人引述外),这一点值得注意。 [CU007, CU009, CU010, CU011, CU016, CU017]
| 指标 | 数值 / 状态 | 分群 | 置信度 | 尽调要求 |
|---|---|---|---|---|
| Net Revenue Retention (NRR) | 未披露 | AI Evaluations 商业客户 | Unknown | 向 LMArena 索取实验室客户购买后第 1–6 个月的 NRR 或 cohort 分析 |
| Gross Revenue Retention (GRR) | 未披露 | AI Evaluations 商业客户 | Unknown | 确认 Sep 2025 上线以来是否有商业客户流失 |
| 月度用户留存(社区) | 未披露;60M 月度对话和持续增长暗示较高 | 社区(免费用户) | 低(推断) | 索取月活用户留存或会话频次数据 |
| 客户流失(商业) | 未披露;上线未满 4 个月,流失数据极少 | AI Evaluations 商业客户 | Unknown | 跟踪 OpenAI/Google/xAI 是否在下个周期续约 |
| 重复提交模型(模型提供方) | 300+ 次发布前测试和 400+ 次公开评估,意味着约 20+ 家实验室持续重复参与 | 模型提供方(免费 + 付费) | 低-中(由平台数据推断) | 获取提交模型超过一次的独立提供方数量 |
| 间接留存信号:排行榜活动 | 根据 changelog,June 2026 每周新增多个模型 | 所有模型提供方 | 中(观察) | 验证活跃模型新增与提供方满意度相关,而不是营销压力 |
| 用户满意度(社区) | 未正式披露;没有发布 NPS 或 CSAT 分数 | 社区用户 | Unknown | 如有发布,查找 NPS 调查或用户评价数据 |
标注「未披露」的数值截至 June 25, 2026 未出现在公开来源中;这不代表数值为零。置信度评级反映可得证据。
[CU016, CU017, CU018]6.5 反向证据:冲突、博弈与集中度风险
LMArena 的客户和模型提供方关系带来有充分记录的结构性风险。第一是利益冲突:付费使用 AI Evaluations、充当商业客户的同一批 AI 实验室(OpenAI、Google、Anthropic 等),也是其模型在公开排行榜上被排名的主体。这天然激励实验室优化 Arena 表现,而它们可以通过预发布测试协议做到这一点。TechCrunch 播客直接提出这个问题:「how a team like theirs can build a neutral benchmark when the companies they're ranking are also their backers.」第二是 Meta Llama 4 Maverick 事件(2025 年 4 月):Meta 向 Arena 提交了一个「chat-optimized experimental」版本的 Maverick,拿到第 2 名,但公开发布模型大约排在第 32。LMArena 更新政策并表示「Meta's interpretation of our policy did not match what we expect from model providers.」该事件说明,依赖实验室提交模型的商业合作结构,会给平台带来实际压力,去容纳损害排行榜完整性的行为。第三,Cohere/Stanford/MIT/AI2 研究(2025 年 4 月)指称,Meta 可在 2025 年 1 月至 3 月间私下测试 27 个模型变体,再选择表现最好者提交;顶级实验室获得不成比例的采样率;研究初步发现分享给 LMArena 后并未被其否认。LMArena 否认具体指控,但宣布了新的采样算法。这些指控尚未经过独立裁定。第四,客户集中:公开具名的付费客户只有三家(OpenAI、Google、xAI),LMArena 的商业收入高度集中;若其中任何一家实验室自建或签约替代评估供应商,对 ARR 的影响可能显著。 [CU005, CU006, CU012, CU019, CU020, CU021]
| 扩张驱动 / 集中度风险 | 影响 | 证据 | 尽调路径 |
|---|---|---|---|
| 客户集中度:3 家具名付费客户 | 高——如果 OpenAI、Google 或 xAI 减少或取消 AI Evaluations 支出,单一客户收入冲击可能很重 | 新闻稿只具名 3 家客户;没有其他具名商业客户 | 确定每家具名客户收入占比;确认未具名商业客户数量 |
| 利益冲突:评分对象 = 付费方 | 高——付费购买评估的实验室也是模型被排名的主体,带来操纵激励和可信度风险 | Meta 操纵事件(Apr 2025);Cohere/Stanford 论文指称存在优先访问 | 结构性防护审计:索取采样算法和发布前测试限制细节 |
| 登陆扩张:一次性评估 → 重复合同 | 中——如果实验室认可评估价值,按量消费模式支持扩张;没有多年合同证据 | CEO 提到「consumption rate」;$30M ARR 快速爬坡暗示重复使用 | 向 LMArena 确认合同结构(现货 vs. 订阅 vs. 年度) |
| 非实验室企业扩张 | 中——潜在市场很大,但目前没有具名的非实验室企业客户 | Series A 新闻稿提到「enterprises」;TechCrunch 文章提到「software engineering, law, medicine」等目标领域 | 识别任何正在推进的非实验室企业试用或试点 |
| 竞争替代风险:实验室内部搭建评估能力 | 中——如果头部实验室(Google、OpenAI)搭出能满足自身需求的内部评估流水线,LMArena 商业需求可能见顶 | Scale AI 的 SEAL Showdown(Mashable 文章)作为竞争性评估产品上线 | 跟踪实验室关系的规模 / 体量随时间增长还是趋稳 |
| 未来操纵事件带来的负面声誉风险 | 高——如果再次出现高关注度操纵事件,社区信任和商业可信度可能同时快速流失 | Meta Llama 4 事件被 The Verge、TechCrunch 等广泛报道;LMArena 更新了政策,但结构性冲突仍在 | 监测新的操纵指控;在下个评估周期复核更新后的采样政策 |
集中度和利益冲突评估基于截至 June 2026 的公开来源证据。
[CU019, CU020, CU021, CU022, CU023, CU024]离散采用里程碑显示,LMArena 社区从 2023 年到 2026 年 1 月的增长轨迹。
数值混合了不同指标(累计投票、数据集提示、MAU)来展示轨迹;彼此不能直接比较。MAU 是所列日期的月活跃用户;投票语料为累计值。240k 是 2024 年 3 月论文发布时的累计投票;截至 2026 年 1 月,累计投票报告为 50M+。
[CU001, CU007, CU008, CU009, CU010]6.6 附录
07风险
7.1 利益冲突与基准完整性风险
LMArena 的核心风险,是独立裁判身份与收入模型之间的结构性冲突。公司向同一批 AI 实验室——OpenAI、Google、xAI——收取付费评估服务费用,而这些实验室的模型又出现在其公开排行榜上。这种双重角色会制造真实或被感知的偏袒激励。来自 Cohere、Stanford、MIT 和 AI2 的学术研究人员在《The Leaderboard Illusion》(arXiv 2504.20879)中记录了这些动态:少数大型提供商可以私下测试多个模型变体,只选择披露最佳分数,并获得不成比例的更多评估对战。Meta 的 Llama 4 Maverick 事件进一步说明,即使没有明确违规,政策缝隙也足以让基准博弈发生:Meta 向 LMArena 提交了一个针对对话性优化、从未公开发布的模型,拿到前二排名,直到差异被公开曝光。LMArena 已在 2026 年 4 月更新政策,但商业依赖与中立感知之间的根本张力仍然存在。包括独立研究员 Gwern 在内的评论者称该排行榜是「a cancer」,并质疑其信号是否仍有科学价值。ucstrategies.com 分析发现,专门为 Arena 偏好调优的模型最多可把分数抬高 100 Elo;SurgeAI 分析则发现,评估者有 52% 的时间不同意 LMArena 投票。一旦可信度叙事永久裂开——例如出现高关注度偏见调查、付费实验室丑闻,或监管机构调查 AI 基准准确性声明——LMArena 面向消费者的排行榜和企业评估收入两端的产品价值主张会同时坍塌。 [CR001, CR002, CR003, CR004, CR005, CR006]
| 风险 / 规则 / 案件 | 司法辖区 | 状态 | 可能性 | 严重性 | 缓释 | 剩余暴露 | 尽调路径 |
|---|---|---|---|---|---|---|---|
| GDPR 数据共享充分性——在没有明确细粒度同意的情况下与 AI 提供方共享用户提示 | EU / EEA | 活跃合规义务;政策于 Sep 2025 更新 | 高 | 关键 | 已更新隐私政策;数据共享前使用 GCP Sensitive Data Protection API | 如果 EU DPA 审计数据共享实践,存在执法风险;训练后模型的删除权存在缺口 | 索取 DPA 往来函件;复核与 AI 提供方的数据处理协议 |
| EU AI Act——GPAI 模型透明度文档要求 | EU | GPAI 义务自 Aug 2025 生效;高风险系统义务于 Aug 2026 生效 | 中 | 高 | 学术论文和开源 Arena-Rank 仓库记录了方法论 | 公开渠道未找到正式 GPAI 文档证明;不清楚 LMArena 是「systemic risk」模型提供方还是下游部署方 | 索取 LMArena 的 EU AI Act 合规备案或自评文件 |
| Digital Services Act——规模扩大后可能触发 VLOP 义务 | EU | 平台达到阈值(>45M EU 月活用户)时适用 | 低-中 | 高 | 根据公开数字,截至 June 2026 平台低于 VLOP 阈值 | 用户规模逼近 VLOP 触发线后,系统性风险评估、独立审计和透明度报告将成为强制要求 | 监测 MAU 披露;聘请 DSA 法律顾问 |
| 美国州隐私法——CPRA、VCDPA、CPA 等 | 美国(多州) | 需要持续合规 | 中 | 中 | 隐私政策承认州法权利章节 | 未找到外部隐私审计或 CPRA 证明;California AG 在 2026 年执法活跃 | 索取隐私审计;验证 CPRA 合规计划 |
| 知识产权——用户提示版权和训练数据来源 | 全球 | 法律未定;AI 版权诉讼在 2025–2026 年全行业活跃 | 中 | 中 | 服务条款授予 LMArena 使用用户内容的许可;提示与提供方共享 | 如果用户提交内容在没有充分许可的情况下用于模型训练,存在版权侵权暴露 | 复核服务条款;审计到提供方训练流水线的数据使用流 |
| 基准准确性主张——FTC 广告真实性风险 | 美国 | 截至运行日期没有已知执法行动 | 低 | 中 | 方法论已开源;学术发表披露了统计局限 | 任何声称 Arena 排名等同于「real-world performance」的商业表述,都可能受到 FTC Section 5 审查 | 监测 FTC AI 指引;复核营销主张措辞 |
所有可能性 / 严重性评级都是作者基于公开证据和适用监管框架的评估;LMArena 没有公开披露官方监管往来函件。各行按严重性排序(关键 → 中)。截至 2026-06-25 的诉讼检索未发现将 Arena Intelligence, Inc. 列为当事方的活跃案件。
[CR007, CR008, CR009, CR010]这张按严重度和可能性交叉的热力图,将八项主要风险映射到五档影响和四档可能性,显示利益冲突与 Goodhart 定律风险占据右上象限。
可能性和影响评级是作者基于公开证据和监管先例作出的评估;LMArena 未披露内部风险台账。
[CR001, CR007, CR012, CR017]7.2 监管、法律与数据隐私风险
LMArena 面对复杂且快速变化的监管环境。法律实体 Arena Intelligence, Inc. d/b/a LMArena 处理来自 150 个国家、500 万月度用户的个人数据,因此触发欧盟《通用数据保护条例》(GDPR)、欧盟 AI Act(通用 AI 模型义务自 2025 年 8 月起生效)、Digital Services Act(规模达到时可能触发 VLOP 门槛),以及包括 California CPRA 在内的一系列美国州隐私法义务。LMArena 自身隐私政策(2025 年 9 月生效)明确提醒用户,prompt 和生成回复可能被公开共享,这带来 GDPR 第 7 条下用户同意充分性的风险。已训练评估模型的删除权挑战(GDPR 第 17 条)是整个 AI 行业已知合规缺口,目前没有明确的执法解释。欧盟 AI Act 关于通用 AI 文档和透明度的要求,也会增加运营负担。LMArena 将用户 prompt 分享给 AI 提供商用于模型改进,这种做法带来数据流义务,必须在各司法辖区用数据处理协议覆盖。欧盟 AI Act 对「高风险」AI 系统的定义尚未明确点名 AI 基准测试平台,但法律、医疗、就业等受监管领域的评估可能引来行业特定审查。截至运行日期,公开记录未发现直接涉及 LMArena 或 Arena Intelligence, Inc. 的诉讼;不过,公司尚未披露其合规认证状态或任何监管问询。用户提交 prompt 被用于训练带来的 IP 风险,在多个司法辖区仍是未定法律问题。 [CR007, CR008, CR009, CR010, CR011]
| 失效模式 | 可能性 | 严重性 | 缓释成熟度 | 剩余暴露 | 未解决缺口 |
|---|---|---|---|---|---|
| AI 提供方 API 撤回——实验室因竞争争议或评分抗议撤销访问 | 中 | 关键 | 低(未披露与提供方的 SLA) | 被撤模型在完整排行榜中留下缺口;若为付费客户则带来收入风险 | 未披露公开 API SLA 或冗余计划 |
| 国家行为体或竞争对手发起协同女巫 / 投票操纵活动 | 中 | 高 | 中等(匿名投票;部分 bot 检测;重加权算法) | 分数污染很难实时发现;事后需要修订方法论 | 没有公开异常检测披露;Battles in Direct 的位置偏差校正显示其具备被动响应能力 |
| 重大数据泄露——社区提示数据集或企业评估结果泄露 | 低-中 | 高 | 低-中(GCP Sensitive Data Protection;标准企业安全) | 用户信任受损;可能触发 GDPR 泄露通知义务;评估 IP 暴露给竞争对手 | 截至运行日期未披露 SOC 2 或 ISO 27001 认证 |
| 云基础设施故障——battle 匹配或 API 路由持续停机 | 低 | 中 | 中等(云提供商 SLA;假设多区域部署但未确认) | 实时评估连续性中断;可能违反企业 SLA | 未找到公开状态页 SLA 或正常运行时间承诺 |
| 方法论过拟合——排名与真实世界模型质量的偏离加速 | 高 | 高 | 中等(开源方法论;学术论文发表;政策更新) | 排行榜失去可信信号;开发者和企业流失 | Leaderboard Illusion 论文后,没有发布评分方法论外部统计审计 |
| 模型蒸馏攻击——对手系统性挖掘 Arena 评估分布 | 中 | 中 | 低(数据按设计部分公开;未披露反爬控制) | 竞争信息泄露;Arena 数据优势被侵蚀 | AI 安全文献记录了蒸馏风险;LMArena 未披露应对措施 |
缓释成熟度评级仅基于公开可得信息。截至 2026-06-25,LMArena 未发布内部安全或基础设施审计结果。各行按严重性排序。
[CR012, CR013, CR014]这张有向无环图展示,根因风险如何传导到收入、用户参与、监管暴露和估值等下游影响。
[CR001, CR003, CR007, CR017]7.3 运营、技术与安全风险
LMArena 的运营风险集中在平台可靠性、数据质量完整性,以及社区生成评估数据集的安全。每月 6000 万次对话流经其基础设施,一旦持续宕机或发生数据泄露,会直接伤害 LMArena 区别于静态基准的实时、连续评估主张。平台高度依赖云基础设施(同时运行多个实时 AI 模型 API 的算力成本),其价格和可用性由第三方控制,其中包括 LMArena 正在评估的同一批 AI 实验室。如果某个提供商撤回 API 访问——当实验室质疑评分或竞争格局变化时,这并非不可行——排行榜会立刻出现缺口。统计完整性是另一项运营风险:LMArena 的 Elo/Bradley-Terry 评分依赖足够均匀且无偏的对战分布;任何系统性 prompt injection、协同投票活动或针对平台用户基础的女巫攻击,都会污染排名信号,而且未必能被实时发现。2026 年 5 月转向「Battles in Direct」后,10% 的 direct-chat 会话生成对战投票,需要事后校正位置偏差和同组织指标偏差,这说明方法论更新会在事后改变模型排名。评估数据集是具有商业价值的资产,如何防范模型蒸馏攻击或未经授权抓取,仍是持续担忧;尤其是领先 AI 实验室有财务动机研究评估分布。约 41 人的员工规模,对于这一量级的平台形成了关键人集中风险。 [CR012, CR013, CR014, CR015, CR016]
| 依赖 | 交易对手 | 角色 | 集中度 | 失效场景 | 严重性 | 缓释 | 剩余暴露 |
|---|---|---|---|---|---|---|---|
| AI 实验室收入——前三大付费实验室(OpenAI、Google、xAI) | OpenAI / Google / xAI 客户 | 主要商业客户 | 高(估计占 $30M ARR 的 >50%) | 客户退出评估合同;排行榜数据冲突 | 关键 | 多实验室客户基础;开源社区中立性承诺 | 如果一家 Tier-1 实验室流失,收入可能断崖式下滑 |
| AI 提供方 API 访问——battle 中所有模型都需要实时 API | 所有主要实验室 | 模型能力提供方 | 高(每家实验室控制自己的 API) | 实验室撤回 API 访问权限;模型从排行榜下线 | 高 | API 访问历史上较稳定;政策承诺纳入公开模型 | 没有合同锁定保障;任何实验室都可短期通知后撤回 |
| 云计算基础设施 | GCP / 主要云服务商 | 计算、存储、数据保护服务 | 高(明确提及 GCP Sensitive Data Protection) | 云服务中断、涨价或条款变化 | 高 | 标准云企业协议;假设已有冗余 | 若依赖单一云,整个平台可能停摆 |
| UC Berkeley / 学术研究管线 | UC Berkeley SkyLab | 人才来源;研究背书;过往资助记录 | 中 | 核心研究人员转向行业;学术合作降温 | 中 | 团队已独立注册公司;研究发表仍在继续 | 若学术切分演变为对立,可信度会受损 |
| 投资人关系 — Felicis(领投)与 UC Investments | Felicis / UC Investments | 资本提供方;战略锚点 | 中(Felicis 的 Peter Deng 曾任 OpenAI 员工) | 若被评测公司与 Arena 排名发生争议,投资人可能出现利益冲突 | 中 | 治理结构未公开披露;Felicis GP 的 OpenAI 背景带来观感风险 | 若投资人冲突公开浮出,独立性认知会受损 |
收入集中度估算基于约 100 个客户和已披露 $30M ARR,并假设早期 B2B SaaS 常见的帕累托分布。按 2026-06-25 的公开信息,LMArena 与 AI 提供商之间尚未披露合同条款或正式 SLA。
[CR017, CR018, CR019]展示 LMArena 对云服务商、AI 实验室 API、学术人才管线和投资方的关键基础设施与商业依赖。
[CR013, CR018]7.4 合作伙伴集中、财务与执行风险
LMArena 的商业模式依赖少数大型 AI 实验室和企业,贡献其 $30M 年化收入的大部分。2026 年初约有 100 家付费客户,头部客户为 OpenAI、Google 和 xAI,收入集中度具有实质性:失去一两家 Tier-1 实验室关系,就可能带来不成比例的收入下滑。同一批实验室既是 LMArena 最好的客户,也是最有动机博弈其排行榜的潜在对手,这一张力没有干净解法。$1.7B 估值对应 $30M ARR 约 57x 收入倍数,定价中包含很高的增长预期:公司必须守住信任、扩展新领域,并向现有客户 upsell。41 名员工的执行产能,相对 LMArena 已承诺的产品面(Agent Arena、WebDev Arena、Vision、Video、Search、Code、Document 排行榜,加商业评估)偏薄。公司起源于 UC Berkeley 研究项目,创始团队的学术关系和发表义务可能分散管理层对商业执行的注意力。现金消耗动态未披露;以 $250M 融资和 41 人团队看,runway 似乎较长,但每月 60M 次对话的基础设施成本很高。未来融资风险中等但并非为零:如果 AI 评估市场未能增长到 2030 年预计的 $3.8B,后续轮次可能会反映失望定价。$382,500 的 Recall Capital-LMArena SEC Form D 备案(2026 年 2 月)反映了二级市场投资者需求,但不能证明公司自身财务健康。 [CR017, CR018, CR019, CR020, CR021]
| 角色 / 职能 | 依赖或缺口 | 可能性 | 严重性 | 缓解措施 | 尽调路径 |
|---|---|---|---|---|---|
| CEO — Anastasios Angelopoulos | 创始 CEO;公开发言人;学术可信度锚点;所有战略关系都经由他推进 | 低-中 | 关键 | 联合创始人 Wei-Lin Chiang 任 CTO;Ion Stoica 任顾问 | 接班计划和关键人条款未公开披露;需索取董事会文件 |
| 统计方法团队 | 核心 Elo / Bradley-Terry 评分能力集中在约 41 人的小型学术团队 | 中 | 高 | 开源 Arena-Rank 仓库;学术论文提供外部校验 | 确认方法团队规模和梯队深度;要求披露员工人数 |
| 企业销售与客户成功 | 尚未公开任命企业销售负责人;商业产品于 2025 年 9 月推出 | 中 | 高 | ARR 快速增至 $30M,说明早期 GTM 已有牵引 | 索取组织架构图;确认 VP Sales 招聘进度 |
| 安全与合规负责人 | 41 人团队中未看到 SOC 2 / GDPR DPO / EU AI Act 合规岗位的公开证据 | 高 | 高 | 已有隐私政策;提及 GCP 数据保护工具 | 核实 DPO 指定情况;索取合规组织结构 |
| 研究科学家留存 | 人才市场竞争激烈;前沿 AI 实验室薪酬为创业公司的 2–3 倍 | 高 | 中 | 股权;使命驱动文化;UC 关系 | 审查期权池规模;留存悬崖日期 |
41 人员工数来自 Latka(2026 年 1 月)。组织架构图未公开。接班和关键人安排未公开披露。各行按严重性排序。
[CR020, CR021]| 风险 | 可监测触发项 | 阈值 / 事件 | 行动含义 |
|---|---|---|---|
| 利益冲突导致可信度坍塌 | Tier-1 AI 媒体报道情绪;批评 Arena 方法论的学术论文数量 | 单季度出现两篇或以上同行评审论文,记录 LMArena 未能反驳的系统性偏差 | 投资逻辑破裂:退出或暂停评估;只有方法论经独立审计后才重新开启 |
| 顶级实验室客户流失 | 与 OpenAI、Google、xAI 的商业评测协议续约情况 | 前三大实验室中任一家公开终止或公开质疑评测结果 | 监控合同续约日期;若续约存疑则升级处理 |
| 监管执法行动 | 披露 EU DPA 询问、FTC 信息要求或州 AG 调查 | 针对 Arena Intelligence, Inc. 的任何正式监管调查启动 | 投资逻辑破裂:冻结新增资本部署,等待法律解决 |
| 开发者社区放弃基准可信度 | GitHub stars、HuggingFace 排行榜引用,以及开发者论坛对 Arena 方法论的提及 | 模型发布公告中的 Arena 引用率同比下降 >30% | 投资逻辑预警:调查根因并评估缓解空间 |
| 收入集中度断崖 | 董事会材料披露前三大客户 ARR 占比 | 单一客户超过 ARR 的 40%,并释放不满信号 | 投资逻辑预警:下一轮融资前必须实现多元化 |
| 关键人物离职 | Angelopoulos 或 Chiang 的公开公告或 LinkedIn 更新 | 任一创始人在 Series A 交割后 24 个月内离职 | 投资逻辑破裂:立即暂停,等待接班安排清晰 |
触发阈值是作者基于同类早期 AI 基础设施公司尽调实践设定的启发式标准。所有触发项都需结合语境判断——单一负面事件不一定单独击穿投资逻辑。
[CR003, CR006, CR017]7.5 附录
08估值
8.1 融资与估值背景
Arena Intelligence, Inc.(d/b/a LMArena)在 2026 年 1 月宣布完成 $150M Series A,投后估值 $1.7B。此前公司于 2025 年 5 月以 $600M 估值完成 $100M 种子轮,约七个月内累计融资达到 $250M。Series A 由 Felicis 和 UC Investments(University of California)领投,Andreessen Horowitz、The House Fund、LDVP、Kleiner Perkins、Lightspeed Venture Partners 和 Laude Ventures 参投。公司年化「consumption run rate」——公司称其等同于 ARR——在 2025 年 12 月超过 $30M,距离 2025 年 9 月推出首个商业产品(AI Evaluations)不到四个月。按 $30M run rate 计算,投后估值 / ARR 倍数约为 57x,而晚期 Series A SaaS 公司通常在 8-15x。不过,四个月内从 $0 ARR(2025 年 5 月)到 $30M ARR(2025 年 12 月)的轨迹极其罕见,投资人定价的是未来增长,而不是当前收入。SEC EDGAR Form D(accession 0002113470-26-000001,2026 年 2 月提交)中「Recall Capital-LMArena」二级基金规模为 $382,500,显示二级市场需求,其隐含估值与 Series A 价格一致。股权结构细节——从种子轮到 Series A 的稀释(种子轮约出售 17%,Series A 约出售 9%)——基于披露的投后估值和轮次规模,隐含 Series A 前企业价值为 $1.55B。未发现可转债、优先清算权或既往 down-round 风险的公开证据;公司八个月内估值 step-up 为 2.83x。 [CV001, CV002, CV003, CV004, CV005, CV041]
| 维度 | 评估 | 备注 |
|---|---|---|
| 推荐 | 跟踪 | 转为买入取决于方法论审计 + 入场价 <30x NTM ARR |
| 置信度 | 中 | 证据足以形成判断;核心风险是结构性的,且仍未解决 |
| 风险评级 | 高 | 利益冲突 + 极高倍数 + 未审计治理 |
| 估值立场 | 偏高 | 57x ARR 倍数位于最高十分位;$1.7B 估值下安全边际不足 |
| 决策含义 | 不按当前价格领投或共同领投;继续跟踪治理改善 | 若 ARR 达到 $57M+,或入场价降至 $1.0–1.2B,则重新评估 |
评估截至 2026-06-25,依据公开披露的融资数据(Series A 投后估值 $1.7B、截至 2025 年 12 月 ARR $30M)和公开风险证据。未审阅内部财务模型、董事会材料或数据室。
[CV001, CV015]| 维度 | 正向投资逻辑 | 反向投资逻辑 | 什么会改变判断 |
|---|---|---|---|
| 市场结构 | 前沿模型快速增多,中立评测基础设施变成必需品;目前没有可信的独立替代者能做到同等规模 | AI 实验室会逐步把评测内化,或组建联盟以绕开第三方费用 | 实验室自建评测达到可比覆盖(5M+ 用户);或 LMArena 失去前三大付费实验室中的 2 家 |
| 网络效应 | 2025 年 5 月以来,5M 月活用户 / 60M 次对话 / 400+ 个模型评测形成数据护城河 | 社区投票可被操纵;Elo 分数会变成优化代理,而不再代表质量 | 2026 年批评 Arena 方法论的学术论文数量翻倍;开发者引用率下滑 |
| 利益冲突 | 透明度、开源方法论和更新后的政策缓解了风险;丑闻之后,社区信任仍然守住 | 同行评审研究已有记录;只要收入依赖被评测实验室,根本张力就无法消除 | 独立审计确认方法论完整性;或来自非实验室客户的收入多元化超过 ARR 的 60% |
| 收入轨迹 | 从 $0 到 $30M ARR 只用 4 个月,表现异常强;说明 AI 评测有强产品市场匹配 | 收入集中在约 100 个客户;前三大实验室很可能占多数;任何流失都会被放大 | 客户 cohort 数据显示前三大集中度 <20%,且 NRR >110% |
| 估值支撑 | 可比 AI 基础设施公司的加权平均倍数可支撑高增长先发者 30-40x ARR | 57x ARR 且未披露盈利路径,相比企业 SaaS 中位数偏贵 | 若到 2026 年底 ARR 增至 $75M+,当前价格隐含倍数会压缩至 23x |
论点综合了公开来源中的证实与反向视角。未审阅公司内部材料。五行对应五个关键投资维度;每个维度都有独立证据支撑。
[CV011, CV012, CV013, CV014]8.2 估值分析——多重视角
收入倍数分析:在 $30M ARR 和 $1.7B 估值下,ARR 倍数为 57x,把 LMArena 放在后期 AI 基础设施融资轮次的最高十分位。AI 评估基础设施的可比公司很少:Weights and Biases 于 2025 年 3 月被 CoreWeave 以 $1.4B 收购(训练 / 评估平台,工具链更宽);Scale AI 的估值已达数十亿美元区间,收入规模也大一个数量级。早期纯 AI 基准测试公司没有直接公开可比。最相关的类比是 AI「卖铲人」基础设施:LMArena 作为中立数据层的定位,类似 Bloomberg 或 S&P Global 服务金融市场——一个具备高切换成本的可信数据特许权。按 The Business Research Company,AI 模型评估市场 2025 年为 $1.86B,预计 2026 年达到 $2.36B,2030 年以 27.3% CAGR 增至 $6.24B。如果 LMArena 到 2028 年拿下 $3B 市场的 5-10%,收入为 $150-300M,按 15-20x 倍数,对应 $2.25-6B 估值——这是公开可比退出的 base-to-bull 区间。增长可持续性:四个月达到 $30M,隐含每月新增 $7.5M 的增长速度。即便未来 12 个月只维持该速度的 50%,也会新增 $45M ARR,到 2026 年底达到 $75M ARR。若倍数压缩到 25x(符合成熟企业 SaaS 特许经营),估值为 $1.875B——大致持平当前。若按 20x 倍数(高增长公开 SaaS 中位数),$75M ARR 对应 $1.5B。因此,当前 $1.7B 估值取决于:(a)维持快速 ARR 增长,(b)守住平台溢价倍数,或(c)两者兼具。下方情景分析更明确地展示结果分布。 [CV006, CV007, CV008, CV009, CV010]
| 情景 | 2026 年底 ARR | 入场收入倍数 | 隐含估值(2028 年退出) | 关键风险 | 概率信号 |
|---|---|---|---|---|---|
| 牛 | $120M+ | 基于 $120M 为 14x | $3–6B(按平台 25-50x ARR) | 可信度保持;前三大实验室续约;方法论审计通过 | 需要持续每月新增 $7.5M ARR,且零可信度事件 |
| 基准 | $60–75M | 基于 $70M 为 23–28x | $1.4–2.1B(按 20-30x ARR) | 轻微信任摩擦;吸收 1–2 轮方法论批评周期 | 需要当前 ARR 增速的 50%;考虑到 $250M 资本和 41 人团队,具备可行性 |
| 熊 | $15–25M(流失) | 基于 $20M 为 68–113x — 不可持续 | $200–500M(平台陷入困境) | 可信度坍塌;顶级实验室客户退出;竞争性评测平台夺走市场 | 由一次重大公开调查或监管执法触发 |
ARR 预测从 2025 年 12 月 $30M 运行率外推,不反映任何内部预测或指引。估值区间使用公开可比 AI 基础设施倍数(The Business Research Company;Latka;二级市场报道)。概率信号列描述把某个情景成立所需的证据,并非正式概率估计。
[CV008, CV009, CV016]| 可比公司 | 类型 | 收入 / ARR | 估值 | 倍数 | 与 LMArena 的相关性 | 局限 |
|---|---|---|---|---|---|---|
| Weights & Biases(CoreWeave 收购,2025 年 3 月) | 私营 → 被收购 | ~$150M ARR(估计) | $1.4B 收购价 | ~9x ARR | AI 开发者平台,具备模型评测和实验跟踪;产品上最直接的参照 | 工具覆盖比 LMArena 更广;收购价不是独立估值;收入未披露 |
| Scale AI | 私营,后期 | $500M+ 收入(2025 年报道) | 报道区间 $14–29B | 28-58x 收入 | 为企业和政府提供数据标注、RLHF、模型评测;相邻市场 | Scale AI 做直接标注劳务;利润率结构不同;收入基础更宽 |
| Hugging Face | 私营 | ~$70M ARR(2024 年估计) | $4.5B(2023 年融资) | ~64x ARR | AI 模型枢纽,带社区评测功能;作为「中立」基础设施,品牌相关 | 不是以基准评测为核心的业务;社区托管才是主产品;估值时间更早 |
| Bloomberg LP | 私营 | ~$6.5B 收入 | 估值 $80–100B | ~13-15x 收入 | 拥有可信中立评分的数据特许经营(如 Bloomberg Intelligence);长期类比 | 规模、资产类别和市场结构差异很大;30+ 年特许经营 |
| S&P Global(评级分部) | 上市公司(SPGI) | 评级分部 ~$4.5B | 市值 $130B+ | ~29x 分部收入 | 受监管认可的中立裁判,具备结构性护城河;数据特许经营的长期类比 | S&P 拥有监管授权和市场锁定;LMArena 处在未监管的基准评测市场 |
| Arize AI(AI 监控) | 私营 | 未披露 | $148M 融资(2024 年) | N/A | 面向生产环境的 AI 模型监控平台;相邻评测市场 | 产品焦点不同(监控 vs 基准评测);规模较小 |
| Aporia Technologies | 私营 | 未披露 | ~$50M 融资(2024 年) | N/A | 负责任 AI 监控与偏差检测;相邻治理市场 | 聚焦合规;没有社区评测组件 |
所有可比财务数据均来自公开报道(Latka、TechCrunch、分析师来源)或可获得的 SEC 文件。私营公司收入为估算或第三方报道近似值。倍数按报道 / 估算数据计算,应视为方向性而非精确值。
[CV007, CV008]展示在不同 ARR 增长情景($45M、$75M、$120M)和退出倍数情景(15x、25x、40x)下,2028 年隐含退出估值如何变化,覆盖从压力情景到牛市情景的区间。
ARR 预测由 2025 年 12 月 $30M 基线外推;退出倍数基于公开 AI 基础设施 SaaS 可比公司。数值单位为 USD 百万美元。不是财务模型输出,仅作示意。
[CV009, CV010, CV016]低 / 基准 / 高三档 2028 年退出估值区间及假设,展示以 $1.7B 入场与示意摊薄入场价对应的回报倍数。
所有数字均为 USD 百万美元(2028 年退出价值)。熊市情景假设可信度崩塌、ARR 流失至 <$25M;基准情景假设持续增长且倍数适度压缩;牛市情景假设突破式增长并维持平台溢价倍数。
[CV016]8.3 投资论点与反论点
投资论点建立在三项耐久的结构性优势上:(1)社区网络效应——5M 月度用户生成 60M 次对话,形成数据护城河,即便投入大量资本,也极难从零复制;(2)关键判断角色的在位优势——排行榜排名会影响 AI 实验室营销、开发者采用和企业采购,为 LMArena 的付费客户创造切换成本;(3)市场时点——随着竞争性前沿模型增多、企业买家面对真实选择复杂度,AI 评估中的真实世界人类偏好数据会越来越有价值,而不是更少。反论点集中在四个担忧:(1)评估收入与被评估主体之间的结构性利益冲突(Leaderboard Illusion 论文已有记录)可能触发信任坍塌,同时摧毁公开排行榜和企业收入;(2)Goodhart 定律动态——当 Arena 成为事实标准,实验室会专门为它优化,削弱信号的真实世界相关性;(3)收入倍数极端,定价中包含既要持续执行、又要保持可信度的增长轨迹;(4)平台依赖风险存在:付费客户本身控制着平台运转所需的 API 访问,形成结构性杠杆失衡。总体看,在合适进入价格下,论点强于反论点;但当前 57x ARR 倍数在冲突未解决时留下的安全边际不足。 [CV011, CV012, CV013, CV014]
| 触发项 | 阈值 / 事件 | 对投资逻辑的传导 | 行动含义 |
|---|---|---|---|
| 方法论独立性破裂 | 同行评审论文或监管询问记录显示,商业收入影响了 LMArena 公开排行榜排名 | 可信度根基坍塌;企业客户退出;估值跌至困境水平 | 立即暂停 / 退出;若无补救,投资逻辑破裂 |
| 顶级实验室客户流失 | OpenAI、Google 或 xAI 任一家公开终止评测合同或移除模型 | 收入断崖($30M ARR 中多数面临风险);社区覆盖出现缺口 | 投资逻辑预警;升级前先评估集中度和替代管线 |
| ARR 增长停滞 | 截至 2025 年 12 月的 $30M ARR 到 2026 年 Q3 未能达到 $50M+ | 意味着产品市场匹配比估值假设更窄;57x 静态倍数无法自洽 | 重新评估投资逻辑;入场价应反映增长停滞倍数(20-25x) |
| 开发者引用率坍塌 | AI 实验室公告中的 Arena 排行榜引用同比下降 >30% | 失去事实标准地位会削弱定价权和获客能力 | 按季度监控;若趋势持续 2 个季度,投资逻辑实质转弱 |
| 监管执法行动 | 针对 Arena Intelligence, Inc. 的任何 EU DPA、FTC 或州 AG 正式调查 | 法律成本、整改负担和声誉损害;可能约束运营 | 立即暂停;重新评级前评估严重性和司法辖区暴露 |
| 关键人物离职 | Anastasios Angelopoulos 或 Wei-Lin Chiang 在 Series A 交割后 18 个月内离职 | 学术可信度锚点和客户关系流失;关键增长拐点面临接班风险 | 投资逻辑破裂;暂停,等待接班和治理清晰 |
触发阈值是作者基于同类 AI 基础设施投资尽调实践设定的监控启发式标准。它们不反映公司指引。
[CV015, CV017, CV018]从市场规模和社区验证出发,经过风险评估和估值证据,最终落到“跟踪,满足条件后买入”的建议。
[CV011, CV012, CV015]8.4 建议、情景与最终尽调要求
总体建议为「track」,并在满足以下条件时具备转向「buy」路径:(a)收到回应 arXiv 2504.20879 发现的独立方法论审计;(b)以低于未来十二个月 ARR 30x 的估值进入(意味着若按当前 $1.7B 价格承诺投资,ARR 至少需达到 $57M);(c)商业合同条款确认,付费实验室客户相对于公开排行榜方法论没有获得偏好评估定价。鉴于利益冲突和倍数压缩风险,风险评级为「high」。对该建议的置信度为「medium」——证据足以形成观点,但关键风险因素具有结构性且未解决。估值立场是「stretched」。公司的轨迹、社区参与度和先发优势真实存在;但为这些优势付出的价格并不便宜。牛市情景要求到 2026 年底 ARR 达到 $75M+ 且可信度不破;熊市情景是可信度坍塌,收入和估值一起压缩至 $200-400M(类似其他行业的基准声誉危机)。持有期较长且有强治理影响力的投资人,如果以 30x forward ARR 或更低价格进入,在 base case 下有可信的风险 / 回报情景。 [CV015, CV016, CV017, CV018, CV019]
| 主题 | 缺失证据 | 为什么重要 | 负责人 / 尽调路径 |
|---|---|---|---|
| 独立方法论审计 | 对 Leaderboard Illusion 论文发现和 LMArena 公开回应做外部统计复现 | 核心可信度资产未经审计;自我证明不足以支撑投资承诺 | 聘请独立学者或咨询机构;把交付结果作为投资条件 |
| 客户集中度与 NRR | 按客户层级拆分 ARR、前三大客户占比、净收入留存率和 cohort 流失数据 | 汇总 $30M ARR 掩盖了极高估值下的集中度风险 | 向数据室索取;若无法提供,考虑延长 90 天尽调 |
| 股权结构与清算优先权 | 完整资本结构表,包括期权池、优先权条款、反稀释条款和 pro-rata 权利 | 57x ARR 意味着持有期较长;清算顺位会显著影响回报情景 | 标准 Series A 尽调;通常可在数据室获得 |
| EU AI Act 合规状态 | EU AI Act(2025 年 8 月)要求的 GPAI 模型文档、数据来源记录和版权政策 | 不合规会带来 EU 市场风险,并引发受监管行业企业客户顾虑 | 索取法律意见;承诺投资前请 EU 律师审阅 |
| 基础设施与 SLA 协议 | 云服务商协议(GCP 合同)、AI 提供商 API 访问协议和企业评测 SLA | 平台可用性和收入连续性依赖未披露的第三方协议 | 在数据室索取;特别关注提供商终止通知期 |
| 员工数与招聘计划 | 组织架构图,列出现有缺口、支撑每名员工 $10M+ ARR 的招聘计划,以及创始人关键人协议 | 41 人支撑 $30M ARR 已很精简;扩至 $100M ARR 需要显著补组织 | 在启动会议中向管理层索取 |
尽调问题按接近当前 $1.7B 估值入场的投资人优先级排序。第 1 和第 2 项被视为买入决策的阻断项。第 3–6 项对投资结构设计和交割后监控有实质影响。
[CV019]围绕七个评估维度给出 IC 可用评分,反映截至运行日期支撑各维度的证据质量和完整度。
分数是作者仅基于可得公开证据按 1–10 分作出的判断,不是算法输出。治理分数因利益冲突未解决、缺少独立审计而受罚;估值分数因 57x ARR 倍数而受罚。
[CV015, CV012, CV013]8.5 附录
免责声明
本报告基于截至 2026-06-25 的公开且可访问来源,不构成投资、法律或会计建议;在作出投资决策前,应补充管理层尽调、客户访谈和一手财务文件。
证据索引
| 编号 | 陈述 | 可信度 | 来源 |
|---|---|---|---|
| CO001 | Chatbot Arena was launched in May 2023 as a public demo by the LMSYS research group at UC Berkeley's Sky Computing Lab. | 高 | SO010, SO011 |
| CO002 | The original Chatbot Arena platform was developed under UC Berkeley's LMSYS (Large Model Systems) research group, operating within the Sky Computing Lab. | 中 | SO010, SO019 |
| CO003 | Anastasios Angelopoulos and Wei-Lin Chiang are co-founders of LMArena (Arena Intelligence Inc.). | 高 | SO001, SO003 |
| CO004 | Ion Stoica, a UC Berkeley CS professor and co-founder of Databricks and Anyscale, is a co-founder of Arena Intelligence Inc. | 高 | SO019, SO004 |
| CO005 | Arena Intelligence Inc. was formally incorporated on April 18, 2025. | 高 | SO020, SO016 |
| CO006 | Anastasios Angelopoulos serves as CEO of Arena Intelligence Inc. | 高 | SO003, SO001 |
| CO007 | LMArena's platform enables users to submit prompts to two anonymous AI models simultaneously, vote on the preferred response, and then see which models they compared — feeding a public leaderboard. | 高 | SO001, SO010, SO011 |
| CO008 | LMArena's headquarters is located in the San Francisco Bay Area; the specific office address is not publicly disclosed. | 低 | SO001, SO019 |
| CO009 | LMArena now operates under the brand name 'Arena'; it was previously branded 'LMArena' and before that 'Chatbot Arena.' | 高 | SO001, SO018 |
| CO010 | LMArena raised a $100 million seed round in May 2025, co-led by Andreessen Horowitz and UC Investments, at a $600 million post-money valuation. | 高 | SO005, SO004, SO006 |
| CO011 | LMArena raised $150 million in Series A financing in January 2026, co-led by Felicis and UC Investments, at a post-money valuation of $1.7 billion. | 高 | SO003, SO004, SO002 |
| CO012 | LMArena's Series A investors include Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed Venture Partners, and Laude Ventures alongside lead investors Felicis and UC Investments. | 高 | SO002, SO003 |
| CO013 | LMArena's total capital raised as of January 2026 is approximately $250 million across seed and Series A rounds. | 高 | SO004, SO003 |
| CO014 | LMArena's $1.7 billion Series A post-money valuation represents nearly triple its $600 million seed-round valuation achieved approximately nine months earlier. | 高 | SO003, SO004 |
| CO015 | LMArena's annualized consumption run rate surpassed $30 million in December 2025, approximately four months after commercial product launch. | 高 | SO003, SO004 |
| CO016 | LMArena launched its commercial AI Evaluations product in September 2025. | 高 | SO007, SO004 |
| CO017 | As of January 2026, LMArena reports more than 5 million monthly users across 150 countries. | 中 | SO003, SO004 |
| CO018 | LMArena processes more than 60 million conversations per month across its evaluation platform. | 中 | SO003, SO004 |
| CO019 | LMArena's community accumulated over 50 million votes across text, vision, web development, search, video, and image modalities by December 2025. | 中 | SO002, SO003 |
| CO020 | LMArena has evaluated more than 400 distinct AI models including both open-weight and proprietary systems since founding. | 中 | SO002, SO003 |
| CO021 | LMArena's pre-commercial research was supported by grants and donations from Google's Kaggle platform, Andreessen Horowitz, and Together AI, primarily in compute resources and cash. | 中 | SO005, SO015 |
| CO022 | LMArena's leaderboard uses a Bradley-Terry model / Elo-based rating system with pairwise crowdsourced comparisons to rank AI models. | 高 | SO011, SO012, SO010 |
| CO023 | LMArena released 145,000 open-source battle data points from expert and occupational evaluation categories in late 2025. | 中 | SO002 |
| CO024 | LMArena's AI Evaluations commercial product includes SLAs with committed delivery timelines, representative feedback samples, and community-grounded performance analytics for model labs and enterprises. | 中 | SO007, SO003 |
| CO025 | LMArena's named commercial customers include OpenAI, Google, and xAI, which use its evaluation services to improve their production models. | 高 | SO003, SO004 |
| CO026 | LMArena has expanded beyond text evaluation to include Search Arena, WebDev Arena, Vision Arena, text-to-image, video generation, and Agent Arena as of June 2026. | 高 | SO024, SO025, SO026, SO009 |
| CO027 | In April 2025, researchers from Cohere, Stanford, MIT, and Ai2 published 'The Leaderboard Illusion,' alleging LMArena systematically allowed certain AI companies to privately test multiple model variants and selectively disclose only top-performing scores. | 高 | SO014, SO016 |
| CO028 | The Leaderboard Illusion paper identified that Meta privately tested at least 27 Llama 4 model variants on Chatbot Arena between January and March 2025, ahead of the public release. | 高 | SO014, SO016 |
| CO029 | The Leaderboard Illusion paper estimated that Google and OpenAI each received approximately 19–20% of all Chatbot Arena battle data, while 83 combined open-weight models received only approximately 29.7% of total data. | 中 | SO014 |
| CO030 | LMArena co-founder Ion Stoica publicly characterized The Leaderboard Illusion paper as containing 'inaccuracies' and 'questionable analysis,' and LMArena invited all model providers to submit more models for testing. | 中 | SO016 |
| CO031 | Critics including researchers at the Allen Institute for AI and King's College London argued that LMArena's user base skews toward technical programmers and AI enthusiasts, making its benchmark unrepresentative of general end-user preferences. | 中 | SO015, SO020 |
| CO032 | LMArena published updated transparency and leaderboard policies as of April 30, 2026, committing to open-sourcing evaluation pipelines and releasing portions of data to support auditing. | 中 | SO008 |
| CO033 | Ion Stoica previously co-founded Databricks (valued at approximately $43 billion) and Anyscale (the Ray computing framework), providing the founding team with deep commercialization experience. | 高 | SO019, SO016 |
| CO034 | LMArena reached a $1.7 billion valuation approximately three years after founding and within seven months of launching its first commercial product. | 中 | SO004, SO003 |
| CO035 | LMArena's commercial business model charges AI labs and enterprises for evaluation services including community-grounded model performance analysis across domains such as software engineering, law, and medicine. | 中 | SO003, SO007 |
| CO036 | LMArena launched Agent Arena in June 2026 with a causal inference methodology using treatment effect estimation to evaluate multi-component AI agents. | 中 | SO024 |
| CO037 | In early April 2025, Meta submitted a specially arena-optimized, unreleased variant of Llama 4 Maverick to Chatbot Arena that ranked 2nd overall; when the standard public release was subsequently scored, it ranked 32nd. | 高 | SO027, SO028 |
| CO038 | LMArena's deduplication system filters approximately 10% of submitted votes and its identity-leak detection pipeline removes fewer than 4% of votes, per July 2025 methodology updates. | 中 | SO009 |
| CO039 | LMArena has not publicly disclosed its total headcount; the company's Series A press release stated funds would be used to expand the technical team. | 中 | SO003, SO002 |
| CO040 | No debt facilities, convertible notes, secondary transactions, or SAFEs have been publicly disclosed for LMArena as of June 2026; financing has been entirely via equity rounds. | 低 | SO003, SO004 |
| CO041 | The Leaderboard Illusion paper estimated that even limited additional Chatbot Arena battle data can yield relative performance gains of up to 112% on Arena-specific benchmarks, demonstrating the value of asymmetric data access. | 中 | SO014 |
| CO042 | LMArena's Chatbot Arena paper (arXiv:2403.04132) reported crowdsourced human votes achieve over 80% agreement with expert rater judgments, establishing methodological validity for the platform's ranking approach. | 高 | SO011, SO012 |
| CO043 | Arena Intelligence Inc. was spun out from UC Berkeley's LMSYS research group with UC Investments (University of California) serving as both lead investor and institutional backer of the spinout; the formal IP licensing terms governing transfer of the Chatbot Arena methodology from the university to the commercial entity are not publicly disclosed. | 中 | SO005, SO020 |
| CM001 | The broad AI model evaluation platform market is sized at $1.86B in 2025 and $2.36B in 2026, implying 27.3% growth into 2026. | 高 | SM001, SM002, SM003 |
| CM002 | The same broad market lens projects the AI model evaluation platform category to reach about $6.24B by 2030 at a 27.3% CAGR. | 中 | SM001, SM003 |
| CM003 | A narrower model evaluation and benchmarking tools lens implies roughly a $0.85B market in 2026 at around 7.3% CAGR, materially below the broad platform estimate. | 中 | SM004 |
| CM004 | Gartner forecasts worldwide AI spending to rise 47% to about $2.59T in 2026, creating a large upstream budget pool for evaluation software. | 高 | SM005, SM013 |
| CM005 | Presenc AI reported that 78% of Global 2000 companies had at least one AI workload in production in Q1 2026, consistent with fast-rising demand for recurring evaluation and governance workflows. | 高 | SM005, SM006 |
| CM006 | LMArena competes in the evaluation software layer that includes public benchmarking, release testing, domain scoring, and workflow-level quality measurement rather than core model training or hosting. | 中 | SM014, SM016, SM017 |
| CM007 | Chatbot Arena introduced LMArena's core pairwise human-preference comparison model, making public leaderboard trust a foundational but not sufficient part of the company's market. | 高 | SM014, SM015 |
| CM008 | Agent Arena broadens LMArena from single-turn chatbot ranking into agent and workflow evaluation, increasing the relevance of multi-step enterprise use cases. | 中 | SM016, SM017 |
| CM009 | The sharp difference between broad and narrow market estimates is best explained by scope: broad reports appear to include wider enterprise evaluation-platform workflows, while narrow reports emphasize benchmarking tools only. | 中 | SM001, SM003, SM004 |
| CM010 | Scale AI launched SEAL Showdown in 2026 with respondents across 100+ countries, 70+ languages, and 200+ professional domains, positioning it as a directly competitive benchmark product. | 中 | SM008, SM009, SM025 |
| CM011 | Scale says GPT-5 tops all SEAL categories while Gemini 2.5 Pro leads most LMArena categories, showing that benchmark outcomes are sensitive to task mix and rater composition. | 中 | SM008, SM009 |
| CM012 | Future AGI's competitor roundup places Galileo, Arize, Patronus, and MLflow in the same evaluation-tool conversation as LMArena, implying a fragmented specialist field. | 中 | SM004, SM007 |
| CM013 | Precedence Research lists major platform and infrastructure vendors such as AWS, Google, Microsoft, IBM, Databricks, Hugging Face, and Arize AI among category participants, implying competition from bundled as well as standalone products. | 中 | SM004, SM007 |
| CM014 | CoreWeave's roughly $1.4B acquisition of Weights & Biases in March 2025 shows strategic M&A appetite around evaluation-adjacent workflows such as experimentation, observability, and model quality. | 中 | SM004, SM021 |
| CM015 | TechCrunch reported that LMArena reached a $1.7B valuation, about $30M ARR, and commercial customers including OpenAI, Google, and xAI four months after launching its product. | 中 | SM012, SM024 |
| CM016 | LMArena says its commercial product targets law, medicine, and engineering workflows, signaling that high-stakes domain evaluation is central to its monetization strategy. | 中 | SM012, SM016 |
| CM017 | Likely budget owners for evaluation software are AI platform leads, model quality owners, or domain product teams rather than generalized IT procurement alone. | 中 | SM007, SM016, SM017 |
| CM018 | Frontier labs typically buy evaluation for model release, trust, and competitive benchmarking, while enterprises buy it for domain reliability, governance, and workflow quality. | 中 | SM016, SM017, SM012 |
| CM019 | The buyer universe is concentrated at the frontier-lab end but much broader across regulated and expertise-heavy enterprise verticals, creating two different procurement motions. | 中 | SM012, SM016, SM007 |
| CM020 | Law, medicine, and engineering are promising beachheads because domain failures there are expensive enough to justify paid evaluation rather than relying on public leaderboard performance alone. | 中 | SM012, SM016 |
| CM021 | A practical LMArena-relevant 2026 SAM is a subset of the broad $2.36B TAM, likely concentrated in frontier labs plus evaluation-heavy enterprise verticals rather than the full surrounding AI-tooling economy. | 低 | SM001, SM004, SM012, SM016 |
| CM022 | A plausible 2026 SAM range for standalone trusted evaluation software relevant to LMArena is about $0.3-0.8B after excluding most bundled hyperscaler, generic MLOps, and non-paid benchmarking activity. | 低 | SM001, SM004, SM016 |
| CM023 | A plausible near-term SOM range for LMArena is about $30-150M because the company has reported roughly $30M ARR and could deepen spend within existing frontier-lab and high-stakes enterprise customers. | 中 | SM012, SM013, SM016 |
| CM024 | The market should be treated as contested because a narrow benchmarking-tools lens can be roughly one-third or less of the broad platform estimate depending on what is included. | 中 | SM001, SM003, SM004 |
| CM025 | Gartner's 2026 forecast includes about $453B of AI software spend and $32.6B of AI model spend, indicating that evaluation budgets need only capture a small fraction of upstream AI activity to support category growth. | 中 | SM005, SM013 |
| CM026 | As AI workloads move into production at large enterprises, evaluation shifts from one-off benchmarking toward recurring regression testing, governance, and release gating. | 中 | SM005, SM006, SM016 |
| CM027 | The rise of agentic AI increases evaluation demand because buyers need to measure multi-step task completion and causal workflow reliability rather than only single-response quality. | 中 | SM016, SM017 |
| CM028 | Benchmark rivalry among frontier labs is itself a demand driver because model providers want external proof points to market releases and defend product claims. | 中 | SM008, SM014, SM015 |
| CM029 | New benchmark launches such as SEAL Showdown confirm that benchmark design is now a contested product category rather than a settled research utility. | 中 | SM008, SM009, SM018 |
| CM030 | Commercializing evaluation in law, medicine, and engineering makes the category more monetizable because quality signals are tied to high-cost business workflows instead of casual consumer usage. | 中 | SM012, SM016 |
| CM031 | Even without exact public compliance budgets, regulated and high-stakes deployments are likely to sustain third-party evaluation demand because internal trust and audit requirements are higher than for generic chat use cases. | 低 | SM016, SM017, SM005 |
| CM032 | The Leaderboard Illusion paper argues that benchmark contamination, hidden sampling choices, and over-interpretation of small score gaps can distort arena-style rankings. | 高 | SM010, SM018, SM019 |
| CM033 | TechCrunch reported in 2024 that Chatbot Arena's user base and interaction style may bias results, weakening its usefulness as a universal benchmark. | 高 | SM010, SM011 |
| CM034 | TechCrunch reported in 2025 that LM Arena faced accusations of helping top labs game its benchmark, creating a direct commercial trust risk for public-score-based products. | 中 | SM018, SM020 |
| CM035 | The Verge's coverage of benchmark gaming around Meta's Llama 4 Maverick suggests that optimization against public benchmarks is ecosystem-wide rather than unique to LMArena. | 中 | SM018, SM020 |
| CM036 | Critics of crowdsourced AI benchmarks argue that rater representativeness and benchmark ethics are material issues, implying that some enterprise buyers may prefer curated private evaluations over open-arena votes. | 中 | SM019, SM011, SM010 |
| CM037 | Standalone vendors like LMArena likely face pricing pressure wherever evaluation is bundled into broader cloud, experimentation, or observability stacks. | 中 | SM004, SM007, SM016 |
| CM038 | Public pricing and average contract values for LMArena and most direct competitors are undisclosed, preventing reliable bottom-up revenue-to-market-share validation from open sources. | 中 | SM007, SM012, SM016 |
| CM039 | No verified public source in this run discloses what share of frontier-lab or enterprise AI budgets is actually allocated to evaluation tooling. | 中 | SM012, SM016 |
| CM040 | Competitor revenue disclosure is sparse for Galileo, Patronus, Arize, Future AGI, and other specialist vendors, making independent startup-share mapping incomplete. | 中 | SM004, SM007 |
| CM041 | LMArena's reported roughly $30M ARR suggests it may already represent a meaningful share of a narrow benchmarking-tools niche while remaining only a small share of the broad AI evaluation platform market. | 中 | SM001, SM004, SM012 |
| CM042 | Because broad and narrow market definitions differ so much, investors should validate how much of LMArena's revenue comes from public benchmarks, agent evaluation, and enterprise domain workflows before relying on TAM-based share math. | 中 | SM001, SM004, SM016, SM017 |
| CM043 | A three-tier lens of $2.36B TAM, about $0.3-0.8B SAM, and about $0.03-0.15B SOM is directionally consistent with the verified evidence but remains an analytical construct rather than a published market model. | 低 | SM001, SM004, SM012, SM016 |
| CM044 | LMArena's buyer-segment matrix is driven more by use-case needs such as benchmark credibility, workflow depth, and domain rigor than by simple company-size segmentation. | 中 | SM007, SM016, SM017, SM008 |
| CM045 | Evaluation spend narrows from open experimentation to recurring governance spend as AI systems move into production, which is why adoption-funnel economics depend on deployment depth rather than benchmark traffic alone. | 中 | SM006, SM016, SM017 |
| CM046 | A low-mid-high range of roughly $0.85B, $2.36B, and $9.57B shows how the category can look small, medium, or strategic depending on whether the lens is narrow 2026 benchmarking, broad 2026 platforms, or longer-horizon strategic infrastructure. | 低 | SM001, SM004 |
| CP001 | LMArena's platform serves more than 5 million monthly users across 150 countries as of January 2026. | 高 | SP007, SP008 |
| CP002 | LMArena generates more than 60 million model-comparison conversations per month as of January 2026. | 高 | SP007, SP008 |
| CP003 | LMArena users span 150 countries according to the company's January 2026 Series A announcement. | 中 | SP008 |
| CP004 | LMArena has evaluated more than 400 public models and conducted more than 300 pre-release tests across multiple modalities as of the two-year anniversary in 2025. | 中 | SP007 |
| CP005 | LMArena has released more than 1.5 million community-contributed prompts as open data for research use, supporting its open-access mission. | 中 | SP007 |
| CP006 | LMArena uses a pairwise comparison approach and Elo-style statistical ranking across blind model battles, as described in the 2024 Chatbot Arena paper by Chiang, Zheng et al. | 高 | SP001, SP023 |
| CP007 | LMArena publicly launched its commercial AI Evaluations enterprise product in September 2025, marking its entry into the paid evaluation-as-a-service market. | 中 | SP008 |
| CP008 | LMArena has established partnerships with OpenAI, Google, Anthropic, Meta, and xAI to make their flagship models available for community evaluation on the platform. | 中 | SP007, SP008 |
| CP009 | FastChat, the open-source platform backing LMArena, has powered over 10 million chat requests for 70+ LLMs, according to the GitHub repository. | 中 | SP011 |
| CP010 | The LMSYS Org at UC Berkeley, which originated LMArena, reports 15+ projects, 79K+ GitHub stars, and 1,000+ contributors across its open-source research portfolio. | 中 | SP023 |
| CP011 | Scale AI operates the Scale GenAI Platform, offering enterprise evaluation, data labeling, and agent deployment services to clients including Meta, Mayo Clinic, and defense agencies. | 中 | SP016, SP017 |
| CP012 | Scale AI's enterprise evaluation product is private and SLA-backed, with no public leaderboard, targeting a different buyer than LMArena's public benchmark community. | 中 | SP016, SP017 |
| CP013 | Scale AI's enterprise engagement pricing reportedly starts at approximately $93,000 per year, with complex projects reaching $400,000 or more, per analyst aggregations. | 低 | SP016 |
| CP014 | The HuggingFace Open LLM Leaderboard uses automated benchmark pipelines to evaluate open-source models on tasks including MMLU, GPQA, ARC, and SWE-Bench, without human preference voting. | 中 | SP014, SP012 |
| CP015 | The HuggingFace Open LLM Leaderboard is freely accessible to model submitters and public readers, with no commercial evaluation service attached. | 中 | SP014 |
| CP016 | Stanford HELM evaluates language models holistically across multiple automated dimensions including accuracy, robustness, calibration, bias, efficiency, and toxicity. | 高 | SP018, SP019 |
| CP017 | Stanford HELM is an academic, non-commercial tool with no enterprise evaluation service; it is freely accessible to any researcher or practitioner. | 中 | SP018 |
| CP018 | EleutherAI's lm-evaluation-harness provides over 60 standardized academic benchmarks and powers the HuggingFace Open LLM Leaderboard. | 中 | SP012 |
| CP019 | EleutherAI's lm-evaluation-harness has no commercial evaluation product and is freely available as open-source software, with the harness running on any infrastructure the user controls. | 中 | SP012 |
| CP020 | OpenAI Evals is an open-source framework for evaluating LLMs that can now be run directly in the OpenAI Dashboard, with a community benchmark registry and enterprise custom eval support. | 中 | SP013 |
| CP021 | BenchLM tracked 261 models across 249 benchmarks as of June 2026, distinguishing verified from provisional rankings, with pricing and speed data included. | 中 | SP020 |
| CP022 | ArtificialAnalysis provides independent, provider-agnostic AI model performance analytics including intelligence, output speed, latency, and cost benchmarking, without a commercial evaluation service. | 中 | SP021 |
| CP023 | None of the major free leaderboard competitors — HuggingFace Open LLM Leaderboard, Stanford HELM, EleutherAI harness, BenchLM, or ArtificialAnalysis — offer an enterprise evaluation-as-a-service product with SLA commitments. | 中 | SP012, SP014, SP018, SP020, SP021 |
| CP024 | LMArena is the only platform that combines large-scale human-preference ranking (5 M+ monthly users) with a commercial enterprise evaluation service in a single brand and infrastructure, as of June 2026. | 中 | SP001, SP007, SP008, SP016 |
| CP025 | LMArena benefits from network effects: a larger and more active user community generates more preference votes, which strengthens the statistical stability of Elo rankings, making the platform more attractive to labs seeking reliable signal. | 中 | SP001, SP007 |
| CP026 | The Elo ranking algorithm used by LMArena is open-source and publicly documented, meaning the mathematical mechanism of ranking can be reproduced by any sufficiently resourced team. | 中 | SP001, SP003 |
| CP027 | OpenAI Evals is built primarily for evaluating OpenAI models, which limits its independence as a neutral tool for cross-provider model comparison. | 中 | SP013 |
| CP028 | Singh et al. (2025) found that Meta tested 27 private LLM variants on Chatbot Arena in the lead-up to the Llama 4 public release, selecting only the best-performing score for public disclosure. | 高 | SP003, SP005 |
| CP029 | Singh et al. (2025) estimated that Google and OpenAI received 19.2% and 20.4% of all Chatbot Arena battle data respectively, while 83 combined open-weight models received only 29.7%. | 高 | SP003, SP005 |
| CP030 | After the gaming incident, Meta's vanilla Maverick model (unoptimized) was ranked approximately 32nd on the LMArena leaderboard, below GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro. | 中 | SP006 |
| CP031 | TechCrunch reported that Chatbot Arena's user base is skewed toward tech and AI professionals; the top questions in the LMSYS-Chat-1M dataset pertain to programming, software bugs, and app design rather than general consumer use. | 中 | SP004 |
| CP032 | Researchers including Yuchen Lin (Allen Institute for AI) and Mike Cook (King's College London) raised concerns that LMArena's evaluation lacks construct validity, meaning it is unclear whether user votes reliably measure real model quality. | 中 | SP004 |
| CP033 | LMArena earns revenue by selling AI evaluation services to the same AI labs (OpenAI, Google, xAI) it ranks on its public leaderboard, creating a structural conflict of interest. | 高 | SP005, SP008 |
| CP034 | LMArena stated in response to the Singh et al. study that it has published information on pre-release testing since March 2024 and that it does not favor any model provider over another. | 中 | SP005 |
| CP035 | LMArena committed to create a new sampling algorithm to address concerns about unequal model battle frequency, acknowledging operational merit in some of the Leaderboard Illusion critiques. | 中 | SP005 |
| CP036 | LMArena has not published a public price list for its AI Evaluations enterprise product; enterprise buyers must contact the sales team for pricing. | 中 | SP008, SP009 |
| CP037 | Scale AI enterprise evaluation engagements reportedly start at approximately $93,000 per year according to third-party analyst aggregations, with complex projects reaching $400,000 or more. | 低 | SP016 |
| CP038 | HuggingFace Open LLM Leaderboard, Stanford HELM, EleutherAI lm-evaluation-harness, BenchLM, and ArtificialAnalysis are all available free of charge with no enterprise evaluation contract required. | 中 | SP012, SP014, SP018, SP020, SP021 |
| CP039 | LMArena's cumulative preference vote dataset (6 M+ votes across 60 M conversations) is the largest publicly known human-preference benchmark dataset for LLMs, with no comparable free alternative. | 中 | SP007, SP008, SP001 |
| CP040 | AI labs cite their Arena Elo scores in press releases and marketing materials, creating a dependency on LMArena's ranking signal for product positioning that raises the switching cost of defecting to an alternative leaderboard. | 中 | SP004, SP005 |
| CP041 | No competitor has replicated LMArena's combination of real-time human-preference data at population scale with a commercial enterprise evaluation product as of June 2026, as inferred from publicly reviewed product surfaces. | 中 | SP007, SP016, SP020, SP021 |
| CI001 | LMArena publicly launched its commercial AI Evaluations enterprise product in September 2025, marking the company's first commercial revenue-generating product. | 高 | SI001, SI002 |
| CI002 | The disclosed $30 million annualized consumption run rate implies roughly $2.5 million of December 2025 monthly revenue at the point LMArena reported the metric. | 高 | SI001, SI002 |
| CI003 | TechCrunch noted that LMArena calls this figure a "consumption rate" — as it "describes its annual recurring revenue (ARR)" — making clear that this is a usage-based annualized run rate, not a contracted ARR with forward booking guarantees. | 高 | SI001, SI002 |
| CI004 | GetLatka estimates approximately 100 enterprise clients for LMArena as of early 2026; this is an analyst aggregation, not a company-confirmed figure. | 低 | SI005 |
| CI005 | Dividing the $30 million annualized consumption run rate by ~100 enterprise clients implies an average contract value of approximately $300,000 per client; this is a derived estimate, not a company-stated figure. | 低 | SI001, SI005 |
| CI006 | OpenAI, Google, and xAI are confirmed as enterprise clients of LMArena's AI Evaluations service, per the January 2026 Series A press release from LMArena. | 高 | SI001, SI004 |
| CI007 | LMArena earns revenue by providing paid AI evaluation services to AI labs and enterprises; the free public leaderboard is not the direct monetized product. | 高 | SI001, SI011 |
| CI008 | LMArena's AI Evaluations product includes comprehensive in-depth evaluations based on community feedback, auditability through representative data samples, and SLA-committed delivery timelines. | 中 | SI011 |
| CI009 | LMArena raised $100 million in a seed round in May 2025 at a post-money valuation of $600 million, led by Andreessen Horowitz and UC Investments, with Lightspeed, Felicis, and Kleiner Perkins also participating. | 高 | SI003, SI004 |
| CI010 | LMArena raised $150 million in a Series A in January 2026 at a post-money valuation of $1.7 billion, led by Felicis and UC Investments with participation from a16z, Kleiner Perkins, Lightspeed, House Fund, LDVP, and Laude Ventures. | 高 | SI001, SI004 |
| CI011 | LMArena has raised $250 million in total across its seed round and Series A, in approximately seven months from May 2025 to January 2026. | 高 | SI002, SI008 |
| CI012 | The LMArena Series A investor group includes UC Investments (managing University of California public funds), which the company's CIO cited as validation of LMArena's role as critical AI evaluation infrastructure. | 中 | SI001 |
| CI013 | An SEC Form D filing (accession 0002113470-26-000001, filed 2026-02-26) documents the formation of Recall Capital-LMArena, a Delaware LLC venture capital feeder fund established to invest in LMArena. | 高 | SI007, SI024 |
| CI014 | LMArena had approximately 41 employees as of the January 2026 Series A announcement, according to TechCrunch's reporting. | 中 | SI002 |
| CI015 | LMArena's primary cost categories are inferred to be AI inference compute for serving 60 million monthly conversations, engineering headcount, and community management; gross margin is not publicly disclosed. | 低 | SI001, SI014 |
| CI016 | LMArena operates a pure software and services model with no physical manufacturing, hardware capital expenditure, or project-finance obligations reported publicly. | 中 | SI001, SI011 |
| CI017 | LMArena's gross margins are not publicly disclosed; analyst estimates suggest a range of 40–70%, depending on whether AI inference costs for partner model hosting are subsidized or borne by LMArena directly. | 低 | SI011, SI006 |
| CI018 | With $250 million raised and a 41-person team as of January 2026, LMArena's estimated cash runway is 24–36 months from the Series A close, though no official burn rate has been confirmed. | 低 | SI002, SI005 |
| CI019 | LMArena has not publicly disclosed monthly burn rate, current cash on hand, or debt obligations; these metrics are fully private. | 中 | SI002 |
| CI020 | LMArena stated it will use the Series A funds to operate its platform, expand its technical team, and strengthen its research capabilities. | 中 | SI001 |
| CI021 | LMArena earns revenue from AI labs (OpenAI, Google, xAI) it evaluates and publicly ranks, creating a structural conflict of interest between its commercial incentive and its neutrality claim. | 高 | SI006, SI010 |
| CI022 | CTOL Digital calculated that LMArena's $1.7 billion valuation is approximately 57 times the $30 million annualized consumption run rate, meaning the valuation embeds an expectation that the commercial conflict can be managed indefinitely. | 中 | SI006 |
| CI023 | LMArena's revenue was concentrated in its first four months of commercial operation; customer concentration risk is high given that three confirmed clients (OpenAI, Google, xAI) are the same entities ranked on the public leaderboard. | 中 | SI001, SI006 |
| CI024 | If enterprise clients reduce engagement with LMArena's evaluation service in response to perceived bias in rankings — as documented in the Singh et al. Leaderboard Illusion paper — the revenue model and the free leaderboard's trust premium would be impaired simultaneously. | 中 | SI009, SI010, SI006 |
| CI025 | No independent audit of LMArena's evaluation methodology, conflict-of-interest management policy, or financial controls has been published as of June 2026. | 中 | SI006, SI009 |
| CI026 | LMArena's GTM motion is a freemium funnel: the free public leaderboard attracts AI labs and enterprise technical teams, which then convert to the paid AI Evaluations service. | 中 | SI001, SI011 |
| CI027 | LMArena has stated that all publicly released models will be evaluated under the same methodology regardless of commercial interest, maintaining the free public leaderboard as a trust anchor for the GTM funnel. | 中 | SI011 |
| CI028 | LMArena has not published information about sales cycle length, CAC, payback period, or enterprise win rates for the AI Evaluations product. | 中 | SI002 |
| CI029 | Both LMArena and TechCrunch used the phrase "consumption rate" rather than a standard SaaS "ARR" metric for the $30 million figure, signaling that the revenue structure may be transaction-based rather than subscription-based. | 高 | SI001, SI002 |
| CI030 | The $30 million annualized consumption run rate is an annualized pace of revenue observed in December 2025, not the full-year 2025 realized revenue figure, which would be substantially lower given product launch in September 2025. | 中 | SI001, SI002 |
| CI031 | If LMArena's billing is usage-based (consumption), enterprise clients may reduce evaluation volume in slow periods without formal churn, making the run rate a less reliable indicator of future revenue than contracted ARR would be. | 中 | SI002, SI006 |
| CI032 | LMArena's $1.7 billion post-money Series A valuation implies a revenue multiple of approximately 57x the $30 million annualized consumption run rate, consistent with early-stage high-growth software but requiring significant revenue growth to justify at later stages. | 中 | SI006, SI002 |
| CI033 | UC Investments, which manages investment assets for the University of California system, led both the seed and Series A rounds, providing institutional and academic credibility to the investment thesis. | 中 | SI003, SI004 |
| CI034 | Andreessen Horowitz (a16z) participated in both the May 2025 seed round and the January 2026 Series A, indicating high-conviction early backing from a top venture firm. | 中 | SI003, SI004 |
| CI035 | LMArena has not published a public price list for its AI Evaluations enterprise product; enterprise buyers must contact the team at evaluations@lmarena.ai. | 中 | SI011 |
| CI036 | The Leaderboard Illusion paper's finding of data-access asymmetry creates a reputational risk that could reduce enterprise clients' willingness to pay LMArena for evaluation services if the methodology is perceived as commercially compromised. | 中 | SI009, SI010 |
| CI037 | Bloomberg reported in April 2025 — one month before the formal seed announcement — that Chatbot Arena was becoming a "real company," providing early public evidence of the commercialization timeline. | 中 | SI003 |
| CI038 | LMArena has released 1.5 million+ community prompts and 145K+ battle data points as open data, which supports its academic credibility narrative but generates no direct revenue. | 中 | SI004, SI014 |
| CE001 | LMArena operates a multi-arena AI evaluation platform at arena.ai covering text, code, search, agent, vision, image generation, video generation, image editing, and document modalities. | 高 | SE001, SE002, SE005 |
| CE002 | The core platform mechanic is a pairwise battle where two anonymous models respond to the same prompt and users vote on the preferred output. | 高 | SE001, SE014 |
| CE003 | All Arena text, code, search, and vision leaderboards use the Bradley-Terry statistical model to infer latent skill coefficients from pairwise win/loss battle outcomes. | 高 | SE009, SE014, SE015, SE018 |
| CE004 | LMArena had 5 million or more monthly users as of January 2026. | 高 | SE003, SE005 |
| CE005 | LMArena has logged more than 250 million real conversations across all arenas since its inception. | 中 | SE003 |
| CE006 | LMArena generates more than 2 million preference votes per month from community users. | 中 | SE003 |
| CE007 | LMArena's community spans users in more than 150 countries. | 中 | SE003 |
| CE008 | Agent Arena was launched on June 4, 2026 as LMArena's newest evaluation arena, ranking orchestrator models for autonomous multi-step task completion. | 高 | SE005, SE006 |
| CE009 | Agent Arena ranks models using causal tracing, treating each component selection as a treatment in a multi-intervention randomized controlled trial measuring five behavioral signals. | 中 | SE006 |
| CE010 | WebDev Arena, launched December 2024, collected over 80,000 community votes on AI-generated web applications before being superseded by Code Arena in 2026. | 中 | SE007 |
| CE011 | Search Arena supports 11 models from three providers (Perplexity, Gemini, and OpenAI) and has collected over 24,000 paired multi-turn user interactions. | 中 | SE008, SE017 |
| CE012 | Code Arena was rebuilt from WebDev Arena with a new isolated agentic coding environment, persistent sessions stored in Cloudflare R2, and live preview rendering via CodeMirror 6. | 高 | SE013, SE005 |
| CE013 | Arena-Rank is an open-source Python package (Apache 2.0) published on GitHub under the lmarena organization and installable from PyPI as 'arena-rank'. | 高 | SE009, SE001, SE020, SE023 |
| CE014 | Arena-Rank implements Bradley-Terry ranking with closed-form confidence-interval calculation and a 30x speedup over the historical FastChat implementation. | 高 | SE009, SE001, SE020 |
| CE015 | FastChat (GitHub: lm-sys/FastChat) was the original open-source platform powering Chatbot Arena but is now primarily in maintenance mode, with ranking code migrated to Arena-Rank. | 中 | SE019, SE009 |
| CE016 | LMArena's style-control extension adds length, markdown header count, bold count, and list count as regression covariates in the BT model to isolate substance from formatting effects; response length is the dominant style factor. | 高 | SE010, SE015 |
| CE017 | The Arena-Hard BenchBuilder pipeline (arXiv:2406.11939) automatically extracts hard prompts from live Arena data using a seven-criterion hardness labeler and generates benchmarks that achieve 98.6% agreement with human preference rankings. | 高 | SE016, SE012 |
| CE018 | LMArena defines prompt hardness using seven criteria including domain knowledge, problem-solving complexity, and real-world applicability; approximately 20% of Arena prompts have a hardness score of 6 or higher. | 中 | SE011 |
| CE019 | The lmarena-ai HuggingFace organization hosts multiple p2l (prompt-to-leaderboard) preference models ranging from 135M to 7B parameters and releases public battle datasets. | 中 | SE022 |
| CE020 | LMArena uses GCP's Sensitive Data Protection API to remove personally identifiable information from conversation data before any public release or data sharing. | 中 | SE004 |
| CE021 | LMArena's leaderboard policy was last updated April 30, 2026 and specifies model eligibility criteria, sampling policies, pre-release testing protocols, and data sharing rules. | 高 | SE004, SE005 |
| CE022 | LMArena's sampling policy requires that at least 20% of all battles involve only publicly available models, with remaining capacity available for unreleased or experimental models. | 中 | SE004 |
| CE023 | LMArena allows model providers to test unreleased models anonymously, shares results privately with the provider, and then removes the model before any public listing. | 高 | SE004, SE025 |
| CE024 | In April 2025 Meta submitted an 'experimental, chat-optimized' version of Llama 4 Maverick (not the public release) to LMArena, achieving a #2 ranking; when the public version was scored, it fell to approximately 32nd place. | 高 | SE025, SE024 |
| CE025 | Following the Meta Llama 4 incident, LMArena updated its leaderboard policies to require that pre-release model variants be explicitly labeled as customized and not represent the publicly released model. | 高 | SE025, SE024 |
| CE026 | A paper by researchers from Cohere, Stanford, MIT, and AI2 (April 2025) alleged that Meta, OpenAI, Google, and Amazon received disproportionately high sampling rates and could suppress low-scoring pre-release variants, constituting benchmark gaming. | 高 | SE024, SE025 |
| CE027 | LMArena denied the claims in the Cohere/Stanford study as containing 'inaccuracies and questionable analysis,' asserting that all model providers are allowed to submit more models for testing and that the leaderboard remains fair. | 中 | SE024 |
| CE028 | Independent researchers have noted that LMArena's user base is skewed toward technical and developer prompts, making it less representative of general-population or enterprise preferences. | 中 | SE024 |
| CE029 | Agent Arena analyzes five behavioral signals: confirmed success, praise vs. complaint, steerability, bash-error recovery, and tool hallucination rate. | 中 | SE006 |
| CE030 | In a recent 7-day window, Agent Mode recorded 160,480 agent tasks on LMArena's platform, with code writing as the largest category at 17.5%. | 中 | SE006 |
| CE031 | Agent Mode issued approximately 2 million structured tool calls in one 7-day period, including 936,000 bash calls and 550,000 file-write operations. | 中 | SE006 |
| CE032 | Code Arena uses Cloudflare R2 for persistent session storage and CodeMirror 6 for source-code display and live preview rendering of generated web applications. | 中 | SE013 |
| CE033 | Battles in Direct, launched May 2026, converts 10% of Direct Chat sessions into anonymous pairwise battles and applies position-bias and same-org-indicator corrections in the BT model. | 中 | SE005 |
| CE034 | Arena-Rank's open-source implementation achieves a 30x speedup over the historical FastChat-based BT implementation and uses reweighting to correct for non-uniform model sampling. | 高 | SE009, SE001, SE020 |
| CE035 | Arena-Hard-Auto v0.1 achieves 98.6% agreement with human preference rankings from Chatbot Arena and provides 3x higher model separation than MT-Bench. | 高 | SE016, SE018 |
| CE036 | The Search Arena paper was accepted at ICLR 2026, providing peer-reviewed validation of LMArena's search-augmented LLM evaluation methodology. | 高 | SE017, SE005 |
| CE037 | LMArena has not publicly disclosed SOC 2 Type II, ISO 27001, or any independent third-party security audit for its AI Evaluations commercial product or the arena.ai platform. | 中 | SE004 |
| CE038 | LMArena's help.arena.ai privacy policy exists but lacks enterprise data processing agreement (DPA) terms, and GDPR/CCPA compliance details are not publicly documented. | 中 | SE004 |
| CE039 | LMArena has evaluated over 400 public models and over 300 pre-release model variants across all its arenas since the platform launched in 2023. | 中 | SE005 |
| CE040 | LMArena has released 1.5 million prompts and over 145,000 battle data points publicly for open research as of early 2026. | 中 | SE005 |
| CE041 | Agent Arena sessions average approximately 16.5 structured tool calls, with about 75.6% of sessions using at least one tool in a measured 7-day window. | 中 | SE006 |
| CE042 | Agent Mode wrote 40.3 million lines of code through successful write_file calls in one 7-day window, approximately 1,000 lines per coding session. | 中 | SE006 |
| CU001 | LMArena had 5 million or more monthly active users as of January 2026, spanning more than 150 countries. | 高 | SU001, SU009 |
| CU002 | LMArena generates more than 60 million conversations per month and has accumulated more than 250 million conversations in total as of January 2026. | 高 | SU001, SU009 |
| CU003 | LMArena's community generates more than 2 million preference votes per month. | 中 | SU013 |
| CU004 | LMArena serves two distinct customer populations: a free community of millions of monthly users providing preference votes, and a small paying commercial segment of AI labs and enterprises purchasing AI Evaluations. | 高 | SU001, SU013 |
| CU005 | OpenAI, Google, and xAI are named paying customers of LMArena's AI Evaluations commercial product, drawing on evaluations to improve their models for production use cases. | 高 | SU001, SU009 |
| CU006 | LMArena partnered with OpenAI, Google, and Anthropic to make their flagship models available for community evaluation; Anthropic's paying status as an AI Evaluations subscriber is not separately confirmed. | 高 | SU001, SU009 |
| CU007 | LMArena's community grew approximately 25 times between May 2025 (seed round) and January 2026 (Series A) as reported by the company. | 中 | SU005 |
| CU008 | LMArena had approximately 3 million or more monthly users at the time of its commercial AI Evaluations product launch in September 2025. | 中 | SU013 |
| CU009 | LMArena has evaluated more than 400 public models and more than 300 pre-release model variants across its arenas as of early 2026. | 中 | SU005, SU006 |
| CU010 | The LMArena community released more than 50 million votes and 1.5 million open prompts by January 2026. | 中 | SU005 |
| CU011 | LMArena's AI Evaluations commercial product achieved an annualized consumption run-rate of $30 million in December 2025, less than four months after its September 2025 launch. | 高 | SU001, SU009 |
| CU012 | Meta submitted 27 Llama 4 model variants to LMArena for private pre-release testing between January and March 2025, then publicly disclosed only the score of the highest-performing experimental variant. | 高 | SU016, SU017 |
| CU013 | The publicly released version of Meta Llama 4 Maverick ranked approximately 32nd on the LMArena leaderboard, versus the experimental version which had ranked #2. | 高 | SU007, SU017 |
| CU014 | Anthropic's Claude models were described as currently winning the expert leaderboard for legal and medical use cases in January 2026. | 中 | SU003 |
| CU015 | Perplexity's Sonar models and Google Gemini were the top-ranked models in the Search Arena as of the ICLR 2026 paper, with Perplexity-Sonar-Reasoning-Pro and Gemini-2.5-Pro at the top. | 中 | SU019, SU024 |
| CU016 | LMArena has not publicly disclosed net revenue retention, gross revenue retention, or customer churn rates for its AI Evaluations commercial product. | 中 | SU011 |
| CU017 | The leaderboard changelog shows multiple model additions per week across all arenas in June 2026, serving as an indirect proxy for ongoing model-provider engagement. | 中 | SU015 |
| CU018 | LMArena describes its revenue as 'annualized consumption rate' rather than annual recurring revenue (ARR), implying a consumption-based billing model rather than committed subscriptions. | 中 | SU001 |
| CU019 | The structural conflict of interest in which the same AI labs that pay for AI Evaluations also have their models ranked on the public leaderboard was identified by independent journalists as a key credibility risk. | 高 | SU002, SU016 |
| CU020 | A paper from Cohere, Stanford, MIT, and AI2 alleged in April 2025 that Meta, OpenAI, Google, and Amazon received disproportionately high sampling rates, enabling benchmark gaming; LMArena denied specific claims and announced a new sampling algorithm. | 高 | SU016, SU003 |
| CU021 | LMArena updated its leaderboard policy after the Meta Llama 4 incident and stated that 'Meta's interpretation of our policy did not match what we expect from model providers.' | 高 | SU017, SU025 |
| CU022 | As of June 2026, only three companies (OpenAI, Google, xAI) are publicly named as paying customers of LMArena's AI Evaluations service; no enterprise non-lab customers have been publicly named. | 高 | SU001, SU009 |
| CU023 | LMArena's revenue is highly concentrated in three named AI lab customers; departure of any one of these would represent a material revenue risk given the $30M ARR run-rate and unknown diversification. | 中 | SU001 |
| CU024 | LMArena's community growth trajectory—25x from May 2025 to January 2026—implies organic conversion of community interest into lab evaluation demand, though the precise pathway from free user to commercial customer is not documented. | 中 | SU023, SU009 |
| CU025 | Prior to commercialization, Google's Kaggle, Andreessen Horowitz, and Together AI had donated compute, cloud credits, and cash to LMSYS as corporate sponsors, creating an indirect financial relationship with leaderboard participants. | 中 | SU016, SU010 |
| CU026 | Felicis General Partner Peter Deng stated in the LMArena Series A press release that LMArena has 'become essential infrastructure for every lab and enterprise' and that Felicis led the round because of LMArena's trustworthy real-world performance signal. | 高 | SU001, SU009 |
| CU027 | Scale AI launched a competing benchmarking product, SEAL Showdown, as a direct rival to LMArena's evaluation platform, representing competitive pressure in the AI evaluation market. | 中 | SU004, SU016 |
| CU028 | OpenTools.AI independently reported that LMArena was 'under fire' for benchmark bias in 2025, reflecting industry-wide concern beyond the single Cohere/Stanford paper. | 低 | SU003 |
| CU029 | PitchBook tracks Arena Intelligence (LMArena) with a confirmed valuation history from the $600M seed in May 2025 to the $1.7B Series A in January 2026. | 高 | SU006, SU009 |
| CU030 | LMArena's original domain lmarena.ai now redirects to arena.ai, reflecting the company's rebranding as it expanded beyond language model evaluation to a multi-arena platform. | 中 | SU007 |
| CU031 | EDGAR full-text search confirms LMArena (entity 'Recall Capital-LMArena') has an active Form D filing for the Series A under CIK 0002113470, providing independent corroboration of the funding event. | 高 | SU008, SU009 |
| CU032 | Anthropic's Claude is described by TechCrunch podcast as currently winning LMArena's expert leaderboard for legal and medical professional use cases as of January 2026, indicating provider engagement depth. | 中 | SU002, SU009 |
| CU033 | LMArena's two-year anniversary blog (April 2025) confirms approximately 41% of battles involve open-source models, indicating a mixed commercial and research user base that constrains pure commercial curation. | 中 | SU024, SU023 |
| CU034 | No independent third-party reviews on G2, Capterra, or Gartner Peer Insights for LMArena's AI Evaluations commercial product are publicly available as of June 2026. | 中 | SU011 |
| CU035 | LMArena's AI Evaluations commercial product targets software engineering, law, medicine, and scientific research as economically valuable industries for paid evaluation services. | 中 | SU001, SU009 |
| CR001 | LMArena earns revenue by selling paid AI evaluation services to AI labs (OpenAI, Google, xAI) whose models simultaneously appear on LMArena's public leaderboard, creating a structural conflict of interest between commercial and evaluation roles. | 高 | SR015, SR011, SR014 |
| CR002 | The Leaderboard Illusion paper (arXiv 2504.20879) found that a small number of providers could privately test multiple model variants and selectively disclose only their best scores, resulting in biased Arena rankings. | 高 | SR002, SR001 |
| CR003 | Meta privately tested 27 Llama-4 model variants on Chatbot Arena in the lead-up to its Llama 4 release, selecting the best-performing variant for its public score. | 高 | SR002, SR001, SR003 |
| CR004 | Meta submitted an "experimental chat version" of Llama 4 Maverick "optimized for conversationality" to LMArena that achieved a top-two leaderboard ranking, but this model was not the same version released publicly. | 高 | SR003, SR004 |
| CR005 | SurgeAI's analysis of 500 LMArena votes found that evaluators disagreed with LMArena's outcomes 52% of the time, with "confidence beats accuracy and formatting beats facts." | 中 | SR006 |
| CR006 | Independent researcher Gwern described LMArena as "a cancer" and questioned whether it is worth running, reflecting significant reputational erosion among technical users. | 中 | SR006 |
| CR007 | Arena Intelligence, Inc. (d/b/a LMArena) processes personal data from users in EU member states and is subject to GDPR compliance obligations including data subject rights, lawful basis for processing, and international data transfer safeguards. | 高 | SR009, SR010 |
| CR008 | LMArena's privacy policy (effective September 2025) explicitly warns users that prompts, votes/ratings, and other user content may be shared publicly and with AI providers as part of the evaluation process. | 高 | SR009, SR013 |
| CR009 | The EU AI Act's general-purpose AI (GPAI) model obligations, which came into force in August 2025, impose transparency documentation, training data summaries, and copyright policy requirements on AI platform operators. | 高 | SR012, SR010 |
| CR010 | No litigation directly naming Arena Intelligence, Inc. or LMArena as a plaintiff or defendant was found in public records as of 2026-06-25. | 中 | SR009 |
| CR011 | LMArena uses GCP's Sensitive Data Protection API to remove personal and sensitive data before sharing conversation data with model providers or publishing it publicly, and GCP is the confirmed primary cloud infrastructure provider. | 中 | SR015, SR009 |
| CR012 | LMArena's leaderboard methodology update in May 2026 (Battles in Direct) discovered and corrected two new biases: position bias favoring Model A, and an advantage for models sharing an organization with prior conversational context. | 高 | SR029, SR008 |
| CR013 | LMArena processes 60 million monthly user conversations across 150 countries on cloud infrastructure, creating platform reliability, data security, and regulatory compliance obligations at scale. | 高 | SR015, SR016 |
| CR014 | No public disclosure of SOC 2, ISO 27001, or equivalent security certification for Arena Intelligence, Inc. has been found as of the run date. | 中 | SR009 |
| CR015 | Chatbot Arena's Elo/Bradley-Terry scoring methodology relies on sufficient uniformity of battle distribution; coordinated voting campaigns, prompt injection, or sybil attacks could corrupt the ranking signal without immediate detection. | 中 | SR002, SR006 |
| CR016 | AI providers including OpenAI and Google have financial incentives to study and potentially overfit to the Arena evaluation distribution, as documented by the finding that access to Arena data yields up to 112% relative performance gains on Arena Hard. | 中 | SR002, SR005 |
| CR017 | LMArena had approximately 100 paying customers as of early 2026, generating $30M ARR, implying material revenue concentration risk with the top AI lab customers. | 中 | SR026, SR015 |
| CR018 | LMArena publicly confirmed that OpenAI, Google, and xAI draw on its evaluations to improve their models, making these companies simultaneously its best customers and its most motivated potential evaluators of its gaming policies. | 高 | SR015, SR020 |
| CR019 | No contractual SLA or public agreement guaranteeing sustained AI provider API access to LMArena's platform has been publicly disclosed; any lab can withdraw its model. No known instance of API withdrawal has occurred as of the run date. | 中 | SR007 |
| CR020 | LMArena employs approximately 41 people as of January 2026, representing a lean team relative to the platform's operational scope and regulatory obligations. | 中 | SR026 |
| CR021 | The lead investor in LMArena's Series A (Peter Deng of Felicis) previously worked at OpenAI, creating an appearance-of-conflict risk between Felicis's portfolio interest and LMArena's independence claims regarding OpenAI model evaluations. | 中 | SR015, SR025 |
| CR022 | Arena Hard performance — a synthetic benchmark LMArena maintains — can be improved by up to 112% relative gains with additional Arena battle data, according to the Leaderboard Illusion researchers' conservative estimates. | 中 | SR002, SR001 |
| CR023 | The Leaderboard Illusion paper found that Google and OpenAI each received an estimated 19.2% and 20.4% of all Chatbot Arena data, while 83 open-weight models combined received only approximately 29.7% of total data. | 高 | SR002, SR001 |
| CR024 | LMArena updated its sampling policy post-April 2026 to guarantee that at least 20% of all battles involve only publicly available models, and committed to reweighting scoring so that sampling probabilities do not bias Arena scores. | 高 | SR007, SR008 |
| CR025 | The SEC EDGAR filing for "Recall Capital-LMArena a Series of CGF2021 LLC" (Form D, filed 2026-02-26, accession 0002113470-26-000001) is a secondary-market venture fund raising $382,500 specifically to invest in LMArena. | 高 | SR024, SR014 |
| CR026 | LMArena's Agent Arena leaderboard launched on June 4, 2026, expanding into agentic evaluation with behavioral signals like file downloads, disapproval events, retries, and steerability rather than static preference votes alone. | 高 | SR008, SR016 |
| CR027 | US state privacy laws including California's CPRA impose separate data-subject rights and business compliance obligations on LMArena that are distinct from GDPR requirements. | 高 | SR009, SR012 |
| CR028 | The EU AI Act imposes penalties of up to 7% of global annual turnover for prohibited AI practices and up to 3% for other violations, applicable to AI platforms with EU operations. | 高 | SR012, SR010 |
| CR029 | CTOL Digital Solutions noted that LMArena's $1.7 billion valuation implies roughly 57x the company's annualized revenue run rate of $30M, pricing in substantial growth expectations that depend on sustained trust in benchmark neutrality. | 中 | SR011 |
| CR030 | LMArena's commercial evaluation product, AI Evaluations (launched September 2025), provides paid evaluation services to enterprises and AI labs, creating a revenue stream directly tied to the labs it ranks publicly. | 高 | SR017, SR015 |
| CR031 | The AI model evaluation platform market is projected to grow from $1.86 billion in 2025 to $2.36 billion in 2026 and $6.24 billion by 2030, at a 27.3-27.5% CAGR. | 中 | SR023 |
| CR032 | LMArena's open-source Arena-Rank repository and methodology academic papers provide external auditability of the leaderboard scoring approach, which is a partial mitigation against benchmark-integrity allegations. | 高 | SR007, SR016 |
| CR033 | LMArena's Series A (January 2026) achieved a post-money valuation of $1.7 billion, nearly triple the $600 million seed valuation from May 2025, on $250M total capital raised. | 高 | SR014, SR015, SR022 |
| CR034 | The Leaderboard Illusion researchers found that proprietary/closed models are sampled at higher battle rates and have fewer models removed from Arena compared to open-weight alternatives, creating a data access asymmetry. | 高 | SR002, SR018 |
| CR035 | A TechCrunch analysis noted that LMArena's commercial relationships with OpenAI, Google, and Anthropic raise questions about whether the benchmark can be trusted to assess AI models without corporate influence clouding the process. | 高 | SR001, SR011 |
| CR036 | LMArena was previously funded through grants and donations from Google's Kaggle, Andreessen Horowitz, and Together AI — organizations with direct stakes in the models being evaluated. | 高 | SR005, SR022 |
| CR037 | LMArena stated that its commercial evaluation product provides the same methodology to all paying customers without preferential treatment, and that the public leaderboard will always be available freely. | 中 | SR017 |
| CR038 | The LMArena policy (updated April 30, 2026) now requires that if a publicly released model differs from the pre-release version tested on Arena, Arena will remove the model from the leaderboard until it can be re-evaluated under the requirements of this policy. | 高 | SR007, SR008 |
| CR039 | TechCrunch noted that a Llama 4 Maverick vanilla release ranked 32nd on the LMArena leaderboard after the experimental version that ranked second was withdrawn, demonstrating a 30-rank gap attributable to benchmark optimization. | 高 | SR004, SR003 |
| CR040 | At 41 employees and $30M ARR, LMArena's revenue-per-employee ratio is approximately $730K/FTE, suggesting significant infrastructure leverage but thin organizational depth for regulatory compliance, security, and enterprise scale-up. | 中 | SR026, SR015 |
| CR041 | LMArena's 5 million monthly users across 150 countries is substantially below the 45 million EU monthly active user threshold that would trigger VLOP designation under the Digital Services Act, making VLOP obligations unlikely in the near term. | 中 | SR015, SR012 |
| CR042 | LMArena has not disclosed any data breach incidents involving the community evaluation dataset or user prompt data as of the run date, but uses GCP security tools rather than published third-party security certifications. | 中 | SR009, SR015 |
| CV001 | LMArena raised $150 million in a Series A round in January 2026 at a post-money valuation of $1.7 billion, led by Felicis and UC Investments, with participation from a16z, Kleiner Perkins, Lightspeed, The House Fund, LDVP, and Laude Ventures. | 高 | SV001, SV002 |
| CV002 | LMArena's annualized "consumption run rate" surpassed $30 million in December 2025, less than four months after launching its first commercial product (AI Evaluations) in September 2025. | 高 | SV002, SV001 |
| CV003 | LMArena raised a $100 million seed round in May 2025 at a $600 million valuation, bringing total capital raised to $250 million by January 2026 across two rounds in approximately seven months. | 高 | SV001, SV007 |
| CV004 | The Latka database records that LMArena sold approximately 17% at the seed round ($100M / $600M) and approximately 9% at the Series A ($150M / $1.7B), providing an implied pre-money Series A enterprise value of approximately $1.55 billion. | 中 | SV003, SV001 |
| CV005 | At $30M ARR and a $1.7B valuation, LMArena's implied ARR multiple is approximately 57x — placing it in the top decile of AI infrastructure Series A rounds in 2025–2026. | 高 | SV013, SV001 |
| CV006 | The $30M ARR figure represents December 2025 monthly revenue annualized (run rate), not a full-year booked revenue number; the company had fewer than four months of commercial operations as of the Series A announcement. | 高 | SV002, SV008 |
| CV007 | The AI model evaluation platform market was valued at $1.86 billion in 2025 and is projected to reach $2.36 billion in 2026 and $6.24 billion by 2030 at a 27.3–27.5% CAGR, according to The Business Research Company's 2026 market report. | 中 | SV010, SV006 |
| CV008 | Weights & Biases (an AI developer platform with model evaluation capabilities) was acquired by CoreWeave for approximately $1.4 billion in March 2025, providing a transaction comparable for AI evaluation infrastructure valuation. | 中 | SV012, SV011 |
| CV009 | A bull case for LMArena requires $120M+ ARR by 2028 and sustained platform multiple above 25x, implying a $3–6B exit value; this requires sustaining the $7.5M/month new ARR velocity seen in the first four months of commercial operations. | 中 | SV008, SV001 |
| CV010 | A base case for LMArena projects $60–75M ARR by end-2026 (at 50% of current ARR velocity), implying a $1.4–2.1B valuation at a 20–30x ARR multiple — roughly in line with the current $1.7B mark, providing limited upside from today's entry price. | 中 | SV001, SV013 |
| CV011 | LMArena's core competitive advantage is a community data moat: 5 million monthly users across 150 countries generating over 60 million conversations monthly, creating a real-time human preference dataset that is extremely difficult to replicate. | 高 | SV002, SV009 |
| CV012 | LMArena's leaderboard has become a standard reference in AI lab product launches, developer procurement decisions, and media coverage, creating network effects and switching costs that reinforce its incumbent position. | 高 | SV022, SV026 |
| CV013 | The structural conflict of interest — where LMArena earns revenue from the same labs it evaluates — creates an existential risk to the investment thesis; a single high-profile investigation confirming revenue influenced rankings would collapse both leaderboard credibility and enterprise revenue simultaneously. | 高 | SV013, SV016 |
| CV014 | Goodhart's Law dynamic — where labs optimize specifically for Arena rather than for genuine capability improvement — degrades the leaderboard's real-world signal value over time, as documented in the Leaderboard Illusion paper and corroborated by multiple independent analyses. | 高 | SV014, SV016 |
| CV015 | The overall investment recommendation for LMArena is "track" with a conditional buy signal requiring an independent methodology audit and an entry price below 30x next-twelve-months ARR. Risk rating is "high"; valuation stance is "stretched." | 中 | SV013, SV001 |
| CV016 | The bear case for LMArena implies a valuation range of $200–500M following a credibility collapse, driven by ARR churn to below $25M and multiple compression to 10–20x distressed ARR — representing a 70–88% loss from the $1.7B entry. | 中 | SV013, SV014 |
| CV017 | The primary thesis-break trigger is discovery of documentary evidence that commercial revenue relationships influenced LMArena's public leaderboard outcomes; secondary triggers include churn of any top-3 lab customers or ARR failing to reach $50M+ by Q3 2026. | 中 | SV013, SV016 |
| CV018 | Peter Deng, the Felicis general partner leading LMArena's Series A round, previously worked at OpenAI — one of LMArena's paying evaluation customers — creating an appearance-of-conflict risk between Felicis's portfolio interest and LMArena's independence claims regarding OpenAI evaluations. | 高 | SV002, SV022 |
| CV019 | The minimum blocking diligence items before a buy decision are: (1) independent statistical audit of the Arena methodology; (2) customer cohort data showing top-3 customer share below 40% or NRR above 110%; (3) cap table with liquidation preference terms; (4) EU AI Act compliance self-assessment. | 中 | SV013, SV016 |
| CV020 | The SEC Form D filing for "Recall Capital-LMArena a Series of CGF2021 LLC" (accession 0002113470-26-000001, filed February 2026) is a $382,500 secondary venture fund, demonstrating secondary-market demand for LMArena exposure at valuations consistent with the Series A. | 高 | SV015, SV001 |
| CV021 | Long-run structural analogies for LMArena's potential platform value include Bloomberg LP and S&P Global's ratings segment, which operate as neutral data arbiters with structural moats and high customer switching costs — though these analogies require LMArena to first resolve its conflict-of-interest and achieve regulatory equivalence. | 中 | SV006, SV013 |
| CV022 | LMArena's product expansion into Agent Arena (June 2026), WebDev Arena, Search Arena, Vision, Video, and coding leaderboards suggests the company is actively increasing its potential ARR ceiling by expanding beyond text model evaluation. | 高 | SV020, SV009 |
| CV023 | A CTOL analysis noted that LMArena's 57x ARR multiple prices in the assumption that the conflict between being a revenue-generating evaluation service and an independent benchmark arbiter "can be managed indefinitely" — a structural assumption that has not been independently validated. | 高 | SV013, SV001 |
| CV024 | LMArena employs approximately 41 people as of January 2026 with approximately $730K revenue per employee, indicating high capital efficiency but thin organizational depth for maintaining a $1.7B asset at scale. | 中 | SV003, SV002 |
| CV025 | No down-round risk has been evidenced in LMArena's financing history; the company raised at 2.83x step-up from seed ($600M) to Series A ($1.7B) in eight months on genuine commercial traction. | 高 | SV001, SV003 |
| CV026 | The revenue concentration risk at LMArena is heightened by its approximately 100 paying customers, where the top AI labs (OpenAI, Google, xAI) likely represent a disproportionate share of $30M ARR per standard B2B Pareto distributions. | 中 | SV003, SV013 |
| CV027 | The AI model evaluation platform market's 27.3% CAGR projection implies that if LMArena captures a 5-10% market share in a $3B+ market by 2028, its revenue would be $150-300M — sufficient to justify the current valuation at 10-15x revenue multiples characteristic of scaled data platforms. | 中 | SV010, SV013 |
| CV028 | LMArena's Arena Intelligence, Inc. d/b/a structure was incorporated in 2025 in Delaware; the company is headquartered in San Francisco, California, and is a private company with no SEC reporting obligations beyond Form D filings. | 高 | SV021, SV015 |
| CV029 | No public evidence of preferred stock liquidation preferences, anti-dilution ratchets, or convertible notes has been disclosed by LMArena or its investors as of the run date, though these terms are standard in Series A financings and their absence from public disclosure does not indicate they do not exist. | 低 | |
| CV030 | LMArena's investor base (Felicis, UC Investments, a16z, Kleiner Perkins, Lightspeed) includes firms with direct investments in AI labs that are LMArena's paying customers, creating a potential for governance conflicts that could compromise the independence of the evaluation platform. | 中 | SV025, SV007 |
| CV031 | LMArena's prior pre-commercial funding through grants from Google's Kaggle and donations from Andreessen Horowitz and Together AI — organizations with evaluated models on the platform — established a pattern of commercial entanglement that the new corporate structure has not fully resolved. | 高 | SV025, SV007 |
| CV032 | Scale AI, the closest large-scale comparable in AI data labeling and evaluation services, was valued in the $14–29B range on reported revenue substantially larger than LMArena's current $30M ARR, suggesting LMArena's 57x multiple is significantly richer than Scale AI's implied multiple on comparable revenue. | 中 | SV010, SV013 |
| CV033 | LMArena generated 50 million votes, 400+ model evaluations, and added 145,000 open-source battle data points to the community between the seed round (May 2025) and the Series A (January 2026), demonstrating substantial community engagement growth. | 高 | SV009, SV002 |
| CV034 | LMArena's AI Evaluations commercial product provides paid evaluation services including comprehensive in-depth evaluations based on community feedback, auditability through representative data samples, and service-level agreements with committed delivery timelines. | 高 | SV019, SV002 |
| CV035 | No institutional investor has publicly disclosed a reduction in LMArena position, secondary sale, or hedge of LMArena exposure since the Series A close as of the run date; the Recall Capital secondary fund implies continued secondary-market demand. | 中 | SV015, SV001 |
| CV036 | LMArena's Agent Arena leaderboard (launched June 4, 2026) represents a new evaluation domain targeting agentic AI systems, expanding the platform's addressable commercial market beyond text model benchmarking. | 高 | SV020, SV026 |
| CV037 | The minimum entry valuation offering adequate risk/return profile is estimated below $1.0–1.2B (approximately 30x $40M NTM ARR), providing sufficient discount to the current $1.7B mark to compensate for conflict-of-interest, concentration, and multiple-compression risks. | 中 | SV013, SV001 |
| CV038 | LMArena's exit readiness is limited in the near term: the company is 18 months old commercially, has no disclosed profitability path, and an IPO would require 2–3 years of operating history and revenue scale above $150M+ to command strong public market reception; strategic acquisition by a hyperscaler is the most plausible nearer-term exit scenario. | 中 | SV001, SV022 |
| CV039 | LMArena's product expansion into evaluation of models across text, code, vision, video, search, documents, and agents (seven distinct modalities as of June 2026) expands the potential ARR ceiling and reduces concentration risk in the core text model benchmarking segment. | 高 | SV020, SV009 |
| CV040 | The OfficeChai analysis notes that LMArena's AI testing market is estimated at $800–900M in 2025, projected to grow to $3.8 billion by 2032, providing a TAM growth context for the $1.7B valuation. | 中 | SV006, SV010 |
| CV041 | Reuters syndication through U.S. News reported that LMArena's valuation tripled to $1.7 billion in about eight months, underscoring how quickly private-market pricing expanded between the $600M seed round and the January 2026 Series A. | 高 | SV033, SV007 |
| 编号 | 出版方 | 标题 | 引文 |
|---|---|---|---|
| SO001 | Arena Intelligence Inc. | About Arena | Crowdsourced AI Model Evaluation Platform | Created by researchers from UC Berkeley, Arena (formerly LMArena) is a community-powered platform for understanding AI performance in the real world. |
| SO002 | Arena Intelligence Inc. | Fueling the World's Most Trusted AI Evaluation Platform (Series A Blog) | We've raised $150M of Series A funding led by Felicis and UC Investments (University of California). |
| SO003 | PR Newswire / LMArena | LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform | LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product. |
| SO004 | TechCrunch | LMArena lands $1.7B valuation four months after launching its product | LMArena raised a $150 million Series A at a post-money valuation of $1.7 billion. |
| SO005 | TechCrunch | LM Arena, the organization behind popular AI leaderboards, lands $100M | LM Arena... has raised $100 million in a seed funding round that values the organization at $600 million. |
| SO006 | Bloomberg | Popular AI Ranking Website Chatbot Arena Is Becoming a Real Company | |
| SO007 | Arena Intelligence Inc. | New Product: AI Evaluations | This service offers enterprises, model labs, and developers comprehensive evaluation services grounded in real-world human feedback. |
| SO008 | Arena Intelligence Inc. | Arena Leaderboard Policy | |
| SO009 | Arena Intelligence Inc. | Arena Leaderboard Changelog | |
| SO010 | LMSYS / UC Berkeley | Chatbot Arena: New Leaderboard & More Models | We are releasing Chatbot Arena, an open-source evaluation platform for LLMs. |
| SO011 | arXiv / UC Berkeley LMSYS | Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference | Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies. |
| SO012 | arXiv / UC Berkeley | Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena | |
| SO013 | arXiv / UC Berkeley | From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline | |
| SO014 | arXiv / Cohere, Stanford, MIT, Ai2 | The Leaderboard Illusion | We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired. |
| SO015 | TechCrunch | The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark | The evaluation is not reproducible, and the limited data released by LMSYS makes it challenging to study the limitations of models in depth. |
| SO016 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification. |
| SO017 | TechCrunch | Here's why most AI benchmarks tell us so little | |
| SO018 | TechCrunch | The leaderboard 'you can't game,' funded by the companies it ranks | |
| SO019 | Founded.com | How two Berkeley roommates built a $1.7B startup that helps you decide which AI to use | The founders behind AI evaluation platform Arena, formerly known as LMArena, have turned that confusion into a business now worth $1.7 billion. |
| SO020 | Winbuzzer | Experts Challenge Validity and Ethics of Crowdsourced AI Benchmarks Like LMArena | Chatbot Arena hasn't shown that voting for one output over another actually correlates with preferences, however they may be defined. |
| SO021 | Hugging Face | lmarena-ai Organization on Hugging Face | |
| SO022 | Andreessen Horowitz (a16z) | Announcing Our Latest Open Source AI Grants | |
| SO023 | Mashable SEA | LMArena has some competition: Scale AI launches Seal Showdown, a new benchmarking tool | Critics say that LMArena's system favors frontier models from big AI companies like Google, xAI, and OpenAI. |
| SO024 | Arena Intelligence Inc. | Agent Arena: Causal Evaluation of Agents in the Real World | |
| SO025 | Arena Intelligence Inc. | Introducing the Search Arena: Evaluating Search-Enabled AI | |
| SO026 | Arena Intelligence Inc. | WebDev Arena: A Live LLM Leaderboard for Web App Development | |
| SO027 | The Verge | Meta got caught gaming AI benchmarks | Meta's interpretation of our policy did not match what we expect from model providers. |
| SO028 | TechCrunch | Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark | |
| SO029 | arXiv / Search Arena (ICLR 2026) | Search Arena: A Crowd-Sourced, Human-Preference Dataset for Search-Augmented LLMs | |
| SM001 | The Business Research Company | Artificial Intelligence (AI) Model Evaluation Platform Market Report | The AI model evaluation platform market grows from $1.86 billion in 2025 to $2.36 billion in 2026 at a CAGR of 27.3%. |
| SM002 | Yahoo Finance | AI Model Evaluation Platform Market article | |
| SM003 | Research & Markets | AI Model Evaluation Platform Market Report | |
| SM004 | Precedence Research | Model Evaluation and Benchmarking Tools Market | |
| SM005 | Gartner | Gartner forecasts worldwide AI spending to grow 47 percent in 2026 | |
| SM006 | Presenc AI | Enterprise AI adoption statistics 2026 | 78% of Global 2000 companies have at least one AI workload in production in Q1 2026. |
| SM007 | Future AGI | Top 5 LLM Evaluation Tools 2025 | |
| SM008 | Scale AI | SEAL Showdown | SEAL Showdown includes users in 100+ countries, 70+ languages, and 200+ professional domains. |
| SM009 | Mashable SEA | LMArena has some competition: Scale AI launches SEAL Showdown | |
| SM010 | arXiv | The Leaderboard Illusion | |
| SM011 | TechCrunch | The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark | |
| SM012 | TechCrunch | LMArena lands $1.7B valuation four months after launching its product | |
| SM013 | PR Newswire | LMArena raises $150 million to build the world's most trusted AI evaluation platform | |
| SM014 | arXiv | Chatbot Arena paper | |
| SM015 | LMSYS | Chatbot Arena launch post | |
| SM016 | Arena AI | AI Evaluations | |
| SM017 | Arena AI | Agent Arena methodology | |
| SM018 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | |
| SM019 | Winbuzzer | Experts challenge validity and ethics of crowdsourced AI benchmarks like LMArena | |
| SM020 | The Verge | Meta Llama 4 Maverick benchmarks gaming | |
| SM021 | TechCrunch | LM Arena, the organization behind popular AI leaderboards, lands $100M | |
| SM022 | arXiv | Arena-Hard paper | |
| SM023 | arXiv | MT-Bench paper | |
| SM024 | Founded | LMArena founders profile | |
| SM025 | Scale AI | SEAL Showdown | |
| SP001 | arXiv (Wei-Lin Chiang, Lianmin Zheng, et al. — UC Berkeley / LMSYS) | Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference | Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies. |
| SP002 | arXiv (Tianle Li, Wei-Lin Chiang, et al. — UC Berkeley / LMSYS) | From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline | Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20. |
| SP003 | arXiv (Singh et al. — Cohere, Stanford, MIT, Ai2) | The Leaderboard Illusion | Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data. |
| SP004 | TechCrunch | The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark | The distribution of testing data may not accurately reflect the target market's real human users. Moreover, the platform's evaluation process is largely uncontrollable. |
| SP005 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | LM Arena allowed some industry-leading AI companies like Meta, OpenAI, Google, and Amazon to privately test several variants of AI models, then not publish the scores of the lowest performers. |
| SP006 | TechCrunch | Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark | The unmodified Maverick, 'Llama-4-Maverick-17B-128E-Instruct,' was ranked below models including OpenAI's GPT-4o, Anthropic's Claude 3.5 Sonnet, and Google's Gemini 1.5 Pro as of Friday. |
| SP007 | LMArena (Arena Intelligence) | Celebrating Community Impact at LMArena (Two-Year Celebration) | 400+ models have been evaluated across various modalities (Text, Vision, Text-to-Image, WebDev and more!). 300+ evaluations have been pre-release. |
| SP008 | LMArena (Arena Intelligence) | New Product: AI Evaluations | This service offers enterprises, model labs, and developers comprehensive evaluation services grounded in real-world human feedback. |
| SP009 | LMArena (Arena Intelligence) | Arena AI: The Official AI Ranking & LLM Leaderboard | |
| SP010 | LMArena (Arena Intelligence) | Arena Leaderboard | |
| SP011 | LMArena / LMSYS Org (GitHub) | FastChat: An open platform for training, serving, and evaluating large language models | FastChat powers Chatbot Arena (lmarena.ai), serving over 10 million chat requests for 70+ LLMs. |
| SP012 | EleutherAI (GitHub) | lm-evaluation-harness: A framework for few-shot evaluation of language models | Over 60 standard academic benchmarks for LLMs, with hundreds of subtasks and variants implemented. |
| SP013 | OpenAI (GitHub) | openai/evals: Evals is a framework for evaluating LLMs and LLM systems | You can now configure and run Evals directly in the OpenAI Dashboard. |
| SP014 | Hugging Face | Open LLM Leaderboard | |
| SP015 | LMArena (Hugging Face Space) | Arena Leaderboard (HuggingFace) | |
| SP016 | Scale AI | Scale AI — Reliable AI Systems | Benchmarking the frontier of AI capability with expert-level evaluations. |
| SP017 | Scale AI | Scale GenAI Platform | Every agent is built and tested against your specific enterprise standards — your workflows, your rules, your definition of good — before it ever touches production. |
| SP018 | Stanford CRFM | Holistic Evaluation of Language Models (HELM) — Latest | |
| SP019 | Stanford CRFM | Holistic Evaluation of Language Models (HELM) — Classic | |
| SP020 | BenchLM | LLM Leaderboard 2026 — Compare 261 AI Models Across 249 Benchmarks | 261 models · 249 benchmarks. The most comprehensive LLM comparison tool — 249 benchmarks, real pricing, and runtime data in one place. |
| SP021 | Artificial Analysis | AI Model & API Providers Analysis | |
| SP022 | MetaTech.dev | LMArena AI Benchmarking Crisis: Why Model Rankings Are Broken | The problem becomes even more concerning when we look at critical applications... This disconnect between benchmark success and practical reliability is exactly what's wrong with current AI benchmarking approaches. |
| SP023 | LMSYS Org (UC Berkeley) | LMSYS Org — Large Model Systems Organization | The Large Model Systems Organization develops large models and systems that are open, accessible, and scalable. |
| SP024 | U.S. Securities and Exchange Commission (EDGAR) | EDGAR Search Results — Recall Capital-LMArena a Series of CGF2021 LLC | |
| SP025 | U.S. Securities and Exchange Commission (EDGAR) | Form D — Recall Capital-LMArena a Series of CGF2021 LLC (search index entry) | |
| SI001 | LMArena (PR Newswire) | LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform | LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product. |
| SI002 | TechCrunch | LMArena lands $1.7B valuation four months after launching its product | LMArena's annualized "consumption rate" — as the company describes its annual recurring revenue (ARR) — of $30 million as of December, less than four months after launch. |
| SI003 | TechCrunch | LM Arena, the organization behind popular AI leaderboards, lands $100M | LM Arena, a crowdsourced benchmarking project that major AI labs rely on to test and market their AI models, has raised $100 million in a seed funding round that values the organization at $600 million. |
| SI004 | LMArena (Arena Intelligence) | Fueling the World's Most Trusted AI Evaluation Platform (Series A Blog) | This year we saw our community grow by over 25x alongside rapid adoption by AI labs who trust this platform as a gold-standard for evaluating real-world model performance. |
| SI005 | GetLatka | LMArena Revenue 2026: $30M ARR, $1.7B Valuation | In 2026, LMArena's revenue reached $30M. LMArena reached a $1.7B valuation in 2026, set during its Series A round. |
| SI006 | CTOL Digital | LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges | The company claims over $30 million in annualized revenue from selling evaluation services to these labs, launching its commercial product only in September 2025. At 57 times that run rate, the valuation prices in not just growth, but the assumption that this inherent conflict can be managed indefinitely. |
| SI007 | U.S. Securities and Exchange Commission | Form D — Recall Capital-LMArena a Series of CGF2021 LLC | |
| SI008 | Wired | LMArena Raises $150 Million to Evaluate the World's Most Powerful AI Models | |
| SI009 | arXiv (Singh et al. — Cohere, Stanford, MIT, Ai2) | The Leaderboard Illusion | |
| SI010 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | |
| SI011 | LMArena (Arena Intelligence) | New Product: AI Evaluations | LMArena earns revenue by providing paid AI evaluation services to AI labs and enterprises. |
| SI012 | LMArena (Arena Intelligence) | Arena AI: The Official AI Ranking & LLM Leaderboard | |
| SI013 | LMArena (Arena Intelligence) | Arena Leaderboard | |
| SI014 | LMArena (Arena Intelligence) | Celebrating Community Impact at LMArena (Two-Year Celebration) | |
| SI015 | MetaTech.dev | LMArena AI Benchmarking Crisis: Why Model Rankings Are Broken | |
| SI016 | LMSYS Org (UC Berkeley) | LMSYS Org — Large Model Systems Organization | |
| SI017 | BenchLM | LLM Leaderboard 2026 — Compare 261 AI Models Across 249 Benchmarks | |
| SI018 | LMArena / LMSYS Org (GitHub) | FastChat: An open platform for training, serving, and evaluating large language models | |
| SI019 | Scale AI | Scale AI — Reliable AI Systems | |
| SI020 | Scale AI | Scale GenAI Platform | |
| SI021 | arXiv (Wei-Lin Chiang, Lianmin Zheng et al.) | Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference | |
| SI022 | TechCrunch | The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark | |
| SI023 | TechCrunch | Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark | |
| SI024 | U.S. Securities and Exchange Commission (EDGAR) | EDGAR Search Results — Recall Capital-LMArena a Series of CGF2021 LLC | |
| SI025 | U.S. Securities and Exchange Commission (EDGAR search index) | SEC EDGAR Full-Text Search — Form D filings mentioning LMArena | |
| SI026 | VentureBeat | LMArena raises $150M at $1.7B valuation to build AI evaluation infrastructure | |
| SI027 | Bloomberg | UC Berkeley AI Benchmarking Group Chatbot Arena Raises $100 Million | |
| SI028 | BusinessWire | LMArena Raises $150 Million, Achieves $1.7B Valuation | |
| SI029 | StartupWired | AI startup LMArena triples valuation to $1.7B in 2026 | |
| SE001 | Arena Intelligence Inc. | Arena AI: The Official AI Ranking & LLM Leaderboard | |
| SE002 | Arena Intelligence Inc. | About Arena — Crowdsourced AI Model Evaluation Platform | |
| SE003 | Arena Intelligence Inc. | New Product: AI Evaluations | LMArena has already logged 250M+ real conversations, 2M+ monthly votes, and has 3M+ monthly users. |
| SE004 | Arena Intelligence Inc. | Arena Leaderboard Policy | |
| SE005 | Arena Intelligence Inc. | Leaderboard Changelog | |
| SE006 | Arena Intelligence Inc. | Agent Arena: Causal Evaluation of Agents in the Real World | The methodology powering the Agent Arena Leaderboard is different from our previous arenas. Rather than pairwise votes, rankings are calculated using a methodology we call causal tracing. |
| SE007 | Arena Intelligence Inc. | WebDev Arena: A Live LLM Leaderboard for Web App Development | |
| SE008 | Arena Intelligence Inc. | Introducing the Search Arena: Evaluating Search-Enabled AI | |
| SE009 | Arena Intelligence Inc. | Arena-Rank: Open Sourcing the Leaderboard Methodology | Arena-Rank, an open-source Python package for ranking that powers the LMArena leaderboard! |
| SE010 | Arena Intelligence Inc. | Does Style Matter in AI Evaluations? | |
| SE011 | Arena Intelligence Inc. | Introducing Hard Prompts Category in Chatbot Arena | |
| SE012 | Arena Intelligence Inc. | The Arena-Hard Pipeline | |
| SE013 | Arena Intelligence Inc. | The Next Stage of AI Coding Evaluation Is Here | Record: Every model action (file creation, edit, or execution) is logged and versioned. Snapshots are stored in Cloudflare R2. |
| SE014 | arXiv (Chiang et al., UC Berkeley / LMSYS) | Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference | |
| SE015 | arXiv / NeurIPS 2023 (Zheng et al., UC Berkeley / LMSYS) | Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena | |
| SE016 | arXiv (Li et al., UC Berkeley / LMSYS) | From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline | |
| SE017 | arXiv / ICLR 2026 (Miroyan et al., UC Berkeley) | Search Arena: Analyzing Search-Augmented LLMs | |
| SE018 | NeurIPS 2023 Proceedings (Zheng et al.) | Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023) | |
| SE019 | GitHub (lm-sys) | FastChat: An Open Platform for Training, Serving, and Evaluating LLMs | |
| SE020 | GitHub (lmarena) | arena-rank: Source Code of Arena Leaderboard Methodology | |
| SE021 | GitHub (lmarena org) | Arena GitHub Organization | |
| SE022 | HuggingFace / lmarena-ai (via Wayback) | lmarena-ai (Arena) Organization on HuggingFace | |
| SE023 | PyPI | arena-rank PyPI package | |
| SE024 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification. |
| SE025 | The Verge | Meta got caught gaming AI benchmarks | Meta's interpretation of our policy did not match what we expect from model providers. |
| SE026 | LMSYS (UC Berkeley SkyLab) | Chatbot Arena: New features and a Elo Rating System (original launch blog) | |
| SE027 | U.S. Securities and Exchange Commission | SEC Form D: Recall Capital-LMArena (Arena Intelligence Series A vehicle) | |
| SU001 | PRNewswire (Arena Intelligence Inc.) | LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform | LMArena works with leading AI labs and enterprises, including OpenAI, Google and xAI, all drawing on LMArena's evaluations to improve their models for production use cases. |
| SU002 | TechCrunch Podcast | The PhD students who became the judges of the AI industry | how a team like theirs can build a neutral benchmark when the companies they're ranking are also their backers |
| SU003 | OpenTools.AI | LM Arena Under Fire: Allegations of Benchmark Bias Stir AI Industry | |
| SU004 | Scale AI | SEAL Showdown: Scale's AI Evaluation Platform | |
| SU005 | AI Wiki | LMArena.org — AI Wiki | |
| SU006 | PitchBook | Arena Intelligence Company Profile | |
| SU007 | Arena Intelligence Inc. | LMArena (original domain) — redirects to arena.ai | |
| SU008 | U.S. Securities and Exchange Commission (EDGAR) | EDGAR Full-Text Search: LMArena Form D filings | |
| SU009 | TechCrunch | LMArena lands $1.7B valuation four months after launching its product | That trajectory, and the startup's popularity, were enough for VCs to pile in for the Series A. |
| SU010 | TechCrunch | LM Arena, the organization behind popular AI leaderboards, lands $100M | |
| SU011 | Arena Intelligence Inc. | Arena AI: The Official AI Ranking & LLM Leaderboard | |
| SU012 | Arena Intelligence Inc. | About Arena — Crowdsourced AI Model Evaluation Platform | |
| SU013 | Arena Intelligence Inc. | New Product: AI Evaluations | LMArena has already logged 250M+ real conversations, 2M+ monthly votes, and has 3M+ monthly users. |
| SU014 | Arena Intelligence Inc. | Arena Leaderboard Policy | |
| SU015 | Arena Intelligence Inc. | Leaderboard Changelog | |
| SU016 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. |
| SU017 | The Verge | Meta got caught gaming AI benchmarks | Meta's interpretation of our policy did not match what we expect from model providers. |
| SU018 | arXiv (Chiang et al., UC Berkeley / LMSYS) | Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference | |
| SU019 | arXiv / ICLR 2026 (Miroyan et al.) | Search Arena: Analyzing Search-Augmented LLMs | |
| SU020 | GitHub (lm-sys) | FastChat: An Open Platform for Training, Serving, and Evaluating LLMs | |
| SU021 | HuggingFace / lmarena-ai (via Wayback) | lmarena-ai (Arena) Organization on HuggingFace | |
| SU022 | Arena Intelligence Inc. | Agent Arena: Causal Evaluation of Agents in the Real World | In a sample of the heaviest real sessions we saw: a live sports-TV schedule site, an autonomous-underwater-vehicle autopilot, a self-hosted movie-watchlist app. |
| SU023 | Arena Intelligence Inc. | Fueling the World's Most Trusted AI Evaluation Platform (Series A blog) | This year we saw our community grow by over 25x alongside rapid adoption by AI labs who trust this platform as a gold-standard for evaluating real-world model performance. |
| SU024 | Arena Intelligence Inc. | Celebrating Community Impact at LMArena (Two-Year Anniversary) | |
| SU025 | TechCrunch | Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark | The release version of Llama 4 has been added to LMArena after it was found out they cheated, but you probably didn't see it because you have to scroll down to 32nd place. |
| SU026 | Discord (LMArena Community) | Arena Discord Community Server | |
| SU027 | Tech in Asia | a16z, Lightspeed back $150M Series A of AI model evaluator LMArena | |
| SU028 | U.S. Securities and Exchange Commission (EDGAR) | EDGAR Full-Text Search: Arena Intelligence Form D 2025-2026 | |
| SR001 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | "Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification." — Sara Hooker, Cohere VP of AI Research |
| SR002 | arXiv (Cohere, Stanford, MIT, AI2) | The Leaderboard Illusion (arXiv:2504.20879) | "We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired." |
| SR003 | The Verge | Meta got caught gaming AI benchmarks | "Meta's interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that 'Llama-4-Maverick-03-26-Experimental' was a customized model to optimize for human preference." |
| SR004 | TechCrunch | Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark | |
| SR005 | TechCrunch | The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark | "Companies can continually optimize their models to better align with the LMSYS user distribution, possibly leading to unfair competition and a less meaningful evaluation." |
| SR006 | UCStrategies | AI Benchmarks Are a Game Now — And the Industry Is Cheating to Win | "SurgeAI analyzed 500 LMArena votes and disagreed with 52%, finding that 'confidence beats accuracy and formatting beats facts.'" |
| SR007 | LMArena | Arena Leaderboard Policy (Last Updated April 30, 2026) | |
| SR008 | LMArena | Leaderboard Changelog | |
| SR009 | Arena Intelligence, Inc. | LMArena Privacy Policy (Previous Version, Effective 2025-09-05) | "Arena Intelligence, Inc. d/b/a LMArena provides a platform for using, comparing, rating, testing, evaluating, and ranking third-party AI models." |
| SR010 | European Commission (Your Europe) | Data protection under GDPR — Your Europe | |
| SR011 | CTOL Digital Solutions | LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges | "LMArena has positioned itself as the independent arbiter of model performance, yet it derives revenue from the same AI labs it evaluates — OpenAI, Google, and xAI among them." |
| SR012 | Didit | AI Compliance in the LLM Era: Regulatory Guide 2026 | |
| SR013 | LMArena | Arena AI: The Official AI Ranking & LLM Leaderboard (Homepage) | |
| SR014 | TechCrunch | LMArena lands $1.7B valuation four months after launching its product | |
| SR015 | PR Newswire | LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform | |
| SR016 | LMArena | Fueling the World's Most Trusted AI Evaluation Platform (Series A blog) | |
| SR017 | LMArena | New Product: AI Evaluations | |
| SR018 | ByteIota | LMArena Raises $150M at $1.7B Valuation in 4 Months | |
| SR019 | StartupWired | AI Startup LMArena Triples Valuation to $1.7B in 2026 | |
| SR020 | The AI Insider | LMArena Secures $150M to Build the World's Most Trusted AI Evaluation Platform | |
| SR021 | OfficeChai | AI Evaluation Platform LMArena Raises Series A At Valuation Of $1.7 Billion | "LMArena, formally known as Arena Intelligence Inc., was founded in 2025 by Anastasios N. Angelopoulos (CEO), Wei-Lin Chiang (CTO), and Ion Stoica (Advisor)." |
| SR022 | TechCrunch | LM Arena, the organization behind popular AI leaderboards, lands $100M | |
| SR023 | The Business Research Company | AI Model Evaluation Platform Market Size and Trends Report 2026 | |
| SR024 | U.S. Securities and Exchange Commission | Form D: Recall Capital-LMArena a Series of CGF2021 LLC (EDGAR filing 0002113470-26-000001) | |
| SR025 | TechCrunch | The leaderboard 'you can't game,' funded by the companies it ranks (video) | |
| SR026 | Latka | LMArena Revenue 2026: $30M ARR, $1.7B Valuation | |
| SR027 | CTOL Digital Solutions | LMArena Raises $150 Million — conflict-of-interest analysis | |
| SR028 | ByteIota | LMArena — open-source model bias and methodological flaws | |
| SR029 | LMArena | Leaderboard Changelog — Battles in Direct Update (May 12, 2026) | "We observed two new biases in the Battles in Direct voting data and corrected for them in the Bradley-Terry fit: the first is a position bias favoring Model A; the second is an advantage given to models that share an organization with the prior turns of context." |
| SR030 | StartupWired | LMArena Triples Valuation — Risks and Challenges Ahead section | "As usage grows, so do demands for data security, fairness, and governance. Any misstep could damage credibility, which forms the core of LMArena's value proposition." |
| SV001 | TechCrunch | LMArena lands $1.7B valuation four months after launching its product | "The startup bolted out of the gate as a commercial venture with a $100 million seed round in May at a $600 million valuation. This new round means it raised $250 million in about seven months." |
| SV002 | PR Newswire | LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform | "LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product." |
| SV003 | Latka | LMArena Revenue 2026: $30M ARR, $1.7B Valuation | |
| SV004 | StartupWired | AI Startup LMArena Triples Valuation to $1.7B in 2026 | |
| SV005 | The AI Insider | LMArena Secures $150M to Build the World's Most Trusted AI Evaluation Platform | |
| SV006 | OfficeChai | AI Evaluation Platform LMArena Raises Series A At Valuation Of $1.7 Billion | "The company is positioning itself in what analysts estimate is an $800-900 million AI testing market in 2025, projected to grow to $3.8 billion by 2032." |
| SV007 | TechCrunch | LM Arena, the organization behind popular AI leaderboards, lands $100M | |
| SV008 | ByteIota | LMArena Raises $150M at $1.7B Valuation in 4 Months | "Four months to $30 million: In May 2025, LMArena raised $100M at $600M; by September, launched AI Evaluations; three months later, hit $30M annualized run rate." |
| SV009 | LMArena | Fueling the World's Most Trusted AI Evaluation Platform (Series A blog) | "Since announcing our $100M Seed round last year in May, LMArena has grown far faster than we imagined. In a matter of months, the community has contributed 50 million votes." |
| SV010 | The Business Research Company | AI Model Evaluation Platform Market Size and Trends Report 2026 | "AI Model Evaluation Platform market size has reached $1.86 billion in 2025; expected to grow to $6.24 billion in 2030 at a CAGR of 27.5%." |
| SV011 | Research and Markets | AI Model Evaluation Platform Market Report 2026 | |
| SV012 | The Business Research Company | Human-In-The-Loop AI Market Size and Drivers Report 2026 | "In March 2025, CoreWeave Inc. acquired Weights & Biases Inc. for approximately $1.4 billion." |
| SV013 | CTOL Digital Solutions | LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges | "At 57 times that run rate, the valuation prices in not just growth, but the assumption that this inherent conflict can be managed indefinitely." |
| SV014 | UCStrategies | AI Benchmarks Are a Game Now — And the Industry Is Cheating to Win | |
| SV015 | U.S. Securities and Exchange Commission | Form D: Recall Capital-LMArena a Series of CGF2021 LLC (EDGAR accession 0002113470-26-000001) | "Recall Capital-LMArena a Series of CGF2021 LLC — Pooled Investment Fund / Venture Capital Fund — Amount Sold: $382,500" |
| SV016 | arXiv (Cohere, Stanford, MIT, AI2) | The Leaderboard Illusion (arXiv:2504.20879) | |
| SV017 | TechCrunch | Study accuses LM Arena of helping top AI labs game its benchmark | |
| SV018 | LMArena | Arena Leaderboard Policy (Last Updated April 30, 2026) | |
| SV019 | LMArena | New Product: AI Evaluations | |
| SV020 | LMArena | Leaderboard Changelog | |
| SV021 | Arena Intelligence, Inc. | LMArena Privacy Policy (Previous Version, Effective 2025-09-05) | "Arena Intelligence, Inc. d/b/a LMArena provides a platform for using, comparing, rating, testing, evaluating, and ranking third-party AI models." |
| SV022 | TechCrunch | The leaderboard 'you can't game,' funded by the companies it ranks (video) | |
| SV023 | The Verge | Meta got caught gaming AI benchmarks | |
| SV024 | TechCrunch | Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark | |
| SV025 | TechCrunch | The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark | |
| SV026 | Grokipedia | Arena (LMArena) — Grokipedia overview | "As of March 5, 2026, the top models on the text leaderboard are claude-opus-4-6 (1504 Elo), gemini-3.1-pro-preview (1500 Elo); total 5,430,034 votes collected across the platform." |
| SV027 | TLDL | AI Company Rankings 2026: Revenue, Funding & Valuation Data for 2,000+ Companies | "Private funding for AI startups topped $150 billion over the trailing twelve months; foundation model companies raising $80 billion in 2025." |
| SV028 | Wellows | 85 Hottest AI Startups to Watch in 2026 [By Valuation, Funding, & Growth] | "Anysphere (Cursor): AI coding assistant, $29.3B valuation, $1B ARR; Harvey: legal AI; LMArena reached a $1.7 billion valuation in under four months." |
| SV029 | AgentMarketCap | LMArena's $1.7B Valuation in 4 Months: Why AI Evaluation Is the New Data Labeling | "Scale AI was valued at $7B in 2021. In June 2025, Meta acquired a 49% stake for $14.3 billion — the largest VC transaction of 2025 — implicitly valuing Scale above $29 billion." |
| SV030 | Axis Intelligence Research | AI Copyright Lawsuits 2026: Status Tracker — Updated Monthly | "$50 billion+: Cumulative legal exposure across all active AI copyright and related IP cases; Anthropic's confirmed settlement in Bartz v. Anthropic: $1.5 billion covering ~482,000 works." |
| SV031 | U.S. Securities and Exchange Commission (EDGAR) | EDGAR Company Search: Recall Capital-LMArena a Series of CGF2021 LLC (CIK 0002113470) | |
| SV032 | Copyright Alliance | AI Copyright Lawsuit Developments in 2025: A Year in Review | "Anthropic's $1.5B settlement in Bartz v. Anthropic required payment of approximately $3,000 for each of the 482,460 books downloaded from pirate libraries — the first publicly confirmed pricing benchmark for AI training on pirated content." |
| SV033 | U.S. News & World Report / Reuters | AI Startup LMArena Triples Its Valuation to $1.7 Billion in Latest Fundraise | "LMArena said on Tuesday its valuation had tripled to $1.7 billion in about eight months, following a new funding round where it raised $150 million." |