Startup Diligence
Diligence report AI infrastructure / model evaluation Series A 2026-06-25

LMArena

Category-defining evaluation franchise with real momentum, but current valuation already prices in exceptional growth and governance repair

LMArena appears to be the category leader in AI evaluation, but the $1.7B Series A price leaves little margin of safety until governance, neutrality, and customer-concentration questions are resolved.

Cover facts

Founded 01
2023 [CO001]
Latest round 02
150 USD M [CO011]
Post-money valuation 03
1700 USD M [CO011, CV001]
Total capital raised 04
250 USD M [CO013]
Monthly users 06
5 M+ [CO017]
Monthly conversations 07
60 M+ [CO018]
Models evaluated 08
400+ [CO024]

Company profile

LMArena grew out of the UC Berkeley LMSYS / Chatbot Arena research effort launched in 2023 and commercialized as Arena Intelligence Inc. in April 2025. Its core product is a blind, pairwise human-preference evaluation platform where users compare anonymous AI model outputs, generating a public leaderboard and a proprietary stream of real-world preference data. The company monetizes through paid AI evaluation services for model labs, enterprises, and developers, including the AI Evaluations product launched in September 2025 with auditability, representative sample reporting, and service-level agreements. LMArena's traction is unusual: public materials cite 5M monthly users, 60M monthly conversations, 50M votes since the seed round, and a $30M annualized consumption run rate by December 2025. The same business also carries a structural tension because the most important customers and investors overlap with the labs whose models the platform evaluates.

Website
lmarena.ai
Founded
2023-05-03
Founders
Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica
Founding location
UC Berkeley / San Francisco Bay Area, California
Headquarters
San Francisco Bay Area, California, USA
Product
Crowdsourced AI model evaluation platform spanning text, coding, search, vision, image, video, and agent benchmarks, plus a paid enterprise evaluation product that packages community-grounded testing, representative battle samples, leaderboard infrastructure, and SLA-backed analysis.
Customers
Frontier AI labs, enterprise AI buyers, product teams, and developers that need comparative model evaluation for release decisions, procurement, and product quality measurement.
Business model
Evaluation-as-a-service: paid custom and private model evaluations, enterprise benchmarking, and related platform services layered on top of the public Arena leaderboard and open evaluation datasets.
Stage
Series A
Funding status
$100M seed at $600M post-money in May 2025 followed by a $150M Series A at $1.7B post-money in January 2026, bringing total disclosed capital raised to about $250M.
[CO001, CO005, CO006, CO007, CO011, CO013, CO017, CO018]

Executive summary

Top strengths

  • The public Arena leaderboard has become a de facto reference point for frontier-model launches, giving LMArena an unusual data and distribution moat.
  • Community scale is real: 5M monthly users, 60M monthly conversations, and 50M votes create an evaluation dataset that is hard for new entrants to reproduce quickly.
  • The company commercialized rapidly, launching AI Evaluations in September 2025 and reaching a $30M annualized run rate by December 2025.
  • Product expansion across search, coding, multimodal, and agent benchmarks increases the plausible ARR ceiling beyond text-model comparison alone.
  • Tier-1 investor support and strong academic roots improve credibility with AI labs and enterprise buyers despite limited operating history.

Top risks

  • Structural conflict of interest: major paying labs and investors overlap with the entities whose models LMArena publicly evaluates.
  • Benchmark integrity has already been challenged by the Leaderboard Illusion paper and the Meta Maverick incident, raising real credibility risk.
  • The $1.7B valuation implies roughly 57x the disclosed $30M annualized run rate, leaving little room for execution misses or multiple compression.
  • Revenue quality is still opaque because the company discloses a run-rate proxy rather than audited revenue, gross margin, or customer-concentration data.
  • Governance disclosure is thin for a company in an adjudication role: public evidence does not yet show board composition, independence controls, or detailed cap-table terms.

Open gaps

  • Top-customer concentration and net revenue retention are not publicly disclosed.
  • Gross margin, compute cost structure, and burn rate remain unavailable from public sources.
  • No independent statistical audit has publicly resolved the neutrality concerns raised in 2025.
  • Board composition, investor rights, and liquidation-preference details are still private.

Contents

Chapter 01

01Company Overview

1.1 Identity and Product

LMArena, now operating under the brand "Arena," is an AI model evaluation platform that enables users to compare frontier AI models through anonymous head-to-head battles. A user submits a prompt to two unlabeled models simultaneously, votes on the better response, and only then sees which models were judged. Aggregated votes power a public leaderboard using a Bradley-Terry / Elo-based ranking algorithm, generating continuous, real-world performance comparisons across large language models and multimodal AI systems. The corporate entity, Arena Intelligence Inc., was incorporated on April 18, 2025, transitioning the Chatbot Arena research project into a commercial enterprise. Headquarters is in the San Francisco Bay Area (specific address undisclosed), consistent with UC Berkeley origins. Primary websites operate at arena.ai and lmarena.ai. LMArena's commercial product, AI Evaluations, launched in September 2025, providing enterprises, model labs, and developers with community-grounded performance analytics, representative feedback samples, and service-level agreements. Named commercial customers include OpenAI, Google, and xAI. The business model is evaluation-as-a-service: clients pay for systematic assessments of model performance across domains such as software engineering, law, and medicine. The platform has expanded well beyond text-only LLM comparison to include Search Arena, WebDev Arena, Vision Arena, text-to-image, text-to-video, and Agent Arena (launched June 2026). This multimodal expansion broadens LMArena's evaluable surface and addressable customer base. As of June 2026, the Arena brand is the de facto public leaderboard for frontier AI models, with its rankings cited in product announcements, investor materials, and academic publications by every major AI lab. [CO001, CO002, CO005, CO007, CO008, CO009]

FO002: LMArena Business Architecture Flow

How LMArena's community inputs, platform infrastructure, data products, and commercial outputs connect.

[CO007, CO017, CO020, CO024, CO025, CO026]
FO003: LMArena Snapshot KPIs

Key performance indicators as of January 2026 across scale, financial, and community dimensions.

Values are company-reported as of January 2026; independent audits unavailable for a private company. ARR stated as annualized consumption run rate, not GAAP revenue.

[CO007, CO011, CO013, CO015, CO017, CO018]

1.2 Founders and Leadership

LMArena was co-founded by three people with UC Berkeley affiliations. Anastasios Angelopoulos, CEO, was a postdoctoral researcher in Statistics and Machine Learning at UC Berkeley and is the primary public spokesperson. Wei-Lin Chiang, co-founder, completed his PhD in distributed systems at Berkeley and was instrumental in building the FastChat serving framework that underlies Chatbot Arena. Ion Stoica, a UC Berkeley Computer Science professor and serial entrepreneur (co-founder of Databricks and Anyscale), joined as co-founder, providing commercialization experience and industry credibility. Original Chatbot Arena research involved additional contributors including Lianmin Zheng and faculty advisors Michael Jordan and Joseph Gonzalez. The founding team demonstrates strong founder-market fit: Angelopoulos and Chiang co-authored the key academic papers establishing the evaluation methodology, and Stoica's track record at Databricks provides operational depth rarely seen in academic spinouts. Key-person dependency is elevated: Angelopoulos serves as both scientific lead and CEO. No non-founder C-suite executives have been publicly named. No board composition or independent director disclosures have been made, typical for early-stage private companies. No leadership departures or governance changes have been disclosed since incorporation. However, the 2025 benchmark integrity controversy revealed a tension between Stoica (who publicly disputed the Leaderboard Illusion findings) and external researchers, highlighting how the founding team's credibility is central to the platform's legitimacy. [CO003, CO004, CO005, CO006, CO021, CO030]

Leadership and founder table
NameRoleBackgroundFounder-Market FitKey-Person Risk
Anastasios AngelopoulosCEO, Co-founderUC Berkeley postdoc, Statistics/ML; co-author of Chatbot Arena and Arena-Hard papers (arXiv:2403.04132, 2406.11939)Very high — technical architect of evaluation methodology and primary public voice of the companyVery high — dual role as scientific authority and CEO; departure would affect both product credibility and investor confidence
Wei-Lin ChiangCo-founderUC Berkeley PhD, distributed systems; built FastChat serving framework underlying Chatbot Arena; co-author of core methodology papersHigh — original platform architect; deep technical executionHigh — core engineering expertise; limited public profile suggests key operational role
Ion StoicaCo-founderUC Berkeley CS professor; co-founder of Databricks (~$43B valuation) and Anyscale (Ray framework); serial entrepreneurHigh — commercialization experience and venture-building credibilityModerate — advisory / board-level role likely; prior company-building track record reduces single-point dependency
Michael JordanFaculty advisor (LMSYS paper)UC Berkeley ML pioneer; foundational work in probabilistic ML, Bayesian methods, and statistical learning theoryReputational — academic prestige lends legitimacy to methodology claimsLow — advisory contribution; no operational role
Joseph E. GonzalezFaculty advisor (LMSYS paper)UC Berkeley Systems+ML faculty; co-developed Ray distributed computing framework; co-author of core papersTechnical — systems expertise relevant to infrastructure scalingLow — advisory; no operational role

Board composition, non-founder C-suite (CTO, CFO, VP Sales/BD), and investor-appointed directors are not publicly disclosed. Enumeration is partial based on public-record sources only.

[CO003, CO004, CO005, CO006, CO033]

1.3 Funding History and Investors

LMArena's pre-commercial funding came from grants and donations, including contributions from Google's Kaggle platform, Andreessen Horowitz, and Together AI, primarily in the form of compute resources and cash to support research infrastructure. In May 2025, following incorporation, LMArena raised a $100 million seed round co-led by Andreessen Horowitz and UC Investments (the University of California's endowment), at a $600 million post-money valuation. Lightspeed Venture Partners, Felicis, and Kleiner Perkins also participated. In January 2026, the company raised a $150 million Series A at a $1.7 billion post-money valuation — nearly triple the seed valuation — co-led by Felicis and UC Investments, with Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed, and Laude Ventures participating. Total capital raised stands at approximately $250 million as of January 2026. UC Investments' dual role as University of California endowment manager and institutional backer of a UC Berkeley spinout, and Andreessen Horowitz's portfolio overlap with models on the Arena leaderboard (including Mistral), represent potential conflicts of interest noted by external critics. No secondary transactions or debt facilities have been publicly disclosed. [CO010, CO011, CO012, CO013, CO014, CO021]

Stakeholder or investor map
StakeholderRoleRound(s)ImportanceDiligence Ask
Felicis VenturesLead investor, Series ASeries ACo-lead investor; GP Peter Deng quoted in official press release; likely board seat or observer rightsBoard seat composition; pro-rata rights; full-ratchet provisions in term sheet
UC Investments (University of California)Lead investor, Seed + Series ASeed + Series ADual-role backer: endowment investor and institutional sponsor of UC Berkeley spinout; structural conflict of interestConflict-of-interest protocol between endowment role and academic IP origination; information rights from UC Berkeley affiliation; IP licensing terms
Andreessen Horowitz (a16z)InvestorSeed + Series ATier-1 VC; participated both rounds; a16z portfolio overlaps with Arena leaderboard model providers (Mistral investment)Portfolio conflict disclosures; Mistral proximity to Arena evaluations; potential for preferential access to pre-release model testing
Kleiner PerkinsInvestorSeed + Series AEstablished VC; two-round participation signals conviction; possible board observer rightsObserver seat terms; governance rights; anti-dilution provisions
Lightspeed Venture PartnersInvestorSeed + Series AActive early-stage tech investor; two-round participationInformation rights; pro-rata in future rounds; governance structure
The House FundInvestorSeries AUC Berkeley-affiliated fund; related-party relationship to founding institutionRelated-party transaction disclosure; financial terms vs. arms-length investors; governance overlap
LDVPInvestorSeries ATechnology VC; limited public information on terms or governance roleFund focus; prior AI portfolio companies; governance participation level
Laude VenturesInvestorSeries ASmaller round participant; no additional public information availableFund background; governance participation; pro-rata rights

Individual investment amounts per investor are not publicly disclosed. UC Investments' dual role as both endowment investor and UC Berkeley institutional backer is a structural conflict requiring diligence. No secondary shareholders, convertible notes, or SAFEs have been disclosed publicly.

[CO010, CO011, CO012]

1.4 Traction Metrics and Financial Indicators

By January 2026, LMArena reported over 5 million monthly users across 150 countries and 60 million conversations per month. The community accumulated more than 50 million votes across text, vision, web development, search, video, and image modalities, and participated in evaluating more than 400 distinct AI models. LMArena also released 145,000 open-source battle data points from expert and occupational evaluation categories. On the financial side, the annualized consumption run rate — the company's ARR proxy — surpassed $30 million in December 2025, approximately four months after commercial launch. This is stated as an annualized run rate rather than realized GAAP revenue, and independently audited figures are unavailable. The rapid ramp from zero to $30 million annualized in four months is exceptional, but the structure of the commercial service (SLA-bound deliverables with community raters) implies potential labor cost constraints on gross margin that have not been disclosed. LMArena has not disclosed headcount, burn rate, or gross margin, which limits financial diligence from publicly available information alone. Deduplication filters remove approximately 10% of submitted votes, and identity-leak detection removes fewer than 4% of all votes as of July 2025 methodology updates, indicating active data quality management. [CO015, CO017, CO018, CO019, CO020, CO023]

Snapshot KPI table
MetricValue / StatusDateConfidenceGap / Diligence Path
Post-Series A valuation$1.7 billion (post-money)Jan 2026HighPre-money not disclosed; no secondary market pricing available
Total capital raised~$250 millionJan 2026HighConfirmed by press release and multiple news sources
Seed valuation$600 million (post-money)May 2025HighAnnounced by company; confirmed by TechCrunch and Bloomberg
ARR (consumption run rate)$30 million+Dec 2025MediumAnnualized run-rate proxy; not audited GAAP revenue; no mid-2026 update
Monthly active users5 million+Jan 2026MediumCompany-reported; 'active' definition unspecified; no independent audit
Monthly conversations60 million+Jan 2026MediumCompany-reported; no third-party verification
Total votes accumulated50 million+Dec 2025MediumIncludes deprecated model battles; breakdown unavailable
Models evaluated400+Dec 2025MediumIncludes deprecated and private-test models; exact active count unknown
Countries served150Jan 2026MediumReflects user geography; not registered entities in 150 countries
HeadcountUndisclosedJun 2026LowNot publicly disclosed; Series A earmarked for technical team expansion
Gross margin / burn rateUndisclosedJun 2026LowPrivate company; no financial disclosures; labor cost structure unknown

Values from company press releases, TechCrunch reporting, and investor announcements. "Consumption run rate" is LMArena's terminology for annualized ARR proxy; not equivalent to recognized GAAP revenue. Null / Undisclosed fields reflect absent public disclosure, not a zero value.

[CO011, CO013, CO014, CO015, CO017, CO018]

1.5 Milestones and Adverse Events

LMArena's trajectory spans approximately three years from research demo to unicorn. The original Chatbot Arena launched in May 2023. Academic publications followed rapidly, culminating in the Chatbot Arena paper (arXiv:2403.04132) establishing its Bradley-Terry methodology. By 2024 the platform had become the de facto reference benchmark, cited by every major AI lab. The first material adverse event came in early 2025 when Meta tested at least 27 private Llama 4 model variants on Chatbot Arena, submitting an arena-optimized version that scored near the top while the publicly released version ranked 32nd. LMArena apologized and updated leaderboard policies. In April 2025, a peer-reviewed paper titled "The Leaderboard Illusion" (Cohere, Stanford, MIT, Ai2) formally documented alleged systematic bias in LMArena's private testing practices. LMArena co-founder Ion Stoica publicly disputed the findings as "inaccuracies." LMArena introduced new sampling algorithms and published updated transparency policies in response. The company incorporated in April 2025, raised $100 million in May 2025, launched its commercial product in September 2025, and closed the Series A in January 2026, achieving unicorn status. No regulatory investigations, litigation, sanctions, or enforcement actions have been publicly disclosed as of June 2026. [CO001, CO002, CO005, CO016, CO027, CO028]

Milestone table
DateEventTypeAmount / StatusParticipantsImplication
May 2023Chatbot Arena launched as public research demofoundingVolunteer research projectUC Berkeley LMSYS team (Angelopoulos, Chiang, et al.)First public crowdsourced LLM leaderboard; rapid organic adoption within tech community
Jun 2023MT-Bench and Chatbot Arena NeurIPS paper submittedproductResearch publication (arXiv:2306.05685)Zheng, Chiang, Angelopoulos, et al.LLM-as-Judge concept established; 30K conversations and 3K expert votes released publicly
Mar 2024Chatbot Arena formal platform paper publishedproductResearch publication (arXiv:2403.04132)Chiang, Zheng, et al.; 240K+ votes amassedBradley-Terry/Elo methodology formalized; cited by all major AI labs in product announcements
Jun 2024Arena-Hard-Auto and BenchBuilder pipeline paper publishedproductResearch publication (arXiv:2406.11939)Li, Chiang, et al.Automated benchmark curation at ~$20/run; 98.6% correlation with human preferences demonstrated
Jan–Mar 2025Meta tested ≥27 Llama 4 model variants privately on Chatbot ArenaadverseUndisclosed private testing; optimized-only variant submittedMeta AI, LMArenaFirst major integrity controversy; the publicly released Llama 4 Maverick ranked 32nd vs. 2nd for arena-optimized version
Apr 18, 2025Arena Intelligence Inc. incorporatedfoundingCompany formationAngelopoulos, Chiang, StoicaFormal transition from academic project to commercial entity; fundraising process begins
Apr 29, 2025The Leaderboard Illusion paper publishedadversePeer-reviewed paper (arXiv:2504.20879)Singh, Hooker, et al. (Cohere, Stanford, MIT, Ai2)Formal academic challenge to benchmark integrity; data access asymmetries documented; major credibility risk
May 2025$100M seed round closed at $600M valuationfinancing$100M raised; $600M post-moneya16z, UC Investments, Lightspeed, Felicis, Kleiner PerkinsFirst major venture round; unicorn-adjacent at seed; signals VC conviction on AI evaluation market
Sep 2025AI Evaluations commercial product launchedproductCommercial launchLMArena team; OpenAI, Google, xAI as anchor customersRevenue generation begins; validates B2B evaluation-as-a-service model
Dec 2025ARR consumption run rate surpassed $30Mscale$30M annualized run rateLMArena commercial teamSub-four-month ramp from zero to $30M annualized; rapid enterprise adoption signal
Jan 2026$150M Series A at $1.7B valuationfinancing$150M raised; $1.7B post-moneyFelicis (lead), UC Investments (co-lead), a16z, LDVP, Kleiner Perkins, Lightspeed, The House Fund, Laude VenturesUnicorn status achieved; total raised reaches ~$250M; ~7 months from product launch to Series A
Apr 2026Transparency policy updated; Arena-Rank open-sourcedgovernancePolicy document published at arena.ai/blog/policy/Arena teamResponse to ongoing benchmark integrity criticism; sampling rules and data-sharing commitments formalized
Jun 2026Agent Arena launched with causal inference methodologyproductNew product verticalArena teamExtends evaluation TAM into agentic AI using treatment effect estimation; first major post-Series-A product release

Event dates derived from press releases and publication timestamps; day-level precision varies. Adverse events are included per diligence scope. Amount/Status reflects primary-source figures; undisclosed amounts are labeled as such.

[CO001, CO002, CO005, CO010, CO011, CO016]
FO001: LMArena Corporate Milestone Timeline

Key milestones from the 2023 research launch through June 2026, including financing events, product launches, and adverse events.

Dates approximated from press releases and publication timestamps; day-level precision varies by source.

[CO001, CO005, CO010, CO011, CO015, CO016]

1.6 Exhibits

Chapter 02

02Market Analysis

2.1 Market Boundary and Definition

LMArena should be analyzed as part of the AI evaluation and benchmarking software layer rather than as a general model-infrastructure or observability company. The included market spans human-preference benchmarking, automated regression testing, domain-specific evaluation workflows for regulated or expertise-heavy use cases, and newer agent-evaluation products that score multi-step task success instead of only single-turn chat quality. Excluded from the serviceable market are foundation-model training compute, generic MLOps orchestration, broad developer tooling, and open-source evaluation scripts that are not sold as managed workflow software. That boundary matters because some publishers size a broad evaluation-platform category, while others size a narrow benchmarking-tools niche, producing multi-fold TAM divergence. LMArena’s own commercial messaging emphasizes law, medicine, and engineering evaluation, which suggests its monetizable market is defined more by high-stakes validation spend than by public leaderboard traffic alone.[CM006, CM007, CM008, CM009, CM016, CM018]

Market definition table
CategoryIncluded spendExcluded spendPrimary buyersWhy it matters to LMArena
Public benchmark and leaderboard operationsHuman-preference benchmarking, side-by-side model comparison, public trust signals, benchmark sponsorshipCore training compute, raw inference spend, generic traffic monetizationFrontier labs, model API vendors, benchmark sponsorsThis is the reputational wedge that created LMArena's market visibility
Enterprise model evaluation workflowsRegression testing, eval datasets, domain scoring, release gating, QA dashboardsGeneric BI, ticketing, or unrelated developer productivity toolsAI platform teams, model quality leads, domain product ownersThis is where recurring software spend is most likely to persist beyond public leaderboard usage
Agent and workflow evaluationMulti-step task success scoring, causal traces, workflow benchmarks, tool-use reliabilityGeneral agent orchestration without measurement, generic copilotsApplied AI teams building agents into workflowsAgent Arena expands the market beyond chatbot ranking into execution reliability
Regulated and high-stakes vertical validationLegal, medical, engineering, and compliance-sensitive evaluation projectsConsumer entertainment chat ranking with no business-critical decision pathDomain leaders, risk owners, quality teamsLMArena explicitly markets these domains as monetizable beachheads
Bundled platform evaluationEvaluation features embedded in clouds, MLOps suites, and experiment platformsN/AAWS, Google, Microsoft, Databricks, MLflow usersThis spend exists in the ecosystem but is only partly accessible to a standalone vendor
Adjacent but excluded infrastructureTraining data pipelines, foundation-model hosting, inference serving, generic observabilityAll of these remain outside the scoped evaluation software marketInfra teams and CTO budgetsExcluding these prevents overstatement of the addressable market

Boundary rows are intentionally scoped to monetizable evaluation software rather than all AI tooling; bundled platform evaluation is included as context but only partly addressable by LMArena.

[CM006, CM007, CM008, CM009, CM016, CM018]

2.2 Market Sizing and Contested Estimates

The cleanest broad 2026 market anchor in the verified source set is the AI model evaluation platform market report showing $1.86B in 2025 growing to $2.36B in 2026 at 27.3% CAGR, and $6.24B by 2030. That lens likely includes enterprise evaluation software sold across model labs, application developers, and compliance-heavy organizations. A narrower lens from Precedence Research implies a materially smaller benchmarking-tools segment, around $0.85B in 2026, because it appears to exclude some bundled or adjacent workflows and emphasizes evaluation plus benchmarking tools rather than the full platform layer. Gartner’s 2026 AI spending forecast and Presenc AI’s production-adoption survey both support the idea that evaluation budgets can expand quickly from here, but they do not resolve the boundary problem. The practical conclusion is that LMArena’s real monetizable market is likely much smaller than the broad TAM yet still large enough to support a meaningful standalone company if it captures trusted, workflow-embedded spend.[CM001, CM002, CM003, CM004, CM005, CM021]

TAM/SAM/SOM or sizing lens table
SourceYearScopeMetricValueGrowth / outlookInterpretationKey limitation
The Business Research Company2026Global AI model evaluation platforms2025-2026 market size$1.86B (2025) to $2.36B (2026)27.3% CAGR; $6.24B by 2030Best verified broad platform TAM anchor in this chapterMethodology details are summarized, not fully disclosed
Yahoo Finance / Research & Markets syndication2026Global AI model evaluation platforms2026 market size$2.36B in 202627.3% CAGRIndependent syndication broadly corroborates the broad-market figureSyndicated press coverage, not original model workbook
Research & Markets2026Global AI model evaluation platforms2030 forecast$6.24B by 203027.3% CAGRConfirms strong forward growth if broad definition holdsStill a broad category with unclear segment decomposition
Precedence Research2026Model evaluation and benchmarking toolsNarrow 2026 lens~$0.85B in 2026~7.3% CAGRUseful lower bound for a narrower benchmarking-tools categoryNot directly comparable to broad platform TAM; category appears narrower
Precedence Research2025-2034Model evaluation and benchmarking toolsLong-range forecast$9.57B by ~2034Long-horizon growth marketShows category importance and strategic value, including M&A contextForecast horizon and scope differ from 2026 broad-market sources
Gartner2026Worldwide AI economyAI software spend$453B AI software in 2026Part of $2.59T worldwide AI spend, +47% yoyUpstream budget pool supporting evaluation purchases is very largeNot a direct evaluation-market measure
Gartner2026Worldwide AI economyAI models spend$32.6B AI models in 2026Rapid expansion alongside software spendHelpful proxy for frontier-lab budget availabilityStill indirect; does not isolate third-party evaluation vendors
Author estimate from verified sources2026LMArena-relevant standalone evaluation SAM / SOMAnalytical rangeSAM ~$0.3-0.8B; SOM ~$0.03-0.15BDerived from broad TAM, narrow lens, and reported ARRUseful decision range for diligence, not a publisher-backed estimateNo public pricing, budget-share, or customer-count detail to validate precisely

This table mixes publisher numbers and one clearly labeled analytical estimate; users should compare definitions before treating rows as directly additive or contradictory.

[CM001, CM002, CM003, CM004, CM005, CM021]
FM001: Market sizing lens

Three-tier lens from broad AI evaluation platforms to LMArena-relevant SAM and near-term SOM.

SAM and SOM are analytical ranges derived from the verified broad TAM, narrow benchmarking lens, and reported ARR; they are not publisher-issued numbers.

[CM001, CM022, CM023, CM043]
FM002: Market estimate range

Low, midpoint, and high market lenses showing how category definition changes the apparent size of LMArena's market.

All rows are expressed in USD billions; the high row combines 2030-2034 directional category endpoints rather than a single year.

[CM001, CM003, CM046]

2.3 Buyer Segmentation and Procurement

LMArena’s buyer map has two distinct centers of gravity. First are frontier labs and model API vendors that need credible third-party benchmarking, launch validation, and competitive signaling before or after major releases. Second are regulated or expertise-heavy enterprises that need domain-specific evaluation for legal, medical, and engineering workflows where failure costs are high and public consumer benchmarks are insufficient. These segments likely purchase differently: labs fund evaluation from central model, safety, or research-platform budgets, while enterprises often buy through AI-platform leaders, product owners, or risk and quality teams. Competition is also mixed. Scale AI, Arize, Galileo, Patronus, MLflow, and large cloud or platform vendors all attack adjacent pieces of the stack, and CoreWeave’s acquisition of Weights & Biases shows that experiment, observability, and evaluation workflows are increasingly converging in enterprise procurement.[CM010, CM011, CM012, CM013, CM014, CM015]

Segment / buyer map
SegmentBuyerUserPayerWorkflowBudget ownerAdoption trigger
Frontier model labsOpenAI, Google, xAI, Anthropic-like labsEvaluation researchers, release managers, safety teamsCentral model-development budgetBenchmark new models before and after release; compare against rivalsVP model quality / research platform leadCompetitive release cadence and need for credible external proof
Model API vendors and platform providersHosted model platforms and AI cloudsPlatform PMs, trust teams, GTM teamsProduct or platform budgetUse external and internal evaluations to support enterprise sales and launch claimsPlatform GM or product leadNeed to differentiate model quality in crowded API markets
Legal AI vendors and enterprisesLegal workflow teams, counsel-tech buyersAttorneys, reviewers, AI product managersBusiness-unit or innovation budgetValidate domain accuracy, citation quality, and workflow reliabilityHead of legal innovation or AI platform leadHigh cost of hallucinations in legal workflows
Medical and health-related AI buyersClinical AI teams, medical documentation vendorsClinicians, quality teams, model validatorsProduct, compliance, or clinical-ops budgetEvaluate safety, terminology accuracy, and failure thresholdsChief medical AI lead or quality ownerPatient-safety and compliance stakes make validation spend easier to justify
Engineering copilots and industrial knowledge workflowsEngineering software teams and applied-AI groupsEngineers, analysts, technical reviewersR&D or product engineering budgetTest task completion, tool use, and domain correctnessHead of applied AI or engineering systemsNeed for measurable productivity gains before broad rollout
Benchmark ecosystem partnersSponsors, evaluators, and adjacent tool vendorsMarket-facing research, developer-relations, trust teamsMarketing, product, or ecosystem budgetUse benchmarks to shape narrative, partnership motions, or integrated workflowsProduct marketing or ecosystem leadNeed to anchor category credibility and public comparability

Buyer rows reflect the highest-likelihood paid segments supported by verified product messaging and media reporting rather than a complete customer roster.

[CM010, CM012, CM013, CM014, CM015, CM016]
FM003: Buyer / segment map

Shows how major buyer segments evaluate LMArena against different purchasing criteria.

Ordinal labels summarize the relative importance of each criterion for each segment, not survey scores.

[CM015, CM017, CM018, CM044]

2.4 Growth Drivers and Adoption Constraints

Several forces support above-market growth for trusted evaluation vendors. Global AI software and model spending are surging in 2026, enterprise AI has moved into production at a high share of large companies, and the shift toward agents increases the need for workflow-level evaluation rather than static prompt testing. Benchmark rivalry itself also stimulates demand, because model providers want external proof points and enterprises want independent quality signals. Yet the category also has real constraints. Benchmark gaming allegations, questions about crowdsourced-rater bias, and skepticism about public leaderboards all undermine willingness to pay for scores that are not clearly tied to production outcomes. Standalone vendors may also face pricing pressure from bundled cloud and MLOps tools, while limited public pricing disclosure makes it hard to separate durable software demand from novelty-driven experimentation.[CM026, CM027, CM028, CM029, CM030, CM031]

Growth drivers and constraints table
FactorTypeTimingImplicationDiligence ask
Enterprise AI production adoption at scaledriver2026-nowMore production workloads create recurring need for regression, governance, and release testingWhat portion of production AI teams currently buy third-party evaluation rather than build in-house?
Exploding upstream AI software and model spenddriver2026-2030Evaluation budgets can grow as a small but expanding percentage of a much larger AI stackCan management show expansion of spend per existing customer as AI programs mature?
Agentic AI and workflow automationdriver2026-2030Multi-step agents need workflow and causal evaluation, expanding beyond chatbot rankingHow much revenue is tied to agent evaluation versus classic leaderboard workflows?
Benchmark rivalry among frontier labsdrivercurrentLaunch competition increases demand for credible third-party measurement and narrative controlWhich buyer segment uses LMArena primarily for external signaling versus internal QA?
Benchmark gaming allegationsconstraintcurrentTrust erosion can make public scores less monetizable unless tied to controlled enterprise workflowsWhat anti-gaming controls and audit trails does LMArena offer paying customers?
Crowdsourced-rater bias and representativeness criticismconstraintcurrentOpen-arena votes may not satisfy regulated buyers that need domain-grounded evaluationWhat share of enterprise evaluations use curated expert raters or private datasets?
Bundling pressure from clouds and MLOps platformsconstraint2026-2028Standalone vendors may face lower pricing power if evaluation becomes a feature not a categoryWhere does LMArena win against bundled MLflow, hyperscaler, or observability workflows?
Sparse public pricing and contract disclosureconstraintcurrentOutside investors cannot independently translate TAM into forecastable revenue captureCan management disclose ACV bands, renewal rates, and seat or usage expansion dynamics?

Drivers and constraints are directional, not weighted; several could be positive for category demand while negative for standalone vendor economics at the same time.

[CM026, CM027, CM028, CM029, CM030, CM031]
FM004: Adoption funnel or value-chain map

Five-stage funnel from experimentation to recurring governance spend for evaluation software.

Values are indexed to show narrowing from experimentation to durable recurring spend; they are not measured conversion rates.

[CM026, CM027, CM045]

2.5 Evidence Gaps and Contradictions

The main diligence problem is not source scarcity but source mismatch. Public market reports disagree on the category boundary; company and media sources disclose valuation and anecdotal customer names but not enough contract detail to model share; competitor-revenue data are fragmentary; and no verified source in this run discloses what percentage of frontier-lab or enterprise AI budgets is actually spent on evaluation software. That means the chapter can support a credible range-based view but not a single precise SAM or market-share conclusion. Investors should preserve the contradiction instead of forcing a point estimate: LMArena may already be large relative to a narrow benchmarking-tools niche, while still tiny relative to the broader AI evaluation platform opportunity. Resolving that gap requires customer, pricing, and budget-intensity evidence that is not public in the verified source set.[CM003, CM024, CM038, CM039, CM040, CM041]

2.6 Exhibits

Chapter 03

03Competitors

3.1 Competitive landscape: direct benchmarkers, automated leaderboards, data labelers, and build-it-yourself substitutes

LMArena operates in the AI model evaluation and benchmarking space, where no single competitor replicates its full stack. The addressable competitive set has four tiers. First, human-preference platforms: LMArena is the dominant public platform for crowdsourced pairwise model comparison; no comparable free alternative exists at anything close to its 5 million monthly user scale or 60 million monthly conversation volume as of January 2026. Second, automated academic leaderboards: the HuggingFace Open LLM Leaderboard (powered by EleutherAI lm-evaluation-harness), Stanford HELM, and BenchLM aggregate standardized benchmark scores such as MMLU, GPQA Diamond, and SWE-Bench across hundreds of models without human preference voting. These tools are free and open-source but measure task accuracy, not holistic user satisfaction. Third, enterprise evaluation vendors: Scale AI's GenAI Platform offers customized evaluation, fine-tuning, and data-labeling pipelines for enterprise and government clients at $93 K–$400 K+ per engagement; it does not host a public leaderboard. Fourth, status-quo substitutes: AI labs can build internal evaluation pipelines using EleutherAI's harness or OpenAI Evals and run their own test sets, avoiding third-party dependency entirely. Likely entrants include any major cloud provider building a benchmark surface to differentiate its model marketplace, or a large AI lab spinning up an in-house neutral evaluation arm. The meaningful competitive dimensions are human-preference scale, enterprise revenue potential, third-party independence, and methodology credibility.[CP001, CP002, CP003, CP011, CP012, CP014]

Competitor profile table
competitorcategoryscale / fundingtarget segmentdifferentiationlimitation
LMArenaHuman-preference leaderboard + enterprise evaluation$250 M raised, $1.7 B valuation (Jan 2026); 5 M monthly usersAI labs (benchmark marketing), enterprises (model selection), researchersLargest public human-preference dataset; cross-modal; real-time Elo rankingsRevenue from same labs it evaluates; user base skewed to tech professionals
Scale AI GenAI PlatformEnterprise AI evaluation, data labeling, fine-tuning$13.8 B valuation (2024); DoD, Meta, Mayo Clinic customers reportedEnterprises and government seeking private, SLA-backed evaluation pipelinesLargest RLHF data-labeling operation; custom private evaluation; GPU-cluster scaleNo public leaderboard; not neutral toward specific models; pricing opaque
HuggingFace Open LLM LeaderboardAutomated academic benchmark aggregator (open-source)Backed by HuggingFace ($4.5 B valuation); free public toolML researchers, open-source developers, model release teamsReproducible benchmarks; open-weight focus; EleutherAI harness poweredNo human preference; tech-only coverage; no enterprise evaluation service
Stanford HELMMulti-dimensional academic evaluation frameworkStanford CRFM (academic); free tool; no commercial productAI researchers and policy audiences needing multi-axis model assessmentCovers accuracy, calibration, robustness, bias, efficiency simultaneouslyPeriodic rather than continuous; no human preference; no enterprise revenue
EleutherAI lm-evaluation-harnessOpen-source evaluation frameworkCommunity-funded (EleutherAI nonprofit); free; GitHub stars ~70 K+ML researchers, leaderboard operators, enterprise data teams building custom evals60+ standardized academic benchmarks; forkable; underpins HF leaderboardNo human preference; no commercial service; no UI/leaderboard product
OpenAI EvalsLLM evaluation framework (open-source + dashboard)OpenAI (backed internally; no separate funding); free frameworkEnterprises and OpenAI customers building custom eval pipelinesIntegrates directly with OpenAI models; dashboard for enterprise use casesModel-provider-owned; limited independence; focused on OpenAI family
BenchLMAutomated LLM leaderboard aggregatorUnknown funding; free public toolEnterprise and developer model selection; 2026 benchmark tracking261 models, 249 benchmarks; separate verified vs provisional rankings; includes price/speedNo human preference; newer platform with lower brand recognition than HF/Arena
ArtificialAnalysisIndependent AI model and API performance analyticsUnknown funding; free public toolEnterprise API buyers comparing speed, throughput, cost, and intelligenceProvider-agnostic latency, throughput, cost, and intelligence benchmarkingNo human preference; no evaluation-as-a-service; limited model breadth vs HF
[CP001, CP002, CP011, CP012, CP013, CP014]
FP001: Competitive positioning map

LMArena holds a differentiated position on both human-preference scale and enterprise service depth; free academic tools cluster in the open/research quadrant; Scale AI occupies the enterprise-only zone.

Scores are ordinal evidence-backed judgments derived from public product surfaces, papers, and pricing pages; not precisely measured metrics.

[CP024, CP025, CP011, CP012, CP014, CP015]

3.2 Feature and capability comparison: human preference versus automated benchmarks

The core capability divide in AI evaluation is between human-preference leaderboards (LMArena, its predecessors at LMSYS) and automated benchmark pipelines (HuggingFace Open LLM Leaderboard, Stanford HELM, EleutherAI lm-evaluation-harness, OpenAI Evals). LMArena's Elo-based ranking system, described in the 2024 Chatbot Arena paper, aggregates user-submitted pairwise votes across any task the user brings, creating a continuously updated signal that reflects real usage patterns rather than curated test sets. In practice this means LMArena is more sensitive to style, fluency, and helpfulness signals that users perceive in everyday tasks — coding, writing, answering questions — but less reproducible than automated benchmarks because two users can legitimately prefer opposite outputs. Automated leaderboards are reproducible and task-specific. Stanford HELM evaluates across accuracy, calibration, robustness, bias, and efficiency dimensions on a common harness, favoring researchers who need controlled measurements. The HuggingFace Open LLM Leaderboard tracks primarily open-source and open-weight models on standardized academic tests. EleutherAI's harness underlies both and is freely forkable. BenchLM aggregates 249 benchmarks across 261 models as of June 2026 and separately tracks price and speed. ArtificialAnalysis provides independent API throughput, latency, and cost benchmarking. Scale AI's enterprise evaluation product fills a different gap: custom, private, SLA-backed evaluation against a client's specific production use case rather than a public leaderboard — a model that does not compete with LMArena for brand recognition but may compete for enterprise budget. LMArena's key differentiators are network effects (more users → more votes → more statistically stable rankings), cross-modal coverage (text, vision, image generation, video, web dev as of 2025–2026), and the breadth of pre-release partnerships with frontier labs. Its weaknesses are methodological reproducibility, user-base bias toward technically sophisticated early adopters, and the inherent tension between selling evaluation services to the same labs it ranks. No competitor combines the public leaderboard surface with the enterprise paid evaluation service in a single brand, which is both LMArena's structural moat and its conflict-of-interest risk.[CP006, CP007, CP008, CP009, CP013, CP015]

Feature / capability matrix
buying criterionLMArenaScale AIHuggingFace Open LLM LeaderboardStanford HELMEleutherAI HarnessBenchLM / ArtificialAnalysis
Human-preference / Elo rankingstrong (unique at scale)absent (enterprise custom, not Elo)absentabsentabsentabsent
Real-time / continuous updatesstrong (live battles 24/7)unknown (SLA-gated, not public)partial (batch releases)periodic (not continuous)on-demand (researcher-run)partial (updated regularly)
Cross-modal evaluation (vision, image, video)strong (text, vision, image, video, web dev as of 2026)unknown (custom; not public)partial (multimodal in progress)partial (limited visual benchmarks)partial (multimodal prototype)partial (image understanding category)
Open-source / academic benchmark coveragepartial (Arena-Hard, MT-Bench derivatives)absent (private pipelines)strong (MMLU, GPQA, ARC, SWE-Bench)strong (multi-axis academic suite)strong (60+ benchmarks)strong (249 benchmarks)
Enterprise evaluation service (paid, SLA-backed)strong (AI Evaluations product launched Sep 2025)strong (core product; multi-year contracts)absentabsentabsentabsent
Independent / non-provider-ownedmedium (earns revenue from evaluated labs; academic roots)medium (data-labeling clients overlap with evaluation clients)strong (HuggingFace is neutral platform)strong (Stanford academic)strong (nonprofit)strong (independent analytics)
Public leaderboard / transparent methodologystrong (open methodology papers; public Elo scores)absent (enterprise-only; private results)strong (open benchmarks, reproducible)strong (multi-axis public results)strong (open-source; reproducible)strong (public leaderboard with separate verified/provisional rankings)
Developer signal (GitHub stars / community)strong (FastChat 38 K+ stars; platform-native community)low (enterprise-first brand; limited open-source presence)strong (HF spaces community; broad ML ecosystem)medium (academic adoption; researcher citation)strong (70 K+ GitHub stars; powers HF leaderboard)medium (growing; no open-source repo)

Capability levels (strong/medium/partial/absent/unknown) are ordinal judgments derived from publicly accessible product surfaces, papers, and benchmark documentation as of June 2026; cells marked unknown reflect private enterprise products where the capability exists but is not publicly verified.

[CP006, CP007, CP008, CP013, CP015, CP017]
Pricing / packaging comparison
platformpricing modellist price / contract rangeincluded capabilitiespublic pricing availabilityimplication for buyer
LMArena AI EvaluationsEnterprise contract (usage-based consumption)Not publicly disclosed; annualized consumption run rate $30 M across ~100 enterprise clients implies avg ~$300 K/client (estimated, not stated)Custom evaluation panels, community feedback data, SLA delivery, analyticsNone (contact sales)Enterprise can pay for priority evaluation but list price is opaque
Scale AI GenAI PlatformEnterprise custom contractAnalyst reports indicate $93 K–$400 K+ per engagementData labeling, model fine-tuning, agent deployment, evaluation pipelinesNone (contact sales)Higher price but private, confidential results; no public leaderboard dependency
HuggingFace Open LLM LeaderboardFree (HuggingFace-subsidized)$0Open benchmark submission, public scores, reproducible test setsFull publicNo cost but results are public; no customization for specific enterprise use cases
Stanford HELMFree (Stanford CRFM-subsidized)$0Multi-axis automated evaluation, public resultsFull publicAcademic rigor, multi-axis coverage, but no human preference or commercial SLA
EleutherAI lm-evaluation-harnessFree (open-source)$0 (infrastructure costs self-hosted)60+ benchmark tasks; self-hosted or cloud-runFull publicMaximum control but requires engineering to operate; no UI or leaderboard
OpenAI EvalsFree framework; pay-per-API for GPT-4o judging$0 for framework; API costs for judge model runs (~$5–30/M output tokens)Custom eval templates, dashboard, community eval registryFull public (framework); API pricing publicDeep OpenAI integration but model-provider-owned; limited for non-OpenAI model comparison
BenchLM / ArtificialAnalysisFree (ad-supported analytics sites)$0LLM leaderboard, benchmark aggregation, speed/cost/intelligence comparisonsFull publicGood for model selection research; no enterprise service or custom evaluation

LMArena per-client average ($300 K) is an estimate derived by dividing the $30 M annualized consumption run rate by ~100 clients; this is not a company-stated figure. Scale AI pricing is from third-party analyst aggregations, not a public price sheet.

[CP036, CP037, CP038]
FP002: Feature breadth / capability map

LMArena leads uniquely on human-preference scale and cross-modal real-time coverage; Scale AI leads on private enterprise depth; free tools lead on reproducibility and open-source signal.

Ordinal capability levels (strong/medium/partial/absent/unknown) are evidence-based judgments from publicly reviewed product surfaces and academic papers; cells marked absent reflect confirmed absence of the capability in the publicly reviewed product surface.

[CP006, CP007, CP017, CP018, CP019, CP023]

3.3 Moat durability, switching costs, and adverse competitive evidence

LMArena's competitive durability rests on three factors: a crowdsourced preference dataset that is the largest of its kind, community trust reinforced by academic roots (UC Berkeley LMSYS Org, FastChat open-source infrastructure), and deep citation network effects — frontier labs cite Arena rankings in marketing materials and investor communications, raising the cost of defection. Switching costs for AI labs are real: if Arena ceases to be the market's signal, their existing benchmark investments lose marketing value. This creates a mutual dependency between LMArena and its client labs that provides revenue stability but also structural independence risk. The adverse evidence on moat is credible and material. The 2025 Leaderboard Illusion paper by Singh et al. (Cohere, Stanford, MIT, Ai2) quantifies data access asymmetry: Google and OpenAI received an estimated 19.2 % and 20.4 % of total Arena data respectively, while 83 combined open-weight models received only 29.7 %. Meta tested 27 private model variants before the Llama 4 release, selecting only the top scorer, which TechCrunch confirmed then ranked 32nd once the vanilla Maverick was re-submitted. LMArena denied the characterization of bias but committed to algorithm changes, signaling the critique had operational merit. A second adverse dynamic is Goodhart's Law: as Arena rankings drive procurement decisions, labs will optimize for Arena-specific patterns, degrading the signal quality and increasing the likelihood a well-resourced competitor can credibly claim superior methodology. The Elo algorithm underlying Arena is open-source and replicable by any well-funded team. What is not easily replicated is the community. Scale AI has a larger enterprise footprint and deeper data-labeling expertise, but has not built a public preference leaderboard. The real displacement threat is internal build: an AI lab that decides Arena's conflict-of-interest problem is unsolvable could fund a competing neutral body. The free academic tools (HELM, HuggingFace) already serve that credibility function for research use cases, keeping a ceiling on how much LMArena can claim scientific neutrality.[CP028, CP029, CP030, CP034, CP035, CP036]

Moat durability / competitive risk register
moat claimthreatseveritymitigation / diligence ask
Human-preference dataset is largest and most cited in the field (6 M+ votes, 60 M conversations)A well-funded lab or consortium could build a competing human-preference platform with equivalent user incentives within 18–24 monthshighConfirm whether LMArena's dataset is proprietary or community-owned; assess data-sharing agreements with labs
Network effects: top labs release models on Arena because community demands itLabs that feel gamed or commercially disadvantaged may defect and fund a rival neutral bodymediumTrack whether any major lab has publicly reduced Arena engagement or launched competing initiative
Academic credibility (UC Berkeley/LMSYS origin; 9 published papers)The Leaderboard Illusion paper (Singh et al., 2025) erodes credibility; further studies could accelerate reputational damagehighReview LMArena's methodological response to selective-disclosure critique; confirm algorithm reform implementation
Revenue from enterprise AI Evaluations creates paid switching costs for lab clientsRevenue from the same labs it evaluates creates structural conflict of interest that undermines independence moathighRequest a structural independence policy (firewalls, editorial independence) from management; review whether top clients can influence rankings
Open-source FastChat infrastructure and open data releases sustain developer goodwillFree automated benchmark tools (EleutherAI, HF, HELM) serve credibility function for academic users, capping LMArena's academic moatlowEvaluate whether LMArena retains academic partnership (papers, citations) as enterprise revenue grows
Platform position as de-facto standard used for model launch benchmarkingGoodhart's Law: as Arena becomes the target, labs over-optimize for Arena-specific patterns, degrading signal quality and opening methodological attacksmediumMonitor whether future model releases specifically cite Arena scores in press releases, and track any public evidence of Arena-specific tuning

Severity (high/medium/low) reflects qualitative judgment based on available public evidence of threat materiality as of June 2026; not a numerical risk score.

[CP028, CP029, CP030, CP033, CP034, CP035]
FP003: Moat / readiness KPIs

LMArena's moat is anchored on data and community scale; the primary vulnerabilities are methodological integrity and commercial conflict of interest.

[CP001, CP002, CP004, CP023, CP029, CP033]

3.4 Exhibits

Chapter 04

04Financials

4.1 Revenue model, pricing, and commercial traction

LMArena's revenue model is a freemium-to-enterprise funnel. The free public leaderboard — which serves 5 million monthly users across 150 countries and generates 60 million model-comparison conversations per month — acts as the community and trust-building layer. Enterprises, model labs, and AI developers pay for LMArena's commercial AI Evaluations product (launched September 2025), which provides custom evaluation panels grounded in real user feedback, representative data samples, and SLA-committed delivery timelines. LMArena has not published a public price list; all enterprise contracts are negotiated directly. Three confirmed enterprise client categories are AI labs (OpenAI, Google, xAI cited explicitly in the January 2026 Series A press release), software enterprises, and regulated professional verticals (law, medicine, scientific research). The company's annualized consumption run rate surpassed $30 million in December 2025 — a pace set less than four months after product launch. Revenue recognition is an important caveat: LMArena and TechCrunch both describe this figure as a "consumption run rate," not a realized annual figure. TechCrunch notes the company "describes its annual recurring revenue (ARR)" as a consumption rate, reflecting usage-based billing rather than upfront contracted SaaS ARR. With approximately 100 enterprise clients (per GetLatka aggregation), the implied average contract value is approximately $300,000 per client — an estimate derived from dividing the run rate by the client count rather than a company-disclosed figure. Actual ACV distribution, contract duration, and renewal rates are not publicly available.[CI001, CI002, CI003, CI004, CI005, CI006]

Revenue streams table
streammechanismunitcurrent value / statusqualitydiligence ask
AI Evaluations enterprise servicePaid custom evaluation panels using community human feedback, delivered under SLAEnterprise contract (consumption-based)$30 M annualized consumption run rate as of Dec 2025 (company-stated); ~100 clients (analyst estimate)medium (consumption run rate ≠ contracted ARR; no churn, NRR, or retention data)Obtain contracted ACV distribution, renewal rates, NRR, and whether billing is time-boxed or usage-based
Data and analytics licensing (potential)Selling aggregated preference data or model performance insights to enterprises and researchersPer-dataset or subscription (unconfirmed)Not confirmed as a separate revenue line; company has released free datasetslow (no confirmed revenue from this stream)Confirm whether any commercial data licensing agreements exist beyond the enterprise evaluation product
Partnership evaluation fees (potential)Fees from AI labs for early/priority access to pre-release model evaluation slotsPer-evaluation or included in enterprise contracts (unconfirmed structure)Not separately disclosed; may be bundled into AI Evaluations contractslow (no separate public disclosure; cannot disaggregate from enterprise line)Confirm whether pre-release evaluation slots carry separate pricing or are bundled; clarify revenue recognition

Current value for AI Evaluations is the "annualized consumption run rate" as of December 2025; this is not a realized 12-month revenue figure. Customer count (~100) is from GetLatka aggregation, not company-disclosed.

[CI001, CI002, CI003, CI004, CI007, CI008]
Pricing / monetization table
product / tierpricing modellist price / contract rangeincluded capabilitiespublic price availabilitysource
Free public leaderboardFree (community-subsidized)$0Head-to-head model battles, Elo rankings, open data releases, multi-modal coverageFully publicarena.ai (official)
AI Evaluations — enterpriseConsumption-based contract (contact sales)Not publicly disclosed; ~$300 K implied avg (analyst estimate, not confirmed by company)Custom evaluation panels, community feedback data, SLA delivery, analytics dashboardsNone (contact evaluations@lmarena.ai)arena.ai/blog/ai-evaluations/ (official)
Open data for researchFree$01.5 M+ community prompts, 145 K+ battle data points (open-source)Fully public (HuggingFace)arena.ai/blog/two-year-celebration/ (official)
Pre-release model evaluations (partner labs)Bundled or negotiated fee (unclear)Not publicly disclosedPriority evaluation slot, pre-release score visibilityNoneinferred from TC and PRNewswire coverage
Academic / open-source tierFlexible pricing per company commitmentUndisclosed (discount from enterprise rate)Access to evaluation community and basic analyticsNone (company commitment to support nonprofits)arena.ai/blog/ai-evaluations/ (official)
[CI006, CI007, CI008, CI035]
FI001: Revenue model bridge

Free community activity drives preference-data value, which converts to enterprise evaluation revenue through the AI Evaluations product.

Revenue, margin, and reinvestment figures are approximate; gross margin is unconfirmed. Run rate is annualized consumption rate (company-stated) rather than contracted ARR.

[CI001, CI002, CI007, CI008, CI030]

4.2 Cost structure, unit economics, and capital adequacy

LMArena's cost structure has three primary categories inferred from public evidence. First, compute and infrastructure: serving 60 million conversations per month across frontier models requires substantial cloud compute. While the public leaderboard uses model APIs from partner labs (which may be provided at reduced or no cost under partnership agreements), the enterprise evaluation product likely incurs real inference costs at scale. Second, engineering and research headcount: the company had 41 employees as of the January 2026 Series A announcement. At typical Silicon Valley all-in compensation for a technical startup, this implies annual headcount costs of roughly $15–25 million. Third, community management, data annotation quality assurance, and go-to-market costs that are impossible to size from public information. Gross margins are not disclosed; a software-like model with low marginal cost per evaluation could achieve 60–80% gross margins, but high compute costs per conversation could compress this to 40–60%. No third-party estimate with strong foundation exists. Capital adequacy appears strong in the near term. The company has raised $250 million total across its May 2025 seed round ($100 million at $600 million valuation) and January 2026 Series A ($150 million at $1.7 billion valuation). With a lean team of 41 employees and first commercial revenue only established in September 2025, the cash burn is likely well below $30 million annually as of early 2026. Rough estimates suggest a runway of at least 18–36 months without further fundraising, but no official burn rate or cash position has been disclosed. LMArena stated it will use the Series A funds to expand its technical team, strengthen research capabilities, and build new platform features — all of which will increase future burn. The Recall Capital-LMArena feeder fund (SEC Form D, filed February 2026) confirms the capital-formation activity but does not disclose LMArena's own balance sheet. No debt or project-finance obligations have been reported publicly.[CI009, CI010, CI011, CI012, CI013, CI014]

Unit economics table
metricvalue / nullconfidencewhy it mattersdiligence ask
Average contract value (ACV)~$300 K (estimated: $30 M run rate ÷ ~100 clients)low (derived estimate; neither input is confirmed by company)Drives revenue scalability and sales efficiency assumptionsConfirm actual ACV distribution, enterprise contract terms, and whether client count is accurate
Gross marginn/a (not disclosed)Key to underwriting; software-like models should exceed 60%; heavy compute could compress to 40-60%Request gross margin in management accounts; confirm whether partner API costs are subsidized
Net revenue retention (NRR)n/a (not disclosed)For consumption-based SaaS, NRR distinguishes growing accounts from one-time evaluationsRequest cohort-level NRR; clarify whether $30 M run rate reflects expanding or stable accounts
Customer acquisition cost (CAC)n/a (not disclosed)Informs unit economics viability and GTM efficiencyRequest CAC by segment (AI labs vs. enterprise); confirm whether free leaderboard is primary acquisition channel
Sales cycle lengthn/a (not disclosed)Enterprise evaluation contracts for large AI labs may involve multi-month procurementRequest average time from initial contact to signed contract; flag if lab budget cycles affect timing
Payback periodn/a (not disclosed)Depends on ACV and CAC; without either confirmed, cannot estimateDerive from ACV and CAC once confirmed; flag if payback exceeds 18 months
Monthly burn ratenull (estimated $1.5–2.5 M/month based on ~41 employees at $300 K fully-loaded avg + infrastructure)low (back-of-envelope; not company-disclosed)Sets minimum cash requirements and fundraising triggerRequest actual monthly P&L; confirm infrastructure costs and any one-time charges
Cash runway (from Jan 2026)null (estimated 24–36 months at estimated burn vs. $250 M raised)low (depends on unconfirmed burn)Determines if next fundraise is near-term necessity or opportunisticObtain cash position at Series A close and current monthly cash outflows

All null values represent metrics not publicly disclosed; estimates (ACV, burn, runway) are back-of-envelope derivations for diligence framing only. Low-confidence values must be confirmed before underwriting.

[CI004, CI005, CI014, CI015, CI017, CI018]
Capital adequacy table
itemvaluesource / confidenceimplication
Seed round (May 2025)$100 M at $600 M post-money valuationhigh (TechCrunch, PRNewswire confirmed)Established commercial runway; led by a16z and UC Investments
Series A (Jan 2026)$150 M at $1.7 B post-money valuationhigh (PRNewswire official press release, TechCrunch confirmed)Primary near-term capital base; use: team expansion, research, platform features
Total raised$250 Mhigh (sum of two confirmed rounds; no bridge or convertible note publicly disclosed)Substantial capital for a lean 41-person team; runway likely 24–36+ months at current scale
Monthly burn estimate$1.5–2.5 M/month (back-of-envelope)low (not disclosed; derived from headcount and infrastructure assumptions)At $2 M/month burn, $250 M implies 10+ years without revenue — conservative; actual burn may be higher
Cash on hand (Jan 2026)Not disclosedn/aNeed management confirmation; relevant for planning headcount expansion and product build-out
Debt / project-finance obligationsNone publicly reportedmedium (no public filings or announcements found)Clean capital structure assumed; verify at due diligence
Next-round triggerNot disclosed; no public statements about next fundraise timelinen/aAt current burn and run rate trajectory, Series B likely 18–30 months from Jan 2026
Planned use of Series A fundsExpand technical team, strengthen research capabilities, build new features (company-stated)medium (company-stated; no budget breakdown)Headcount will increase burn; need updated projections post-expansion

Monthly burn is a rough estimate; actual burn depends on headcount growth pace, infrastructure scaling, and data quality investment. Series B timeline estimate is illustrative, not a company statement.

[CI009, CI010, CI011, CI012, CI013, CI018]
FI002: Unit economics bridge

Enterprise contracts flow through panel assembly and inference cost pools before generating gross profit; all margin inputs remain private.

ACV ($300 K) is an estimate. Gross margin is unknown; the range reflects software-like upper bound vs compute-heavy lower bound. All unit economics require management disclosure to confirm.

[CI005, CI015, CI016, CI017, CI028]
FI004: Capital intensity / cash-flow map

$250 M raised against estimated annual burn of $20–45 M implies strong near-term capital adequacy, but cost assumptions are unconfirmed.

All cost items are estimates derived from headcount and infrastructure benchmarks; the $225 M net cash estimate is illustrative, not a company-stated figure. Revenue is the run rate (annualized consumption), not the realized 2025 figure.

[CI009, CI010, CI011, CI014, CI018, CI019]

4.3 Financial risks, conflict of interest, and diligence blockers

The most material financial risk is structural: LMArena earns revenue from the same AI labs — OpenAI, Google, xAI — that it publicly ranks on its leaderboard. CTOL Digital documented this paradox explicitly after the January 2026 funding: the $1.7 billion valuation is "at 57 times that run rate, pricing in not just growth, but the assumption that this inherent conflict can be managed indefinitely." If enterprise clients come to believe that LMArena's rankings are influenced by commercial relationships (as the Leaderboard Illusion paper alleged for data-access asymmetry), the evaluation service's pricing power erodes and the free leaderboard's trust advantage — which is the core marketing asset — collapses simultaneously. This dual exposure is unusual: most SaaS companies face customer churn risk; LMArena faces credibility churn risk where the same event that loses a client also damages the free product that generates new clients. A second risk is revenue concentration: with approximately 100 enterprise clients and a run-rate dominated by three major AI labs, any single-client defection from OpenAI, Google, or xAI could materially reduce revenue. A third risk is the "consumption rate" framing: if billing is usage-based, revenue in any given month may not be predictable, and the $30 million annualized figure reflects a peak-month extrapolation rather than a contracted forward obligation. Fourth, the valuation at 57x run rate implies the capital markets are pricing in both high revenue growth (revenue must exceed $100 million to justify a 17x multiple, a common late-stage benchmark) and durable gross margins — both of which are unverifiable from public data. Any slowdown in enterprise adoption or loss of methodological trust could trigger a significant valuation reset that complicates future fundraising or exit optionality.[CI021, CI022, CI023, CI024, CI025, CI032]

Public financial gaps table
missing private metricimpact on analysisexact diligence path
Realized 2025 revenue (not run-rate)Cannot verify whether $30 M annualized run rate translates to $8–10 M actual 2025 revenue (4 months of product) or a higher/lower number depending on ramp shapeRequest monthly revenue actuals from product launch (September 2025) through December 2025
NRR / account expansion dataWithout NRR, cannot distinguish growing accounts (bullish) from one-time pilot evaluations (bearish for durability)Obtain cohort-level NRR for AI lab clients and enterprise clients separately; confirm if any have already churned
Gross margin (P&L level)Gross margin drives whether $30 M run rate implies a scalable, high-margin business or a service delivery model with limited leverageRequest gross profit line from management accounts; confirm treatment of partner API costs and infrastructure costs
ACV and contract termsCannot verify $300 K implied ACV or whether pricing is time-boxed, usage-based, or event-basedObtain a sample de-identified contract and ACV distribution across client tiers
Customer concentration (top-3 revenue share)If three lab clients (OpenAI, Google, xAI) represent 50%+ of revenue, any single defection is a material eventRequest revenue breakdown by customer tier; flag any single client representing >20% of revenue
Headcount and burn as of June 2026Company had 41 employees at Series A (January 2026); burn will have increased meaningfully if expansion plans are underwayRequest current headcount, monthly payroll, and infrastructure cost run rate as of most recent quarter

Each gap represents a material variable needed to underwrite the revenue quality and financial sustainability claims implied by the $1.7 B valuation.

[CI003, CI004, CI017, CI019, CI023, CI025]
FI003: Financial estimate range

Key financial variables have wide uncertainty bands; the 57x valuation multiple is confirmed, but the revenue and cost inputs needed to validate it are largely private.

All ranges except the valuation multiple are estimates derived from headcount, typical SV compensation, infrastructure benchmarks, and the disclosed run-rate figure. The mid-point of the valuation multiple (57x) is calculated from confirmed $1.7 B valuation and $30 M run rate. Ranges reflect uncertainty in both inputs and the distinction between run rate and realized revenue.

[CI002, CI011, CI014, CI017, CI018, CI032]

4.4 Exhibits

Chapter 05

05Product & Technology

5.1 Product Portfolio and Arena Modalities

LMArena's core user-facing product is the Arena platform at arena.ai, which presents battles between two anonymized AI models for a given prompt and records the community's preference vote. As of June 2026 the platform supports ten distinct evaluation arenas: Text (general chat), Code (agentic web-development and coding), Search (retrieval-augmented generation), Agent (autonomous multi-step task completion), Vision (multimodal understanding), Text-to-Image (image generation), Text-to-Video (video generation), Image Edit (image-editing models), Document (long-form document analysis), and a separate Direct Chat mode where users interact with a chosen model without the pairwise battle structure. The Agent Arena is the newest and most technically complex arena, launched June 4, 2026, which ranks orchestrator models on causal treatment effects derived from 2M+ tool calls per week. Code Arena, rebuilt from the earlier WebDev Arena (launched December 2024), operates as a live, isolated coding environment where models autonomously create, modify, and execute files through structured tool calls, with outputs persisted in Cloudflare R2 and displayed via CodeMirror 6. Since May 2026 LMArena also includes "Battles in Direct," routing 10% of Direct Chat sessions into anonymous pairwise battles to increase daily vote volume. The commercial product, AI Evaluations (launched September 2025), is a contract evaluation service sold to AI labs and enterprises, offering in-depth evaluations with representative feedback samples and SLA-backed delivery timelines. All public model evaluations are available free of charge on the public leaderboard; the commercial product adds private, confidential, SLA-governed evaluation runs. [CE001, CE002, CE008, CE012, CE029, CE032]

Product Module and Arena Asset Matrix
Arena / ModuleUser / BuyerStatus / MaturityEvaluation MethodKey DifferentiationDiligence Gap
Text Arena (Chatbot Arena)General users, researchers, model providersLive / Mature (since 2023)Bradley-Terry pairwise votes250M+ real conversations; style-controlled rankingsNo independent reliability audit of BT estimates
Code ArenaDevelopers, enterprises, model providersLive / Growing (rebuilt 2026)Pairwise votes on generated web apps; agentic tool-call environmentPersistent sessions, CodeMirror 6 live preview, Cloudflare R2 snapshotsFunctional correctness vs. preference divergence not quantified
Agent ArenaPower users, enterprises, model providersLive / Early (launched Jun 2026)Causal tracing: multi-signal RCT frameworkNovel methodology; 2M+ weekly tool calls; first causal agent leaderboardSignal selection and causal model assumptions not peer-reviewed yet
Search ArenaResearchers, model providers, enterprisesLive / GrowingBradley-Terry; citation-style randomizationICLR 2026 paper; 24k+ battles; grounding quality analysisOnly 3 providers supported; geolocation feature limited
WebDev ArenaDevelopers, model providersLive / Stable (since Dec 2024)Bradley-Terry on web app preference votes80k+ votes; topic modeling analysis; CSS/JS/HTML generationPredecessor to Code Arena; methodological refresh pending
Vision ArenaMultimodal users, model providersLive / StableBradley-Terry pairwise votes on image-understanding tasksCovers major multimodal modelsCoverage limited relative to text arena; image fidelity metrics not included
Text-to-Image ArenaCreative users, model providersLive / GrowingPairwise preference votes on generated imagesCovers Flux, Midjourney, DALL-E, etc.Aesthetic preference vs. prompt adherence not separated
Text-to-Video ArenaCreative users, model providersLive / EarlyPairwise preference votes on generated video clipsEarly-stage; wan2.7-t2v, gemini-omni-flash added in 2026Limited vote volume; temporal coherence not separately scored
Document ArenaEnterprise users, model providersLive / GrowingBradley-Terry pairwise votes on document tasksLong-context evaluation; added to several leaderboards in 2026Specific document-type breakdowns not published
AI Evaluations (commercial)AI labs, enterprisesGA since Sep 2025Private evaluation with SLA delivery; community-grounded feedback$30M ARR run-rate in Dec 2025; OpenAI, Google, xAI cited as customersSLA terms, security disclosures, DPA not public
Direct ChatAll usersLive / StableNo evaluation; free model accessNo-paywall access to frontier modelsNot revenue-generating; cost center for inference
AI Evaluations (API/pipeline)Enterprises, developersRoadmap / plannedProgrammatic evaluation submissionWould unlock self-serve evaluation at scaleNo timeline announced; pre-launch risk

Status and vote counts derived from official LMArena blog posts, leaderboard changelog, and press releases through June 2026. Diligence gaps are researcher inference.

[CE001, CE002, CE008, CE012, CE029, CE032]
FE001: LMArena Product Architecture Stack

Layered view of LMArena's platform from data collection through ranking engine to evaluation products.

Layer boundaries inferred from blog posts and GitHub repos; internal service decomposition not publicly disclosed.

[CE001, CE002, CE003, CE009, CE013]

5.2 Technical Architecture and Evaluation Methodology

All Arena leaderboards use the Bradley-Terry (BT) pairwise comparison model, which infers a latent skill coefficient for each model from win/loss outcomes. LMArena's Arena-Rank Python package, open-sourced under Apache 2.0 and published on GitHub and PyPI, implements the BT fitting with closed-form confidence-interval calculation and a 30x speedup over the historical FastChat implementation. The package also supports reweighting to correct for non-uniform sampling, meaning models with fewer battles are not penalized. Style-control extensions layer additional covariates (token length, markdown header count, markdown bold count, markdown list count) into the BT regression to separate substance from stylistic formatting effects; coefficient estimation shows that length is the dominant style factor. For the Agent Arena, ranking uses causal tracing rather than pairwise votes: the arena treats each component selection as a treatment in a multi-intervention randomized controlled trial, aggregating multiple behavioral signals (confirmed success, praise/complaint, steerability, bash-error recovery, tool hallucination) into a single net-improvement estimate per model. The Arena-Hard pipeline (BenchBuilder, arXiv:2406.11939) automates benchmark construction by extracting hard prompts from the live Arena dataset via a seven-criterion hardness labeler; Arena-Hard-Auto v0.1 achieves 98.6% agreement with human preference rankings at a cost of approximately $20. The Search Arena methodology, published as a dataset and paper accepted at ICLR 2026 (arXiv:2506.05334), extends the BT model to search-augmented LLM systems across 11+ models from Perplexity, Gemini, and OpenAI. Prior to sharing any conversation data, LMArena uses GCP's Sensitive Data Protection API to remove personally identifiable information. Since July 2025, methodology updates are publicly logged in the Leaderboard Changelog. [CE003, CE009, CE013, CE014, CE015, CE016]

Technology and Operating Architecture
Layer / ComponentRoleDependency / ImplementationRisk
Arena-Rank (ranking engine)Computes BT coefficients, confidence intervals, style-controlled scores for all arenasOpen-source Python package (Apache 2.0); GitHub lmarena/arena-rank; PyPI arena-rankMethodology open = reproducible but also replicable by rivals
Bradley-Terry pairwise modelCore statistical model; infers latent skill from win/loss battlesCustom Python implementation; 30x speedup over FastChat baseline; closed-form CIsModel assumptions (IID battles, no temporal drift) may degrade under heavy sampling manipulation
Style-control extensionSeparates substance from formatting effects in BT regressionAdditive covariate regression; length, markdown headers, bold, list countsOnly four style covariates; richer semantic style not yet captured
Causal tracing (Agent Arena)Ranks agent components via multi-intervention RCT; measures net treatment effectsCustom statistical framework combining 5 behavioral signalsNovel methodology; causal identification assumptions unvalidated by third parties
BenchBuilder / Arena-Hard pipelineAuto-generates hard benchmark subsets from live Arena data using LLM-labelerarXiv:2406.11939; GPT-4-based hardness scorer; 7 hardness criteriaDepends on GPT-4 for labeling; bias from labeler bleeds into benchmark quality
Data sanitization (GCP Sensitive Data Protection API)PII removal before public data releaseGoogle Cloud Platform API; called pre-release of any shared datasetReliance on single cloud vendor; no independent PII audit published
Web platform (arena.ai)Hosts battles, voting UI, leaderboards, Direct Chat, Agent Mode, Code ArenaJavaScript front-end; server architecture not disclosedSingle point of failure for community trust; uptime SLA not public
Cloudflare R2 + CodeMirror 6 (Code Arena)Stores persistent code sessions; renders live web app previewsCloudflare R2 for snapshots; CodeMirror 6 for code display; streaming frontendData residency and retention policy not disclosed
FastChat (legacy)Historical platform for serving and training; original Chatbot Arena backendGitHub lm-sys/FastChat; primarily maintenance mode as of 2025Actively deprioritized; migration risk for external users still on FastChat
p2l models (HuggingFace)Prompt-to-leaderboard preference models; used for internal evaluation researchlmarena-ai org on HuggingFace; 0.1B-7B parameter rangeNot publicly documented for production use; purpose and deployment not clear

Architecture inferred from official blog posts, GitHub repos, PyPI package, and academic papers. Server infrastructure details are not publicly disclosed.

[CE003, CE009, CE013, CE014, CE015, CE016]
FE002: User Workflow for LMArena Text and Code Battle

End-to-end flow from user prompt submission through vote recording to leaderboard update.

Flow reconstructed from official blog posts and policy page; internal system architecture not disclosed.

[CE002, CE016, CE020, CE021]
FE003: Critical Dependency Map

Key platform, data, regulatory, and partner dependencies for LMArena's product.

Dependency relationships inferred from public documentation; relative criticality is researcher assessment.

[CE013, CE014, CE019, CE020, CE032]

5.3 Deployment, Data Infrastructure, and Developer Ecosystem

The Arena-Rank ranking engine is published as a pip-installable Python package (pip install arena-rank) under the lmarena GitHub organization, which also hosts the list repository. The FastChat repository (lm-sys/FastChat, now primarily in maintenance mode) originally powered Chatbot Arena and remains a reference for the Vicuna model weights. The lmarena-ai HuggingFace organization hosts p2l (prompt-to-leaderboard) model variants used internally for preference modeling, and releases public battle datasets (e.g., arena-human-preference-140k) to support external research. As of June 2026 the changelog records ongoing model additions at a pace of several new models per week across leaderboards. The Code Arena's persistent sessions are built on Cloudflare R2 storage with CodeMirror 6 for source views and live rendering of generated web applications. Agent Mode sessions averaged approximately 16.5 structured tool calls per session, with roughly 75.6% of sessions using at least one tool; the highest-use sessions ran very long chains in a single week, covering coding, file creation, and web synthesis tasks. In one measured 7-day window, Agent Mode wrote 40.3 million lines of code across successful write_file calls, approximately 1,000 lines per coding session. Battles in Direct, integrated since May 2026, converts 10% of direct chat sessions into pairwise battles, with position-bias and same-org-indicator corrections applied to the BT fit. No API is publicly documented for third-party consumption of Arena scores or raw vote data, though open datasets are released periodically on HuggingFace. [CE013, CE014, CE015, CE019, CE030, CE031]

Key User Workflow and Use-Case Table
User JobLegacy WorkflowLMArena SolutionMeasurable BenefitKnown Limitation
Compare LLM quality before selecting for productionRun internal A/B tests; rely on static benchmarks (MMLU, HumanEval)Text/Code Arena pairwise battle vote; see real-user preference at scaleStatistically calibrated ranking with 95% CIs; style-controlled and hard-prompt subsetsRankings reflect general user preferences, not enterprise-specific use cases
Benchmark agentic coding models for real workloadsStatic code-correctness benchmarks; synthetic test suitesCode Arena: isolated agent environment; live web app generation; persistent sessionsAgentic, multi-turn behavior captured; shareable runsCorrectness vs. aesthetic preference not separated; limited programming-language diversity
Evaluate search-augmented LLMs for retrieval tasksSimpleQA static benchmarks; manual annotationSearch Arena: crowdsourced votes on multi-turn search-RAG outputs24k+ paired interactions; citation-count analysis; ICLR 2026 validationOnly 3 providers currently; academic/domain-specific coverage limited
Measure autonomous agent performance for business tasksManual human rater sessions; tool-specific benchmarksAgent Arena: causal tracing across 2M+ weekly tool calls; 5-signal leaderboardRCT-based causal treatment effects; interpretable per-signal breakdownsMethodology not yet peer-reviewed; signal set will evolve
Get confidential model evaluation with SLA for pre-launchProprietary red-teaming or vendor-specific eval teamsAI Evaluations: private evaluation with SLA, representative feedback, data samplesGrounded in real-world human preferences; certified delivery timelinePricing, SLA terms, and security details not publicly disclosed

Workflows are synthesized from LMArena blog posts, academic papers, and TechCrunch reporting. Benefit claims are company-stated unless marked as researcher inference.

[CE002, CE009, CE012, CE029]
FE004: Product Maturity and Capability Map

Maturity of each evaluation arena across four capability dimensions as of June 2026.

Maturity ratings are researcher judgments based on available public evidence as of June 2026. Vote volume tiers: High = 250M+; Medium = 24k-80k; Low = <24k.

[CE002, CE008, CE009, CE035, CE036]

5.4 Differentiation, IP, and Research Advantage

LMArena's core differentiation is the combination of real-world, in-the-wild data at scale with methodological transparency. Unlike static academic benchmarks (MMLU, HumanEval) or automated LLM-as-judge pipelines, Arena captures genuine multi-turn user interactions across diverse languages, topics, and skill levels. Published research validates this approach: the original Chatbot Arena paper (arXiv:2403.04132) showed >80% agreement with expert raters; MT-Bench (NeurIPS 2023) demonstrated that LLM-as-a-judge can achieve human-level inter-rater agreement. The Arena-Hard benchmark (arXiv:2406.11939) achieves 3× higher model separation than MT-Bench at a fraction of the cost. The Search Arena ICLR 2026 paper (arXiv:2506.05334) extends the methodology to retrieval-augmented systems, demonstrating preference correlations with citation count and cited source types. LMSYS (the precursor nonprofit) published the foundational evaluation methodology papers that underpin the current company's differentiation; those papers are widely cited across the AI industry. The HuggingFace organization releases preference datasets (140k+ labeled battle pairs) enabling external researchers to audit rankings independently. The company's 9 published papers and 15+ blog posts covering evaluation methodology constitute a recognized body of IP. However, methodological assets are openly shared, which means rivals can replicate the algorithm while the company retains the data-network advantage (scale of battles) and community trust as the lasting moat. [CE003, CE016, CE035, CE036, CE037, CE039]

Roadmap, Releases, and Development Stage
Date / PeriodFeature / MilestoneStatusImplicationSource
May 2023Chatbot Arena (Text Arena) public launchCompletedEstablished crowdsourced evaluation as viable methodologylmsys.org/blog/2023-05-03-arena/
Jun 2023MT-Bench multi-turn question set and LLM-as-a-judge paper (NeurIPS 2023)CompletedAcademic validation of pairwise evaluation methodologyarXiv:2306.05685
Mar 2024Chatbot Arena technical paper published (arXiv:2403.04132)Completed240k+ votes; credibility established for industry citationarXiv:2403.04132
Apr 2024Arena-Hard pipeline (BenchBuilder) publishedCompletedAutomated benchmark creation; 98.6% agreement with human rankingsarXiv:2406.11939
Dec 2024WebDev Arena launchedCompleted; superseded by Code ArenaFirst coding-specific real-world evaluation; 80k+ votes collectedarena.ai/blog/webdev-arena/
Sep 2025AI Evaluations commercial product launchedGA; $30M ARR by Dec 2025Revenue engine; SLA-governed enterprise and lab evaluation servicearena.ai/blog/ai-evaluations/
Mar 2026Search Arena paper accepted at ICLR 2026 (arXiv:2506.05334)Accepted and presentedPeer-reviewed validation of search-augmented LLM evaluation methodologyarXiv:2506.05334
Mar 2026 (est.)Battles in Direct experiment began (10% of direct sessions)Experiment phaseIncreases vote volume; corrects position bias and same-org biasarena.ai/blog/leaderboard-changelog/
May 2026Battles in Direct votes included in leaderboardsLiveTightens confidence intervals; shifts prompt distribution toward harder queriesarena.ai/blog/leaderboard-changelog/
May 2026Code Arena launched (rebuilt from WebDev Arena)LiveAgentic coding environment; persistent sessions; Cloudflare R2 snapshotsarena.ai/blog/code-arena/
Jun 4, 2026Agent Arena launched with causal-tracing methodologyLive / EarlyFirst RCT-based agent leaderboard; 2M+ weekly tool calls analyzedarena.ai/blog/agent-arena-methodology/
2026 (planned)Additional behavioral signals for Agent Arena; more arenas / modalitiesAnnounced intent; no specific datesPlatform expansion in agentic evaluation; signal set will evolvearena.ai/blog/agent-arena-methodology/

Dates derived from blog posts, academic paper timestamps, and leaderboard changelog. Planned items are based on stated intent, not confirmed release dates.

[CE008, CE012, CE029, CE033, CE035, CE036]

5.5 Trust, Safety, Quality Controls, and Adverse Evidence

LMArena's public leaderboard policy (last updated April 30, 2026) specifies eligibility criteria for model listing, a sampling policy requiring ≥20% of battles to be between publicly available models, a pre-release testing protocol for anonymized models, and a data-sharing policy using GCP's Sensitive Data Protection API to sanitize conversations before release. However, as of June 2026 LMArena has not published SOC 2 Type II or ISO 27001 certifications, no third-party security audit is publicly disclosed, and the help.arena.ai privacy policy page lacks enterprise data processing agreement (DPA) details. The platform has faced two serious credibility challenges. First, in April 2025 Meta submitted an "experimental, chat-optimized" version of Llama 4 Maverick that ranked #2 on the arena but was not the publicly released model; LMArena subsequently updated its policy and rescored the public version, which fell to 32nd place. Second, a paper authored by Cohere, Stanford, MIT, and AI2 in April 2025 alleged that certain large labs (Meta, OpenAI, Google, Amazon) received disproportionately high sampling rates and could selectively suppress low-performing pre-release variants, constituting benchmark gaming; LMArena denied the inaccuracies and pointed to its published sampling transparency. These incidents create ongoing reputational risk and have prompted calls from the research community for tighter pre-release limits, independent audits, and algorithmic transparency. The AI Evaluations commercial product offers SLA commitments but detailed uptime SLA terms and security disclosures are not public. [CE020, CE021, CE022, CE023, CE024, CE025]

Trust, Quality, and Compliance Controls
Control / CertificationStatusScopeGap / Risk
Leaderboard policy (public)Live; last updated April 30, 2026Model eligibility, sampling policy, pre-release testing protocol, data sharing rulesPolicy is self-enforced; no independent auditor or enforcement mechanism
Data sanitization (GCP Sensitive Data Protection API)Active for public data releasesPII removal from conversation data before HuggingFace/public releaseSingle vendor dependency; no published audit of sanitization efficacy
Open-source ranking methodology (Apache 2.0)Live; arena-rank v1 on GitHub/PyPIBradley-Terry engine, reweighting, CIs; powers all current leaderboardsAlgorithm is open but platform data is proprietary; full reproducibility limited
Academic peer review of methodology3 published and 1 ICLR 2026 accepted papersChatbot Arena (arXiv 2403.04132), MT-Bench (NeurIPS 2023), Arena-Hard (arXiv 2406.11939), Search Arena (ICLR 2026)Agent Arena causal-tracing methodology not yet peer-reviewed
Pre-release testing protocolPublished policy since March 2024; updated April 2026Model providers may test anonymized pre-release models; results shared privatelyPreferential access controversy (Cohere/Stanford paper) led to policy update but no independent audit
Enterprise SLA (AI Evaluations)GA since Sep 2025; SLA offeredCommitted delivery timelines for commercial evaluation runsSLA uptime terms, penalties, and security audit scope not publicly disclosed
Privacy policyExists on help.arena.aiUser data handling for community platform interactionsNo enterprise DPA; GDPR/CCPA compliance details not publicly documented
SOC 2 Type II / ISO 27001Not publicly disclosedN/A — would apply to enterprise cloud servicesAbsence of published certifications is a diligence gap for enterprise buyers

Status reflects publicly available documentation as of June 25, 2026. 'Not publicly disclosed' does not confirm absence; LMArena may hold certifications not published.

[CE020, CE021, CE022, CE023, CE024, CE025]

5.6 Exhibits

Chapter 06

06Customers

6.1 Customer Base Segmentation

LMArena's customer base has two fundamentally different segments that must be distinguished carefully. The first is the free community, a global population of 5 million or more monthly users across 150+ countries who interact with the platform at no charge through arena.ai. These users submit prompts, vote on model outputs, and power the collective preference signal that underpins the leaderboards. This community spans developers, researchers, students, knowledge workers, and enthusiasts; per the two-year anniversary blog (April 2025), approximately 41% of battles involve open-source models, suggesting a technically sophisticated user base. The company generates no direct revenue from this segment but derives its core IP (preference data) and brand credibility from it. The second is the commercial customer segment, which consists of AI labs and enterprises that pay for the AI Evaluations service launched in September 2025. Named paying customers confirmed in the January 2026 Series A press release include OpenAI, Google, and xAI. Anthropic and Meta are major model providers on the leaderboard but their paying status as AI Evaluations subscribers is not separately confirmed in public disclosures; the TechCrunch 2026 article notes LMArena "partnered with" OpenAI, Google, and Anthropic for model submissions, a category that encompasses both commercial and non-commercial relationships. Enterprises (non-lab companies using AI Evaluations to benchmark models for their own applications) are a third stated target segment; no named enterprise non-lab customers have been publicly disclosed. [CU001, CU002, CU003, CU004, CU005, CU006]

Customer Segmentation Table
SegmentBuyer / User / PayerUse CaseScaleRevenue / Strategic ValueKey Gap
Community users (free)End users (researchers, devs, enthusiasts)Free model battles; leaderboard access; Direct Chat5M+ monthly users; 150+ countriesNo direct revenue; core IP and brand moat; PR/growth engineUser demographics not publicly broken down by profession or industry
AI lab model providers (non-paying / free)AI labs: Anthropic, Meta, Mistral, DeepSeek, etc.Submit models for free public evaluation to gain leaderboard ranking and market credibility400+ models evaluated; 300+ pre-release testsNo direct revenue; provides model supply for platform valueNot clearly distinguished from paying customers in all communications
AI lab commercial customers (paying)AI labs: OpenAI, Google, xAI (named)Private AI Evaluations with SLA; comprehensive feedback for model improvement3 named; unknown othersRevenue-generating; $30M ARR run-rate Dec 2025 from this segmentCustomer count, contract sizes, and individual contribution not disclosed
Enterprise customers (paying)Enterprises wanting to benchmark AI models for production useAI Evaluations to measure model performance for specific domains (law, medicine, software)Size unknown; no named enterprise customersStated target segment; no public evidence of enterprise non-lab customers yetNo named enterprise customers; pipeline and conversion rate unknown
Academic / research users (free)University researchers, think tanks, open-source contributorsPublic datasets for research; leaderboard as reference; battle data analysisHundreds of research papers cite ArenaNo direct revenue; builds academic credibility and methodology validationNo formal academic partnership program announced

Segments based on official press releases, blog posts, and news coverage. Revenue attribution is estimated from the single disclosed ARR figure; per-segment breakdown is not public.

[CU001, CU002, CU003, CU004, CU005, CU006]
FU001: LMArena Customer Journey Map

Customer segments, adoption surfaces, and key engagement touchpoints from first contact to commercial expansion.

Journey stages inferred from official communications and press coverage; no customer research or NPS data is publicly available.

[CU001, CU002, CU005, CU006]

6.2 Adoption Trajectory and Usage Metrics

LMArena's community adoption has been rapid and well-documented. The platform launched in May 2023 as a UC Berkeley research project; by May 2025 it disclosed the $100M seed round at a $600M valuation with significant traction. By January 2026, at the time of the Series A announcement, LMArena reported 5M+ monthly users (up from an implied ~3M cited in the September 2025 AI Evaluations blog), 60M+ monthly conversations, 250M+ cumulative conversations, and 2M+ monthly votes. The Series A blog post quantified community growth as "25x" since the seed round in May 2025. In terms of model evaluations, LMArena has evaluated 400+ public models and 300+ pre-release model variants. From September 2025 (commercial launch) to December 2025 (less than four months), the annualized revenue consumption run-rate reached $30M, per CEO Anastasios Angelopoulos and confirmed in TechCrunch's January 2026 article and the official press release. The Series A post also reported 50M+ community votes collected, 145k+ open-source battle data points released, and 400+ model evaluations across modalities. Notably, the platform's growth appears primarily pull-driven: the TechCrunch podcast episode notes that LMArena leaderboards "became something of an obsession among model makers," implying low customer-acquisition cost for the community segment. [CU001, CU002, CU003, CU004, CU007, CU008]

Customer Growth and Adoption Trajectory
MetricValueDateSourceConfidenceImplication
Monthly active users3M+Sep 2025arena.ai/blog/ai-evaluations/Medium (company-claimed)Community had already scaled significantly at commercial launch
Monthly active users5M+Jan 2026PRNewswire / arena.ai/blog/series-a/High (multi-source)5x MAU growth in under 2 years; strong organic pull
Monthly conversations60M+Jan 2026PRNewswire / TechCrunch Jan 2026High (multi-source)Each user generates ~12 conversations/month on average
Cumulative conversations250M+Jan 2026PRNewswire / arena.ai/blog/ai-evaluations/Medium (company-claimed)Long tail of historical data; strengthens preference model quality
Monthly votes2M+Sep 2025 est.arena.ai/blog/ai-evaluations/Medium (company-claimed)High engagement rate relative to users; ~0.7 votes/user/month
Community votes total50M+Jan 2026arena.ai/blog/series-a/Medium (company-claimed)Cumulative vote corpus underpins BT model calibration
Models evaluated (public)400+Apr 2025 / Jan 2026arena.ai/blog/two-year-celebration/, arena.ai/blog/series-a/Medium (company-claimed)Breadth of coverage strengthens platform neutrality claim
Pre-release model tests300+Apr 2025arena.ai/blog/two-year-celebration/Medium (company-claimed)Strong lab engagement; also the source of gaming-controversy risk
Geographic reach150+ countriesJan 2026PRNewswireMedium (company-claimed)Global diversity strengthens preference data representativeness
ARR run-rate$30M annualizedDec 2025TechCrunch Jan 2026 / PRNewswireHigh (multi-source)Rapid ramp from zero in Sep 2025; validates commercial model
Community growth rate25x year-over-yearMay 2025–Jan 2026arena.ai/blog/series-a/Low-Medium (company framing)Single data point; metric denominator unclear

Values are company-disclosed except where marked estimated. 'High' confidence requires two independent sources. Month-over-month trajectory is not available in public disclosures.

[CU001, CU002, CU003, CU007, CU008, CU009]
FU002: Adoption and Deployment Funnel

From discovery to community voting to model submission to commercial evaluation: the LMArena engagement funnel.

Monthly visitors is a rough estimate; vote conversion rate is derived from disclosed MAU and monthly vote counts. Commercial customer count is a minimum (only named ones).

[CU001, CU002, CU007, CU008, CU009]

6.3 Named Customer Proof and Production Evidence

Named paying customer evidence for the AI Evaluations commercial product is limited but specific. The January 2026 PR Newswire press release states: "LMArena works with leading AI labs and enterprises, including OpenAI, Google and xAI, all drawing on LMArena's evaluations to improve their models for production use cases." This constitutes direct company-stated customer naming and is the strongest public customer-proof available. Separately, the Felicis partner Peter Deng stated in the same press release: "We're leading this round because LMArena has built the most trusted, reliable, real-world signal of AI performance. They have become essential infrastructure for every lab and enterprise." The company's Agent Arena methodology blog post documents real user sessions — a live sports-TV schedule site, a self-hosted movie watchlist, an autonomous underwater vehicle autopilot — as examples of production-style deployments, though these are community users, not paying enterprise customers. Prior to commercialization, Google's Kaggle, a16z, and Together AI had donated compute and grants to LMSYS (the precursor nonprofit). The lmsys.org historical launch records show the original 2023 partnership structure with UC Berkeley. For model providers that are not confirmed as paying: Anthropic and Meta are frequently cited as having their flagship models on the leaderboard, but the TechCrunch 2025 and 2026 articles consistently name only OpenAI, Google, and Anthropic as "partners" and only OpenAI, Google, and xAI as named commercial customers in the press release. No government procurement records, G2/Capterra reviews, or independent third-party case studies are publicly available for the AI Evaluations service. [CU005, CU006, CU012, CU013, CU014, CU015]

Named Customer Proof Table
Customer / UserSegmentDeployment / Use CaseProduction vs. PilotEvidence of OutcomeEvidence Limitation
OpenAIAI lab (paying commercial)AI Evaluations — evaluating GPT model variants for human preference; flagship models on public leaderboardProduction (named in press release)Named in Jan 2026 Series A PR Newswire press release as a customer drawing on evaluations for production useNo quote from OpenAI directly; company-claimed only
GoogleAI lab (paying commercial)AI Evaluations — evaluating Gemini model variants; flagship models on public leaderboardProduction (named in press release)Named in Jan 2026 PR Newswire press release; Gemini models consistently rank on text and vision leaderboardsNo direct Google statement; company-claimed
xAIAI lab (paying commercial)AI Evaluations — evaluating Grok model variants for human preferenceProduction (named in press release)Named in Jan 2026 PR Newswire press release alongside OpenAI and GoogleNewest of the three named customers; product usage details not disclosed
AnthropicAI lab (model provider; paying status unconfirmed)Claude models on public leaderboard; 'partnered' for model submission per TechCrunch 2026Active production (leaderboard)Claude described as winning expert leaderboard for legal/medical use cases (TechCrunch podcast); partner relationship confirmedPaying AI Evaluations customer status not separately confirmed in any source
MetaAI lab (model provider; paying status unconfirmed; adverse history)Llama models on public leaderboard; 27 pre-release Llama 4 variants tested Jan–Mar 2025Active production (leaderboard)Extensive pre-release testing confirmed by Cohere/Stanford paper; Llama models regularly updated on leaderboardGaming incident: experimental Maverick submitted and ranked #2; public release ranked ~32nd; relationship adversely affected
PerplexityAI lab (model provider; search arena partner)Sonar models on Search Arena leaderboard; citation-style collaborationActive (leaderboard)11 Perplexity model variants evaluated in Search Arena; style randomization done in collaboration with providerNo evidence of AI Evaluations paid subscription

Named paying customers (OpenAI, Google, xAI) sourced from the January 2026 PR Newswire press release. Model provider participation (Anthropic, Meta, Perplexity) sourced from news articles and blog posts but does not confirm AI Evaluations commercial subscription. No independent customer quotes or testimonials are publicly available.

[CU005, CU006, CU012, CU013, CU014, CU015]
FU003: Customer Proof Quality Matrix

Evidence quality across dimensions for each named customer or model provider.

Evidence quality ratings reflect the best available public evidence as of June 2026; 'Unconfirmed' does not mean 'No.'

[CU005, CU006, CU012, CU013, CU014, CU021]

6.4 Retention, Durability, and Repeat Usage

LMArena has not publicly disclosed formal retention metrics such as net revenue retention (NRR), gross revenue retention (GRR), or customer churn for its AI Evaluations product. The platform is too young commercially (launched September 2025) to have multi-year cohort data. For the community segment, the principal retention signal is platform activity: the 25x community growth year-over-year (May 2025 to January 2026 per the Series A blog) and 250M+ cumulative conversations versus 60M/month active generation suggest that most of the cumulative usage reflects a growing ongoing community rather than one-time visits. The AI Evaluations product exhibits a consumption-based model ("annualized consumption rate" rather than subscription ARR), which means that retention depends on labs and enterprises continuously submitting evaluation requests. The rapid ramp from $0 to $30M annualized in under four months implies initial strong lab demand; whether that sustains as labs develop internal evaluation capabilities is a key risk. The leaderboard changelog shows continuous model additions (several per week in June 2026), which serves as an indirect proxy for ongoing model-provider engagement. The lack of any disclosed customer logo wall, case study site, or testimonial page (other than investor quotes from Felicis and UC Investments) is notable for a company at $1.7B valuation and $30M ARR run-rate. [CU007, CU009, CU010, CU011, CU016, CU017]

Retention, Repeat Usage, and Satisfaction Metrics
MetricValue / StatusSegmentConfidenceDiligence Ask
Net Revenue Retention (NRR)Not disclosedAI Evaluations commercial customersUnknownRequest NRR or cohort analysis from LMArena for lab customers month 1–6 post-purchase
Gross Revenue Retention (GRR)Not disclosedAI Evaluations commercial customersUnknownConfirm whether any commercial customers have churned since Sep 2025 launch
Monthly user retention (community)Not disclosed; implied high via 60M monthly conversations and continued growthCommunity (free users)Low (inferred)Request monthly active user retention or session frequency data
Customer churn (commercial)Not disclosed; <4 months since launch so minimal churn data existsAI Evaluations commercial customersUnknownTrack whether OpenAI/Google/xAI renew contracts at next cycle
Repeat model submission (model providers)300+ pre-release tests and 400+ public evaluations implies ongoing repeat engagement from ~20+ labsModel providers (free + paying)Low-Medium (inferred from platform data)Obtain count of unique providers that have submitted models more than once
Indirect retention signal: leaderboard activitySeveral new models added per week in June 2026 per changelogAll model providersMedium (observed)Validate that active model additions correlate with provider satisfaction rather than marketing pressure
User satisfaction (community)Not formally disclosed; no NPS or CSAT scores publishedCommunity usersUnknownLook for NPS survey or user review data if published

Values marked 'Not disclosed' are absent from public sources as of June 25, 2026; they do not confirm zero values. Confidence ratings reflect available evidence.

[CU016, CU017, CU018]

6.5 Adverse Evidence: Conflicts, Gaming, and Concentration Risk

LMArena's customer and model-provider relationships create well-documented structural risks. First, the conflict of interest: the same AI labs (OpenAI, Google, Anthropic, and others) that pay for AI Evaluations and serve as commercial customers are also the entities whose models are ranked on the public leaderboard. This creates an inherent incentive for labs to optimize for Arena performance, which they can do via the pre-release testing protocol. The TechCrunch podcast episode directly raises this: "how a team like theirs can build a neutral benchmark when the companies they're ranking are also their backers." Second, the Meta Llama 4 Maverick incident (April 2025): Meta submitted a "chat-optimized experimental" version of Maverick to the arena that achieved a #2 ranking, but the publicly released model ranked approximately 32nd. LMArena updated its policy and stated "Meta's interpretation of our policy did not match what we expect from model providers." This incident shows that the commercial partnership structure (which depends on labs submitting models) creates practical pressure to accommodate behavior that damages leaderboard integrity. Third, the Cohere/Stanford/MIT/AI2 study (April 2025) alleged that Meta could privately test 27 model variants between January and March 2025 before selecting the highest performer, that sampling rates were disproportionate for top labs, and that the study's preliminary findings were not disputed by LMArena when shared. LMArena denied the specific claims but announced a new sampling algorithm. These allegations have not been independently adjudicated. Fourth, customer concentration: with only three publicly named paying customers (OpenAI, Google, xAI), LMArena's commercial revenue is highly concentrated; if any of these labs develop or contract an alternative evaluation vendor, the impact on ARR could be significant. [CU005, CU006, CU012, CU019, CU020, CU021]

Expansion, Concentration Risk, and Adverse Dynamics
Expansion Driver / Concentration RiskImpactEvidenceDiligence Path
Customer concentration: 3 named paying customersHigh — if OpenAI, Google, or xAI reduce or cancel AI Evaluations spend, single-customer revenue impact could be severeOnly 3 customers named in press release; no other named commercial customersDetermine revenue share of each named customer; confirm # of unnamed commercial customers
Conflict of interest: raters = payersHigh — labs paying for evaluations are also the entities whose models are ranked, creating gaming incentive and credibility riskMeta gaming incident (Apr 2025); Cohere/Stanford paper alleging preferential accessStructural safeguard audit: request details of sampling algorithm and pre-release testing limits
Land-and-expand: one-time evaluation → repeat contractMedium — consumption-based model supports expansion if labs find evaluations valuable; no evidence of multi-year contractsCEO described a 'consumption rate'; rapid $30M ARR ramp suggests repeat usageConfirm contract structure (spot vs. subscription vs. annual) with LMArena
Enterprise non-lab expansionMedium — large potential market but zero named enterprise non-lab customers yetSeries A press release mentions 'enterprises'; TechCrunch article mentions 'software engineering, law, medicine' as target domainsIdentify any non-lab enterprise trials or pilots in progress
Competitive displacement risk: internal evaluation build-out at labsMedium — if top labs (Google, OpenAI) develop internal evaluation pipelines that satisfy their needs, LMArena commercial demand could plateauScale AI's SEAL Showdown (Mashable article) launched as a competing evaluation productTrack whether scale/size of lab relationships grows or stabilizes over time
Adverse reputational risk from future gaming incidentsHigh — a second high-profile gaming incident could rapidly erode community trust and commercial credibility simultaneouslyMeta Llama 4 incident was widely covered by The Verge, TechCrunch, and others; LMArena policy updated but structural conflict persistsMonitor for additional gaming allegations; review updated sampling policy at next evaluation cycle

Concentration and conflict-of-interest assessments are based on publicly available source evidence as of June 2026.

[CU019, CU020, CU021, CU022, CU023, CU024]
FU004: Community Adoption Growth: Key Milestones

Discrete adoption milestones showing LMArena's community growth trajectory from 2023 through January 2026.

Values mix distinct metrics (cumulative votes, dataset prompts, MAU) to show trajectory; not directly comparable. MAU are monthly active users at the stated dates; vote corpus is cumulative. 240k is cumulative votes at the Mar 2024 paper; 50M+ cumulative votes reported by Jan 2026.

[CU001, CU007, CU008, CU009, CU010]

6.6 Exhibits

Chapter 07

07Risks

7.1 Conflict of Interest and Benchmark Integrity Risk

LMArena's cardinal risk is the structural conflict between its independent-arbiter identity and its revenue model. The company charges the same AI laboratories — OpenAI, Google, xAI — whose models appear on its public leaderboard for paid evaluation services. This dual role creates incentives, real or perceived, for preferential treatment. Academic researchers from Cohere, Stanford, MIT, and AI2 documented these dynamics in "The Leaderboard Illusion" (arXiv 2504.20879), finding that a handful of large providers could privately test multiple model variants, selectively disclose only their best scores, and receive disproportionately more evaluation battles. Meta's Llama 4 Maverick episode — in which Meta submitted a conversationality-optimized model to LMArena that it never publicly released, achieving a top-two ranking before the discrepancy was publicly exposed — demonstrated that policy gaps can allow benchmark gaming even without explicit rule violations. LMArena updated its policies in April 2026 but the fundamental tension between commercial dependency and perceived neutrality persists. Commentators including independent researcher Gwern have called the leaderboard "a cancer" and questioned whether its signal retains scientific value. The ucstrategies.com analysis found that models tuned specifically for Arena preferences can inflate scores by up to 100 Elo points, while SurgeAI analysis found evaluators disagreed with LMArena votes 52% of the time. If the credibility narrative fractures permanently — through a high-profile bias investigation, a paying-lab scandal, or a regulatory inquiry into AI benchmark accuracy claims — LMArena's entire product value proposition collapses simultaneously in both its consumer-facing leaderboard and its enterprise evaluation revenue. [CR001, CR002, CR003, CR004, CR005, CR006]

Regulatory / legal risk register
Risk / Rule / CaseJurisdictionStatusLikelihoodSeverityMitigationResidual ExposureDiligence Path
GDPR data-sharing adequacy — user prompts shared with AI providers without explicit granular consentEU / EEAActive compliance obligation; policy updated Sep 2025HighCriticalPrivacy policy updated; GCP Sensitive Data Protection API used before data sharingEnforcement action risk if EU DPA audits data-sharing practices; right-to-erasure gap for trained modelsRequest DPA correspondence; review data-processing agreements with AI providers
EU AI Act — GPAI model transparency documentation requirementsEUIn force since Aug 2025 for GPAI obligations; high-risk system obligations Aug 2026MediumHighDocumentation of methodology in academic papers and open-source Arena-Rank repositoryNo formal GPAI documentation attestation found publicly; unclear whether LMArena is a 'systemic risk' model provider or a downstream deployerRequest LMArena's EU AI Act compliance filing or self-assessment document
Digital Services Act — potential VLOP obligations at scaleEUApplicable if platform reaches thresholds (>45M EU monthly active users)Low-MediumHighPlatform below VLOP threshold as of June 2026 based on public figuresAs user base grows toward VLOP trigger, systemic risk assessments, independent audits, and transparency reporting become mandatoryMonitor MAU disclosures; engage DSA legal counsel
US state privacy laws — CPRA, VCDPA, CPA and othersUSA (multi-state)Ongoing compliance requiredMediumMediumPrivacy policy acknowledges state-law rights sectionNo external privacy audit or CPRA attestation found; California AG enforcement is active in 2026Request privacy audit; verify CPRA compliance program
Intellectual property — user prompt copyright and training data provenanceGlobalUnsettled law; active AI copyright litigation industry-wide 2025–2026MediumMediumTerms of service grant LMArena license to user content; prompts shared with providersCopyright infringement exposure if user-submitted content used in model training without adequate licensingReview terms of service; audit data-use flow to provider training pipelines
Benchmark accuracy claims — FTC advertising truth-in-advertising riskUSANo known enforcement action as of run dateLowMediumMethodology is open-sourced; statistical limitations disclosed in academic publicationsAny commercial claim that Arena rankings equal 'real-world performance' may attract scrutiny under FTC Section 5Monitor FTC AI guidance; review marketing claim language

All likelihood/severity ratings are the author's assessment based on public evidence and applicable regulatory frameworks; no official regulatory correspondence from LMArena has been publicly disclosed. Rows ordered by severity (Critical → Medium). Litigation search as of 2026-06-25 found no active cases naming Arena Intelligence, Inc.

[CR007, CR008, CR009, CR010]
FR001: Risk Heatmap — LMArena Risk Inventory by Impact and Likelihood

Severity-by-likelihood heatmap mapping eight primary risks across five impact bands and four likelihood tiers, illustrating that conflict-of-interest and Goodhart's-Law risks dominate the top-right quadrant.

Likelihood and impact ratings are author assessments based on public evidence and regulatory precedent; no internal risk register has been disclosed by LMArena.

[CR001, CR007, CR012, CR017]

7.2 Regulatory, Legal, and Data Privacy Risks

LMArena operates under a complex and rapidly evolving regulatory environment. The legal entity, Arena Intelligence, Inc. d/b/a LMArena, processes personal data from 5 million monthly users across 150 countries, triggering obligations under the EU General Data Protection Regulation (GDPR), the EU AI Act (general-purpose AI model obligations active since August 2025), the Digital Services Act (potentially VLOP thresholds at scale), and a patchwork of US state privacy laws including California's CPRA. LMArena's own privacy policy (effective September 2025) explicitly warns users that prompts and generated responses may be shared publicly, raising user-consent adequacy risks under GDPR Article 7. The right-to-erasure challenge for trained evaluation models (GDPR Article 17) is a known compliance gap across the AI industry with no settled enforcement interpretation. The EU AI Act's general-purpose AI documentation and transparency requirements impose additional operational overhead. LMArena's practice of sharing user prompts with AI providers for model improvement purposes creates data-flow obligations that must be covered by data processing agreements in each jurisdiction. The EU AI Act's definition of "high-risk" AI system has not explicitly named AI benchmarking platforms, but evaluations in regulated domains (law, medicine, employment) could attract sector-specific scrutiny. No litigation directly involving LMArena or Arena Intelligence, Inc. was found in public records as of the run date; however, the company has not disclosed its compliance attestation status or any regulatory inquiries. IP risk from user-submitted prompts used in training is an unsettled area of law across multiple jurisdictions. [CR007, CR008, CR009, CR010, CR011]

Operational and Security Risk Register
Failure ModeLikelihoodSeverityMitigation MaturityResidual ExposureUnresolved Gap
AI provider API withdrawal — lab pulls access amid competitive dispute or rating protestMediumCriticalLow (no disclosed SLA with providers)Full leaderboard gap for pulled model; revenue risk if paying customerNo public API SLA or redundancy plan disclosed
Coordinated sybil / voting manipulation campaign by state-actor or competitorMediumHighModerate (anonymous voting; some bot-detection; reweighting algorithm)Score corruption difficult to detect in real time; methodology revision required post-hocNo public anomaly-detection disclosure; Battles in Direct position-bias correction shows reactive capability
Major data breach — community prompt dataset or enterprise evaluation results leakedLow-MediumHighLow-Medium (GCP Sensitive Data Protection; standard enterprise security)Loss of user trust; potential GDPR breach notification obligation; evaluation IP exposed to competitorsNo SOC 2 or ISO 27001 certification disclosed as of run date
Cloud infrastructure outage — sustained downtime for battle matching or API routingLowMediumModerate (cloud provider SLAs; multi-region deployment assumed but not confirmed)Real-time evaluation continuity breaks; enterprise SLA breach possibleNo public status page SLA or uptime commitments found
Methodological overfitting — ranking divergence from real-world model quality acceleratesHighHighModerate (open-source methodology; academic paper publications; policy updates)Leaderboard loses credibility signal; developer and enterprise churnNo external statistical audit of scoring methodology published post-Leaderboard Illusion paper
Model distillation attack — adversary systematically mines Arena evaluation distributionMediumMediumLow (data is partially public by design; no anti-scraping controls disclosed)Competitive information leakage; Arena data advantage erodedDistillation risk documented in AI security literature; LMArena has not disclosed countermeasures

Mitigation maturity ratings based on publicly available information only. No internal security or infrastructure audit results have been published by LMArena as of 2026-06-25. Rows ordered by severity.

[CR012, CR013, CR014]
FR002: Risk Transmission Map — How LMArena Risks Flow to Revenue and Credibility

Directed acyclic graph showing causal pathways from root-cause risks to downstream impacts on revenue, user engagement, regulatory exposure, and valuation.

[CR001, CR003, CR007, CR017]

7.3 Operational, Technical, and Security Risks

LMArena's operational risks center on platform reliability, data quality integrity, and the security of its community-generated evaluation dataset. With 60 million monthly conversations flowing through its infrastructure, a sustained outage or data breach would directly damage the real-time, continuous-evaluation proposition that differentiates LMArena from static benchmarks. The platform depends heavily on cloud infrastructure (compute costs for running multiple live AI model APIs simultaneously) whose pricing and availability are controlled by third parties including the same AI labs LMArena evaluates. A provider pulling API access — feasible if a lab disputes a rating or competitive dynamic shifts — would immediately create gaps in the leaderboard. Statistical integrity is another operational risk: LMArena's Elo/Bradley-Terry scoring relies on a sufficiently uniform and unbiased distribution of battles; any systematic prompt injection, coordinated voting campaigns, or sybil attacks on the platform's user base would corrupt the ranking signal without necessarily being detectable in real time. The shift to "Battles in Direct" (May 2026) where 10% of direct-chat sessions generate battle votes required post-hoc bias corrections for position bias and same-organization indicator bias, demonstrating that methodological updates can shift model rankings after the fact. Security of the evaluation dataset — a commercially valuable asset — against model distillation attacks or unauthorized scraping is an ongoing concern, particularly given that leading AI labs have financial incentives to study the evaluation distribution. Employee headcount of ~41 creates key-person concentration for a platform of this scale. [CR012, CR013, CR014, CR015, CR016]

Partner and Dependency Risk Register
DependencyCounterpartyRoleConcentrationFailure ScenarioSeverityMitigationResidual Exposure
AI lab revenue — top-3 paying labs (OpenAI, Google, xAI)OpenAI / Google / xAIPrimary commercial customersHigh (estimated >50% of $30M ARR)Customer exits evaluation contract; leaderboard data conflictCriticalMulti-lab customer base; open-source community neutrality pledgeRevenue cliff if one Tier-1 lab defects
AI provider API access — all models in battles require live APIAll major labsModel capability providersHigh (each lab controls its own API)Lab withdraws API access; model drops from leaderboardHighHistorical track record of stable API access; policy commits to public model inclusionNo contractual lock-in guarantees; any lab can withdraw with short notice
Cloud compute infrastructureGCP / major cloud providerCompute, storage, data protection servicesHigh (GCP Sensitive Data Protection explicitly cited)Cloud outage, pricing increase, or terms changeHighStandard cloud enterprise agreements; redundancy assumedFull platform downtime if single-cloud dependency
UC Berkeley / academic research pipelineUC Berkeley SkyLabTalent source; research legitimacy; grant historyMediumKey researchers depart for industry; academic collaboration winds downMediumTeam has already incorporated independently; research publication continuesCredibility loss if academic separation becomes adversarial
Investor relationships — Felicis (lead) and UC InvestmentsFelicis / UC InvestmentsCapital providers; strategic anchorsMedium (Peter Deng at Felicis was OpenAI veteran)Investor conflict of interest if portfolio companies dispute Arena rankingsMediumGovernance structure not publicly disclosed; Felicis GP's OpenAI background creates appearance riskPerceived independence risk if investor conflict surfaces publicly

Revenue concentration estimates are inferred from customer count (~100) and stated ARR ($30M) assuming Pareto distribution typical of early-stage B2B SaaS. No contractual terms or formal SLA between LMArena and AI providers have been publicly disclosed as of 2026-06-25.

[CR017, CR018, CR019]
FR003: Dependency Map — Critical Platform Dependencies

Illustrates LMArena's critical infrastructure and commercial dependencies on cloud providers, AI lab APIs, academic talent pipeline, and investor stakeholders.

[CR013, CR018]

7.4 Partner Concentration, Financial, and Execution Risks

LMArena's commercial model depends on a small set of large AI labs and enterprises for the majority of its $30M annualized revenue. With approximately 100 paying customers as of early 2026 and the top clients being OpenAI, Google, and xAI, revenue concentration is material: the loss of even one or two Tier-1 lab relationships could represent a disproportionate revenue decline. The same labs are simultaneously LMArena's best customers and its most motivated potential gaming adversaries, creating a tension with no clean resolution. The $1.7 billion valuation implies a ~57x revenue multiple on $30M ARR, pricing in substantial growth expectations that require maintaining trust, expanding into new domains, and upselling existing customers. At 41 employees, execution capacity is thin relative to the product surface area LMArena has already committed to (Agent Arena, WebDev Arena, Vision, Video, Search, Code, Document leaderboards plus commercial evaluations). The company's origins as a UC Berkeley research project mean that founding team academic ties and publication obligations may divert leadership attention from commercial execution. Cash burn dynamics are undisclosed; at $250M raised and 41 headcount, the runway appears long, but infrastructure costs at 60M monthly conversations are substantial. Future funding risk is moderate given market conditions but not zero: if the AI evaluation market fails to grow to the $3.8B projected by 2030, subsequent rounds may price in disappointment. The $382,500 Recall Capital-LMArena SEC Form D filing (February 2026) reflects secondary-market investor demand but is not evidence of the company's own financial health. [CR017, CR018, CR019, CR020, CR021]

People and Execution Risk Register
Role / FunctionDependency or GapLikelihoodSeverityMitigationDiligence Path
CEO — Anastasios AngelopoulosFounding CEO; public spokesperson; academic credibility anchor; all strategic relationships flow through himLow-MediumCriticalCo-founder Wei-Lin Chiang as CTO; Ion Stoica as advisorSuccession plan and key-man clause terms not publicly disclosed; request board documentation
Statistical methodology teamCore Elo/Bradley-Terry scoring expertise concentrated in small academic team of ~41MediumHighOpen-source Arena-Rank repository; academic papers provide external checksIdentify team size and bench depth for methodology; headcount disclosure
Enterprise sales and customer successNo publicly named enterprise sales leadership; commercial product launched Sep 2025MediumHighRapid ARR growth to $30M suggests initial GTM tractionRequest org chart; identify VP Sales hire status
Security and compliance leadershipSOC 2 / GDPR DPO / EU AI Act compliance roles not publicly evidenced at 41 headcountHighHighPrivacy policy is in place; GCP data protection tools citedVerify DPO designation; request compliance org structure
Research scientist retentionCompetitive talent market; frontier AI labs pay 2–3x startup compensationHighMediumEquity; mission-driven culture; UC affiliationReview option pool size; retention cliff dates

Headcount of 41 sourced from Latka (January 2026). Organization chart not publicly available. Succession and key-man provisions are not publicly disclosed. Rows ordered by severity.

[CR020, CR021]
Mitigation and Kill Criteria
RiskMonitorable TriggerThreshold / EventAction Implication
Conflict-of-interest credibility collapseCoverage sentiment in Tier-1 AI press; academic paper count criticizing Arena methodologyTwo or more peer-reviewed papers in a single quarter documenting systematic bias that LMArena does not rebutThesis break: exit or suspend evaluation; reopen only if methodology independently audited
Top-lab customer churnRenewal of commercial evaluation agreements with OpenAI, Google, xAIAny one of the top-3 labs publicly terminates or publicly disputes evaluation resultsMonitor contract renewal dates; escalate if renewal in doubt
Regulatory enforcement actionEU DPA inquiries, FTC requests for information, or state AG investigations disclosedAny formal regulatory investigation opened against Arena Intelligence, Inc.Thesis break: freeze new capital deployment pending legal resolution
Benchmark credibility abandonment by developer communityGitHub stars, HuggingFace leaderboard citations, and developer forum references to Arena methodologyArena citation rate in model launch announcements drops >30% YoYThesis warning: investigate root cause and assess mitigation
Revenue concentration cliffShare of ARR from top-3 customers disclosed in board materialsSingle customer exceeds 40% of ARR and signals dissatisfactionThesis warning: diversification required before next funding round
Key-person departurePublic announcement or LinkedIn update from Angelopoulos or ChiangEither founder departs within 24 months of Series A closeThesis break: immediate hold pending succession clarity

Trigger thresholds are author-defined heuristics based on comparable early-stage AI infrastructure company diligence practice. All triggers are subject to context — a single negative event does not necessarily break the investment thesis in isolation.

[CR003, CR006, CR017]

7.5 Exhibits

Chapter 08

08Valuation

8.1 Financing and Valuation Context

Arena Intelligence, Inc. (d/b/a LMArena) raised $150 million in a Series A round announced January 2026, achieving a post-money valuation of $1.7 billion. This followed a $100 million seed round at a $600 million valuation in May 2025, bringing total capital raised to $250 million in approximately seven months. The Series A was led by Felicis and UC Investments (University of California), with participation from Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed Venture Partners, and Laude Ventures. The company's annualized "consumption run rate" — which it describes as equivalent to ARR — surpassed $30 million in December 2025, less than four months after launching its first commercial product (AI Evaluations) in September 2025. This gives a post-money valuation-to-ARR multiple of approximately 57x on the $30M run rate, compared with 8-15x typically seen for late Series A SaaS companies. However, the trajectory from $0 ARR (May 2025) to $30M ARR (December 2025) in four months is exceptional and investors are pricing forward growth, not current revenue. The SEC EDGAR Form D (accession 0002113470-26-000001, filed February 2026) for the "Recall Capital-LMArena" secondary fund at $382,500 indicates secondary-market demand at implied valuations consistent with the Series A price. Cap table details — dilution from seed to Series A (approximately 17% sold at seed, approximately 9% sold at Series A) imply a pre-money Series A enterprise value of $1.55 billion based on disclosed post-money valuation and round size. No public evidence of convertible notes, preferred liquidation preferences, or prior down-round risk has been found; the company raised at a 2.83x step-up in eight months. [CV001, CV002, CV003, CV004, CV005, CV041]

Recommendation summary table
DimensionAssessmentNotes
RecommendationTrackPath to Buy conditional on methodology audit + entry <30x NTM ARR
ConfidenceMediumSufficient evidence for a view; key risks are structural and unresolved
Risk RatingHighConflict of interest + extreme multiple + unaudited governance
Valuation StanceStretched57x ARR multiple is top-decile; insufficient margin of safety at $1.7B
Decision ImplicationDo not lead or co-lead at current price; track for governance improvementRe-evaluate if ARR reaches $57M+ or entry price drops to $1.0–1.2B

Assessment is as of 2026-06-25 based on publicly disclosed financing data (Series A at $1.7B post-money, $30M ARR as of December 2025) and public risk evidence. No internal financial model, board materials, or data room has been reviewed.

[CV001, CV015]
Thesis / anti-thesis table
DimensionThesis ArgumentAnti-Thesis ArgumentWhat Would Change the View
Market structureNeutral evaluation infrastructure is essential as frontier models proliferate; no credible independent alternative at scaleAI labs will internalize evaluation over time or build coalitions to avoid third-party feesLab-built evaluation reaches comparable coverage (5M+ users); or LMArena loses 2 of top-3 paying labs
Network effects5M monthly users / 60M conversations / 400+ model evaluations since May 2025 create data moatCommunity votes can be gamed; Elo scores become a proxy for optimization, not qualityAcademic paper count critiquing Arena methodology doubles in 2026; developer citation rate drops
Conflict of interestTransparency, open-source methodology, and updated policies mitigate the risk; community trust has held post-scandalDocumented in peer-reviewed research; fundamental tension unresolvable while revenue depends on evaluated labsIndependent audit confirms methodology integrity; or revenue diversification to non-lab customers exceeds 60% of ARR
Revenue trajectory$30M ARR in 4 months from $0 is exceptional; indicates strong product-market fit in AI evaluationRevenue concentrated in ~100 customers; top-3 labs likely represent majority; any churn is disproportionateCustomer cohort data shows <20% concentration in top 3 and NRR >110%
Valuation supportWeighted-average comparable AI infra multiples support 30-40x ARR for high-growth first movers57x ARR with no path to profitability disclosed is expensive relative to enterprise SaaS mediansARR growth to $75M+ by end-2026 would compress the implied multiple to 23x at current price

Arguments represent both confirming and adverse perspectives synthesized from public sources. No internal company materials have been reviewed. Rows represent five key investment dimensions; each has independent supporting evidence.

[CV011, CV012, CV013, CV014]

8.2 Valuation Analysis — Multiple Lenses

Revenue-multiple analysis: At $30M ARR and $1.7B valuation, the ARR multiple is 57x — placing LMArena in the top decile of late-stage AI infrastructure funding rounds. The comparable universe for AI evaluation infrastructure is thin: Weights and Biases was acquired by CoreWeave for $1.4 billion in March 2025 (a training/evaluation platform with broader tooling), Scale AI has been valued in the multi-billion range with revenues an order of magnitude larger. For early-stage pure-play AI benchmarking, no direct public comparables exist. The most relevant analogies are AI "picks and shovels" infrastructure: LMArena's position as a neutral data layer resembles how Bloomberg or S&P Global serves financial markets — a trusted data franchise with high switching costs. The AI model evaluation market was $1.86B in 2025 and is projected to reach $2.36B in 2026 and $6.24B by 2030 at 27.3% CAGR per The Business Research Company. If LMArena captures 5-10% of a $3B market by 2028, revenue of $150-300M at a 15-20x multiple implies a $2.25-6B valuation — the base-to-bull range for a public-comparable exit. Growth sustainability: reaching $30M in four months implies an implied growth rate of $7.5M per month. Sustaining even 50% of this adds $45M ARR in the next 12 months, reaching $75M ARR by end-2026. At a compressed 25x multiple (consistent with a maturing enterprise SaaS franchise), that implies a $1.875B valuation — roughly flat to current. At 20x multiple (public SaaS median for high-growth), $75M ARR yields $1.5B. The current $1.7B valuation thus depends on either (a) sustaining rapid ARR growth, (b) maintaining the premium platform multiple, or (c) both. Scenario analysis below reveals the distribution of outcomes more explicitly. [CV006, CV007, CV008, CV009, CV010]

Bull / base / bear scenario table
ScenarioARR by End-2026Revenue Multiple at EntryImplied Valuation (2028 exit)Key RisksProbability Signal
Bull$120M+14x on $120M$3–6B (at 25-50x ARR for platform)Credibility maintained; top-3 labs renew; methodology audit clearsRequires sustained $7.5M/month new ARR and zero credibility incidents
Base$60–75M23–28x on $70M$1.4–2.1B (at 20-30x ARR)Minor credibility friction; 1–2 methodology criticism cycles absorbedRequires 50% of current ARR velocity; feasible given $250M capital and 41 headcount
Bear$15–25M (churn)68–113x on $20M — unsustainable$200–500M (platform distressed)Credibility collapse; top-lab customer exits; competing evaluation platform captures marketTriggered by a single major public investigation or regulatory enforcement

ARR projections are extrapolations from the $30M December 2025 run rate and do not reflect any internal forecasts or guidance. Valuation ranges use public comparable AI infrastructure multiples (The Business Research Company; Latka; secondary market reporting). Probability signal column describes the evidence required to place a scenario, not a formal probability estimate.

[CV008, CV009, CV016]
Comparable valuation table
ComparableTypeRevenue / ARRValuationMultipleRelevance to LMArenaLimitation
Weights & Biases (CoreWeave acquisition, March 2025)Private → acquired~$150M ARR (estimated)$1.4B acquisition price~9x ARRAI developer platform with model evaluation and experiment tracking; most direct product analogBroader tooling than LMArena; acquisition not a standalone valuation; revenue not disclosed
Scale AIPrivate, late-stage$500M+ revenue (reported 2025)$14–29B reported range28-58x revenueData labeling, RLHF, model evaluation for enterprise and government; adjacent marketScale AI does direct annotation labor; different margin profile; broader revenue base
Hugging FacePrivate~$70M ARR (estimated 2024)$4.5B (2023 round)~64x ARRAI model hub with community evaluation features; brand relevance as 'neutral' infrastructureNot a benchmarking-first business; community hosting is primary product; older valuation
Bloomberg LPPrivate~$6.5B revenueEstimated $80–100B value~13-15x revenueData franchise with trusted neutral scores (e.g., Bloomberg Intelligence); long-run analogyVery different scale, asset class, and market structure; 30+ year franchise
S&P Global (ratings segment)Public (SPGI)Ratings segment ~$4.5BMarket cap $130B+~29x segment revenueRegulatory-blessed neutral arbiter with structural moat; long-run analogy for data franchiseS&P has regulatory mandate and market lock-in; LMArena operates in unregulated benchmarking
Arize AI (AI monitoring)PrivateUndisclosed$148M raised (2024)N/AAI model monitoring platform for production use; adjacent evaluation marketDifferent product focus (monitoring vs benchmarking); smaller scale
Aporia TechnologiesPrivateUndisclosed~$50M raised (2024)N/AResponsible AI monitoring and bias detection; adjacent governance marketCompliance-focused; no community evaluation component

All comparable financials are derived from public reporting (Latka, TechCrunch, analyst sources) or SEC filings where available. Revenue figures for private companies are estimates or reported third-party approximations. Multiples are computed on reported/estimated figures and should be treated as indicative, not precise.

[CV007, CV008]
FV002: Valuation Sensitivity — ARR Growth vs Exit Multiple

Illustrates how implied 2028 exit valuation varies across ARR growth scenarios ($45M, $75M, $120M) and exit multiple scenarios (15x, 25x, 40x), showing the range from stressed to bull outcomes.

ARR projections extrapolated from $30M December 2025 baseline; exit multiples based on public AI infrastructure SaaS comparables. Values in USD millions. Not financial model output; illustrative only.

[CV009, CV010, CV016]
FV003: Valuation / Return Range — LMArena Entry at $1.7B

Low/base/high 2028 exit valuation range with assumptions, illustrating return multiples for a $1.7B entry vs illustrative diluted entry.

All figures in USD millions (2028 exit value). Bear case assumes credibility collapse and ARR churn to <$25M; base case assumes sustained growth and moderate multiple compression; bull case assumes breakout growth and premium platform multiple maintenance.

[CV016]

8.3 Investment Thesis and Anti-Thesis

The investment thesis rests on three durable structural advantages: (1) community network effects — 5M monthly users generating 60M conversations create a data moat that is extremely difficult to replicate from scratch even with significant capital; (2) incumbency in a critical judgment role — leaderboard rankings influence AI lab marketing, developer adoption, and enterprise procurement in a way that creates switching costs for LMArena's paying customers; (3) market timing — real-world human preference data for AI evaluation will become more valuable, not less, as the number of competing frontier models proliferates and enterprise buyers face genuine selection complexity. The anti-thesis centers on four concerns: (1) the structural conflict of interest between evaluation revenue and evaluated subjects (documented by the Leaderboard Illusion paper) could precipitate a trust collapse that simultaneously destroys both the public leaderboard and the enterprise revenue; (2) Goodhart's Law dynamics — as Arena becomes the de facto standard, labs optimize specifically for it, degrading the signal's real-world relevance; (3) the revenue multiple is extreme and prices in a growth trajectory that requires both sustained execution and continued credibility; (4) a platform dependency risk exists where the very AI labs that are paying customers control the API access that makes the platform function, creating a structural leverage imbalance. On balance, the thesis is stronger than the anti-thesis at the right entry price, but the current 57x ARR multiple leaves insufficient margin of safety given the unresolved conflicts. [CV011, CV012, CV013, CV014]

Thesis-break and kill triggers table
TriggerThreshold / EventTransmission to ThesisAction Implication
Methodology independence breachA peer-reviewed paper or regulatory inquiry documents that commercial revenue influenced LMArena's public leaderboard rankingsCredibility foundation collapses; enterprise customers exit; valuation deflates to distressed levelImmediate hold / exit; thesis breaks without remediation
Top-lab customer churnAny of OpenAI, Google, or xAI publicly terminates evaluation contract or removes modelsRevenue cliff (majority of $30M ARR at risk); community coverage gapThesis warning; evaluate concentration and replacement pipeline before escalating
ARR growth stall$30M ARR as of December 2025 fails to reach $50M+ by Q3 2026Implies product-market fit was narrower than priced; 57x static multiple is unjustifiableThesis re-evaluation; entry price should reflect stalled-growth multiple (20-25x)
Developer citation rate collapseArena leaderboard citations in AI lab announcements drop >30% YoYLoss of status as de facto benchmark reduces pricing power and customer acquisitionMonitor quarterly; if trend persists 2 quarters, thesis weakens materially
Regulatory enforcement actionAny EU DPA, FTC, or state AG formal investigation against Arena Intelligence, Inc.Legal costs, remediation overhead, and reputational damage; potential operational constraintImmediate hold; assess severity and jurisdictional exposure before re-rating
Key-person departureAnastasios Angelopoulos or Wei-Lin Chiang departs within 18 months of Series A closeAcademic credibility anchor and customer relationship loss; succession risk at critical growth inflectionThesis break; hold pending succession and governance clarity

Trigger thresholds are author-defined investment monitoring heuristics based on comparable AI infrastructure diligence practice. They do not reflect company guidance.

[CV015, CV017, CV018]
FV001: Recommendation Logic Flow — From Evidence to Conviction

Chain from market scale and community proof through risk assessment and valuation evidence to the final recommendation of "track with conditional buy."

[CV011, CV012, CV015]

8.4 Recommendation, Scenarios, and Final Diligence Asks

The overall recommendation is "track" with a path to "buy" conditional on: (a) receipt of an independent methodology audit responding to arXiv 2504.20879 findings; (b) entry at a valuation below 30x next-twelve-months ARR (implying ARR of at least $57M before committing at the current $1.7B price); and (c) commercial contract terms confirming no preferential evaluation pricing for paying lab customers versus public leaderboard methodology. Risk rating is "high" given the conflict-of-interest and multiple-compression risks. The confidence level in this recommendation is "medium" — sufficient evidence exists to form a view, but the key risk factors are structural and unresolved. Valuation stance is "stretched" at current price. The company's trajectory, community engagement, and first-mover advantages are genuine; the price paid for them is not. Bull case requires $75M+ ARR by end-2026 with maintained credibility; bear case is a credibility collapse in which both revenue and valuation deflate to $200-400M (similar to benchmark reputation crises in other industries). Investors with long hold periods and strong governance influence who enter at 30x forward ARR or below have a credible risk/reward scenario in the base case. [CV015, CV016, CV017, CV018, CV019]

Final diligence asks table
TopicMissing EvidenceWhy It MattersOwner / Diligence Path
Independent methodology auditExternal statistical replication of Leaderboard Illusion paper findings and LMArena's published responsesCore credibility asset is unaudited; self-attestation insufficient for investment commitmentEngage independent academic or consulting firm; condition investment on delivery
Customer concentration and NRRARR breakdown by customer tier, top-3 customer share, net revenue retention rate, and cohort churn dataExtreme valuation concentration risk hidden behind aggregate $30M ARR figureRequest from data room; consider 90-day diligence extension if unavailable
Cap table and liquidation preferencesFull capitalization table with option pool, preference terms, anti-dilution provisions, and pro-rata rights57x ARR implies long hold period; liquidation stack materially affects return scenariosStandard Series A due diligence; typically available in data room
EU AI Act compliance statusGPAI model documentation, data provenance records, and copyright policy required under EU AI Act (August 2025)Non-compliance creates EU market risk and enterprise customer concern in regulated industriesRequest legal opinion; review with EU counsel before commitment
Infrastructure and SLA agreementsCloud provider agreements (GCP contract), AI provider API access agreements, and enterprise evaluation SLAsPlatform availability and revenue continuity depend on undisclosed third-party agreementsRequest in data room; pay particular attention to provider termination notice periods
Headcount and hiring planOrganization chart with current gaps, hiring plan to $10M+ ARR per employee, and key-man agreements for founders41 headcount at $30M ARR is thin; scaling to $100M ARR requires significant organizational buildRequest from management in kick-off meeting

Diligence asks are prioritized for an investor entering at or near the current $1.7B valuation. Items 1 and 2 are considered blocking for a buy decision. Items 3–6 are material for investment structuring and post-close monitoring.

[CV019]
FV004: Investment KPIs — IC-Ready Scorecard

IC-ready scoring across seven evaluation dimensions, reflecting the quality and completeness of evidence supporting each dimension as of the run date.

Scores are author judgments on a 1–10 scale based on available public evidence only. Not an algorithmic output. Governance score penalized for unresolved conflict-of-interest and absence of independent audit. Valuation score penalized for 57x ARR multiple.

[CV015, CV012, CV013]

8.5 Exhibits

Disclaimer

This report is based on public and accessible sources as of 2026-06-25. It is not investment, legal, or accounting advice and should be supplemented with management diligence, customer calls, and primary financial documents before making an investment decision.

Evidence index

Claims
IDStatementConfidenceSources
CO001 Chatbot Arena was launched in May 2023 as a public demo by the LMSYS research group at UC Berkeley's Sky Computing Lab. High SO010, SO011
CO002 The original Chatbot Arena platform was developed under UC Berkeley's LMSYS (Large Model Systems) research group, operating within the Sky Computing Lab. Medium SO010, SO019
CO003 Anastasios Angelopoulos and Wei-Lin Chiang are co-founders of LMArena (Arena Intelligence Inc.). High SO001, SO003
CO004 Ion Stoica, a UC Berkeley CS professor and co-founder of Databricks and Anyscale, is a co-founder of Arena Intelligence Inc. High SO019, SO004
CO005 Arena Intelligence Inc. was formally incorporated on April 18, 2025. High SO020, SO016
CO006 Anastasios Angelopoulos serves as CEO of Arena Intelligence Inc. High SO003, SO001
CO007 LMArena's platform enables users to submit prompts to two anonymous AI models simultaneously, vote on the preferred response, and then see which models they compared — feeding a public leaderboard. High SO001, SO010, SO011
CO008 LMArena's headquarters is located in the San Francisco Bay Area; the specific office address is not publicly disclosed. Low SO001, SO019
CO009 LMArena now operates under the brand name 'Arena'; it was previously branded 'LMArena' and before that 'Chatbot Arena.' High SO001, SO018
CO010 LMArena raised a $100 million seed round in May 2025, co-led by Andreessen Horowitz and UC Investments, at a $600 million post-money valuation. High SO005, SO004, SO006
CO011 LMArena raised $150 million in Series A financing in January 2026, co-led by Felicis and UC Investments, at a post-money valuation of $1.7 billion. High SO003, SO004, SO002
CO012 LMArena's Series A investors include Andreessen Horowitz, The House Fund, LDVP, Kleiner Perkins, Lightspeed Venture Partners, and Laude Ventures alongside lead investors Felicis and UC Investments. High SO002, SO003
CO013 LMArena's total capital raised as of January 2026 is approximately $250 million across seed and Series A rounds. High SO004, SO003
CO014 LMArena's $1.7 billion Series A post-money valuation represents nearly triple its $600 million seed-round valuation achieved approximately nine months earlier. High SO003, SO004
CO015 LMArena's annualized consumption run rate surpassed $30 million in December 2025, approximately four months after commercial product launch. High SO003, SO004
CO016 LMArena launched its commercial AI Evaluations product in September 2025. High SO007, SO004
CO017 As of January 2026, LMArena reports more than 5 million monthly users across 150 countries. Medium SO003, SO004
CO018 LMArena processes more than 60 million conversations per month across its evaluation platform. Medium SO003, SO004
CO019 LMArena's community accumulated over 50 million votes across text, vision, web development, search, video, and image modalities by December 2025. Medium SO002, SO003
CO020 LMArena has evaluated more than 400 distinct AI models including both open-weight and proprietary systems since founding. Medium SO002, SO003
CO021 LMArena's pre-commercial research was supported by grants and donations from Google's Kaggle platform, Andreessen Horowitz, and Together AI, primarily in compute resources and cash. Medium SO005, SO015
CO022 LMArena's leaderboard uses a Bradley-Terry model / Elo-based rating system with pairwise crowdsourced comparisons to rank AI models. High SO011, SO012, SO010
CO023 LMArena released 145,000 open-source battle data points from expert and occupational evaluation categories in late 2025. Medium SO002
CO024 LMArena's AI Evaluations commercial product includes SLAs with committed delivery timelines, representative feedback samples, and community-grounded performance analytics for model labs and enterprises. Medium SO007, SO003
CO025 LMArena's named commercial customers include OpenAI, Google, and xAI, which use its evaluation services to improve their production models. High SO003, SO004
CO026 LMArena has expanded beyond text evaluation to include Search Arena, WebDev Arena, Vision Arena, text-to-image, video generation, and Agent Arena as of June 2026. High SO024, SO025, SO026, SO009
CO027 In April 2025, researchers from Cohere, Stanford, MIT, and Ai2 published 'The Leaderboard Illusion,' alleging LMArena systematically allowed certain AI companies to privately test multiple model variants and selectively disclose only top-performing scores. High SO014, SO016
CO028 The Leaderboard Illusion paper identified that Meta privately tested at least 27 Llama 4 model variants on Chatbot Arena between January and March 2025, ahead of the public release. High SO014, SO016
CO029 The Leaderboard Illusion paper estimated that Google and OpenAI each received approximately 19–20% of all Chatbot Arena battle data, while 83 combined open-weight models received only approximately 29.7% of total data. Medium SO014
CO030 LMArena co-founder Ion Stoica publicly characterized The Leaderboard Illusion paper as containing 'inaccuracies' and 'questionable analysis,' and LMArena invited all model providers to submit more models for testing. Medium SO016
CO031 Critics including researchers at the Allen Institute for AI and King's College London argued that LMArena's user base skews toward technical programmers and AI enthusiasts, making its benchmark unrepresentative of general end-user preferences. Medium SO015, SO020
CO032 LMArena published updated transparency and leaderboard policies as of April 30, 2026, committing to open-sourcing evaluation pipelines and releasing portions of data to support auditing. Medium SO008
CO033 Ion Stoica previously co-founded Databricks (valued at approximately $43 billion) and Anyscale (the Ray computing framework), providing the founding team with deep commercialization experience. High SO019, SO016
CO034 LMArena reached a $1.7 billion valuation approximately three years after founding and within seven months of launching its first commercial product. Medium SO004, SO003
CO035 LMArena's commercial business model charges AI labs and enterprises for evaluation services including community-grounded model performance analysis across domains such as software engineering, law, and medicine. Medium SO003, SO007
CO036 LMArena launched Agent Arena in June 2026 with a causal inference methodology using treatment effect estimation to evaluate multi-component AI agents. Medium SO024
CO037 In early April 2025, Meta submitted a specially arena-optimized, unreleased variant of Llama 4 Maverick to Chatbot Arena that ranked 2nd overall; when the standard public release was subsequently scored, it ranked 32nd. High SO027, SO028
CO038 LMArena's deduplication system filters approximately 10% of submitted votes and its identity-leak detection pipeline removes fewer than 4% of votes, per July 2025 methodology updates. Medium SO009
CO039 LMArena has not publicly disclosed its total headcount; the company's Series A press release stated funds would be used to expand the technical team. Medium SO003, SO002
CO040 No debt facilities, convertible notes, secondary transactions, or SAFEs have been publicly disclosed for LMArena as of June 2026; financing has been entirely via equity rounds. Low SO003, SO004
CO041 The Leaderboard Illusion paper estimated that even limited additional Chatbot Arena battle data can yield relative performance gains of up to 112% on Arena-specific benchmarks, demonstrating the value of asymmetric data access. Medium SO014
CO042 LMArena's Chatbot Arena paper (arXiv:2403.04132) reported crowdsourced human votes achieve over 80% agreement with expert rater judgments, establishing methodological validity for the platform's ranking approach. High SO011, SO012
CO043 Arena Intelligence Inc. was spun out from UC Berkeley's LMSYS research group with UC Investments (University of California) serving as both lead investor and institutional backer of the spinout; the formal IP licensing terms governing transfer of the Chatbot Arena methodology from the university to the commercial entity are not publicly disclosed. Medium SO005, SO020
CM001 The broad AI model evaluation platform market is sized at $1.86B in 2025 and $2.36B in 2026, implying 27.3% growth into 2026. High SM001, SM002, SM003
CM002 The same broad market lens projects the AI model evaluation platform category to reach about $6.24B by 2030 at a 27.3% CAGR. Medium SM001, SM003
CM003 A narrower model evaluation and benchmarking tools lens implies roughly a $0.85B market in 2026 at around 7.3% CAGR, materially below the broad platform estimate. Medium SM004
CM004 Gartner forecasts worldwide AI spending to rise 47% to about $2.59T in 2026, creating a large upstream budget pool for evaluation software. High SM005, SM013
CM005 Presenc AI reported that 78% of Global 2000 companies had at least one AI workload in production in Q1 2026, consistent with fast-rising demand for recurring evaluation and governance workflows. High SM005, SM006
CM006 LMArena competes in the evaluation software layer that includes public benchmarking, release testing, domain scoring, and workflow-level quality measurement rather than core model training or hosting. Medium SM014, SM016, SM017
CM007 Chatbot Arena introduced LMArena's core pairwise human-preference comparison model, making public leaderboard trust a foundational but not sufficient part of the company's market. High SM014, SM015
CM008 Agent Arena broadens LMArena from single-turn chatbot ranking into agent and workflow evaluation, increasing the relevance of multi-step enterprise use cases. Medium SM016, SM017
CM009 The sharp difference between broad and narrow market estimates is best explained by scope: broad reports appear to include wider enterprise evaluation-platform workflows, while narrow reports emphasize benchmarking tools only. Medium SM001, SM003, SM004
CM010 Scale AI launched SEAL Showdown in 2026 with respondents across 100+ countries, 70+ languages, and 200+ professional domains, positioning it as a directly competitive benchmark product. Medium SM008, SM009, SM025
CM011 Scale says GPT-5 tops all SEAL categories while Gemini 2.5 Pro leads most LMArena categories, showing that benchmark outcomes are sensitive to task mix and rater composition. Medium SM008, SM009
CM012 Future AGI's competitor roundup places Galileo, Arize, Patronus, and MLflow in the same evaluation-tool conversation as LMArena, implying a fragmented specialist field. Medium SM004, SM007
CM013 Precedence Research lists major platform and infrastructure vendors such as AWS, Google, Microsoft, IBM, Databricks, Hugging Face, and Arize AI among category participants, implying competition from bundled as well as standalone products. Medium SM004, SM007
CM014 CoreWeave's roughly $1.4B acquisition of Weights & Biases in March 2025 shows strategic M&A appetite around evaluation-adjacent workflows such as experimentation, observability, and model quality. Medium SM004, SM021
CM015 TechCrunch reported that LMArena reached a $1.7B valuation, about $30M ARR, and commercial customers including OpenAI, Google, and xAI four months after launching its product. Medium SM012, SM024
CM016 LMArena says its commercial product targets law, medicine, and engineering workflows, signaling that high-stakes domain evaluation is central to its monetization strategy. Medium SM012, SM016
CM017 Likely budget owners for evaluation software are AI platform leads, model quality owners, or domain product teams rather than generalized IT procurement alone. Medium SM007, SM016, SM017
CM018 Frontier labs typically buy evaluation for model release, trust, and competitive benchmarking, while enterprises buy it for domain reliability, governance, and workflow quality. Medium SM016, SM017, SM012
CM019 The buyer universe is concentrated at the frontier-lab end but much broader across regulated and expertise-heavy enterprise verticals, creating two different procurement motions. Medium SM012, SM016, SM007
CM020 Law, medicine, and engineering are promising beachheads because domain failures there are expensive enough to justify paid evaluation rather than relying on public leaderboard performance alone. Medium SM012, SM016
CM021 A practical LMArena-relevant 2026 SAM is a subset of the broad $2.36B TAM, likely concentrated in frontier labs plus evaluation-heavy enterprise verticals rather than the full surrounding AI-tooling economy. Low SM001, SM004, SM012, SM016
CM022 A plausible 2026 SAM range for standalone trusted evaluation software relevant to LMArena is about $0.3-0.8B after excluding most bundled hyperscaler, generic MLOps, and non-paid benchmarking activity. Low SM001, SM004, SM016
CM023 A plausible near-term SOM range for LMArena is about $30-150M because the company has reported roughly $30M ARR and could deepen spend within existing frontier-lab and high-stakes enterprise customers. Medium SM012, SM013, SM016
CM024 The market should be treated as contested because a narrow benchmarking-tools lens can be roughly one-third or less of the broad platform estimate depending on what is included. Medium SM001, SM003, SM004
CM025 Gartner's 2026 forecast includes about $453B of AI software spend and $32.6B of AI model spend, indicating that evaluation budgets need only capture a small fraction of upstream AI activity to support category growth. Medium SM005, SM013
CM026 As AI workloads move into production at large enterprises, evaluation shifts from one-off benchmarking toward recurring regression testing, governance, and release gating. Medium SM005, SM006, SM016
CM027 The rise of agentic AI increases evaluation demand because buyers need to measure multi-step task completion and causal workflow reliability rather than only single-response quality. Medium SM016, SM017
CM028 Benchmark rivalry among frontier labs is itself a demand driver because model providers want external proof points to market releases and defend product claims. Medium SM008, SM014, SM015
CM029 New benchmark launches such as SEAL Showdown confirm that benchmark design is now a contested product category rather than a settled research utility. Medium SM008, SM009, SM018
CM030 Commercializing evaluation in law, medicine, and engineering makes the category more monetizable because quality signals are tied to high-cost business workflows instead of casual consumer usage. Medium SM012, SM016
CM031 Even without exact public compliance budgets, regulated and high-stakes deployments are likely to sustain third-party evaluation demand because internal trust and audit requirements are higher than for generic chat use cases. Low SM016, SM017, SM005
CM032 The Leaderboard Illusion paper argues that benchmark contamination, hidden sampling choices, and over-interpretation of small score gaps can distort arena-style rankings. High SM010, SM018, SM019
CM033 TechCrunch reported in 2024 that Chatbot Arena's user base and interaction style may bias results, weakening its usefulness as a universal benchmark. High SM010, SM011
CM034 TechCrunch reported in 2025 that LM Arena faced accusations of helping top labs game its benchmark, creating a direct commercial trust risk for public-score-based products. Medium SM018, SM020
CM035 The Verge's coverage of benchmark gaming around Meta's Llama 4 Maverick suggests that optimization against public benchmarks is ecosystem-wide rather than unique to LMArena. Medium SM018, SM020
CM036 Critics of crowdsourced AI benchmarks argue that rater representativeness and benchmark ethics are material issues, implying that some enterprise buyers may prefer curated private evaluations over open-arena votes. Medium SM019, SM011, SM010
CM037 Standalone vendors like LMArena likely face pricing pressure wherever evaluation is bundled into broader cloud, experimentation, or observability stacks. Medium SM004, SM007, SM016
CM038 Public pricing and average contract values for LMArena and most direct competitors are undisclosed, preventing reliable bottom-up revenue-to-market-share validation from open sources. Medium SM007, SM012, SM016
CM039 No verified public source in this run discloses what share of frontier-lab or enterprise AI budgets is actually allocated to evaluation tooling. Medium SM012, SM016
CM040 Competitor revenue disclosure is sparse for Galileo, Patronus, Arize, Future AGI, and other specialist vendors, making independent startup-share mapping incomplete. Medium SM004, SM007
CM041 LMArena's reported roughly $30M ARR suggests it may already represent a meaningful share of a narrow benchmarking-tools niche while remaining only a small share of the broad AI evaluation platform market. Medium SM001, SM004, SM012
CM042 Because broad and narrow market definitions differ so much, investors should validate how much of LMArena's revenue comes from public benchmarks, agent evaluation, and enterprise domain workflows before relying on TAM-based share math. Medium SM001, SM004, SM016, SM017
CM043 A three-tier lens of $2.36B TAM, about $0.3-0.8B SAM, and about $0.03-0.15B SOM is directionally consistent with the verified evidence but remains an analytical construct rather than a published market model. Low SM001, SM004, SM012, SM016
CM044 LMArena's buyer-segment matrix is driven more by use-case needs such as benchmark credibility, workflow depth, and domain rigor than by simple company-size segmentation. Medium SM007, SM016, SM017, SM008
CM045 Evaluation spend narrows from open experimentation to recurring governance spend as AI systems move into production, which is why adoption-funnel economics depend on deployment depth rather than benchmark traffic alone. Medium SM006, SM016, SM017
CM046 A low-mid-high range of roughly $0.85B, $2.36B, and $9.57B shows how the category can look small, medium, or strategic depending on whether the lens is narrow 2026 benchmarking, broad 2026 platforms, or longer-horizon strategic infrastructure. Low SM001, SM004
CP001 LMArena's platform serves more than 5 million monthly users across 150 countries as of January 2026. High SP007, SP008
CP002 LMArena generates more than 60 million model-comparison conversations per month as of January 2026. High SP007, SP008
CP003 LMArena users span 150 countries according to the company's January 2026 Series A announcement. Medium SP008
CP004 LMArena has evaluated more than 400 public models and conducted more than 300 pre-release tests across multiple modalities as of the two-year anniversary in 2025. Medium SP007
CP005 LMArena has released more than 1.5 million community-contributed prompts as open data for research use, supporting its open-access mission. Medium SP007
CP006 LMArena uses a pairwise comparison approach and Elo-style statistical ranking across blind model battles, as described in the 2024 Chatbot Arena paper by Chiang, Zheng et al. High SP001, SP023
CP007 LMArena publicly launched its commercial AI Evaluations enterprise product in September 2025, marking its entry into the paid evaluation-as-a-service market. Medium SP008
CP008 LMArena has established partnerships with OpenAI, Google, Anthropic, Meta, and xAI to make their flagship models available for community evaluation on the platform. Medium SP007, SP008
CP009 FastChat, the open-source platform backing LMArena, has powered over 10 million chat requests for 70+ LLMs, according to the GitHub repository. Medium SP011
CP010 The LMSYS Org at UC Berkeley, which originated LMArena, reports 15+ projects, 79K+ GitHub stars, and 1,000+ contributors across its open-source research portfolio. Medium SP023
CP011 Scale AI operates the Scale GenAI Platform, offering enterprise evaluation, data labeling, and agent deployment services to clients including Meta, Mayo Clinic, and defense agencies. Medium SP016, SP017
CP012 Scale AI's enterprise evaluation product is private and SLA-backed, with no public leaderboard, targeting a different buyer than LMArena's public benchmark community. Medium SP016, SP017
CP013 Scale AI's enterprise engagement pricing reportedly starts at approximately $93,000 per year, with complex projects reaching $400,000 or more, per analyst aggregations. Low SP016
CP014 The HuggingFace Open LLM Leaderboard uses automated benchmark pipelines to evaluate open-source models on tasks including MMLU, GPQA, ARC, and SWE-Bench, without human preference voting. Medium SP014, SP012
CP015 The HuggingFace Open LLM Leaderboard is freely accessible to model submitters and public readers, with no commercial evaluation service attached. Medium SP014
CP016 Stanford HELM evaluates language models holistically across multiple automated dimensions including accuracy, robustness, calibration, bias, efficiency, and toxicity. High SP018, SP019
CP017 Stanford HELM is an academic, non-commercial tool with no enterprise evaluation service; it is freely accessible to any researcher or practitioner. Medium SP018
CP018 EleutherAI's lm-evaluation-harness provides over 60 standardized academic benchmarks and powers the HuggingFace Open LLM Leaderboard. Medium SP012
CP019 EleutherAI's lm-evaluation-harness has no commercial evaluation product and is freely available as open-source software, with the harness running on any infrastructure the user controls. Medium SP012
CP020 OpenAI Evals is an open-source framework for evaluating LLMs that can now be run directly in the OpenAI Dashboard, with a community benchmark registry and enterprise custom eval support. Medium SP013
CP021 BenchLM tracked 261 models across 249 benchmarks as of June 2026, distinguishing verified from provisional rankings, with pricing and speed data included. Medium SP020
CP022 ArtificialAnalysis provides independent, provider-agnostic AI model performance analytics including intelligence, output speed, latency, and cost benchmarking, without a commercial evaluation service. Medium SP021
CP023 None of the major free leaderboard competitors — HuggingFace Open LLM Leaderboard, Stanford HELM, EleutherAI harness, BenchLM, or ArtificialAnalysis — offer an enterprise evaluation-as-a-service product with SLA commitments. Medium SP012, SP014, SP018, SP020, SP021
CP024 LMArena is the only platform that combines large-scale human-preference ranking (5 M+ monthly users) with a commercial enterprise evaluation service in a single brand and infrastructure, as of June 2026. Medium SP001, SP007, SP008, SP016
CP025 LMArena benefits from network effects: a larger and more active user community generates more preference votes, which strengthens the statistical stability of Elo rankings, making the platform more attractive to labs seeking reliable signal. Medium SP001, SP007
CP026 The Elo ranking algorithm used by LMArena is open-source and publicly documented, meaning the mathematical mechanism of ranking can be reproduced by any sufficiently resourced team. Medium SP001, SP003
CP027 OpenAI Evals is built primarily for evaluating OpenAI models, which limits its independence as a neutral tool for cross-provider model comparison. Medium SP013
CP028 Singh et al. (2025) found that Meta tested 27 private LLM variants on Chatbot Arena in the lead-up to the Llama 4 public release, selecting only the best-performing score for public disclosure. High SP003, SP005
CP029 Singh et al. (2025) estimated that Google and OpenAI received 19.2% and 20.4% of all Chatbot Arena battle data respectively, while 83 combined open-weight models received only 29.7%. High SP003, SP005
CP030 After the gaming incident, Meta's vanilla Maverick model (unoptimized) was ranked approximately 32nd on the LMArena leaderboard, below GPT-4o, Claude 3.5 Sonnet, and Gemini 1.5 Pro. Medium SP006
CP031 TechCrunch reported that Chatbot Arena's user base is skewed toward tech and AI professionals; the top questions in the LMSYS-Chat-1M dataset pertain to programming, software bugs, and app design rather than general consumer use. Medium SP004
CP032 Researchers including Yuchen Lin (Allen Institute for AI) and Mike Cook (King's College London) raised concerns that LMArena's evaluation lacks construct validity, meaning it is unclear whether user votes reliably measure real model quality. Medium SP004
CP033 LMArena earns revenue by selling AI evaluation services to the same AI labs (OpenAI, Google, xAI) it ranks on its public leaderboard, creating a structural conflict of interest. High SP005, SP008
CP034 LMArena stated in response to the Singh et al. study that it has published information on pre-release testing since March 2024 and that it does not favor any model provider over another. Medium SP005
CP035 LMArena committed to create a new sampling algorithm to address concerns about unequal model battle frequency, acknowledging operational merit in some of the Leaderboard Illusion critiques. Medium SP005
CP036 LMArena has not published a public price list for its AI Evaluations enterprise product; enterprise buyers must contact the sales team for pricing. Medium SP008, SP009
CP037 Scale AI enterprise evaluation engagements reportedly start at approximately $93,000 per year according to third-party analyst aggregations, with complex projects reaching $400,000 or more. Low SP016
CP038 HuggingFace Open LLM Leaderboard, Stanford HELM, EleutherAI lm-evaluation-harness, BenchLM, and ArtificialAnalysis are all available free of charge with no enterprise evaluation contract required. Medium SP012, SP014, SP018, SP020, SP021
CP039 LMArena's cumulative preference vote dataset (6 M+ votes across 60 M conversations) is the largest publicly known human-preference benchmark dataset for LLMs, with no comparable free alternative. Medium SP007, SP008, SP001
CP040 AI labs cite their Arena Elo scores in press releases and marketing materials, creating a dependency on LMArena's ranking signal for product positioning that raises the switching cost of defecting to an alternative leaderboard. Medium SP004, SP005
CP041 No competitor has replicated LMArena's combination of real-time human-preference data at population scale with a commercial enterprise evaluation product as of June 2026, as inferred from publicly reviewed product surfaces. Medium SP007, SP016, SP020, SP021
CI001 LMArena publicly launched its commercial AI Evaluations enterprise product in September 2025, marking the company's first commercial revenue-generating product. High SI001, SI002
CI002 The disclosed $30 million annualized consumption run rate implies roughly $2.5 million of December 2025 monthly revenue at the point LMArena reported the metric. High SI001, SI002
CI003 TechCrunch noted that LMArena calls this figure a "consumption rate" — as it "describes its annual recurring revenue (ARR)" — making clear that this is a usage-based annualized run rate, not a contracted ARR with forward booking guarantees. High SI001, SI002
CI004 GetLatka estimates approximately 100 enterprise clients for LMArena as of early 2026; this is an analyst aggregation, not a company-confirmed figure. Low SI005
CI005 Dividing the $30 million annualized consumption run rate by ~100 enterprise clients implies an average contract value of approximately $300,000 per client; this is a derived estimate, not a company-stated figure. Low SI001, SI005
CI006 OpenAI, Google, and xAI are confirmed as enterprise clients of LMArena's AI Evaluations service, per the January 2026 Series A press release from LMArena. High SI001, SI004
CI007 LMArena earns revenue by providing paid AI evaluation services to AI labs and enterprises; the free public leaderboard is not the direct monetized product. High SI001, SI011
CI008 LMArena's AI Evaluations product includes comprehensive in-depth evaluations based on community feedback, auditability through representative data samples, and SLA-committed delivery timelines. Medium SI011
CI009 LMArena raised $100 million in a seed round in May 2025 at a post-money valuation of $600 million, led by Andreessen Horowitz and UC Investments, with Lightspeed, Felicis, and Kleiner Perkins also participating. High SI003, SI004
CI010 LMArena raised $150 million in a Series A in January 2026 at a post-money valuation of $1.7 billion, led by Felicis and UC Investments with participation from a16z, Kleiner Perkins, Lightspeed, House Fund, LDVP, and Laude Ventures. High SI001, SI004
CI011 LMArena has raised $250 million in total across its seed round and Series A, in approximately seven months from May 2025 to January 2026. High SI002, SI008
CI012 The LMArena Series A investor group includes UC Investments (managing University of California public funds), which the company's CIO cited as validation of LMArena's role as critical AI evaluation infrastructure. Medium SI001
CI013 An SEC Form D filing (accession 0002113470-26-000001, filed 2026-02-26) documents the formation of Recall Capital-LMArena, a Delaware LLC venture capital feeder fund established to invest in LMArena. High SI007, SI024
CI014 LMArena had approximately 41 employees as of the January 2026 Series A announcement, according to TechCrunch's reporting. Medium SI002
CI015 LMArena's primary cost categories are inferred to be AI inference compute for serving 60 million monthly conversations, engineering headcount, and community management; gross margin is not publicly disclosed. Low SI001, SI014
CI016 LMArena operates a pure software and services model with no physical manufacturing, hardware capital expenditure, or project-finance obligations reported publicly. Medium SI001, SI011
CI017 LMArena's gross margins are not publicly disclosed; analyst estimates suggest a range of 40–70%, depending on whether AI inference costs for partner model hosting are subsidized or borne by LMArena directly. Low SI011, SI006
CI018 With $250 million raised and a 41-person team as of January 2026, LMArena's estimated cash runway is 24–36 months from the Series A close, though no official burn rate has been confirmed. Low SI002, SI005
CI019 LMArena has not publicly disclosed monthly burn rate, current cash on hand, or debt obligations; these metrics are fully private. Medium SI002
CI020 LMArena stated it will use the Series A funds to operate its platform, expand its technical team, and strengthen its research capabilities. Medium SI001
CI021 LMArena earns revenue from AI labs (OpenAI, Google, xAI) it evaluates and publicly ranks, creating a structural conflict of interest between its commercial incentive and its neutrality claim. High SI006, SI010
CI022 CTOL Digital calculated that LMArena's $1.7 billion valuation is approximately 57 times the $30 million annualized consumption run rate, meaning the valuation embeds an expectation that the commercial conflict can be managed indefinitely. Medium SI006
CI023 LMArena's revenue was concentrated in its first four months of commercial operation; customer concentration risk is high given that three confirmed clients (OpenAI, Google, xAI) are the same entities ranked on the public leaderboard. Medium SI001, SI006
CI024 If enterprise clients reduce engagement with LMArena's evaluation service in response to perceived bias in rankings — as documented in the Singh et al. Leaderboard Illusion paper — the revenue model and the free leaderboard's trust premium would be impaired simultaneously. Medium SI009, SI010, SI006
CI025 No independent audit of LMArena's evaluation methodology, conflict-of-interest management policy, or financial controls has been published as of June 2026. Medium SI006, SI009
CI026 LMArena's GTM motion is a freemium funnel: the free public leaderboard attracts AI labs and enterprise technical teams, which then convert to the paid AI Evaluations service. Medium SI001, SI011
CI027 LMArena has stated that all publicly released models will be evaluated under the same methodology regardless of commercial interest, maintaining the free public leaderboard as a trust anchor for the GTM funnel. Medium SI011
CI028 LMArena has not published information about sales cycle length, CAC, payback period, or enterprise win rates for the AI Evaluations product. Medium SI002
CI029 Both LMArena and TechCrunch used the phrase "consumption rate" rather than a standard SaaS "ARR" metric for the $30 million figure, signaling that the revenue structure may be transaction-based rather than subscription-based. High SI001, SI002
CI030 The $30 million annualized consumption run rate is an annualized pace of revenue observed in December 2025, not the full-year 2025 realized revenue figure, which would be substantially lower given product launch in September 2025. Medium SI001, SI002
CI031 If LMArena's billing is usage-based (consumption), enterprise clients may reduce evaluation volume in slow periods without formal churn, making the run rate a less reliable indicator of future revenue than contracted ARR would be. Medium SI002, SI006
CI032 LMArena's $1.7 billion post-money Series A valuation implies a revenue multiple of approximately 57x the $30 million annualized consumption run rate, consistent with early-stage high-growth software but requiring significant revenue growth to justify at later stages. Medium SI006, SI002
CI033 UC Investments, which manages investment assets for the University of California system, led both the seed and Series A rounds, providing institutional and academic credibility to the investment thesis. Medium SI003, SI004
CI034 Andreessen Horowitz (a16z) participated in both the May 2025 seed round and the January 2026 Series A, indicating high-conviction early backing from a top venture firm. Medium SI003, SI004
CI035 LMArena has not published a public price list for its AI Evaluations enterprise product; enterprise buyers must contact the team at evaluations@lmarena.ai. Medium SI011
CI036 The Leaderboard Illusion paper's finding of data-access asymmetry creates a reputational risk that could reduce enterprise clients' willingness to pay LMArena for evaluation services if the methodology is perceived as commercially compromised. Medium SI009, SI010
CI037 Bloomberg reported in April 2025 — one month before the formal seed announcement — that Chatbot Arena was becoming a "real company," providing early public evidence of the commercialization timeline. Medium SI003
CI038 LMArena has released 1.5 million+ community prompts and 145K+ battle data points as open data, which supports its academic credibility narrative but generates no direct revenue. Medium SI004, SI014
CE001 LMArena operates a multi-arena AI evaluation platform at arena.ai covering text, code, search, agent, vision, image generation, video generation, image editing, and document modalities. High SE001, SE002, SE005
CE002 The core platform mechanic is a pairwise battle where two anonymous models respond to the same prompt and users vote on the preferred output. High SE001, SE014
CE003 All Arena text, code, search, and vision leaderboards use the Bradley-Terry statistical model to infer latent skill coefficients from pairwise win/loss battle outcomes. High SE009, SE014, SE015, SE018
CE004 LMArena had 5 million or more monthly users as of January 2026. High SE003, SE005
CE005 LMArena has logged more than 250 million real conversations across all arenas since its inception. Medium SE003
CE006 LMArena generates more than 2 million preference votes per month from community users. Medium SE003
CE007 LMArena's community spans users in more than 150 countries. Medium SE003
CE008 Agent Arena was launched on June 4, 2026 as LMArena's newest evaluation arena, ranking orchestrator models for autonomous multi-step task completion. High SE005, SE006
CE009 Agent Arena ranks models using causal tracing, treating each component selection as a treatment in a multi-intervention randomized controlled trial measuring five behavioral signals. Medium SE006
CE010 WebDev Arena, launched December 2024, collected over 80,000 community votes on AI-generated web applications before being superseded by Code Arena in 2026. Medium SE007
CE011 Search Arena supports 11 models from three providers (Perplexity, Gemini, and OpenAI) and has collected over 24,000 paired multi-turn user interactions. Medium SE008, SE017
CE012 Code Arena was rebuilt from WebDev Arena with a new isolated agentic coding environment, persistent sessions stored in Cloudflare R2, and live preview rendering via CodeMirror 6. High SE013, SE005
CE013 Arena-Rank is an open-source Python package (Apache 2.0) published on GitHub under the lmarena organization and installable from PyPI as 'arena-rank'. High SE009, SE001, SE020, SE023
CE014 Arena-Rank implements Bradley-Terry ranking with closed-form confidence-interval calculation and a 30x speedup over the historical FastChat implementation. High SE009, SE001, SE020
CE015 FastChat (GitHub: lm-sys/FastChat) was the original open-source platform powering Chatbot Arena but is now primarily in maintenance mode, with ranking code migrated to Arena-Rank. Medium SE019, SE009
CE016 LMArena's style-control extension adds length, markdown header count, bold count, and list count as regression covariates in the BT model to isolate substance from formatting effects; response length is the dominant style factor. High SE010, SE015
CE017 The Arena-Hard BenchBuilder pipeline (arXiv:2406.11939) automatically extracts hard prompts from live Arena data using a seven-criterion hardness labeler and generates benchmarks that achieve 98.6% agreement with human preference rankings. High SE016, SE012
CE018 LMArena defines prompt hardness using seven criteria including domain knowledge, problem-solving complexity, and real-world applicability; approximately 20% of Arena prompts have a hardness score of 6 or higher. Medium SE011
CE019 The lmarena-ai HuggingFace organization hosts multiple p2l (prompt-to-leaderboard) preference models ranging from 135M to 7B parameters and releases public battle datasets. Medium SE022
CE020 LMArena uses GCP's Sensitive Data Protection API to remove personally identifiable information from conversation data before any public release or data sharing. Medium SE004
CE021 LMArena's leaderboard policy was last updated April 30, 2026 and specifies model eligibility criteria, sampling policies, pre-release testing protocols, and data sharing rules. High SE004, SE005
CE022 LMArena's sampling policy requires that at least 20% of all battles involve only publicly available models, with remaining capacity available for unreleased or experimental models. Medium SE004
CE023 LMArena allows model providers to test unreleased models anonymously, shares results privately with the provider, and then removes the model before any public listing. High SE004, SE025
CE024 In April 2025 Meta submitted an 'experimental, chat-optimized' version of Llama 4 Maverick (not the public release) to LMArena, achieving a #2 ranking; when the public version was scored, it fell to approximately 32nd place. High SE025, SE024
CE025 Following the Meta Llama 4 incident, LMArena updated its leaderboard policies to require that pre-release model variants be explicitly labeled as customized and not represent the publicly released model. High SE025, SE024
CE026 A paper by researchers from Cohere, Stanford, MIT, and AI2 (April 2025) alleged that Meta, OpenAI, Google, and Amazon received disproportionately high sampling rates and could suppress low-scoring pre-release variants, constituting benchmark gaming. High SE024, SE025
CE027 LMArena denied the claims in the Cohere/Stanford study as containing 'inaccuracies and questionable analysis,' asserting that all model providers are allowed to submit more models for testing and that the leaderboard remains fair. Medium SE024
CE028 Independent researchers have noted that LMArena's user base is skewed toward technical and developer prompts, making it less representative of general-population or enterprise preferences. Medium SE024
CE029 Agent Arena analyzes five behavioral signals: confirmed success, praise vs. complaint, steerability, bash-error recovery, and tool hallucination rate. Medium SE006
CE030 In a recent 7-day window, Agent Mode recorded 160,480 agent tasks on LMArena's platform, with code writing as the largest category at 17.5%. Medium SE006
CE031 Agent Mode issued approximately 2 million structured tool calls in one 7-day period, including 936,000 bash calls and 550,000 file-write operations. Medium SE006
CE032 Code Arena uses Cloudflare R2 for persistent session storage and CodeMirror 6 for source-code display and live preview rendering of generated web applications. Medium SE013
CE033 Battles in Direct, launched May 2026, converts 10% of Direct Chat sessions into anonymous pairwise battles and applies position-bias and same-org-indicator corrections in the BT model. Medium SE005
CE034 Arena-Rank's open-source implementation achieves a 30x speedup over the historical FastChat-based BT implementation and uses reweighting to correct for non-uniform model sampling. High SE009, SE001, SE020
CE035 Arena-Hard-Auto v0.1 achieves 98.6% agreement with human preference rankings from Chatbot Arena and provides 3x higher model separation than MT-Bench. High SE016, SE018
CE036 The Search Arena paper was accepted at ICLR 2026, providing peer-reviewed validation of LMArena's search-augmented LLM evaluation methodology. High SE017, SE005
CE037 LMArena has not publicly disclosed SOC 2 Type II, ISO 27001, or any independent third-party security audit for its AI Evaluations commercial product or the arena.ai platform. Medium SE004
CE038 LMArena's help.arena.ai privacy policy exists but lacks enterprise data processing agreement (DPA) terms, and GDPR/CCPA compliance details are not publicly documented. Medium SE004
CE039 LMArena has evaluated over 400 public models and over 300 pre-release model variants across all its arenas since the platform launched in 2023. Medium SE005
CE040 LMArena has released 1.5 million prompts and over 145,000 battle data points publicly for open research as of early 2026. Medium SE005
CE041 Agent Arena sessions average approximately 16.5 structured tool calls, with about 75.6% of sessions using at least one tool in a measured 7-day window. Medium SE006
CE042 Agent Mode wrote 40.3 million lines of code through successful write_file calls in one 7-day window, approximately 1,000 lines per coding session. Medium SE006
CU001 LMArena had 5 million or more monthly active users as of January 2026, spanning more than 150 countries. High SU001, SU009
CU002 LMArena generates more than 60 million conversations per month and has accumulated more than 250 million conversations in total as of January 2026. High SU001, SU009
CU003 LMArena's community generates more than 2 million preference votes per month. Medium SU013
CU004 LMArena serves two distinct customer populations: a free community of millions of monthly users providing preference votes, and a small paying commercial segment of AI labs and enterprises purchasing AI Evaluations. High SU001, SU013
CU005 OpenAI, Google, and xAI are named paying customers of LMArena's AI Evaluations commercial product, drawing on evaluations to improve their models for production use cases. High SU001, SU009
CU006 LMArena partnered with OpenAI, Google, and Anthropic to make their flagship models available for community evaluation; Anthropic's paying status as an AI Evaluations subscriber is not separately confirmed. High SU001, SU009
CU007 LMArena's community grew approximately 25 times between May 2025 (seed round) and January 2026 (Series A) as reported by the company. Medium SU005
CU008 LMArena had approximately 3 million or more monthly users at the time of its commercial AI Evaluations product launch in September 2025. Medium SU013
CU009 LMArena has evaluated more than 400 public models and more than 300 pre-release model variants across its arenas as of early 2026. Medium SU005, SU006
CU010 The LMArena community released more than 50 million votes and 1.5 million open prompts by January 2026. Medium SU005
CU011 LMArena's AI Evaluations commercial product achieved an annualized consumption run-rate of $30 million in December 2025, less than four months after its September 2025 launch. High SU001, SU009
CU012 Meta submitted 27 Llama 4 model variants to LMArena for private pre-release testing between January and March 2025, then publicly disclosed only the score of the highest-performing experimental variant. High SU016, SU017
CU013 The publicly released version of Meta Llama 4 Maverick ranked approximately 32nd on the LMArena leaderboard, versus the experimental version which had ranked #2. High SU007, SU017
CU014 Anthropic's Claude models were described as currently winning the expert leaderboard for legal and medical use cases in January 2026. Medium SU003
CU015 Perplexity's Sonar models and Google Gemini were the top-ranked models in the Search Arena as of the ICLR 2026 paper, with Perplexity-Sonar-Reasoning-Pro and Gemini-2.5-Pro at the top. Medium SU019, SU024
CU016 LMArena has not publicly disclosed net revenue retention, gross revenue retention, or customer churn rates for its AI Evaluations commercial product. Medium SU011
CU017 The leaderboard changelog shows multiple model additions per week across all arenas in June 2026, serving as an indirect proxy for ongoing model-provider engagement. Medium SU015
CU018 LMArena describes its revenue as 'annualized consumption rate' rather than annual recurring revenue (ARR), implying a consumption-based billing model rather than committed subscriptions. Medium SU001
CU019 The structural conflict of interest in which the same AI labs that pay for AI Evaluations also have their models ranked on the public leaderboard was identified by independent journalists as a key credibility risk. High SU002, SU016
CU020 A paper from Cohere, Stanford, MIT, and AI2 alleged in April 2025 that Meta, OpenAI, Google, and Amazon received disproportionately high sampling rates, enabling benchmark gaming; LMArena denied specific claims and announced a new sampling algorithm. High SU016, SU003
CU021 LMArena updated its leaderboard policy after the Meta Llama 4 incident and stated that 'Meta's interpretation of our policy did not match what we expect from model providers.' High SU017, SU025
CU022 As of June 2026, only three companies (OpenAI, Google, xAI) are publicly named as paying customers of LMArena's AI Evaluations service; no enterprise non-lab customers have been publicly named. High SU001, SU009
CU023 LMArena's revenue is highly concentrated in three named AI lab customers; departure of any one of these would represent a material revenue risk given the $30M ARR run-rate and unknown diversification. Medium SU001
CU024 LMArena's community growth trajectory—25x from May 2025 to January 2026—implies organic conversion of community interest into lab evaluation demand, though the precise pathway from free user to commercial customer is not documented. Medium SU023, SU009
CU025 Prior to commercialization, Google's Kaggle, Andreessen Horowitz, and Together AI had donated compute, cloud credits, and cash to LMSYS as corporate sponsors, creating an indirect financial relationship with leaderboard participants. Medium SU016, SU010
CU026 Felicis General Partner Peter Deng stated in the LMArena Series A press release that LMArena has 'become essential infrastructure for every lab and enterprise' and that Felicis led the round because of LMArena's trustworthy real-world performance signal. High SU001, SU009
CU027 Scale AI launched a competing benchmarking product, SEAL Showdown, as a direct rival to LMArena's evaluation platform, representing competitive pressure in the AI evaluation market. Medium SU004, SU016
CU028 OpenTools.AI independently reported that LMArena was 'under fire' for benchmark bias in 2025, reflecting industry-wide concern beyond the single Cohere/Stanford paper. Low SU003
CU029 PitchBook tracks Arena Intelligence (LMArena) with a confirmed valuation history from the $600M seed in May 2025 to the $1.7B Series A in January 2026. High SU006, SU009
CU030 LMArena's original domain lmarena.ai now redirects to arena.ai, reflecting the company's rebranding as it expanded beyond language model evaluation to a multi-arena platform. Medium SU007
CU031 EDGAR full-text search confirms LMArena (entity 'Recall Capital-LMArena') has an active Form D filing for the Series A under CIK 0002113470, providing independent corroboration of the funding event. High SU008, SU009
CU032 Anthropic's Claude is described by TechCrunch podcast as currently winning LMArena's expert leaderboard for legal and medical professional use cases as of January 2026, indicating provider engagement depth. Medium SU002, SU009
CU033 LMArena's two-year anniversary blog (April 2025) confirms approximately 41% of battles involve open-source models, indicating a mixed commercial and research user base that constrains pure commercial curation. Medium SU024, SU023
CU034 No independent third-party reviews on G2, Capterra, or Gartner Peer Insights for LMArena's AI Evaluations commercial product are publicly available as of June 2026. Medium SU011
CU035 LMArena's AI Evaluations commercial product targets software engineering, law, medicine, and scientific research as economically valuable industries for paid evaluation services. Medium SU001, SU009
CR001 LMArena earns revenue by selling paid AI evaluation services to AI labs (OpenAI, Google, xAI) whose models simultaneously appear on LMArena's public leaderboard, creating a structural conflict of interest between commercial and evaluation roles. High SR015, SR011, SR014
CR002 The Leaderboard Illusion paper (arXiv 2504.20879) found that a small number of providers could privately test multiple model variants and selectively disclose only their best scores, resulting in biased Arena rankings. High SR002, SR001
CR003 Meta privately tested 27 Llama-4 model variants on Chatbot Arena in the lead-up to its Llama 4 release, selecting the best-performing variant for its public score. High SR002, SR001, SR003
CR004 Meta submitted an "experimental chat version" of Llama 4 Maverick "optimized for conversationality" to LMArena that achieved a top-two leaderboard ranking, but this model was not the same version released publicly. High SR003, SR004
CR005 SurgeAI's analysis of 500 LMArena votes found that evaluators disagreed with LMArena's outcomes 52% of the time, with "confidence beats accuracy and formatting beats facts." Medium SR006
CR006 Independent researcher Gwern described LMArena as "a cancer" and questioned whether it is worth running, reflecting significant reputational erosion among technical users. Medium SR006
CR007 Arena Intelligence, Inc. (d/b/a LMArena) processes personal data from users in EU member states and is subject to GDPR compliance obligations including data subject rights, lawful basis for processing, and international data transfer safeguards. High SR009, SR010
CR008 LMArena's privacy policy (effective September 2025) explicitly warns users that prompts, votes/ratings, and other user content may be shared publicly and with AI providers as part of the evaluation process. High SR009, SR013
CR009 The EU AI Act's general-purpose AI (GPAI) model obligations, which came into force in August 2025, impose transparency documentation, training data summaries, and copyright policy requirements on AI platform operators. High SR012, SR010
CR010 No litigation directly naming Arena Intelligence, Inc. or LMArena as a plaintiff or defendant was found in public records as of 2026-06-25. Medium SR009
CR011 LMArena uses GCP's Sensitive Data Protection API to remove personal and sensitive data before sharing conversation data with model providers or publishing it publicly, and GCP is the confirmed primary cloud infrastructure provider. Medium SR015, SR009
CR012 LMArena's leaderboard methodology update in May 2026 (Battles in Direct) discovered and corrected two new biases: position bias favoring Model A, and an advantage for models sharing an organization with prior conversational context. High SR029, SR008
CR013 LMArena processes 60 million monthly user conversations across 150 countries on cloud infrastructure, creating platform reliability, data security, and regulatory compliance obligations at scale. High SR015, SR016
CR014 No public disclosure of SOC 2, ISO 27001, or equivalent security certification for Arena Intelligence, Inc. has been found as of the run date. Medium SR009
CR015 Chatbot Arena's Elo/Bradley-Terry scoring methodology relies on sufficient uniformity of battle distribution; coordinated voting campaigns, prompt injection, or sybil attacks could corrupt the ranking signal without immediate detection. Medium SR002, SR006
CR016 AI providers including OpenAI and Google have financial incentives to study and potentially overfit to the Arena evaluation distribution, as documented by the finding that access to Arena data yields up to 112% relative performance gains on Arena Hard. Medium SR002, SR005
CR017 LMArena had approximately 100 paying customers as of early 2026, generating $30M ARR, implying material revenue concentration risk with the top AI lab customers. Medium SR026, SR015
CR018 LMArena publicly confirmed that OpenAI, Google, and xAI draw on its evaluations to improve their models, making these companies simultaneously its best customers and its most motivated potential evaluators of its gaming policies. High SR015, SR020
CR019 No contractual SLA or public agreement guaranteeing sustained AI provider API access to LMArena's platform has been publicly disclosed; any lab can withdraw its model. No known instance of API withdrawal has occurred as of the run date. Medium SR007
CR020 LMArena employs approximately 41 people as of January 2026, representing a lean team relative to the platform's operational scope and regulatory obligations. Medium SR026
CR021 The lead investor in LMArena's Series A (Peter Deng of Felicis) previously worked at OpenAI, creating an appearance-of-conflict risk between Felicis's portfolio interest and LMArena's independence claims regarding OpenAI model evaluations. Medium SR015, SR025
CR022 Arena Hard performance — a synthetic benchmark LMArena maintains — can be improved by up to 112% relative gains with additional Arena battle data, according to the Leaderboard Illusion researchers' conservative estimates. Medium SR002, SR001
CR023 The Leaderboard Illusion paper found that Google and OpenAI each received an estimated 19.2% and 20.4% of all Chatbot Arena data, while 83 open-weight models combined received only approximately 29.7% of total data. High SR002, SR001
CR024 LMArena updated its sampling policy post-April 2026 to guarantee that at least 20% of all battles involve only publicly available models, and committed to reweighting scoring so that sampling probabilities do not bias Arena scores. High SR007, SR008
CR025 The SEC EDGAR filing for "Recall Capital-LMArena a Series of CGF2021 LLC" (Form D, filed 2026-02-26, accession 0002113470-26-000001) is a secondary-market venture fund raising $382,500 specifically to invest in LMArena. High SR024, SR014
CR026 LMArena's Agent Arena leaderboard launched on June 4, 2026, expanding into agentic evaluation with behavioral signals like file downloads, disapproval events, retries, and steerability rather than static preference votes alone. High SR008, SR016
CR027 US state privacy laws including California's CPRA impose separate data-subject rights and business compliance obligations on LMArena that are distinct from GDPR requirements. High SR009, SR012
CR028 The EU AI Act imposes penalties of up to 7% of global annual turnover for prohibited AI practices and up to 3% for other violations, applicable to AI platforms with EU operations. High SR012, SR010
CR029 CTOL Digital Solutions noted that LMArena's $1.7 billion valuation implies roughly 57x the company's annualized revenue run rate of $30M, pricing in substantial growth expectations that depend on sustained trust in benchmark neutrality. Medium SR011
CR030 LMArena's commercial evaluation product, AI Evaluations (launched September 2025), provides paid evaluation services to enterprises and AI labs, creating a revenue stream directly tied to the labs it ranks publicly. High SR017, SR015
CR031 The AI model evaluation platform market is projected to grow from $1.86 billion in 2025 to $2.36 billion in 2026 and $6.24 billion by 2030, at a 27.3-27.5% CAGR. Medium SR023
CR032 LMArena's open-source Arena-Rank repository and methodology academic papers provide external auditability of the leaderboard scoring approach, which is a partial mitigation against benchmark-integrity allegations. High SR007, SR016
CR033 LMArena's Series A (January 2026) achieved a post-money valuation of $1.7 billion, nearly triple the $600 million seed valuation from May 2025, on $250M total capital raised. High SR014, SR015, SR022
CR034 The Leaderboard Illusion researchers found that proprietary/closed models are sampled at higher battle rates and have fewer models removed from Arena compared to open-weight alternatives, creating a data access asymmetry. High SR002, SR018
CR035 A TechCrunch analysis noted that LMArena's commercial relationships with OpenAI, Google, and Anthropic raise questions about whether the benchmark can be trusted to assess AI models without corporate influence clouding the process. High SR001, SR011
CR036 LMArena was previously funded through grants and donations from Google's Kaggle, Andreessen Horowitz, and Together AI — organizations with direct stakes in the models being evaluated. High SR005, SR022
CR037 LMArena stated that its commercial evaluation product provides the same methodology to all paying customers without preferential treatment, and that the public leaderboard will always be available freely. Medium SR017
CR038 The LMArena policy (updated April 30, 2026) now requires that if a publicly released model differs from the pre-release version tested on Arena, Arena will remove the model from the leaderboard until it can be re-evaluated under the requirements of this policy. High SR007, SR008
CR039 TechCrunch noted that a Llama 4 Maverick vanilla release ranked 32nd on the LMArena leaderboard after the experimental version that ranked second was withdrawn, demonstrating a 30-rank gap attributable to benchmark optimization. High SR004, SR003
CR040 At 41 employees and $30M ARR, LMArena's revenue-per-employee ratio is approximately $730K/FTE, suggesting significant infrastructure leverage but thin organizational depth for regulatory compliance, security, and enterprise scale-up. Medium SR026, SR015
CR041 LMArena's 5 million monthly users across 150 countries is substantially below the 45 million EU monthly active user threshold that would trigger VLOP designation under the Digital Services Act, making VLOP obligations unlikely in the near term. Medium SR015, SR012
CR042 LMArena has not disclosed any data breach incidents involving the community evaluation dataset or user prompt data as of the run date, but uses GCP security tools rather than published third-party security certifications. Medium SR009, SR015
CV001 LMArena raised $150 million in a Series A round in January 2026 at a post-money valuation of $1.7 billion, led by Felicis and UC Investments, with participation from a16z, Kleiner Perkins, Lightspeed, The House Fund, LDVP, and Laude Ventures. High SV001, SV002
CV002 LMArena's annualized "consumption run rate" surpassed $30 million in December 2025, less than four months after launching its first commercial product (AI Evaluations) in September 2025. High SV002, SV001
CV003 LMArena raised a $100 million seed round in May 2025 at a $600 million valuation, bringing total capital raised to $250 million by January 2026 across two rounds in approximately seven months. High SV001, SV007
CV004 The Latka database records that LMArena sold approximately 17% at the seed round ($100M / $600M) and approximately 9% at the Series A ($150M / $1.7B), providing an implied pre-money Series A enterprise value of approximately $1.55 billion. Medium SV003, SV001
CV005 At $30M ARR and a $1.7B valuation, LMArena's implied ARR multiple is approximately 57x — placing it in the top decile of AI infrastructure Series A rounds in 2025–2026. High SV013, SV001
CV006 The $30M ARR figure represents December 2025 monthly revenue annualized (run rate), not a full-year booked revenue number; the company had fewer than four months of commercial operations as of the Series A announcement. High SV002, SV008
CV007 The AI model evaluation platform market was valued at $1.86 billion in 2025 and is projected to reach $2.36 billion in 2026 and $6.24 billion by 2030 at a 27.3–27.5% CAGR, according to The Business Research Company's 2026 market report. Medium SV010, SV006
CV008 Weights & Biases (an AI developer platform with model evaluation capabilities) was acquired by CoreWeave for approximately $1.4 billion in March 2025, providing a transaction comparable for AI evaluation infrastructure valuation. Medium SV012, SV011
CV009 A bull case for LMArena requires $120M+ ARR by 2028 and sustained platform multiple above 25x, implying a $3–6B exit value; this requires sustaining the $7.5M/month new ARR velocity seen in the first four months of commercial operations. Medium SV008, SV001
CV010 A base case for LMArena projects $60–75M ARR by end-2026 (at 50% of current ARR velocity), implying a $1.4–2.1B valuation at a 20–30x ARR multiple — roughly in line with the current $1.7B mark, providing limited upside from today's entry price. Medium SV001, SV013
CV011 LMArena's core competitive advantage is a community data moat: 5 million monthly users across 150 countries generating over 60 million conversations monthly, creating a real-time human preference dataset that is extremely difficult to replicate. High SV002, SV009
CV012 LMArena's leaderboard has become a standard reference in AI lab product launches, developer procurement decisions, and media coverage, creating network effects and switching costs that reinforce its incumbent position. High SV022, SV026
CV013 The structural conflict of interest — where LMArena earns revenue from the same labs it evaluates — creates an existential risk to the investment thesis; a single high-profile investigation confirming revenue influenced rankings would collapse both leaderboard credibility and enterprise revenue simultaneously. High SV013, SV016
CV014 Goodhart's Law dynamic — where labs optimize specifically for Arena rather than for genuine capability improvement — degrades the leaderboard's real-world signal value over time, as documented in the Leaderboard Illusion paper and corroborated by multiple independent analyses. High SV014, SV016
CV015 The overall investment recommendation for LMArena is "track" with a conditional buy signal requiring an independent methodology audit and an entry price below 30x next-twelve-months ARR. Risk rating is "high"; valuation stance is "stretched." Medium SV013, SV001
CV016 The bear case for LMArena implies a valuation range of $200–500M following a credibility collapse, driven by ARR churn to below $25M and multiple compression to 10–20x distressed ARR — representing a 70–88% loss from the $1.7B entry. Medium SV013, SV014
CV017 The primary thesis-break trigger is discovery of documentary evidence that commercial revenue relationships influenced LMArena's public leaderboard outcomes; secondary triggers include churn of any top-3 lab customers or ARR failing to reach $50M+ by Q3 2026. Medium SV013, SV016
CV018 Peter Deng, the Felicis general partner leading LMArena's Series A round, previously worked at OpenAI — one of LMArena's paying evaluation customers — creating an appearance-of-conflict risk between Felicis's portfolio interest and LMArena's independence claims regarding OpenAI evaluations. High SV002, SV022
CV019 The minimum blocking diligence items before a buy decision are: (1) independent statistical audit of the Arena methodology; (2) customer cohort data showing top-3 customer share below 40% or NRR above 110%; (3) cap table with liquidation preference terms; (4) EU AI Act compliance self-assessment. Medium SV013, SV016
CV020 The SEC Form D filing for "Recall Capital-LMArena a Series of CGF2021 LLC" (accession 0002113470-26-000001, filed February 2026) is a $382,500 secondary venture fund, demonstrating secondary-market demand for LMArena exposure at valuations consistent with the Series A. High SV015, SV001
CV021 Long-run structural analogies for LMArena's potential platform value include Bloomberg LP and S&P Global's ratings segment, which operate as neutral data arbiters with structural moats and high customer switching costs — though these analogies require LMArena to first resolve its conflict-of-interest and achieve regulatory equivalence. Medium SV006, SV013
CV022 LMArena's product expansion into Agent Arena (June 2026), WebDev Arena, Search Arena, Vision, Video, and coding leaderboards suggests the company is actively increasing its potential ARR ceiling by expanding beyond text model evaluation. High SV020, SV009
CV023 A CTOL analysis noted that LMArena's 57x ARR multiple prices in the assumption that the conflict between being a revenue-generating evaluation service and an independent benchmark arbiter "can be managed indefinitely" — a structural assumption that has not been independently validated. High SV013, SV001
CV024 LMArena employs approximately 41 people as of January 2026 with approximately $730K revenue per employee, indicating high capital efficiency but thin organizational depth for maintaining a $1.7B asset at scale. Medium SV003, SV002
CV025 No down-round risk has been evidenced in LMArena's financing history; the company raised at 2.83x step-up from seed ($600M) to Series A ($1.7B) in eight months on genuine commercial traction. High SV001, SV003
CV026 The revenue concentration risk at LMArena is heightened by its approximately 100 paying customers, where the top AI labs (OpenAI, Google, xAI) likely represent a disproportionate share of $30M ARR per standard B2B Pareto distributions. Medium SV003, SV013
CV027 The AI model evaluation platform market's 27.3% CAGR projection implies that if LMArena captures a 5-10% market share in a $3B+ market by 2028, its revenue would be $150-300M — sufficient to justify the current valuation at 10-15x revenue multiples characteristic of scaled data platforms. Medium SV010, SV013
CV028 LMArena's Arena Intelligence, Inc. d/b/a structure was incorporated in 2025 in Delaware; the company is headquartered in San Francisco, California, and is a private company with no SEC reporting obligations beyond Form D filings. High SV021, SV015
CV029 No public evidence of preferred stock liquidation preferences, anti-dilution ratchets, or convertible notes has been disclosed by LMArena or its investors as of the run date, though these terms are standard in Series A financings and their absence from public disclosure does not indicate they do not exist. Low
CV030 LMArena's investor base (Felicis, UC Investments, a16z, Kleiner Perkins, Lightspeed) includes firms with direct investments in AI labs that are LMArena's paying customers, creating a potential for governance conflicts that could compromise the independence of the evaluation platform. Medium SV025, SV007
CV031 LMArena's prior pre-commercial funding through grants from Google's Kaggle and donations from Andreessen Horowitz and Together AI — organizations with evaluated models on the platform — established a pattern of commercial entanglement that the new corporate structure has not fully resolved. High SV025, SV007
CV032 Scale AI, the closest large-scale comparable in AI data labeling and evaluation services, was valued in the $14–29B range on reported revenue substantially larger than LMArena's current $30M ARR, suggesting LMArena's 57x multiple is significantly richer than Scale AI's implied multiple on comparable revenue. Medium SV010, SV013
CV033 LMArena generated 50 million votes, 400+ model evaluations, and added 145,000 open-source battle data points to the community between the seed round (May 2025) and the Series A (January 2026), demonstrating substantial community engagement growth. High SV009, SV002
CV034 LMArena's AI Evaluations commercial product provides paid evaluation services including comprehensive in-depth evaluations based on community feedback, auditability through representative data samples, and service-level agreements with committed delivery timelines. High SV019, SV002
CV035 No institutional investor has publicly disclosed a reduction in LMArena position, secondary sale, or hedge of LMArena exposure since the Series A close as of the run date; the Recall Capital secondary fund implies continued secondary-market demand. Medium SV015, SV001
CV036 LMArena's Agent Arena leaderboard (launched June 4, 2026) represents a new evaluation domain targeting agentic AI systems, expanding the platform's addressable commercial market beyond text model benchmarking. High SV020, SV026
CV037 The minimum entry valuation offering adequate risk/return profile is estimated below $1.0–1.2B (approximately 30x $40M NTM ARR), providing sufficient discount to the current $1.7B mark to compensate for conflict-of-interest, concentration, and multiple-compression risks. Medium SV013, SV001
CV038 LMArena's exit readiness is limited in the near term: the company is 18 months old commercially, has no disclosed profitability path, and an IPO would require 2–3 years of operating history and revenue scale above $150M+ to command strong public market reception; strategic acquisition by a hyperscaler is the most plausible nearer-term exit scenario. Medium SV001, SV022
CV039 LMArena's product expansion into evaluation of models across text, code, vision, video, search, documents, and agents (seven distinct modalities as of June 2026) expands the potential ARR ceiling and reduces concentration risk in the core text model benchmarking segment. High SV020, SV009
CV040 The OfficeChai analysis notes that LMArena's AI testing market is estimated at $800–900M in 2025, projected to grow to $3.8 billion by 2032, providing a TAM growth context for the $1.7B valuation. Medium SV006, SV010
CV041 Reuters syndication through U.S. News reported that LMArena's valuation tripled to $1.7 billion in about eight months, underscoring how quickly private-market pricing expanded between the $600M seed round and the January 2026 Series A. High SV033, SV007
Sources
IDPublisherTitleQuote
SO001 Arena Intelligence Inc. About Arena | Crowdsourced AI Model Evaluation Platform Created by researchers from UC Berkeley, Arena (formerly LMArena) is a community-powered platform for understanding AI performance in the real world.
SO002 Arena Intelligence Inc. Fueling the World's Most Trusted AI Evaluation Platform (Series A Blog) We've raised $150M of Series A funding led by Felicis and UC Investments (University of California).
SO003 PR Newswire / LMArena LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product.
SO004 TechCrunch LMArena lands $1.7B valuation four months after launching its product LMArena raised a $150 million Series A at a post-money valuation of $1.7 billion.
SO005 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M LM Arena... has raised $100 million in a seed funding round that values the organization at $600 million.
SO006 Bloomberg Popular AI Ranking Website Chatbot Arena Is Becoming a Real Company
SO007 Arena Intelligence Inc. New Product: AI Evaluations This service offers enterprises, model labs, and developers comprehensive evaluation services grounded in real-world human feedback.
SO008 Arena Intelligence Inc. Arena Leaderboard Policy
SO009 Arena Intelligence Inc. Arena Leaderboard Changelog
SO010 LMSYS / UC Berkeley Chatbot Arena: New Leaderboard & More Models We are releasing Chatbot Arena, an open-source evaluation platform for LLMs.
SO011 arXiv / UC Berkeley LMSYS Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies.
SO012 arXiv / UC Berkeley Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
SO013 arXiv / UC Berkeley From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
SO014 arXiv / Cohere, Stanford, MIT, Ai2 The Leaderboard Illusion We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired.
SO015 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark The evaluation is not reproducible, and the limited data released by LMSYS makes it challenging to study the limitations of models in depth.
SO016 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification.
SO017 TechCrunch Here's why most AI benchmarks tell us so little
SO018 TechCrunch The leaderboard 'you can't game,' funded by the companies it ranks
SO019 Founded.com How two Berkeley roommates built a $1.7B startup that helps you decide which AI to use The founders behind AI evaluation platform Arena, formerly known as LMArena, have turned that confusion into a business now worth $1.7 billion.
SO020 Winbuzzer Experts Challenge Validity and Ethics of Crowdsourced AI Benchmarks Like LMArena Chatbot Arena hasn't shown that voting for one output over another actually correlates with preferences, however they may be defined.
SO021 Hugging Face lmarena-ai Organization on Hugging Face
SO022 Andreessen Horowitz (a16z) Announcing Our Latest Open Source AI Grants
SO023 Mashable SEA LMArena has some competition: Scale AI launches Seal Showdown, a new benchmarking tool Critics say that LMArena's system favors frontier models from big AI companies like Google, xAI, and OpenAI.
SO024 Arena Intelligence Inc. Agent Arena: Causal Evaluation of Agents in the Real World
SO025 Arena Intelligence Inc. Introducing the Search Arena: Evaluating Search-Enabled AI
SO026 Arena Intelligence Inc. WebDev Arena: A Live LLM Leaderboard for Web App Development
SO027 The Verge Meta got caught gaming AI benchmarks Meta's interpretation of our policy did not match what we expect from model providers.
SO028 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SO029 arXiv / Search Arena (ICLR 2026) Search Arena: A Crowd-Sourced, Human-Preference Dataset for Search-Augmented LLMs
SM001 The Business Research Company Artificial Intelligence (AI) Model Evaluation Platform Market Report The AI model evaluation platform market grows from $1.86 billion in 2025 to $2.36 billion in 2026 at a CAGR of 27.3%.
SM002 Yahoo Finance AI Model Evaluation Platform Market article
SM003 Research & Markets AI Model Evaluation Platform Market Report
SM004 Precedence Research Model Evaluation and Benchmarking Tools Market
SM005 Gartner Gartner forecasts worldwide AI spending to grow 47 percent in 2026
SM006 Presenc AI Enterprise AI adoption statistics 2026 78% of Global 2000 companies have at least one AI workload in production in Q1 2026.
SM007 Future AGI Top 5 LLM Evaluation Tools 2025
SM008 Scale AI SEAL Showdown SEAL Showdown includes users in 100+ countries, 70+ languages, and 200+ professional domains.
SM009 Mashable SEA LMArena has some competition: Scale AI launches SEAL Showdown
SM010 arXiv The Leaderboard Illusion
SM011 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark
SM012 TechCrunch LMArena lands $1.7B valuation four months after launching its product
SM013 PR Newswire LMArena raises $150 million to build the world's most trusted AI evaluation platform
SM014 arXiv Chatbot Arena paper
SM015 LMSYS Chatbot Arena launch post
SM016 Arena AI AI Evaluations
SM017 Arena AI Agent Arena methodology
SM018 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark
SM019 Winbuzzer Experts challenge validity and ethics of crowdsourced AI benchmarks like LMArena
SM020 The Verge Meta Llama 4 Maverick benchmarks gaming
SM021 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SM022 arXiv Arena-Hard paper
SM023 arXiv MT-Bench paper
SM024 Founded LMArena founders profile
SM025 Scale AI SEAL Showdown
SP001 arXiv (Wei-Lin Chiang, Lianmin Zheng, et al. — UC Berkeley / LMSYS) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference Chatbot Arena has emerged as one of the most referenced LLM leaderboards, widely cited by leading LLM developers and companies.
SP002 arXiv (Tianle Li, Wei-Lin Chiang, et al. — UC Berkeley / LMSYS) From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline Arena-Hard-Auto provides 3x higher separation of model performances compared to MT-Bench and achieves 98.6% correlation with human preference rankings, all at a cost of $20.
SP003 arXiv (Singh et al. — Cohere, Stanford, MIT, Ai2) The Leaderboard Illusion Providers like Google and OpenAI have received an estimated 19.2% and 20.4% of all data on the arena, respectively. In contrast, a combined 83 open-weight models have only received an estimated 29.7% of the total data.
SP004 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark The distribution of testing data may not accurately reflect the target market's real human users. Moreover, the platform's evaluation process is largely uncontrollable.
SP005 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark LM Arena allowed some industry-leading AI companies like Meta, OpenAI, Google, and Amazon to privately test several variants of AI models, then not publish the scores of the lowest performers.
SP006 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark The unmodified Maverick, 'Llama-4-Maverick-17B-128E-Instruct,' was ranked below models including OpenAI's GPT-4o, Anthropic's Claude 3.5 Sonnet, and Google's Gemini 1.5 Pro as of Friday.
SP007 LMArena (Arena Intelligence) Celebrating Community Impact at LMArena (Two-Year Celebration) 400+ models have been evaluated across various modalities (Text, Vision, Text-to-Image, WebDev and more!). 300+ evaluations have been pre-release.
SP008 LMArena (Arena Intelligence) New Product: AI Evaluations This service offers enterprises, model labs, and developers comprehensive evaluation services grounded in real-world human feedback.
SP009 LMArena (Arena Intelligence) Arena AI: The Official AI Ranking & LLM Leaderboard
SP010 LMArena (Arena Intelligence) Arena Leaderboard
SP011 LMArena / LMSYS Org (GitHub) FastChat: An open platform for training, serving, and evaluating large language models FastChat powers Chatbot Arena (lmarena.ai), serving over 10 million chat requests for 70+ LLMs.
SP012 EleutherAI (GitHub) lm-evaluation-harness: A framework for few-shot evaluation of language models Over 60 standard academic benchmarks for LLMs, with hundreds of subtasks and variants implemented.
SP013 OpenAI (GitHub) openai/evals: Evals is a framework for evaluating LLMs and LLM systems You can now configure and run Evals directly in the OpenAI Dashboard.
SP014 Hugging Face Open LLM Leaderboard
SP015 LMArena (Hugging Face Space) Arena Leaderboard (HuggingFace)
SP016 Scale AI Scale AI — Reliable AI Systems Benchmarking the frontier of AI capability with expert-level evaluations.
SP017 Scale AI Scale GenAI Platform Every agent is built and tested against your specific enterprise standards — your workflows, your rules, your definition of good — before it ever touches production.
SP018 Stanford CRFM Holistic Evaluation of Language Models (HELM) — Latest
SP019 Stanford CRFM Holistic Evaluation of Language Models (HELM) — Classic
SP020 BenchLM LLM Leaderboard 2026 — Compare 261 AI Models Across 249 Benchmarks 261 models · 249 benchmarks. The most comprehensive LLM comparison tool — 249 benchmarks, real pricing, and runtime data in one place.
SP021 Artificial Analysis AI Model & API Providers Analysis
SP022 MetaTech.dev LMArena AI Benchmarking Crisis: Why Model Rankings Are Broken The problem becomes even more concerning when we look at critical applications... This disconnect between benchmark success and practical reliability is exactly what's wrong with current AI benchmarking approaches.
SP023 LMSYS Org (UC Berkeley) LMSYS Org — Large Model Systems Organization The Large Model Systems Organization develops large models and systems that are open, accessible, and scalable.
SP024 U.S. Securities and Exchange Commission (EDGAR) EDGAR Search Results — Recall Capital-LMArena a Series of CGF2021 LLC
SP025 U.S. Securities and Exchange Commission (EDGAR) Form D — Recall Capital-LMArena a Series of CGF2021 LLC (search index entry)
SI001 LMArena (PR Newswire) LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product.
SI002 TechCrunch LMArena lands $1.7B valuation four months after launching its product LMArena's annualized "consumption rate" — as the company describes its annual recurring revenue (ARR) — of $30 million as of December, less than four months after launch.
SI003 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M LM Arena, a crowdsourced benchmarking project that major AI labs rely on to test and market their AI models, has raised $100 million in a seed funding round that values the organization at $600 million.
SI004 LMArena (Arena Intelligence) Fueling the World's Most Trusted AI Evaluation Platform (Series A Blog) This year we saw our community grow by over 25x alongside rapid adoption by AI labs who trust this platform as a gold-standard for evaluating real-world model performance.
SI005 GetLatka LMArena Revenue 2026: $30M ARR, $1.7B Valuation In 2026, LMArena's revenue reached $30M. LMArena reached a $1.7B valuation in 2026, set during its Series A round.
SI006 CTOL Digital LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges The company claims over $30 million in annualized revenue from selling evaluation services to these labs, launching its commercial product only in September 2025. At 57 times that run rate, the valuation prices in not just growth, but the assumption that this inherent conflict can be managed indefinitely.
SI007 U.S. Securities and Exchange Commission Form D — Recall Capital-LMArena a Series of CGF2021 LLC
SI008 Wired LMArena Raises $150 Million to Evaluate the World's Most Powerful AI Models
SI009 arXiv (Singh et al. — Cohere, Stanford, MIT, Ai2) The Leaderboard Illusion
SI010 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark
SI011 LMArena (Arena Intelligence) New Product: AI Evaluations LMArena earns revenue by providing paid AI evaluation services to AI labs and enterprises.
SI012 LMArena (Arena Intelligence) Arena AI: The Official AI Ranking & LLM Leaderboard
SI013 LMArena (Arena Intelligence) Arena Leaderboard
SI014 LMArena (Arena Intelligence) Celebrating Community Impact at LMArena (Two-Year Celebration)
SI015 MetaTech.dev LMArena AI Benchmarking Crisis: Why Model Rankings Are Broken
SI016 LMSYS Org (UC Berkeley) LMSYS Org — Large Model Systems Organization
SI017 BenchLM LLM Leaderboard 2026 — Compare 261 AI Models Across 249 Benchmarks
SI018 LMArena / LMSYS Org (GitHub) FastChat: An open platform for training, serving, and evaluating large language models
SI019 Scale AI Scale AI — Reliable AI Systems
SI020 Scale AI Scale GenAI Platform
SI021 arXiv (Wei-Lin Chiang, Lianmin Zheng et al.) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
SI022 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark
SI023 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SI024 U.S. Securities and Exchange Commission (EDGAR) EDGAR Search Results — Recall Capital-LMArena a Series of CGF2021 LLC
SI025 U.S. Securities and Exchange Commission (EDGAR search index) SEC EDGAR Full-Text Search — Form D filings mentioning LMArena
SI026 VentureBeat LMArena raises $150M at $1.7B valuation to build AI evaluation infrastructure
SI027 Bloomberg UC Berkeley AI Benchmarking Group Chatbot Arena Raises $100 Million
SI028 BusinessWire LMArena Raises $150 Million, Achieves $1.7B Valuation
SI029 StartupWired AI startup LMArena triples valuation to $1.7B in 2026
SE001 Arena Intelligence Inc. Arena AI: The Official AI Ranking & LLM Leaderboard
SE002 Arena Intelligence Inc. About Arena — Crowdsourced AI Model Evaluation Platform
SE003 Arena Intelligence Inc. New Product: AI Evaluations LMArena has already logged 250M+ real conversations, 2M+ monthly votes, and has 3M+ monthly users.
SE004 Arena Intelligence Inc. Arena Leaderboard Policy
SE005 Arena Intelligence Inc. Leaderboard Changelog
SE006 Arena Intelligence Inc. Agent Arena: Causal Evaluation of Agents in the Real World The methodology powering the Agent Arena Leaderboard is different from our previous arenas. Rather than pairwise votes, rankings are calculated using a methodology we call causal tracing.
SE007 Arena Intelligence Inc. WebDev Arena: A Live LLM Leaderboard for Web App Development
SE008 Arena Intelligence Inc. Introducing the Search Arena: Evaluating Search-Enabled AI
SE009 Arena Intelligence Inc. Arena-Rank: Open Sourcing the Leaderboard Methodology Arena-Rank, an open-source Python package for ranking that powers the LMArena leaderboard!
SE010 Arena Intelligence Inc. Does Style Matter in AI Evaluations?
SE011 Arena Intelligence Inc. Introducing Hard Prompts Category in Chatbot Arena
SE012 Arena Intelligence Inc. The Arena-Hard Pipeline
SE013 Arena Intelligence Inc. The Next Stage of AI Coding Evaluation Is Here Record: Every model action (file creation, edit, or execution) is logged and versioned. Snapshots are stored in Cloudflare R2.
SE014 arXiv (Chiang et al., UC Berkeley / LMSYS) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
SE015 arXiv / NeurIPS 2023 (Zheng et al., UC Berkeley / LMSYS) Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
SE016 arXiv (Li et al., UC Berkeley / LMSYS) From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
SE017 arXiv / ICLR 2026 (Miroyan et al., UC Berkeley) Search Arena: Analyzing Search-Augmented LLMs
SE018 NeurIPS 2023 Proceedings (Zheng et al.) Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (NeurIPS 2023)
SE019 GitHub (lm-sys) FastChat: An Open Platform for Training, Serving, and Evaluating LLMs
SE020 GitHub (lmarena) arena-rank: Source Code of Arena Leaderboard Methodology
SE021 GitHub (lmarena org) Arena GitHub Organization
SE022 HuggingFace / lmarena-ai (via Wayback) lmarena-ai (Arena) Organization on HuggingFace
SE023 PyPI arena-rank PyPI package
SE024 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification.
SE025 The Verge Meta got caught gaming AI benchmarks Meta's interpretation of our policy did not match what we expect from model providers.
SE026 LMSYS (UC Berkeley SkyLab) Chatbot Arena: New features and a Elo Rating System (original launch blog)
SE027 U.S. Securities and Exchange Commission SEC Form D: Recall Capital-LMArena (Arena Intelligence Series A vehicle)
SU001 PRNewswire (Arena Intelligence Inc.) LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform LMArena works with leading AI labs and enterprises, including OpenAI, Google and xAI, all drawing on LMArena's evaluations to improve their models for production use cases.
SU002 TechCrunch Podcast The PhD students who became the judges of the AI industry how a team like theirs can build a neutral benchmark when the companies they're ranking are also their backers
SU003 OpenTools.AI LM Arena Under Fire: Allegations of Benchmark Bias Stir AI Industry
SU004 Scale AI SEAL Showdown: Scale's AI Evaluation Platform
SU005 AI Wiki LMArena.org — AI Wiki
SU006 PitchBook Arena Intelligence Company Profile
SU007 Arena Intelligence Inc. LMArena (original domain) — redirects to arena.ai
SU008 U.S. Securities and Exchange Commission (EDGAR) EDGAR Full-Text Search: LMArena Form D filings
SU009 TechCrunch LMArena lands $1.7B valuation four months after launching its product That trajectory, and the startup's popularity, were enough for VCs to pile in for the Series A.
SU010 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SU011 Arena Intelligence Inc. Arena AI: The Official AI Ranking & LLM Leaderboard
SU012 Arena Intelligence Inc. About Arena — Crowdsourced AI Model Evaluation Platform
SU013 Arena Intelligence Inc. New Product: AI Evaluations LMArena has already logged 250M+ real conversations, 2M+ monthly votes, and has 3M+ monthly users.
SU014 Arena Intelligence Inc. Arena Leaderboard Policy
SU015 Arena Intelligence Inc. Leaderboard Changelog
SU016 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others.
SU017 The Verge Meta got caught gaming AI benchmarks Meta's interpretation of our policy did not match what we expect from model providers.
SU018 arXiv (Chiang et al., UC Berkeley / LMSYS) Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
SU019 arXiv / ICLR 2026 (Miroyan et al.) Search Arena: Analyzing Search-Augmented LLMs
SU020 GitHub (lm-sys) FastChat: An Open Platform for Training, Serving, and Evaluating LLMs
SU021 HuggingFace / lmarena-ai (via Wayback) lmarena-ai (Arena) Organization on HuggingFace
SU022 Arena Intelligence Inc. Agent Arena: Causal Evaluation of Agents in the Real World In a sample of the heaviest real sessions we saw: a live sports-TV schedule site, an autonomous-underwater-vehicle autopilot, a self-hosted movie-watchlist app.
SU023 Arena Intelligence Inc. Fueling the World's Most Trusted AI Evaluation Platform (Series A blog) This year we saw our community grow by over 25x alongside rapid adoption by AI labs who trust this platform as a gold-standard for evaluating real-world model performance.
SU024 Arena Intelligence Inc. Celebrating Community Impact at LMArena (Two-Year Anniversary)
SU025 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark The release version of Llama 4 has been added to LMArena after it was found out they cheated, but you probably didn't see it because you have to scroll down to 32nd place.
SU026 Discord (LMArena Community) Arena Discord Community Server
SU027 Tech in Asia a16z, Lightspeed back $150M Series A of AI model evaluator LMArena
SU028 U.S. Securities and Exchange Commission (EDGAR) EDGAR Full-Text Search: Arena Intelligence Form D 2025-2026
SR001 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark "Only a handful of [companies] were told that this private testing was available, and the amount of private testing that some [companies] received is just so much more than others. This is gamification." — Sara Hooker, Cohere VP of AI Research
SR002 arXiv (Cohere, Stanford, MIT, AI2) The Leaderboard Illusion (arXiv:2504.20879) "We find that undisclosed private testing practices benefit a handful of providers who are able to test multiple variants before public release and retract scores if desired."
SR003 The Verge Meta got caught gaming AI benchmarks "Meta's interpretation of our policy did not match what we expect from model providers. Meta should have made it clearer that 'Llama-4-Maverick-03-26-Experimental' was a customized model to optimize for human preference."
SR004 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SR005 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark "Companies can continually optimize their models to better align with the LMSYS user distribution, possibly leading to unfair competition and a less meaningful evaluation."
SR006 UCStrategies AI Benchmarks Are a Game Now — And the Industry Is Cheating to Win "SurgeAI analyzed 500 LMArena votes and disagreed with 52%, finding that 'confidence beats accuracy and formatting beats facts.'"
SR007 LMArena Arena Leaderboard Policy (Last Updated April 30, 2026)
SR008 LMArena Leaderboard Changelog
SR009 Arena Intelligence, Inc. LMArena Privacy Policy (Previous Version, Effective 2025-09-05) "Arena Intelligence, Inc. d/b/a LMArena provides a platform for using, comparing, rating, testing, evaluating, and ranking third-party AI models."
SR010 European Commission (Your Europe) Data protection under GDPR — Your Europe
SR011 CTOL Digital Solutions LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges "LMArena has positioned itself as the independent arbiter of model performance, yet it derives revenue from the same AI labs it evaluates — OpenAI, Google, and xAI among them."
SR012 Didit AI Compliance in the LLM Era: Regulatory Guide 2026
SR013 LMArena Arena AI: The Official AI Ranking & LLM Leaderboard (Homepage)
SR014 TechCrunch LMArena lands $1.7B valuation four months after launching its product
SR015 PR Newswire LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform
SR016 LMArena Fueling the World's Most Trusted AI Evaluation Platform (Series A blog)
SR017 LMArena New Product: AI Evaluations
SR018 ByteIota LMArena Raises $150M at $1.7B Valuation in 4 Months
SR019 StartupWired AI Startup LMArena Triples Valuation to $1.7B in 2026
SR020 The AI Insider LMArena Secures $150M to Build the World's Most Trusted AI Evaluation Platform
SR021 OfficeChai AI Evaluation Platform LMArena Raises Series A At Valuation Of $1.7 Billion "LMArena, formally known as Arena Intelligence Inc., was founded in 2025 by Anastasios N. Angelopoulos (CEO), Wei-Lin Chiang (CTO), and Ion Stoica (Advisor)."
SR022 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SR023 The Business Research Company AI Model Evaluation Platform Market Size and Trends Report 2026
SR024 U.S. Securities and Exchange Commission Form D: Recall Capital-LMArena a Series of CGF2021 LLC (EDGAR filing 0002113470-26-000001)
SR025 TechCrunch The leaderboard 'you can't game,' funded by the companies it ranks (video)
SR026 Latka LMArena Revenue 2026: $30M ARR, $1.7B Valuation
SR027 CTOL Digital Solutions LMArena Raises $150 Million — conflict-of-interest analysis
SR028 ByteIota LMArena — open-source model bias and methodological flaws
SR029 LMArena Leaderboard Changelog — Battles in Direct Update (May 12, 2026) "We observed two new biases in the Battles in Direct voting data and corrected for them in the Bradley-Terry fit: the first is a position bias favoring Model A; the second is an advantage given to models that share an organization with the prior turns of context."
SR030 StartupWired LMArena Triples Valuation — Risks and Challenges Ahead section "As usage grows, so do demands for data security, fairness, and governance. Any misstep could damage credibility, which forms the core of LMArena's value proposition."
SV001 TechCrunch LMArena lands $1.7B valuation four months after launching its product "The startup bolted out of the gate as a commercial venture with a $100 million seed round in May at a $600 million valuation. This new round means it raised $250 million in about seven months."
SV002 PR Newswire LMArena Raises $150 Million to Build the World's Most Trusted AI Evaluation Platform "LMArena's annualized consumption run rate surpassed $30 million in December, less than four months after launching its AI evaluation product."
SV003 Latka LMArena Revenue 2026: $30M ARR, $1.7B Valuation
SV004 StartupWired AI Startup LMArena Triples Valuation to $1.7B in 2026
SV005 The AI Insider LMArena Secures $150M to Build the World's Most Trusted AI Evaluation Platform
SV006 OfficeChai AI Evaluation Platform LMArena Raises Series A At Valuation Of $1.7 Billion "The company is positioning itself in what analysts estimate is an $800-900 million AI testing market in 2025, projected to grow to $3.8 billion by 2032."
SV007 TechCrunch LM Arena, the organization behind popular AI leaderboards, lands $100M
SV008 ByteIota LMArena Raises $150M at $1.7B Valuation in 4 Months "Four months to $30 million: In May 2025, LMArena raised $100M at $600M; by September, launched AI Evaluations; three months later, hit $30M annualized run rate."
SV009 LMArena Fueling the World's Most Trusted AI Evaluation Platform (Series A blog) "Since announcing our $100M Seed round last year in May, LMArena has grown far faster than we imagined. In a matter of months, the community has contributed 50 million votes."
SV010 The Business Research Company AI Model Evaluation Platform Market Size and Trends Report 2026 "AI Model Evaluation Platform market size has reached $1.86 billion in 2025; expected to grow to $6.24 billion in 2030 at a CAGR of 27.5%."
SV011 Research and Markets AI Model Evaluation Platform Market Report 2026
SV012 The Business Research Company Human-In-The-Loop AI Market Size and Drivers Report 2026 "In March 2025, CoreWeave Inc. acquired Weights & Biases Inc. for approximately $1.4 billion."
SV013 CTOL Digital Solutions LMArena Raises $150 Million to Rank AI Models While Selling Evaluation Services to the Labs It Judges "At 57 times that run rate, the valuation prices in not just growth, but the assumption that this inherent conflict can be managed indefinitely."
SV014 UCStrategies AI Benchmarks Are a Game Now — And the Industry Is Cheating to Win
SV015 U.S. Securities and Exchange Commission Form D: Recall Capital-LMArena a Series of CGF2021 LLC (EDGAR accession 0002113470-26-000001) "Recall Capital-LMArena a Series of CGF2021 LLC — Pooled Investment Fund / Venture Capital Fund — Amount Sold: $382,500"
SV016 arXiv (Cohere, Stanford, MIT, AI2) The Leaderboard Illusion (arXiv:2504.20879)
SV017 TechCrunch Study accuses LM Arena of helping top AI labs game its benchmark
SV018 LMArena Arena Leaderboard Policy (Last Updated April 30, 2026)
SV019 LMArena New Product: AI Evaluations
SV020 LMArena Leaderboard Changelog
SV021 Arena Intelligence, Inc. LMArena Privacy Policy (Previous Version, Effective 2025-09-05) "Arena Intelligence, Inc. d/b/a LMArena provides a platform for using, comparing, rating, testing, evaluating, and ranking third-party AI models."
SV022 TechCrunch The leaderboard 'you can't game,' funded by the companies it ranks (video)
SV023 The Verge Meta got caught gaming AI benchmarks
SV024 TechCrunch Meta's vanilla Maverick AI model ranks below rivals on a popular chat benchmark
SV025 TechCrunch The AI industry is obsessed with Chatbot Arena, but it might not be the best benchmark
SV026 Grokipedia Arena (LMArena) — Grokipedia overview "As of March 5, 2026, the top models on the text leaderboard are claude-opus-4-6 (1504 Elo), gemini-3.1-pro-preview (1500 Elo); total 5,430,034 votes collected across the platform."
SV027 TLDL AI Company Rankings 2026: Revenue, Funding & Valuation Data for 2,000+ Companies "Private funding for AI startups topped $150 billion over the trailing twelve months; foundation model companies raising $80 billion in 2025."
SV028 Wellows 85 Hottest AI Startups to Watch in 2026 [By Valuation, Funding, & Growth] "Anysphere (Cursor): AI coding assistant, $29.3B valuation, $1B ARR; Harvey: legal AI; LMArena reached a $1.7 billion valuation in under four months."
SV029 AgentMarketCap LMArena's $1.7B Valuation in 4 Months: Why AI Evaluation Is the New Data Labeling "Scale AI was valued at $7B in 2021. In June 2025, Meta acquired a 49% stake for $14.3 billion — the largest VC transaction of 2025 — implicitly valuing Scale above $29 billion."
SV030 Axis Intelligence Research AI Copyright Lawsuits 2026: Status Tracker — Updated Monthly "$50 billion+: Cumulative legal exposure across all active AI copyright and related IP cases; Anthropic's confirmed settlement in Bartz v. Anthropic: $1.5 billion covering ~482,000 works."
SV031 U.S. Securities and Exchange Commission (EDGAR) EDGAR Company Search: Recall Capital-LMArena a Series of CGF2021 LLC (CIK 0002113470)
SV032 Copyright Alliance AI Copyright Lawsuit Developments in 2025: A Year in Review "Anthropic's $1.5B settlement in Bartz v. Anthropic required payment of approximately $3,000 for each of the 482,460 books downloaded from pirate libraries — the first publicly confirmed pricing benchmark for AI training on pirated content."
SV033 U.S. News & World Report / Reuters AI Startup LMArena Triples Its Valuation to $1.7 Billion in Latest Fundraise "LMArena said on Tuesday its valuation had tripled to $1.7 billion in about eight months, following a new funding round where it raised $150 million."