Startup Diligence
Diligence report AI inference infrastructure / MaaS Late-stage private (Series B / B+ financing; HKEX applicant) 2026-07-22

SiliconFlow

China-based AI inference infrastructure company with real scale and strategic relevance, but still negative public-cloud economics, thin public retention visibility, and a valuation that requires evidence-sensitive price discipline.

SiliconFlow looks like a real and increasingly important China-based inference platform, but the disclosed June 2026 price already asks investors to underwrite future margin improvement, enterprise durability, and supply / compliance execution that the public record still does not fully prove.

Cover facts

Latest disclosed round 01
296 USD million [CO036]
Public valuation anchor 02
1200 USD million [CV040]
2025 revenue 03
55.33 RMB million [CI011]
2025 blended gross margin 04
-24 pct [CI012]
Registered users 05
10M+ latest practicable date [CO013]
Enterprise customers 06
13000+ latest practicable date [CO017]
Founded 07
2023-08-29 [CO001]
Post-2025 cash consideration 08
1480 RMB million [CO035]

Company profile

SiliconFlow is a Beijing-founded AI inference infrastructure company established in August 2023. Public evidence describes a multi-surface business: pay-as-you-go model APIs, reserved or dedicated inference capacity, acceleration services, and private deployment for enterprise or compliance-sensitive workloads. By mid-2026 the company had reached meaningful adoption scale, with more than 10 million registered users, more than 13,000 enterprise customers in filing-era disclosure, and a public model library above 200 models. The company also disclosed seven funding rounds through June 2026, including a Series B/B+ sequence that lifted filing-disclosed post-money valuation to RMB7.74 billion while independent media described the round as roughly US$296 million at about a US$1.2 billion valuation. SiliconFlow is strategically relevant within the China inference stack, but the public record still shows weak public-cloud economics, supplier dependence, and incomplete disclosure on retention and concentration.

Website
siliconflow.cn
Founded
2023-08-29
Founders
Yuan Jinhui
Founding location
Beijing, China
Headquarters
Beijing, China
Product
SiliconFlow delivers an API-first inference platform with usage-based access to large language, multimodal, image, audio, and video models, alongside reserved instances, inference-acceleration services, and private-deployment options for enterprise workloads.
Customers
Developers, AI builders, enterprise platform teams, regulated or state-linked institutions, and telecom / infrastructure partners that need model access, predictable inference performance, private deployment, or domestic-chip adaptation.
Business model
Usage-based serverless APIs, reserved or dedicated capacity, and private-deployment solutions. Public evidence suggests a two-line business mix between public cloud services and on-premise deployment, with the latter currently carrying materially better gross margins.
Stage
Late-stage private; June 2026 Series B / B+ financing and active HKEX Chapter 18C filing
Funding status
The public record supports a rapid funding climb from angel through Series B+ between December 2023 and June 2026. The HKEX filing shows post-money valuation steps up to RMB7.74 billion and aggregate post-2025 cash consideration of about RMB1.48 billion, while external media described the June 2026 financing as roughly US$296 million at a US$1.2 billion valuation.
[CO001, CO023, CO026, CO032, CO033, CO034, CO035, CO036]

Executive summary

Top strengths

  • SiliconFlow has genuine strategic relevance in AI inference: broad model access, multiple deployment modes, and meaningful enterprise/infrastructure use cases rather than a simple demo API surface.
  • Public adoption signals are substantial for a company founded in 2023, including more than 10 million registered users, more than 13,000 enterprise customers in filing-era disclosure, and token throughput large enough to support dedicated-capacity case studies.
  • The company appears well positioned for China-specific inference demand because it combines model distribution, private deployment, and domestic-chip adaptation narratives.
  • The June 2026 financing syndicate and HKEX filing process suggest that SiliconFlow has real capital-market relevance and time to continue executing.

Top risks

  • Public-cloud economics are still weak: 2025 blended gross margin was negative and the public- cloud line remained deeply loss-making, indicating that scale has not yet solved quality of growth.
  • Supplier and compute dependence remain material because the company rents or leases critical infrastructure rather than fully controlling its own semiconductor supply chain.
  • Regulatory and compliance complexity is meaningful across AI labeling, privacy, identity, telecom licensing, and domestic/international surface management.
  • Customer breadth is public, but retention, concentration, and revenue-quality metrics remain too thin to support a high-conviction underwriting case.
  • The disclosed valuation can still prove expensive if current usage scale does not convert into higher-margin dedicated/private deployments and durable enterprise retention.

Open gaps

  • Current 2026 revenue run-rate, product-line mix, and margin trend by serverless, dedicated, and private deployment.
  • Enterprise retention, logo concentration, spend concentration, and API-to-dedicated expansion metrics.
  • Current cash balance, monthly burn, and financing scenario analysis after the June 2026 round.
  • Security / assurance package: uptime history, incident history, SOC 2 / ISO or equivalent evidence.
  • Supplier concentration terms and contingency plans under export-control or capacity-shock scenarios.

Contents

Chapter 01

01Company Overview

1.1 Identity, Legal Footprint, and Business Model

SiliconFlow's public identity is already more complex than a generic API startup. The China-facing domain presents Beijing SiliconFlow as an AI-capability provider with a concrete address in Haidian, a visible ICP filing, and a Beijing value-added telecom permit, while the global terms and privacy pages are issued by SiliconFlow Technology Pte. Ltd. and explicitly route mainland-China users back to siliconflow.cn. The HKEX filing adds a third layer: a PRC joint-stock company with a Beijing registered office and a Hong Kong place of business, but with headquarters, senior management, and operations still centered outside Hong Kong. In practical terms, the company is China-rooted, globally marketed, and already structuring itself for cross-border capital-market access. The underlying business model is clearer than the legal map. Across the China homepage, global homepage, GitHub profile, and API documentation, SiliconFlow consistently markets itself as an inference infrastructure layer rather than a model lab or end-application vendor. It sells standardized access to models through large-model APIs, dedicated or reserved instances, inference acceleration services, and private deployment. The docs emphasize rapid onboarding, API-key self-service, and an OpenAI-style surface; the marketing pages emphasize speed, cost control, multimodal coverage, and predictable usage-based delivery. That combination is important because it shows the company is trying to sit between model suppliers, compute suppliers, and downstream developers or enterprises as a neutral token-supply platform.[CO001, CO002, CO003, CO004, CO005, CO006]

Snapshot KPI table
MetricValue / statusDate / periodConfidenceGap or caveat
Legal establishment2023-08-29 PRC LLC; converted to joint-stock company on 2026-06-222023-08 to 2026-06HighConversion is clear; ultimate post-IPO structure still subject to listing completion
Registered HQHaidian District, BeijingCurrent public recordHighCountry-by-country office list beyond Beijing and HK is not fully public
Global legal entity.com terms issued by SiliconFlow Technology Pte. Ltd.Current public recordHighPublic entity map between PRC and Singapore vehicles is still incomplete
Latest capital-market stageHKEX Chapter 18C applicant / pre-commercial listing process active2026-06 to 2026-07HighListing is not yet approved or completed
Latest disclosed valuation rungRMB 7.74B post-money (Series B+)2026-06-09 / 2026-06-18 sequenceHighExternal headlines often refer instead to an over-RMB2B Series B
Post-2025 financing cash received~RMB 1.48B across A+, B, and B+After 2025-12-31HighDoes not include earlier pre-2026 rounds
Registered users10.28M+2026-04-30HighPoint-in-time metric; current run-date level undisclosed
Enterprise customers13,000+ servedLatest practicable date before 2026-06-30 filingHighNo public split by paying status, cohort, or concentration
Model coverage170+ in filing; 200+ on July 2026 model library2026-06 to 2026-07HighPublic website appears newer than filing snapshot
Token throughput578.5B average daily; 1,071.4B peak daily2026-04HighUsage scale does not directly imply healthy unit economics
Overseas commercializationMonthly overseas revenue exceeded US$1M2026-05HighNo public geographic or product mix disclosed
Current headcountNot publicly supportable from retained sourcesAs of 2026-07-22HighRequires management confirmation

Combines filing-backed operating metrics with public-website and financing context. Where filing and media snapshots differ, the table preserves the discrepancy rather than smoothing it away.

[CO001, CO002, CO003, CO004, CO005, CO013]
FO002: SiliconFlow Company Snapshot Logic

Flow diagram linking SiliconFlow's legal footprint, platform architecture, customer base, investor set, and listing trajectory.

1.2 Leadership, Governance, and Control Signals

Public disclosure on leadership is better than a typical private startup because of the HKEX filing, but it still leaves important diligence gaps. Dr. Yuan Jinhui is clearly the central figure: founder, chairman, CEO, general manager, and financial controller. The filing also identifies Liu Juncheng as CTO and executive director, Zeng Hua as commercialization lead and executive director, and Chen Yingjie as the non-executive director linked to Alibaba's strategic-investment arm. Beyond the formal board, the history and employee-incentive sections point to a broader founding and core-team nucleus that includes Zhao Zhen and Hu Jian, both tied to the earlier OneFlow network. That matters because SiliconFlow appears to be less a lightly staffed API aggregator than a technically opinionated team with a pre-existing system-software lineage. The biographies reinforce that interpretation. Yuan came through Microsoft China and OneFlow, Liu also spent years at OneFlow in AI-system-software roles, and Zeng adds commercialization experience from Microsoft China, Baidu, and JD. Chen Yingjie adds a strategic-capital bridge to Alibaba. The board structure after listing is straightforward on paper—three executive directors, one non-executive, and three independent non-executives—but the public record still lacks details on control rights, founder dilution, secondary activity, and board-economics terms. One additional wrinkle is that third-party databases do not perfectly agree on the co-founder slate, which makes the filing the authoritative source but also shows how noisy startup databases remain for a fast-moving China AI company.[CO022, CO023, CO024, CO025, CO026, CO027]

Leadership and founder table
Person / groupCurrent roleRelevant prior backgroundWhy it mattersKey-person or disclosure note
Dr. Yuan JinhuiFounder, chairman, CEO, general manager, financial controllerMicrosoft China lead researcher; founded OneFlowTechnical and strategic center of gravity for inference-platform thesisVery high dependence on Yuan across product, finance, and capital markets
Liu JunchengExecutive director, CTOAI system-software and OneFlow R&D backgroundOwns core technology execution and architecture continuityHigh dependence on core-system-software talent
Zeng HuaExecutive director, deputy general managerMicrosoft China, Baidu, JD commercialization backgroundAdds enterprise sales and go-to-market credibilityImportant for translating technical scale into monetization
Chen YingjieNon-executive directorAlibaba strategic-investment managing director; ex-PwC; former XPeng and MiniMax board rolesRepresents strategic-capital and ecosystem signal from a major platform investorBoard seat and rights are not fully disclosed in public materials
Zhao ZhenChief operating officer / co-founder per filing historyFormer COO of OneFlowOperational continuity from prior team networkNot a board member in public filing; external profile coverage is sparse
Hu Jian and broader core teamDeputy GM / core R&D and employee-incentive participantsCore team overlaps with OneFlow-era talent baseShows deeper bench than a two-founder startup narrativePublic biographies for several core operators remain limited
Independent director slateWu Chuan, Li Dan, Hou Hong (proposed INEDs)Academic, accounting, and strategy backgroundsSuggests listing-readiness and added governance depthIndependent slate is filing-disclosed, but committee practice is not yet proven

The HKEX filing materially improves leadership visibility, but several operational leaders and the economics of board control remain only partially disclosed.

[CO022, CO023, CO024, CO025, CO026, CO027]

1.3 Funding History and Capital-Market Story

SiliconFlow's capital formation has been unusually compressed. The HKEX filing documents seven financing rounds from December 2023 through June 2026, with disclosed capital injections stepping from RMB47.2 million at angel to RMB65.84 million at angel+, RMB71.48 million at pre-A, RMB285.99 million at Series A, RMB220 million at Series A+, RMB520 million at Series B, and RMB740 million at Series B+. The corresponding post-money valuations rose from RMB280.0 million to RMB7.74 billion in roughly two and a half years. That is not just a fundraising story; it is a signal that investors rapidly repriced the company as the AI-token middle layer became investable in China. The June 2026 headlines need careful handling. External media consistently described SiliconFlow as completing an over-RMB2 billion Series B, backed by a strategic syndicate spanning Trip.com or Ctrip, JinkoSolar, Kingdee, Unicom-linked investors, Biren, NIO Capital, SenseTime, GGV, and others, with China Renaissance advising. But the HKEX filing then presents a more granular 2026 sequence of A+, B, and B+ financings, with aggregate post-2025 cash consideration of about RMB1.48 billion and valuation steps from RMB3.12 billion to RMB5.02 billion and then RMB7.74 billion. Public databases also show unreconciled totals, so the right conclusion is not that one source is necessarily wrong, but that the public capital history should be treated as directionally clear and transaction-label details as still needing data-room reconciliation.[CO032, CO033, CO034, CO035, CO036, CO037]

Stakeholder or investor map
Stakeholder / investorRole or round contextStrategic relevanceWhat public record showsDiligence ask
Alibaba / Hangzhou DuoxiangShareholder and board-linked strategic investorPotential cloud, distribution, and ecosystem leverageAlibaba-linked entity is a substantial shareholder; Chen Yingjie sits on the boardConfirm exact ownership, information rights, and commercial relationship scope
Trip.com / Ctrip Strategic InvestmentNamed June 2026 strategic investorTravel, enterprise-demand, and capital-markets signalRepeatedly named in June 2026 financing coverageVerify whether Trip.com invested at B, B+, or both and whether commercial pilots exist
SenseTime strategic investmentNamed June 2026 investor and existing AI-ecosystem peerSignals model and AI-infra ecosystem alignmentNamed in multiple June 2026 funding reportsCheck whether the relationship is purely financial or includes workload, model, or channel cooperation
JinkoSolar HoldingNamed June 2026 investorLinks AI infrastructure to power and datacenter economics36Kr frames Jinko as an energy-to-compute strategic partnerDetermine whether this includes actual infrastructure contracts or only capital
Unicom-linked capitalNamed June 2026 investor groupTelecom and computing-network integration may aid enterprise deploymentFunding coverage cites Unicom Xinwo and Unicom Capital vehiclesClarify customer or network resource commitments
Biren TechnologyNamed June 2026 strategic investorDomestic-chip adaptation is central to SiliconFlow's differentiationFunding coverage frames Biren around domestic-chip inference collaborationCheck binding commercial terms, exclusivity, or co-marketing obligations
NIO Capital and other financial investorsGrowth capital in late-stage roundsAdds financial sponsorship beyond strategic corporatesNamed across financing coverage and CB Insights investor tableRequest full cap table and pro-rata or liquidation rights
China RenaissanceExclusive adviser on June 2026 financing; later HKEX overall coordinatorBridges private financing and public-listing preparationNamed in Caixin Global and HKEX coordinator announcementDetermine whether adviser economics or process dependencies create pressure for rapid listing

Public information strongly supports a strategically dense investor base, but not the control rights or commercial obligations attached to those investors.

[CO025, CO026, CO033, CO036, CO037, CO047]

1.4 Scale, Traction, and Geographic Reach

The filing and corroborating coverage show that SiliconFlow has already achieved real usage scale, even if its commercial density is still evolving. As of April 30, 2026, the company disclosed more than 10 million registered users, average daily throughput of 578.5 billion tokens, peak daily throughput above one trillion tokens, and more than 13,000 enterprise customers by the latest practicable date. The platform supported over 170 models in the filing and more than 200 on the public model library by July 2026. Those are unusually large infrastructure-side operating metrics for a company incorporated only in August 2023, and they help explain why investors were willing to fund it aggressively despite immature profitability. At the same time, the footprint evidence is broad but not exhaustive. Public pages firmly place the operating center in Beijing, and the filing confirms a Hong Kong business address as part of the IPO process. The official and third-party materials also suggest overseas commercialization is no longer hypothetical: the filing says overseas monthly revenue exceeded US$1 million in May 2026, while June coverage repeatedly described millions of dollars in overseas monthly revenue and global platform traction. What is still missing is a clean public breakdown by country, office, headcount, or overseas revenue mix. The company therefore looks scaled in usage, scaled enough to matter in enterprise adoption, and only partially disclosed in organizational footprint.[CO013, CO014, CO015, CO016, CO017, CO018]

FO003: SiliconFlow Snapshot KPIs

Key maturity, scale, and caution indicators for SiliconFlow as of the 2026-07-22 run date.

1.5 Milestones, Compliance Signals, and Adverse Context

The operating chronology is coherent. SiliconFlow says it started inference-engine R&D in August 2023, launched public-cloud MaaS in May 2024, shipped DeepSeek inference on Huawei Ascend in February 2025, launched private MaaS in September 2025, released the Elastic GPU scheduler in April 2026, converted into a joint-stock company in June 2026, filed its HKEX application proof on June 30, and then added China Renaissance as an overall coordinator on July 14. That timeline shows a company moving in lockstep across productization, commercial expansion, financing, and listing preparation. The adverse frame is just as important. HKEX financial disclosure shows 2025 revenue scaling rapidly to RMB55.3 million, but gross margin swung negative at -24.0%. KrASIA and Hello China Tech sharpen the point: SiliconFlow appears to be renting compute, distributing third-party models, and competing in a price war where developer credits and below-cost token supply can accelerate adoption faster than profitability. Public-cloud economics look structurally tougher than on-premise deployment economics, which means the company's headline growth does not yet settle the question of durable monetization. In other words, the company overview supports a serious infrastructure platform with elite strategic backers, but not yet a fully de-risked business model.[CO039, CO040, CO042, CO043, CO044, CO045]

Milestone table
DateEventTypeAmount / statusParticipantsImplication
2023-08-29Company established in BeijingfoundingPRC LLC formedYuan-led founding teamCreates the legal base for later product, financing, and listing work
2023-08-01Inference-engine R&D beginsproductEngine development startedFounding engineering teamShows the company began at the systems layer rather than at the application layer
2024-05-01Public-cloud MaaS launchesproductCommercial public-cloud service liveSiliconFlowMarks transition from R&D project to external platform
2025-02-01DeepSeek services on Huawei Ascend launchproductDomestic-chip token service milestoneSiliconFlow, DeepSeek, Huawei AscendStrengthens domestic-chip and heterogeneous-compute positioning
2025-07-10Series A consideration fully paidfinancing~RMB285.99M Series APrivate investors per filingMoves valuation to RMB2.286B and funds broader commercialization
2025-09-01Private MaaS launchesproductPrivate deployment offering liveSiliconFlow enterprise teamAdds higher-touch compliance-oriented revenue path
2026-04-01Elastic GPU launchesproductHeterogeneous scheduling engine liveSiliconFlowSupports the token-factory efficiency thesis
2026-06-17Series B consideration fully paidfinancing~RMB520M Series B2026 investor syndicatePushes disclosed post-money valuation to RMB5.02B
2026-06-18Series B+ consideration paidfinancing~RMB740M Series B+Late-2026 investorsPushes disclosed post-money valuation to RMB7.74B just before listing
2026-06-30HKEX application proof filedgovernanceChapter 18C pre-commercial filingSiliconFlow, Huatai, HaitongCreates the first deep public disclosure set
2026-07-14China Renaissance added as overall coordinatorgovernanceListing process expandedSiliconFlow, China RenaissanceSignals active continuation of IPO preparation

This is the single chronology of record for the overview chapter. Dates come from the filing or directly from the cited June-July 2026 announcements; month-only milestones use the first day of the disclosed month as an anchor.

[CO001, CO003, CO032, CO033, CO034, CO042]
FO001: SiliconFlow Company Milestone Timeline

Chronology of SiliconFlow's formation, product launches, financing steps, and listing milestones from August 2023 through July 2026.

1.6 Exhibits

Chapter 02

02Market Analysis

2.1 Market Boundary, Included Spend, and Substitutes

SiliconFlow is not selling a generic 'AI' product; it sits in the inference middleware layer that turns third-party models and rented or customer-owned compute into metered token output. The June 2026 HKEX filing, SiliconFlow's models catalog, and its onboarding docs all point to the same market boundary: public-cloud serverless token APIs, dedicated inference instances, and private deployments that let enterprises or developers consume many models through one interface. That means the included spend is not model training, semiconductor design, or end-user AI SaaS seats. It is the spend tied to model serving, token distribution, orchestration, and adjacent tooling that reduces the operational work of adopting AI in production. The most important substitutes are therefore not only other startups. Closed-ecosystem hyperscaler platforms such as Amazon Bedrock, Microsoft Foundry, and Alibaba Cloud Model Studio compete for the same model-access and agent-building budgets, while independent inference platforms such as Together and Fireworks compete on developer portability, speed, and price. In many buyer journeys the status quo is a mix of direct model-lab APIs, internal self-hosting, or an incumbent cloud already trusted for security and procurement. SiliconFlow's market matters because it sits between those options: more open than a single-cloud stack, but more productized than a do-it-yourself inference layer.[CM001, CM002, CM003, CM004, CM005, CM006]

Market definition table
Segment / categoryIncluded spendExcluded spendBuyer / payerWhy it matters
Public-cloud MaaS / token APIsServerless token calls, dedicated instances, model access and routingModel training, chips, end-user AI SaaS seatsDevelopers, startups, product teamsThis is SiliconFlow's core public-cloud revenue path
Private / on-prem MaaSDeployment software and enterprise-specific token factories in customer environmentsGeneric IT outsourcing or unrelated cloud migrationRegulated enterprises, SOEs, finance, public sectorCaptures buyers that need compliance and supply control
Developer distribution layerBYOK usage inside apps, SDKs, orchestration tools, preintegrated marketplacesClosed single-app subscriptionsDevelopers and tool operatorsReduces switching friction and broadens demand capture
Hyperscaler managed model platformsEnterprise governance, agent tooling, model catalogs, routing, guardrailsConsumer AI assistantsCloud platform owners, CIO budgetsPrimary substitute set for enterprise-grade procurement
AI inference infrastructure contextUnderlying inference hardware, cloud capacity, and serving softwareTraining-only infrastructure and non-AI computeCSPs, infrastructure buyersUseful for TAM context but too broad for a direct SiliconFlow SAM

Boundary anchored on SiliconFlow's public-cloud and private-deployment positioning, then widened to the substitute set buyers actually compare during procurement.

[CM001, CM002, CM003, CM004, CM005, CM006]
FM004: Adoption funnel or value-chain map

Inference-market value chain from model and compute supply through platforms, tools, and enterprise production deployment.

[CM004, CM006, CM007, CM022, CM027, CM028]

2.2 Market Sizing Through China MaaS and Global Inference Lenses

No single public number cleanly captures SiliconFlow's opportunity, so the defensible approach is to preserve multiple lenses rather than force one TAM. On the China demand lens, IDC says enterprise MaaS token consumption jumped from 114 trillion tokens in 2024 to 1,944 trillion in 2025 and projects roughly 40,000 trillion in 2026, while public-cloud MaaS revenue reaches RMB3.07 billion in 2025 and RMB18.6 billion in 2026. On the filing lens, Frost & Sullivan's industry work inside the HKEX prospectus shows China's token-supply market growing 1,602.6% from 2024 to 2025 and reaching 53.2 quintillion tokens by 2030. Those two China-centric lenses broadly agree on explosive expansion, but they do not use identical definitions or baselines, which is exactly why they should be triangulated instead of averaged. On the broader global lens, third-party analyst pages from MarketsandMarkets, Grand View Research, and Fortune Business Insights all cluster the AI inference market around roughly USD100 billion in 2024-2025, with forecast endpoints ranging from roughly USD254 billion by 2030 to more than USD312 billion by 2034. That larger global figure is useful context, but it overstates SiliconFlow's directly reachable near-term market because much of it includes hardware and hyperscaler infrastructure spend far above an independent token-distribution layer. The better underwriting logic is stacked: global inference infrastructure for context, China public-cloud MaaS for monetizable demand, and SiliconFlow's 1.5% 2025 throughput share as a signal of present competitive relevance rather than a direct revenue-share bridge.[CM009, CM010, CM011, CM012, CM013, CM014]

TAM/SAM/SOM or sizing lens table
Lens / publisherGeographyValueCAGR / growthMethodology cueConfidenceLimitation
IDC public-cloud MaaS revenueChinaRMB3.07B (2025) to RMB18.6B (2026)~6.1x y/yEnterprise MaaS revenue on public cloudMediumRevenue lens excludes private deployment
IDC token consumptionChina1,944T tokens (2025); ~40,000T (2026)~16x in 2025; ~20x in 2026Token throughput / usage lensHighUsage is not the same as monetized revenue
Frost & Sullivan via HKEXChina2,426.3T tokens total market in 2025; 53.2 quintillion by 20301,602.6% growth 2024-2025; 638.3% CAGR 2025-2030Token supply market throughputMediumDefinition differs from IDC MaaS lens
SiliconFlow share signal via HKEXChina1.5% throughput share in 2025; rank #4 overall, #1 independentn/aCompany ranking within token-supply platformsHighShare is by throughput, not revenue
MarketsandMarkets AI inferenceGlobalUSD106.15B (2025) to USD254.98B (2030)19.2%Broad AI inference infrastructure marketMediumToo broad for direct SiliconFlow SAM
Grand View AI inferenceGlobalUSD97.24B (2024) to USD253.75B (2030)17.5%Broad AI inference infrastructure marketMediumOne year earlier base than other reports
Fortune Business Insights AI inferenceGlobalUSD103.73B (2025) to USD312.64B (2034)12.98%Broad AI inference infrastructure marketMediumLonger forecast window increases dispersion

Use these lenses together, not as additive components. China MaaS demand, global AI-inference infrastructure, and SiliconFlow share signals answer different questions.

[CM009, CM010, CM011, CM012, CM013, CM014]
FM001: Market sizing lens

A layered view from broad global AI-inference infrastructure down to SiliconFlow's current share signal in China token supply.

Layers are directional and not additive because the underlying units and market definitions differ.

[CM011, CM012, CM015, CM016, CM017, CM018]
FM002: Market estimate range

Low/base/high source-backed ranges for the broad global AI-inference market in consistent USD billions.

USD billions. Base-year draws from 2024-2025 report pages; forecast endpoint uses nearest published 2030 or 2034 figure from retained sources.

[CM016, CM017, CM018, CM019]

2.3 Buyer Segmentation, Workflow Owners, and Budget Paths

The filing makes clear that SiliconFlow's market spans several buyer archetypes rather than one homogenous user. At the low-friction end are individual developers and startups that want commitment-free model access, rapid onboarding, prepaid usage, and the ability to keep existing tools through BYOK or compatible APIs. At the higher-control end are large enterprises and institutions that care about dedicated performance, stable supply, low latency, private deployment, or supply assurance. The same pattern shows up across comparable platforms: Bedrock markets to startups and global enterprises; Microsoft Foundry markets a unified resource plane for agents, governance, and fleet-wide controls; Alibaba's Token Plan converts inference into seat-based team productivity spend; and Together and Fireworks emphasize OpenAI-compatible APIs that reduce migration work. That mix implies multiple budget owners. For self-serve API usage the first budget can live with a developer lead, startup founder, or a team productivity manager. As deployments become larger or regulated, ownership migrates toward the CIO, platform engineering leader, procurement, or security-compliance gatekeepers. The adoption path usually runs from a quick serverless experiment to heavier usage, then to dedicated instances or private deployment once cost, stability, or compliance start to matter. SiliconFlow's advantage is that it can participate in more than one stage of that path; its challenge is that every stage brings stronger competition from incumbent clouds and lower tolerance for product or supply instability.[CM022, CM023, CM024, CM025, CM026, CM027]

Segment / buyer map
SegmentBuyerUserPayer / budget ownerWorkflowAdoption triggerImplication
Individual developersDeveloper lead / founderEngineerCredit balance or small team budgetPrototype and direct API useNeed fast model access with no commitmentSelf-serve onboarding matters more than contracts
AI-native startups and toolsCTO / product leadApplication teamProduct or infra budgetEmbed multi-model inference into appsNeed portability and speed to ship featuresBYOK and pre-integrations reduce friction
Enterprise AI platform teamsCIO / platform engineeringInternal product teamsCloud / platform budgetStandardize model access across business unitsNeed governance, routing, and observabilityCompete directly with incumbent clouds
Regulated enterprises / institutionsCIO + security/complianceBusiness units and knowledge workersEnterprise procurementDedicated or private deploymentNeed stability, latency, or data-control assurancesHigher-value but longer sales cycle
Team productivity / coding organizationsEngineering manager or team adminDevelopersSeat/subscription budgetCredits-based daily AI usageNeed predictable spend and team controlsSubscription packaging can expand budget pool

Buyer map combines SiliconFlow's own segmentation with how AWS, Microsoft, Alibaba, Together, and Fireworks package inference for different budget owners.

[CM022, CM023, CM024, CM025, CM026, CM027]
FM003: Buyer / segment map

Segments mapped by budget owner, buying trigger, operational requirement, and migration path.

[CM022, CM023, CM024, CM025, CM026, CM027]

2.4 Growth Drivers, Adoption Constraints, and Valuation Relevance

The most important tailwinds are visible in both macro and platform-level sources. Stanford's 2026 AI Index says frontier-model progress did not plateau in 2025 and that organizational adoption reached 88%, while IDC argues China's MaaS competition is no longer just about cheap tokens but about the combined package of price, performance, and toolchain support. The official product pages of Bedrock, Foundry, Alibaba, Together, and Fireworks all reinforce that shift: they sell routing, observability, governance, privacy controls, dedicated throughput, caching, and model evaluation as essential features. That matters for SiliconFlow because the market is maturing from raw API access into a more operationally demanding layer where orchestration and enterprise fit are increasingly monetizable. The constraints are equally material. IDC still ranks performance, security and compliance, answer quality, platform availability, and cost effectiveness among the top enterprise selection factors. Fortune's market page highlights hardware cost and integration complexity, while the filing and independent analyses show what this means economically for a neutral platform: public-cloud growth can be real even when the vendor is still renting compute, subsidizing developers, and operating at negative gross margins. Hello China Tech adds a sharper warning by describing a 2023-2026 price war in model APIs and compute-backed token supply. For valuation, that means the market can be enormous and still not translate cleanly into attractive independent-economics businesses unless the platform can protect pricing, improve utilization, and win higher-value enterprise workloads over time.[CM029, CM030, CM031, CM032, CM033, CM034]

Growth drivers and constraints table
Driver / constraintDirectionTimingEvidenceWhy it mattersDiligence ask
Frontier-model progress and broad AI adoptionDriverNowStanford AI Index 2026More capable models create more inference demandHow much of new usage stays open/multi-model?
Multimodal and agent workloadsDriverNow to medium termIDC + hyperscaler product pagesRaises token intensity and expands use casesWhich workloads are highest value for SiliconFlow?
Price + performance + toolchain competitionDriver and filterNowIDCWinners need more than low token pricesHow differentiated is SiliconFlow tooling?
Governance, privacy, and enterprise controlsDriverNowAWS + Microsoft + AlibabaEnterprise buyers increasingly require these featuresAre SiliconFlow controls comparable enough to win large accounts?
Developer portability and OpenAI compatibilityDriverNowSiliconFlow + Together + FireworksMakes adoption easier and lowers migration costDoes easy switching help SiliconFlow or hurt moat?
Hardware cost and integration complexityConstraintNowFortune Business InsightsInference scale remains operationally expensiveWhat portion of cost can SiliconFlow structurally remove?
Price wars and subsidized user acquisitionConstraintNowHKEX + Hello China Tech + KrASIAMarket growth may outrun profit-pool growthWhen do unit economics inflect positive?
Compute supply and heterogeneous orchestrationConstraint / differentiatorNow to medium termHKEX + 36KrAccess to supply can shape margins and reliabilityHow durable are supplier relationships and chip coverage?

Several factors are two-sided: they enlarge the market while simultaneously raising the execution bar for independent platforms.

[CM029, CM030, CM031, CM032, CM033, CM034]

2.5 Conflicting Estimates and Remaining Diligence Gaps

A disciplined market chapter should preserve what is still unknown. The first gap is definitional: IDC's China MaaS framing, Frost & Sullivan's token-supply framing, and the global AI-inference reports are all useful, but they are not measuring the same basket of spend or activity. The second gap is company-specific: no retained public source isolates SiliconFlow's serviceable revenue share by geography, by customer segment, or by model category, and none shows the retention or switching dynamics that would turn gross throughput into a durable share thesis. The practical conclusion is that market evidence is strongest at the level of demand direction and weakest at the level of precise independent-platform profit pools. SiliconFlow clearly participates in a fast-growing market; it is much harder to prove from public sources alone how much of that growth will accrue to neutral token platforms versus hyperscalers, model labs, or private deployments inside large enterprises. Investors should therefore treat market size as supportive context, but reserve underwriting conviction for evidence on unit economics, segment mix, and defensibility in later chapters.[CM021, CM041, CM042]

2.6 Exhibits

Chapter 03

03Competitors

3.1 Competitive Landscape by Class

SiliconFlow's competitive set spans at least four classes. First are independent open inference platforms such as Together AI, Fireworks, and OpenRouter, which compete on model breadth, developer portability, speed, and pricing rather than on ownership of a proprietary frontier model. Second are incumbent cloud platforms such as Amazon Bedrock, Microsoft Foundry, and Alibaba Cloud Model Studio, which sell the same broad job—enterprise access to many models—but bundle it into larger governance, networking, identity, and procurement stacks. Third are direct model-lab APIs and single-vendor ecosystems, which can bypass neutral platforms entirely when a customer wants one preferred model rather than a marketplace or broker. Fourth is internal build: engineering teams can self-host open models or stitch together vendor APIs when they believe cost, control, or performance justify the complexity. SiliconFlow's own filing is clear that it wants to sit in the open, neutral middle. It explicitly contrasts itself with closed ecosystems, claims first place among independent ecosystem token-supply platforms in China, and positions itself as a connective layer across heterogeneous compute and multiple model families. That means the most relevant competitor question is not simply who has the biggest model catalog; it is which vendor class wins as buyers progress from experimentation to production. In early developer adoption, independent platforms and brokers can look interchangeable. In regulated or very large deployments, the balance often shifts toward the providers with stronger governance, procurement reach, and guaranteed supply.[CP001, CP002, CP003, CP004, CP005, CP006]

Competitor profile table
CompetitorClassScale / funding signalTarget segmentDifferentiationLimitation / diligence note
SiliconFlowIndependent token-supply platformHKEX applicant; 1.5% China throughput share in 2025; RMB7.74B latest disclosed valuation rungDevelopers, startups, enterprises, private deploymentsChina-local neutrality, heterogeneous chips, 200+ models, domestic-chip adaptationMuch smaller than hyperscalers; public-cloud economics still negative in filing
Together AIIndependent open-model platformPrivate company; retained sources here emphasize product rather than current funding or revenueDevelopers and AI-native builders scaling from serverless to dedicated GPUsOpen-source focus, strong serverless catalog, dedicated endpoints, speed claimsNo retained independent source here establishes current scale or margins
Fireworks AIIndependent open-model platformFireworks says it processes 40T+ tokens/day; retained sources here do not establish current funding or revenueDevelopers, enterprises needing training/inference, on-demand GPU usersSpecialized training + inference, serverless and on-demand, fine-tuning, performance positioningThroughput claim is vendor-authored and not independently reconciled here
OpenRouterBroker / router / aggregatorUnified API across hundreds of models; retained sources emphasize routing rather than disclosed scaleDevelopers and agent builders optimizing cost, latency, or fallback behaviorProvider routing, cost/latency sorting, multi-homing, BYOK-friendly designRelies on external providers rather than owning core compute supply
Amazon BedrockHyperscaler managed model platformAWS platform serving 100,000+ organizations on BedrockEnterprise builders from startups to global enterprisesGovernance, guardrails, model choice, batch/priority options, procurement strengthLess neutral than an independent broker and tied to AWS estate
Microsoft FoundryHyperscaler managed agent/model platformMicrosoft positions Foundry as unified control plane with 1,900+ to 11,000+ model access surfacesEnterprise platform teams and agent buildersRBAC, policy, observability, governance, model catalog, managed computePricing pages do not always expose token prices cleanly in retained text
Alibaba Cloud Model StudioRegional hyperscaler / model platformAlibaba combines Qwen ownership with third-party models and regional endpointsTeams and enterprises across China and international regionsOpenAI compatibility, multimodal Qwen stack, regional deployments, team credit plansPlatform economics and realized discounting versus list pricing remain unclear

Profile rows mix public-company incumbents, open independent peers, and a routing broker because buyers evaluate all of them against the same model-access job.

[CP001, CP002, CP003, CP004, CP005, CP006]
FP001: Competitive positioning map

Ordinal positioning by openness / portability and enterprise control / procurement strength.

Ordinal 1-10 scores synthesized from retained product, pricing, and governance evidence; not a third-party benchmark.

[CP001, CP013, CP014, CP015, CP016, CP017]

3.2 Capability Breadth and Pricing Competition

Capability overlap is substantial. SiliconFlow, Together, Fireworks, OpenRouter, Alibaba, and other rivals all present some mix of OpenAI-compatible APIs, multi-model access, and low-friction onboarding. SiliconFlow advertises 200-plus models; OpenRouter says hundreds of models through a unified API; Bedrock says 100-plus foundation models; Microsoft Foundry markets 1,900-plus models in documentation and 11,000-plus models on its pricing surface; Alibaba packages Qwen plus DeepSeek, Kimi, GLM, and other third-party models; Together and Fireworks each combine serverless and higher-control deployment paths. This is a market where feature parity on basic access is becoming table stakes. Pricing and packaging are therefore increasingly strategic. Together and Fireworks both offer per-token serverless pricing with cached-token and batch discounts, then move heavier users toward dedicated hardware. AWS Bedrock layers token pricing, batch discounts, and provisioned or reserved capacity. Alibaba mixes pay-as-you-go rates with regional discounting and credit-based team plans. OpenRouter turns pricing itself into part of the product by routing across providers and sorting for price, latency, or throughput. The net result is that buyers can often find comparable base-model access across multiple vendors, but the all-in value proposition still differs meaningfully once deployment mode, rate limits, governance, routing, and performance controls enter the decision.[CP010, CP011, CP012, CP013, CP014, CP015]

Feature / capability matrix
Buying criterionSiliconFlowTogetherFireworksOpenRouterAWS BedrockMicrosoft FoundryAlibaba Model Studio
OpenAI-compatible APIYesYesYesDrop-in OpenAI SDK pathn/a in retained sourceYes via OpenAI()/project endpointYes
Broad multi-model catalog200+ modelsYes100+ open text + vision + moreHundreds of models100+ models1,900+ to 11,000+ modelsQwen + DeepSeek/Kimi/GLM + multimodal
Dedicated / reserved capacityDedicated instances + private deploymentDedicated model inferenceOn-demand deploymentsRoutes to providers rather than own dedicated fleetProvisioned / reserved tiersManaged compute / provisioned throughputRegional deployment scopes and team plans
Routing / fallback controlsModel selection + BYOK channelsModel choice; serverless to dedicatedServing paths and tieringExplicit provider routing and fallbacksPrompt routingModel router and unified control planeRegional endpoints and billing controls
Enterprise governance / controlsPartially public from docs and filingSome docs; limited retained trust evidenceUsage metrics and dashboardsRouting/data controls but lighter enterprise stackGuardrails, privacy, complianceRBAC, network, policy, observabilityData privacy statement and monitoring

Unsupported cells are phrased conservatively and limited to retained-source evidence only.

[CP010, CP011, CP012, CP013, CP014, CP015]
Pricing / packaging comparison
VendorExample unit / packagingIllustrative retained pricingDiscount / control leverImplication
SiliconFlowPer-model inference APIPublic model library exposes per-model prices; exact realized rates not equal to marginBYOK and tool integrationsCompetes on breadth and convenience, but realized economics are undisclosed
TogetherServerless per token; dedicated per GPU-minuteQwen 3.7 Max $1.25 input / $3.75 output per 1M tokens; DeepSeek V4 Pro $1.74 / $3.48Cached-input discounts and batch discounts; dedicated cheaper at high utilizationStrong direct comparable for open-model token pricing
FireworksServerless per token; on-demand GPU-hour; fine-tuningDeepSeek V4 Pro $1.74 / $0.145 cached / $3.48; H100 on-demand $7.00/hrPriority/Fast tiers, cached-token discount, batch 50% of standardCompetes on both token pricing and higher-control infrastructure
OpenRouterProvider-routed per-token APIRates vary by underlying provider; router can sort by price, throughput, or latencyFallbacks, max_price, preferred latency/throughput, ZDRTurns routing logic into part of the commercial offer
AWS BedrockPer-token plus batch / provisioned optionsClaude Opus 4.8 $6 input / $30 output per 1M; DeepSeek v3.2 $0.62 / $1.85 in listed regions50% batch discount; standard / priority / reserved tiersEnterprise incumbent can span premium and low-cost models
Microsoft FoundryServerless / managed compute / provisionedManaged GPU families listed; many token prices obscured as $- in retained surfaceACU pre-purchase plans and resource-wide governanceCommercial model is broader platform contract, not just one API rate
Alibaba Model StudioPay-as-you-go plus credits subscriptionqwen3.7-max list price $2.5 input / $7.5 output with temporary discounts; team seats from $30 to $200 per monthNight/day discounts, free quota, shared credit packsAggressive regional discounting and seat packaging widen the competitive set

Illustrative prices are list or posted rates captured on 2026-07-22 and should not be read as realized net pricing.

[CP018, CP019, CP020, CP021, CP022, CP023]
FP002: Feature breadth / capability map

Capability comparison emphasizing where overlap is high and where meaningful divergence remains.

Evidence-backed qualitative labels derived from retained official docs and pricing pages only.

[CP010, CP011, CP012, CP013, CP014, CP015]

3.3 Switching Costs, Lock-in, and Distribution Power

For independent inference platforms, switching costs look structurally low at the API surface and much higher at the surrounding workflow layer. The common use of OpenAI-compatible endpoints across SiliconFlow, Together, Fireworks, OpenRouter, and Alibaba means that many developers can test or swap providers with limited code change. OpenRouter's routing controls and BYOK logic go even further by encouraging multi-homing instead of hard lock-in. SiliconFlow's own BYOK and preselected-tool distribution serve a similar goal: they make adoption easier, but they also make competitive displacement easier if another provider offers better cost, throughput, or reliability. Incumbents answer that weakness with distribution power. Bedrock and Foundry do not just sell tokens; they sell inference inside larger enterprise control planes with IAM, policies, guardrails, observability, and existing procurement relationships. Alibaba can similarly lean on regional cloud infrastructure, Qwen ownership, and packaged subscriptions. These surrounding advantages can create more durable attachment than model access alone. SiliconFlow therefore needs to win where openness, neutral routing across models, local chip heterogeneity, or China-specific deployment fit outweighs the convenience of staying inside a hyperscaler's broader estate.[CP024, CP025, CP026, CP027, CP028, CP029]

3.4 Moat Durability and Adverse Competitive Risks

The adverse evidence is unusually important here because the entire sector is moving toward commoditization on core API access. Hello China Tech argues mainstream model-API prices have fallen more than 90% since 2023, while the filing shows that SiliconFlow's public-cloud business prioritized share and user acquisition over near-term profitability. In China, the top three token suppliers by throughput are still hyperscaler divisions, and SiliconFlow's 1.5% share leaves it materially smaller than the largest incumbents. That does not invalidate the business, but it does mean investors should be careful about treating usage growth as equivalent to durable competitive advantage. The strongest moat candidate in retained sources is not proprietary model ownership or hard customer lock-in. It is SiliconFlow's neutral position in the China stack: broad open-model access, heterogeneous chip support including domestic compute, and the ability to serve both self-serve and private-deployment workloads. The problem is that similar neutrality and portability are also selling points for Together, Fireworks, and OpenRouter, while hyperscalers can mimic many API features and underwrite them with deeper balance sheets. Competitive durability therefore depends less on surface feature checklists and more on supply access, operational efficiency, trust, and the ability to convert price-sensitive developer traffic into harder-to-displace enterprise accounts.[CP030, CP031, CP032, CP033, CP034, CP035]

Moat durability / competitive risk register
Moat claimThreatSeverityWhy it mattersMitigation / diligence ask
China-local neutral platformHyperscalers replicate API surface and underprice trafficHighLow switching costs can erase feature-only advantagesTest enterprise win reasons beyond raw price
Heterogeneous chip orchestrationCompetitors improve multi-chip support or secure preferred supplyHighSupply access and unit economics are central to reliability and marginVerify supplier contracts and domestic-chip depth
Broad model accessCatalog breadth commoditizes quickly across routers and cloudsMediumHundreds of models alone do not create durable lock-inMeasure actual usage concentration by model family
Developer distribution through toolsBYOK and compatibility also enable multi-homing away from SiliconFlowMediumChannels boost top-of-funnel but may weaken retentionRequest cohort retention by acquisition channel
Enterprise/private deployment pathIncumbent clouds can bundle governance, network, and procurementHighBundling can overpower neutral-platform advantages in big accountsProbe security posture and procurement wins against clouds
Independent positioningPrice wars and negative-margin public cloud compress the whole categoryHighMarket growth may not translate into profit-pool growthBenchmark gross margins and promotional-credit discipline versus peers

The risk register focuses on durability, not feature checklists.

[CP029, CP030, CP031, CP032, CP033, CP034]
FP003: Moat / readiness KPIs

Compact view of what matters most in competitive durability for SiliconFlow.

[CP002, CP024, CP029, CP032, CP034, CP042]

3.5 Exhibits

Chapter 04

04Financials

4.1 Revenue Model and Pricing Architecture

SiliconFlow's filing shows two economically different businesses under one brand. Public cloud-based services include serverless token services and dedicated instances; on-premise deployment solutions install inference software in customer environments. Serverless services are prepaid, low-priced, consumption-driven products aimed at developers and smaller customers, while dedicated instances and private deployments serve larger buyers that need stability, reserved capacity, or compliance control. The public models page and API docs make the commercial surface visible: list pricing is usage-based at the model level, access is API-key driven, and the company relies on rapid onboarding and broad model choice to drive adoption. The problem is that list pricing is only the top of the revenue story. Peer pricing pages from Together, Fireworks, AWS, and Alibaba show the same market logic: cached tokens, batch jobs, dedicated capacity, and fine-tuning or deployment charges all change realized economics relative to simple per-token list rates. SiliconFlow therefore operates in a market where customers are trained to expect pay-as-you-go flexibility and frequent discounts, while heavy users can often migrate toward more cost-efficient dedicated capacity. That makes revenue quality highly sensitive to customer mix, effective discounts, and whether higher-value enterprise workloads eventually outweigh low-ARPU developer traffic.[CI001, CI002, CI003, CI004, CI005, CI006]

Revenue streams table
StreamMechanismUnitCurrent value / statusRevenue qualityDiligence ask
Serverless token servicesPrepaid, pay-as-you-go token consumptionTokens / API callsRMB14.3M revenue in 2025Low current quality: high volume, low spend densityNeed cohort retention and effective realized price per token
Dedicated instancesReserved computing capacity on public cloudInstance / reserved capacityRMB15.0M revenue in 2025Better than pure serverless but still compute-rental exposedNeed utilization and contract-length disclosure
Public cloud totalServerless + dedicated instancesRMB revenueRMB29.261M, 52.9% of 2025 revenueGrowth engine but negative marginNeed line-level margin bridge and enterprise mix
On-premise deploymentSoftware / deployment inside customer environmentProject revenueRMB26.069M, 47.1% of 2025 revenueHigher quality on margin, lower scalabilityNeed sales-cycle length and repeatability data
Model-level list pricingAPI price menu on public model libraryPer-token / per-unit list pricePublicly visible list prices on platformNot equal to realized pricing or gross marginNeed discount and promotion policy by cohort

Revenue mix is filing-backed; pricing posture is supplemented by public product surfaces and peer list-pricing evidence.

[CI001, CI002, CI003, CI004, CI005, CI009]
Pricing / monetization table
Vendor / surfacePrice / unit / contractList vs realizedDiscount / control leverSource / implication
SiliconFlow models pageUsage-based model pricingList onlyModel choice and credits affect realized spendConfirms pay-as-you-go positioning, not realized revenue
Together serverlessPer-token pricingList onlyCached-input and batch discounts; migrate to dedicatedIllustrates market pressure toward lower effective unit cost
Together dedicatedPer-GPU-minute / hourList onlyAutoscaling and reserved capacityDedicated can be cheaper at high utilization
Fireworks serverlessPer-token pricingList onlyPriority / Fast tiers, cached-input discount, batch 50% of standardOperational attributes alter effective economics
AWS BedrockPer-token plus batch / provisioned optionsList only50% batch discount; provisioned throughputLarge incumbents can span price segments
Alibaba Model StudioPer-token plus seat / credit planList onlyRegional discounting, free quota, subscription seatsPackaging broadens addressable budgets but muddies realized pricing

Official pricing pages are list pricing only; none of them reveal realized contract economics or contribution margins.

[CI006, CI007, CI008, CI010]
FI001: Revenue model bridge

How token consumption and deployment choices convert into different revenue-quality profiles.

[CI001, CI002, CI003, CI004, CI039]

4.2 Cost Structure and Unit Economics

The filing is unusually clear about what hurts current economics. In 2025 SiliconFlow generated RMB55.33 million of revenue but recorded cost of sales of RMB68.632 million, implying a negative gross profit and a blended gross margin of -24.0%. The public-cloud line was much worse: its gross loss margin was -119.0% in 2025 after -271.6% in 2024, while on-premise deployment remained high-margin at 82.5%. Compute rental dominated the cost structure, accounting for 86.9% of cost of sales, and the filing explicitly says the company prioritized market share, user acquisition, and ecosystem building over immediate profitability. Independent commentary sharpens the unit-economics picture. Hello China Tech and KrASIA both note that SiliconFlow is effectively a middle layer renting compute and packaging third-party models into a price-war environment. The best public proxy for monetization quality is not total registered users but spend density. Serverless paying accounts rose from 2,455 to 716,000 in one year, yet serverless token revenue was only RMB14.3 million, implying extremely small average annual spend per account. More than 64% of 2025 sales and marketing expense went to promotional compute credits. That mix can create large usage numbers quickly, but it does not yet prove attractive CAC payback or durable gross-margin recovery.[CI011, CI012, CI013, CI014, CI015, CI016]

Unit economics table
MetricValue / statusConfidenceWhy it mattersDiligence ask
2025 revenueRMB55.33MHighTop-line scale anchorConfirm 2026 run-rate and revenue recognition cadence
2025 blended gross margin-24.0%HighShows scale has not yet fixed economicsBridge public-cloud vs on-prem mix effects
2025 public-cloud gross loss margin-119.0%HighCore growth engine still loss-makingNeed per-product margin path and utilization targets
2025 on-prem gross margin82.5%HighHigher-quality but lower-scale revenue streamAssess repeatability and ceiling of this line
Compute rental share of COGS86.9%HighSupply cost dominates economicsNeed supplier contracts and pricing roadmap
Serverless paying accounts716,000 at 2025 year-endHighDemonstrates acquisition scaleNeed active spend distribution by cohort
Approx. serverless revenue per paying account~RMB20/year using year-end accounts as rough divisorMediumSuggests very low spend densityNeed average monthly active payer and revenue by decile
Promotional credits share of S&M>64% in 2025HighCustomer acquisition may be subsidy-heavyNeed CAC payback and promo-credit conversion

Uses public figures and simple cautionary proxies; the RMB20/account figure is explicitly approximate and not a management KPI.

[CI011, CI012, CI013, CI014, CI015, CI016]
FI002: Unit economics bridge

Why scale has not yet translated into attractive public-cloud unit economics.

[CI011, CI013, CI015, CI018, CI020, CI040]
FI003: Financial estimate range

Source-backed range view for key financial inputs and rough adequacy signals.

Range mixes business-line and balance-sheet anchors to show spread in economics; not a single management scenario model.

[CI012, CI014, CI023, CI026]

4.3 Capital Adequacy and Financing Dependency

Capital adequacy improved meaningfully after the 2026 financing sequence, but the business still looks financing-dependent rather than self-funding. At the end of 2025 SiliconFlow held RMB171 million of cash and cash equivalents plus RMB100 million of time deposits, while net cash used in operating activities was RMB172 million for the year. On that simple backward-looking lens, the company would not have had comfortable standalone runway without fresh capital. The post-2025 A+, B, and B+ financings disclosed in the filing changed that picture by injecting roughly RMB1.48 billion in cash consideration after year-end, which is why the question shifts from immediate survival to how efficiently that new capital can be converted into better unit economics. The use-of-proceeds and supplier sections still point to meaningful dependency on external compute supply and continuing commercialization spend. The company does not buy chips directly; it mainly leases computing resources through partners, and purchases are concentrated among a handful of major suppliers. That means liquidity needs are tied not just to R&D burn but to compute-availability strategy, promotional usage credits, and the mix between serverless traffic and higher-quality enterprise work. The retained public record does not support a clean current runway calculation as of 2026-07-22, so investors should treat capital adequacy as improved but still linked to operating execution rather than solved by the June 2026 rounds alone.[CI023, CI024, CI025, CI026, CI027, CI028]

Capital adequacy table
ItemValue / statusWhy it mattersConfidenceDiligence ask
Cash and cash equivalentsRMB171M at 2025 year-endImmediate liquidity base before 2026 financingsHighConfirm current unrestricted cash
Time depositsRMB100M at 2025 year-endAdds liquidity cushion but may not equal instant operating cashHighClarify maturity and restrictions
Operating cash burnRMB172M net cash used in operations in 2025Shows pre-2026 business was not self-fundingHighProvide monthly 2026 burn trend
Post-2025 financing cash consideration~RMB1.48B across A+, B, and B+Materially extends runway relative to year-end cashHighConfirm cash received net of fees and restricted uses
Compute-supply dependencyLeases compute rather than buying chips; supplier concentration is highRunway depends on supplier terms as much as cashMediumNeed payables, prepayment, and contract structure
Current runway as of 2026-07-22Not publicly supportable from retained sourcesPrevents precise underwritingHighManagement should provide cash, burn, and runway at run date

Historical funding chronology lives in Chapter 1; this table focuses on forward adequacy and dependency.

[CI023, CI024, CI025, CI026, CI027, CI028]
FI004: Capital intensity / cash-flow map

Cash requirements are shaped by compute procurement, commercialization subsidies, and financing support.

[CI024, CI025, CI026, CI027, CI028, CI029]

4.4 Public Gaps and Financial Verdict

The public record supports a strong financial caution signal but not a full underwriting model. We have revenue, gross margin, some mix data, customer-count proxies, supplier concentration, and a post-2025 financing bridge. We do not have clean net revenue retention, cohort margins, realized discount rates, customer concentration, monthly burn as of run date, or a management-backed timetable for public-cloud breakeven. Because official pricing pages across the market are list prices, not realized rates, they are useful for competitive context but not enough to infer actual SiliconFlow contribution margins. The financial verdict is therefore mixed. Revenue growth and user acquisition show demand is real. On-premise deployments appear economically healthier than public cloud. But the current core growth engine—public-cloud token supply—still looks subsidy-heavy, compute-rental-heavy, and margin-negative. For investors, the most important diligence blocker is not whether there is a large market, but whether SiliconFlow can translate its token-factory scale into materially better revenue quality and structurally improved gross margins before the next capital-reliance cycle begins.[CI031, CI032, CI033, CI034, CI035, CI036]

Public financial gaps table
Missing private metricImpactWhy missing mattersExact diligence path
Net revenue retention / gross retentionMaterialWithout retention data, user growth may mask churn or weak monetizationRequest retention by developer, enterprise, and private-deployment cohorts
Realized effective price per tokenMaterialList prices do not reveal discounting or promotion dependenceAsk for realized net pricing by top model families
Customer concentrationMaterialLarge enterprise concentration can distort revenue qualityRequest top-customer revenue share and contract terms
Current monthly burnMaterialPost-2026 financing adequacy cannot be underwritten from 2025 numbers aloneAsk for latest monthly cash burn and cash balance
Public-cloud breakeven timelineImportantCore growth engine valuation depends on margin inflection timingRequest management operating plan and utilization milestones
Sales efficiency / CAC paybackImportantPromotional credits may be masking expensive acquisitionRequest payback by acquisition channel and customer type

Every missing field is directly tied to underwriting, not curiosity.

[CI031, CI032, CI033, CI034, CI035, CI036]

4.5 Exhibits

Chapter 05

05Product & Technology

5.1 Product Definition and Module Map

SiliconFlow's product is best understood as a multilayer inference delivery platform. The Chinese homepage, English docs introduction, models catalog, and API references all present a workflow where developers or enterprise teams pick a model family, obtain an API key, and then call hosted inference endpoints using usage-based pricing. That core serverless surface is surrounded by adjacent offerings—reserved instances for stable enterprise workloads, private deployment for regulated or data-sensitive buyers, and acceleration services for customers that want faster inference on self-developed or open-source models. In other words, the company is selling access, orchestration, optimization, and deployment modes around model inference, not a single proprietary frontier model. This module mix matters strategically. It lets SiliconFlow serve different buyer jobs with one control plane: fast experimentation for developers, higher-SLA reserved capacity for enterprises, and private or hybrid deployment for customers that cannot stay on a shared public cloud. The breadth is a strength because it reduces the number of separate vendors a customer must test. It is also a complexity risk because the company must keep model inventory current, maintain heterogeneous serving paths, and explain which delivery mode actually fits each workload.[CE001, CE002, CE003, CE004, CE005, CE006]

Product module / asset matrix
Module / assetPrimary userStatus / maturityDifferentiationDiligence gap
Serverless model APIsDevelopers, startups, enterprise buildersGA / publicly documentedBroad model catalog with pay-as-you-go accessNeed realized uptime / latency / error-rate metrics
Reserved instancesEnterprise platform teamsGA / homepage-promotedDedicated capacity and cost optimization for core inference workloadsNeed SLA terms and utilization economics
Inference acceleration serviceModel builders, infra teamsGA / homepage-promotedPerformance optimization for self-developed or open-source modelsNeed independent benchmarks outside company claims
Private deploymentRegulated or data-sensitive enterprisesGA / homepage-promotedBYOC, isolation, and private deployment optionsNeed deployment references and security attestations
Models catalog / control planeAll usersGA / publicly visibleSingle surface across text, speech, image, video, and multimodal modelsNeed lifecycle / deprecation policy and migration tooling detail

Status is public-surface verified, not independently audited production maturity.

[CE001, CE002, CE003, CE004, CE006]
Workflow / use-case table
User jobCurrent workflowSiliconFlow solutionMeasurable benefitLimitation
Prototype an AI app quicklySelect model, create API key, call endpointServerless API catalogFast experimentation without GPU ops burdenCost / performance may change as model roster changes
Run stable enterprise inferenceMove from bursty demand to predictable loadReserved instancesCapacity control and stronger performance predictabilityNeed public SLA and support-package details
Deploy in private environmentKeep data or workloads off shared public cloudPrivate deployment / BYOCMore control over privacy and compliance postureImplementation complexity and reference depth unclear
Optimize custom or open-source model servingTune runtime for speed and costAcceleration serviceLower latency and better hardware utilizationIndependent benchmark coverage is limited
Switch across model families for specific tasksCompare model capabilities by taskBroad catalog across modalitiesReduced integration switching costPlatform still depends on upstream model availability

Use-case mapping focuses on customer workflow, not company marketing categories.

[CE005, CE007, CE008, CE009]
FE001: Product architecture map

SiliconFlow combines customer-facing APIs, deployment modes, optimization software, and underlying compute into one inference-delivery stack.

Layer values are qualitative weights representing relative functional scope within the public architecture, not resource allocation or revenue mix.

[CE001, CE003, CE010, CE022]

5.2 Architecture and Developer Workflow

The public technical surface is unmistakably API-first. The chat-completions reference exposes model selection, streaming, context-window controls, tool calling, and tracing headers. The broader docs introduction describes pay-as-you-go API access across text, image, speech, video, vector, reranking, and multimodal models. The workflow implied by these sources is simple for users: register, create an API key, choose a model from the catalog, integrate against a largely OpenAI-style request pattern, and then decide later whether to remain on shared serverless, migrate to reserved instances, or move into private deployment. That low-friction path is central to why SiliconFlow can aggregate so many models under one platform. Under the hood, the platform is not only an API broker. SiliconFlow repeatedly emphasizes self-developed efficient operators, optimization frameworks, acceleration engines, dynamic scaling, monitoring, and fault tolerance. The open-source OneDiff project provides the clearest external engineering proof that the company actually builds inference-optimization software rather than only packaging other vendors' models. OneDiff focuses on acceleration libraries, compiler backends, and optimized kernels for diffusion models, which does not prove the entire platform architecture but does support the broader claim that inference performance engineering is an in-house competency. The technical story is therefore more credible than pure marketing copy, though still stronger on acceleration tooling than on independently audited reliability metrics.[CE010, CE011, CE012, CE013, CE014, CE015]

Technology / operating architecture table
Layer / componentRoleDependencyRisk
API gateway and authEntry point for application callsDeveloper credentials and request managementCustomer-facing outages or auth failures hit all modules
Model catalog / routing layerMaps requests to available hosted modelsUpstream model providers and version changesFrequent model churn can force migration work
Inference acceleration layerImproves latency / throughput / costInternal optimization software and GPU-specific tuningPerformance claims may not generalize across workloads
Compute orchestration and scalingAllocates shared or dedicated capacityLeased GPU resources and autoscaling logicCapacity shortages or cost spikes affect margins and reliability
Monitoring / traceabilityOperational debugging and service assuranceLogging, trace ids, observability stackLack of public SLO reporting reduces outside verification

Architecture is public-source grounded and avoids unsupported hidden-layer speculation.

[CE010, CE011, CE012, CE014, CE023]
FE002: Customer workflow / operating flow

Typical path from evaluation to scaled production use on SiliconFlow.

[CE005, CE011, CE013, CE017]
FE003: Critical dependency map

The product relies on upstream model, compute, and open-source optimization dependencies.

[CE014, CE018, CE022, CE023, CE025]

5.3 Differentiation, Roadmap, and Dependencies

SiliconFlow's main product differentiation is combinational: breadth of model access, fast integration, multiple deployment modes, and a performance/cost optimization narrative. The homepages cite 10x+ speed improvement for language models, 1-second image generation, and double-digit cost savings for several scenarios; the docs emphasize large model coverage; and the product menu spans public cloud, reserved instances, acceleration services, and private deployment. That bundle is more differentiated than a simple model marketplace because it gives the company room to compete on service quality and deployment flexibility, not only on catalog breadth. At the same time, the dependency map is substantial. SiliconFlow depends on upstream model providers staying available, on continued access to leased GPU capacity, on developer trust in the API abstraction, and on its own ability to keep pace with a rapidly changing model roster. The API docs explicitly warn that model availability and capabilities will change over time and that some services are still being updated. The Kimi K3 launch post shows the platform can add new models quickly, but it also underlines that roadmap execution is partly a model-onboarding race rather than only a deep proprietary R&D race. That creates a product moat that is real in operations and integration quality, but potentially more fragile than the moats of model creators or hyperscalers with first-party compute and distribution.[CE020, CE021, CE022, CE023, CE024, CE025]

Trust / quality / compliance table
Control / metricStatusScopeGap
BYOC deployment supportPublicly claimedEnterprise / private deploymentsNeed architecture review and customer reference
Compute / network / storage isolationPublicly claimedData security controlsNeed third-party attestations
Privacy policyPublicly availableInternational platform data handlingNeed DPA / subprocessors / retention detail
Separate terms for China vs international surfacePublicly availableJurisdiction and contracting structureNeed legal-entity-by-customer mapping
Trace identifiers in API responsesPublicly documentedSupport and troubleshootingNeed public incident / status history
Industry-standard compliance claimMarketing-level claimEnterprise positioningNeed named certifications and audit reports

Public trust signals exist, but assurance evidence is incomplete.

[CE029, CE030, CE031, CE032, CE033]
Roadmap / release / development-stage table
Date / stageFeature / milestoneStatusImplicationSource
2026-07-22 currentExtensive model catalog with multimodal coverageLiveBreadth supports one-platform positioningDocs / models surfaces
2026-07-22 currentReasoning and thinking-mode parameters for supported modelsLiveAPI surface keeps pace with newer model behaviorsChat API reference
2026-07-22 currentReserved instance and private deployment offeringsLiveSignals enterprise packaging beyond developer APIsChinese homepage
2026-07-22 currentRapid onboarding of Kimi K3 / GLM-5.2 and similar launchesLive / recentRoadmap execution partly depends on fast partner/model integrationOfficial blog / homepage
2024 open-source release cadenceOneDiff acceleration library releasesHistorical but relevantShows continuing engineering work in inference optimizationGitHub releases

Roadmap is inferred from public release behavior; no full changelog or committed forward roadmap was found.

[CE020, CE021, CE024, CE026, CE028]
FE004: Product maturity / capability map

Publicly visible maturity and verification quality across SiliconFlow capability areas.

[CE020, CE027, CE034, CE035]

5.4 Trust, Safety, Security, and Quality Controls

SiliconFlow does publish meaningful trust and compliance signals, but they are not yet the same as comprehensive enterprise assurance. The Chinese homepage cites data isolation across compute, network, and storage, support for BYOC deployment, and compliance with industry standards and regulatory requirements. The privacy policy and terms confirm that the company has distinct China and international service surfaces, which matters for data handling and jurisdictional separation. The public API surface also includes trace identifiers that aid issue troubleshooting. Together these are real operational signals that the platform is built for supportability rather than only demo usage. However, the retained source set does not provide a public SOC 2 report, ISO certification list, public uptime history, security incident history, detailed DPA package, or benchmarked error-rate / latency SLOs. The gap is important because SiliconFlow is selling critical inference infrastructure to enterprise workloads. For diligence, the relevant question is not whether the company knows security matters—the website clearly says it does—but whether it can furnish the third-party attestations, reliability reports, and governance artifacts that large enterprise buyers will expect before standardizing on the platform.[CE029, CE030, CE031, CE032, CE033, CE034]

5.5 Exhibits

Chapter 06

06Customers

6.1 Customer Segmentation and Adoption Surfaces

SiliconFlow's customer base is not one homogeneous SaaS list. The public sources point to at least four practical segments: developers building directly on the API, enterprise platform teams using reserved or private deployments, telecom / compute partners that embed the inference stack into regional infrastructure, and community projects that route external user demand through SiliconFlow-hosted models. The QQ financing coverage and independent reports give the broadest top-line adoption markers—more than 10 million users and 10,000 enterprise customers—while the filing and case studies show that enterprise use is concentrated in compute-intensive, infrastructure-like scenarios rather than lightweight chatbot experiments alone. The adoption surfaces also matter because they imply different monetization and durability profiles. Community and developer integrations such as MindSearch, Continue, and Cline prove API relevance and onboarding ease, but they are weak proxies for high-ARPU contracts. In contrast, Guizhou Mobile cooperation, reserved-instance case studies, and private-deployment stories show deeper operational embedding, but often without public contract values or renewal terms. That split means SiliconFlow clearly has demand breadth, yet investors still need to separate proof of usage from proof of durable, high-quality revenue.[CU001, CU002, CU003, CU004, CU005, CU006]

Customer segmentation table
SegmentBuyer / user / payerUse caseScale signalRevenue / strategic valueGap
Developers and AI buildersUser=developer; payer=individual/teamPrototype and launch AI apps via APINamed integrations and community guides are abundantBroad top-of-funnel demand and ecosystem relevanceNeed active-paid-developer count and revenue share
Enterprise platform teamsBuyer=IT/AI platform lead; user=internal app teamsReserved instances, private deployment, coding or knowledge workloadsAnonymous case studies plus reserved-instance exampleLikely higher ARPU and stronger embeddingNeed named logo roster and contract duration
Telecom / compute partnersBuyer=regional infra operator; user=industry customersToken factory / inference infrastructure co-buildGuizhou Mobile partnership is explicitStrategic distribution and supply leverageNeed commercial terms and scale of live workloads
Regulated / state-linked institutionsBuyer=央企 / public-sector style institutions国产化 private deployment and cost/performance optimizationAirline and energy央企 cases show patternImportant for trust and domestic moat narrativeNamed customer list mostly withheld
Community tools / open-source appsUser=end developers or end users of third-party toolsSearch/RAG, coding, translation, agent workflowsMindSearch, Continue, Cline, and broader usercase hubHigh awareness and repeated usage surfaceNot enough evidence on monetization or exclusivity

Segments distinguish buyer, user, and payer roles rather than treating all “customers” as equal.

[CU001, CU003, CU004, CU005, CU006]
Customer growth / adoption trajectory table
MetricValueDateSourceConfidenceImplicationMissing denominator
Users served10M+ users2026-06-16QQ financing coverage / official disclosureMediumVery broad adoption surfaceUnknown monthly active share
Enterprise customers10,000+ enterprise customers2026-06-16QQ financing coverage / official disclosureMediumEnterprise reach is materially larger than a small pilot rosterUnknown production-grade share
Serverless paying accounts716,000 year-end accounts2025-12-31IPO filing / coverageMediumSelf-serve monetization exists at scaleUnknown active monthly payer rate
Reserved-instance example demandSingle customer at 100B daily tokens2026Official case studyMediumSome workloads are large enough to justify dedicated capacityAnonymous customer limits segmentation
Guizhou Mobile strategic upgradeDeepened 2026 agreement after 2025 cooperation2026-06-10Official partnership announcementHighInstitutional relationship appears multi-phase rather than one-offCommercial value undisclosed
Community usercase breadthMultiple named integration guides across search, coding, translation, and RAG2026-07-22Official docs usercase hubHighAPI relevance spans many practitioner workflowsUnknown how many convert into paid production use

This table mixes broad customer-reach markers with deeper deployment signals to show funnel layers, not a single conversion chain.

[CU002, CU007, CU010, CU013, CU020, CU029]
FU001: Customer journey map

SiliconFlow customer motion starts with easy API access and can expand into dedicated or joint-infrastructure relationships.

Stages are supported by retained sources but public conversion rates between stages are unavailable.

[CU003, CU013, CU020, CU032]
FU002: Adoption / deployment funnel

Public proof narrows from broad user and enterprise counts to the much smaller subset with deep deployment evidence.

The funnel mixes counts from different layers of proof. The final zero means no public retention cohort disclosure, not zero retention.

[CU002, CU007, CU011, CU022]

6.2 Named Customer and User Proof

The strongest public named proof comes from two channels. First are enterprise and ecosystem partnerships, especially Guizhou Mobile, where SiliconFlow publicly describes a deepened strategic collaboration around inference frameworks, token services, and joint operating systems. Second are developer-facing user cases where named tools or open-source projects publish integration guides around SiliconFlow APIs. MindSearch, Continue, and Cline are especially relevant because they show the platform being used in search/RAG and coding-agent workflows that are naturally token-intensive and repeat usage-heavy. These proofs are not all equal. Guizhou Mobile is a named institutional counterparty with executive quotes and a concrete operating scope, which is much stronger evidence than a how-to guide. The MindSearch, Continue, and Cline pages prove awareness and practical integration, but they do not disclose paid conversion, scale, or exclusive commitment. Anonymous enterprise case studies in aviation, energy, and coding agents add depth on deployment patterns and outcomes, yet their anonymity limits concentration analysis. The correct reading is therefore that named proof exists and spans both developer and enterprise surfaces, but only part of it qualifies as strong commercial evidence.[CU010, CU011, CU012, CU013, CU014, CU015]

Named customer proof table
Customer / projectSegmentDeployment / use caseProduction vs pilotOutcome / proofLimitation
Guizhou MobileTelecom / regional AI infrastructureInference framework deployment, token services, joint operationsProduction / strategic cooperation evidence stronger than pilotOfficial signed agreement with executive quotes and multi-workstream scopeRevenue, deployment count, and renewal terms undisclosed
MindSearchDeveloper / search-RAG projectIntegration with SiliconFlow API and deployment guidePractical integration proofOfficial how-to guide with config and deployment steps, including HuggingFace Space pathDoes not prove paid production scale or exclusivity
ContinueDeveloper / IDE coding assistantVS Code / JetBrains integration with SiliconFlow-hosted modelsPractical integration proofOfficial guide maps real model selection, caching, and verification workflowNo disclosed conversion or spend data
ClineDeveloper / coding agentOpenAI-compatible API integration for coding-agent workflowsPractical integration proofOfficial guide details base URL, model IDs, and multi-mode setupNo customer-count or revenue disclosure

Named proof spans one institutional partner and multiple developer-tool integrations; this is useful evidence of adoption breadth, but it is not a substitute for a named enterprise revenue roster.

[CU011, CU014, CU015, CU016, CU017, CU018]
FU003: Customer proof matrix

Institutional partnership proof is stronger than classic retention visibility; developer-tool proof is broad but lighter on commercial weight.

Labels reflect the combination of naming, deployment detail, quantified outcomes, and revenue-proximity.

[CU014, CU018, CU019, CU023, CU030, CU033]

6.3 Durability, Expansion, and Concentration

Public sources support an expansion story more convincingly than a retention story. The reserved-instance case shows one lifecycle from pay-as-you-go usage toward locked dedicated capacity as token demand grows. The Guizhou Mobile partnership points to platform embedding deeper into regional compute and service infrastructure. The enterprise case studies show SiliconFlow moving beyond public-cloud API calls into private deployment,国产 chip adaptation, and broader operational integration. Those patterns are consistent with land-and-expand behavior: start with easy API access, then deepen into reserved instances, private deployment, or joint operations as usage becomes mission-critical. What is missing is classic SaaS durability evidence. The retained public set does not disclose NRR, GRR, churn, average contract length, renewal rates, top-customer concentration, or segment-level revenue contribution. The official and independent record also does not separate how many of the 10,000 enterprise customers are truly production-grade, how many are small-budget experimentation accounts, or what percentage of demand is mediated through a handful of strategic partners. That means customer breadth is real, but customer quality and concentration remain unresolved diligence questions.[CU020, CU021, CU022, CU023, CU024, CU025]

Retention / repeat usage / satisfaction table
MetricValue / nullSegmentConfidenceDiligence ask
Public NRRnullAll paying customersHighRequest NRR by developer, enterprise, and private-deployment cohorts
Public GRR / logo retentionnullAll paying customersHighRequest logo retention and renewal rates by segment
Contract durationnullEnterprise / telecomHighRequest average term and renewal options for reserved and private deals
Repeat usage proofQualitative only: some workloads move from pay-as-you-go to reserved capacityEnterprise heavy usersMediumProvide conversion rate from API spend to reserved-instance deployment
Satisfaction / reference depthMixed: strong workflow detail, weak public formal testimonialsDevelopers and enterprise buyersMediumRequest NPS / CSAT and customer references
Public adverse churn evidenceNo robust churn dataset in retained corpusAll segmentsMediumProvide churn reasons and lost-logo analysis

Null means no public disclosure in retained sources, not that the metric is irrelevant.

[CU021, CU022, CU023, CU024, CU030]
Expansion and concentration risk table
Expansion driverConcentration riskImpactDiligence path
API to reserved-instance upgrade pathUnknown whether expansion depends on a small set of whalesLarge revenue concentration could hide beneath broad user countsRequest spend deciles and expansion cohorts
Private deployment for regulated enterprisesNamed proof is thin because key case studies are anonymousHard to assess renewal risk and sales-cycle durabilityRequest named references under NDA
Telecom / infra partnershipsStrategic partners may become concentrated routes to volume growthCould create bargaining-power and dependency issuesRequest partner contribution to bookings and compute supply
Community-tool integrationsHigh awareness may not equal durable monetizationCan inflate usage without proving high-LTV customersRequest paid conversion from community integrations
Cross-industry央企 adoption narrativeOfficial story is strong, but exact vertical mix is vagueCould overstate diversification if revenue is concentratedRequest revenue by vertical and top ten accounts

The main uncertainty is not whether demand exists, but where revenue concentration hides inside the demand base.

[CU025, CU026, CU027, CU028, CU031]
FU004: Retention / repeat cohort visibility proxy

Public sources show customer formation and some expansion clues, but not true retention percentages.

This is a visibility proxy. “1” marks one clear public expansion pattern rather than a rate; zeros mean no public percentage or cohort disclosure.

[CU021, CU022, CU024, CU034]

6.4 Customer Verdict and Gaps

The overall customer verdict is positive on relevance and mixed on durability. SiliconFlow has enough public proof to show that developers actively integrate the API, that institutional buyers are willing to use the company in significant inference contexts, and that at least some workloads become large enough to justify reserved capacity or private deployment. This is more than superficial logo collection. The customer narrative is supported by case-study detail, integration depth, and specific workflow framing. However, the public evidence is still not strong enough to underwrite concentration risk, expansion efficiency, or retention quality. Community user cases are useful but not equivalent to paid production logos. Anonymous enterprise stories prove pattern, not precise account value. Even Guizhou Mobile is a strategic partnership rather than a disclosed revenue contract. The diligence priority is therefore straightforward: obtain segmented active-customer counts, retained cohorts, top-customer exposure, and expansion conversion from serverless into reserved or private deployments.[CU029, CU030, CU031, CU032, CU033, CU034]

6.5 Exhibits

Chapter 07

07Risks

7.1 Regulatory and Legal Risk

SiliconFlow sits inside multiple overlapping compliance regimes. Its China terms and privacy policy require real-name verification for many users, restrict certain application domains, commit the company to AI-generated-content labeling obligations, and place user-information storage within mainland China for domestic operations. These are not generic website boilerplate points: they directly affect onboarding friction, enterprise contracting, product design, and logging obligations. The terms explicitly state the service is not suitable for automatic control, medical information services, psychological counseling, and critical information infrastructure scenarios, which narrows some of the highest-consequence deployment categories unless customized controls exist outside the public surface. The broader Chinese policy stack is becoming denser rather than lighter. The AI labeling measures effective in 2025 require explicit and implicit labels for generated synthetic content, with logging, metadata, and platform-distribution responsibilities. SiliconFlow's domestic terms already acknowledge those rules and prohibit users from deleting or tampering with labels. That alignment is positive, but it also means compliance risk is operational: if SiliconFlow's tools or enterprise customers mis-handle labeling, metadata, or downstream distribution, the company may face administrative scrutiny even when model output generation itself is legal. The legal risk is therefore less about a single banned activity and more about the burden of continuously implementing a moving regulatory stack across domestic and international surfaces.[CR001, CR002, CR003, CR004, CR005, CR006]

Regulatory / legal risk register
Rule / caseJurisdictionStatusLikelihoodSeverityMitigationResidual exposureDiligence path
AI-generated content labeling obligationsChinaEffective / activeMedium-highHighTerms acknowledge labeling duties; provider can add labels and retain logsOperational mis-implementation risk remainsRequest labeling architecture and compliance audit
Real-name verification and identity checksChinaActive contractual / legal requirementMediumMedium-highDomestic onboarding and verification procedures in privacy/termsFriction and privacy risk remain for some usersRequest KYC flow and exception handling
Telecom / internet service licensingChinaActiveMediumHighICP / telecom filings visible on domestic siteScope of license fit for evolving services is not publicly tested hereObtain counsel memo mapping licenses to current services
Sensitive-domain exclusions (medical, psych, auto-control, CII)China / customer contractActive contractual limitationMediumHighPublic terms narrow scope unless custom arrangements existEnterprise misuse could create legal or reputational issuesRequest sector-specific deployment controls and approval process
Cross-border legal-surface splitChina vs internationalActiveMediumMediumSeparate China and international terms/privacy surfacesOperational confusion or inconsistent controls can still occurReview entity, data-flow, and contracting map

Rows are ordered by investment relevance rather than by comprehensive legal taxonomy.

[CR001, CR002, CR003, CR004, CR005, CR006]
FR001: Risk heatmap

Highest residual risk clusters around regulation, compute dependency, and low-visibility customer economics.

[CR006, CR015, CR024, CR029, CR037]

7.2 Operational, Security, and Quality Risk

Operationally, SiliconFlow looks more like infrastructure than like a simple application layer. That raises the cost of outages, latency swings, bad deployments, and compliance failures. Case studies and product pages emphasize private deployment, heterogeneous chip adaptation, reserved instances, monitoring, fault tolerance, and cost optimization—all signals that customers are placing meaningful workloads on the platform. But the public assurance surface remains incomplete. The retained corpus does not show public uptime reports, detailed SLOs, SOC 2 or ISO attestations, or a public incident archive. That makes it hard for investors to judge whether operational maturity matches the company's market position. Security and content-governance exposure also interact. The domestic terms say the service should not be used for certain sensitive domains and require users to comply with wide content restrictions; the privacy policy explains that SiliconFlow collects logs and identity information for compliance and security purposes; and labeling-related rules preserve logs for six months in some cases. These controls can help reduce misuse, but they also expand the surface area for privacy, content, and access-control failures. For an inference platform, the relevant risk is not only whether the model hallucinates, but whether the platform can reliably enforce account controls, data handling, labeling, and operational isolation at scale while still shipping fast enough for the market.[CR012, CR013, CR014, CR015, CR016, CR017]

Operational / quality / security risk register
Failure modeLikelihoodSeverityMitigation maturityResidual exposureUnresolved gap
Platform outage / latency instability on critical workloadsMediumHighMediumHighNo public uptime/SLO record in retained corpus
Security-control gap versus enterprise expectationsMediumHighLow-MediumHighNo public SOC 2 / ISO / incident archive identified
Privacy / log-handling failureMediumHighMediumMedium-highDomestic privacy policy shows significant data-handling duties
Labeling / metadata enforcement failureMediumMedium-highMediumMedium-highNeed proof of pipeline-level explicit and implicit labeling controls
Quality or model-lifecycle disruption from rapid model churnMedium-highMediumMediumMediumAPI docs explicitly warn model availability can change

This register focuses on infrastructure-like failure modes rather than abstract model-risk discourse.

[CR012, CR013, CR014, CR015, CR016, CR017]
FR002: Risk transmission map

Shows how regulation and operations transmit into revenue quality, margin, financing, and valuation.

[CR007, CR016, CR030, CR037]

7.3 Dependency, Financial, and Competitive Risk

The most acute non-regulatory risk is external dependency. SiliconFlow does not own a first-party frontier model franchise or a first-party global hyperscale cloud. Its model breadth depends on upstream model partners staying relevant and available; its gross margin depends on access to external compute; and its supply resilience depends on both domestic chip adaptation and continued ability to source or lease advanced semiconductor capacity. The company's own filing and independent analyses show compute rental dominating cost of sales and major supplier concentration remaining high. That makes the business vulnerable to both cost spikes and strategic bargaining from suppliers or strategic channels. Geopolitics amplifies the problem. U.S. export-control guidance continues to tighten around advanced computing semiconductors and diversion to PRC-linked entities, while China simultaneously increases domestic governance over AI-generated content and internet services. SiliconFlow may adapt with domestic chips, private deployment, and token-factory operating models, but those are mitigations, not immunity. Competitive pressure compounds the exposure: the same market that validates SiliconFlow also trains buyers to compare it against hyperscalers and other inference providers on price, reliability, and integration convenience. If price war, labeling compliance costs, and compute scarcity all intensify together, SiliconFlow's weakest link becomes margin and capital adequacy rather than simple demand generation.[CR022, CR023, CR024, CR025, CR026, CR027]

Partner / dependency risk register
DependencyCounterparty / classRoleConcentrationFailure scenarioSeverityMitigationResidual exposure
Upstream model providersModel labs / OSS ecosystemsCatalog breadth and user demandMedium-highModel access or competitiveness weakens quicklyHighMulti-model routing and fast onboardingStill lacks first-party model control
Leased compute suppliersGPU / cloud / infrastructure partnersCore service deliveryHighCost spikes or capacity constraints compress margin and reliabilityHighDomestic-chip adaptation and partner networkSupply bargaining power remains external
Strategic telecom / infra partnersGuizhou Mobile and similar channelsDistribution plus compute coordinationMediumPartner-led growth becomes concentrated or politically exposedMedium-highJoint operations and regional ecosystem strategyEconomic contribution mix is not public
Developer ecosystem surfacesContinue / Cline / MindSearch and community toolsDemand generation and usage growthLow-MediumAwareness fails to convert into paid durable accountsMediumLow-friction API and many modelsCommercial conversion data absent
Regulatory dependenciesCAC / MIIT / local compliance regimesOperating permission and content governanceHighRule changes force workflow redesign or stricter auditsHighTerms, privacy controls, labeling alignmentRules continue to evolve

Dependencies mix commercial, technical, and regulatory counterparties because all can transmit into growth and margin.

[CR022, CR023, CR024, CR025, CR032]
People / execution risk register
Role / functionDependency or gapLikelihoodSeverityMitigationDiligence path
Founder / senior technical leadershipInference infrastructure and regulatory navigation remain founder-heavyMediumMedium-highRecent financing broadens resourcesAssess second-line operating bench
Compliance / legal operationsMust map fast-moving AI, privacy, telecom, and labeling rules into productsMedium-highHighPublic terms show awarenessRequest compliance org chart and escalation process
Enterprise delivery / supportPrivate deployment and reserved instances require strong implementation disciplineMediumHighCase studies imply experienceRequest deployment timelines, staffing, and support SLAs
Supplier / capacity planningMargin and service quality depend on accurate capacity and supplier managementHighHighPartner relationships and domestic adaptationRequest procurement governance and contingency plans

Execution risk is tied more to operating depth than to raw headcount size.

[CR026, CR027, CR028, CR033]
FR003: Dependency map

SiliconFlow depends on regulators, compute suppliers, model providers, and enterprise channels at the same time.

[CR023, CR024, CR025, CR027, CR032]

7.4 Mitigations, Monitoring, and Kill Triggers

The retained evidence does show mitigation paths. SiliconFlow has distinct domestic and international legal surfaces, real-name and content controls, BYOC / private-deployment options, domestic-chip optimization narratives, reserved instances, and joint operations with compute partners like Guizhou Mobile. Those mechanisms can reduce some of the most obvious risks: private deployment can help customers with data localization concerns, labeling obligations are now reflected in terms, and dedicated capacity can stabilize heavy workloads. The question is whether execution quality will keep pace with regulatory complexity and commercial scaling. From an investment perspective, the right kill criteria are observable. A negative regulatory event around AI-content labeling or telecom compliance, evidence of export-control-induced compute bottlenecks, a failure to improve public-cloud margin, or proof that enterprise customer concentration is much higher than implied by broad user counts would each materially damage the thesis. In contrast, the thesis strengthens if SiliconFlow can show audited security controls, stable supplier diversification, clear enterprise retention, and continued migration of high-volume accounts into higher-quality dedicated or private-deployment contracts.[CR032, CR033, CR034, CR035, CR036, CR037]

Mitigation and kill criteria table
RiskMonitorable triggerThreshold / eventAction implication
China regulatory compliance driftRegulator notice, enforcement, or required remediationFormal action tied to labeling, privacy, or telecom compliancePause / re-underwrite compliance readiness
Compute-supply shockCapacity shortage, material supplier loss, or cost spikeService degradation or gross-margin re-deteriorationRe-cut downside case and runway assumptions
Customer-quality shortfallWeak enterprise retention or whale concentrationRetention materially below expectation or top-customer share unexpectedly highReduce conviction and pressure-test valuation
Operational maturity gapSecurity incident or inability to furnish enterprise assurance artifactsMajor incident or failed enterprise auditDelay or avoid underwriting enterprise moat
Price-war / competition escalationPersistent margin pressure with no mix improvementPublic-cloud losses fail to narrow despite scaleTreat scale as low-quality growth

Kill criteria are designed to be observable in diligence or early post-investment monitoring.

[CR034, CR035, CR036, CR037, CR038, CR039]

7.5 Exhibits

Chapter 08

08Valuation

8.1 Recommendation, Risk Rating, and Price Discipline

SiliconFlow clears the threshold for strategic relevance but not for conviction buying at the publicly described June 2026 round price. The company has several features investors want: a real product, a large market, developer adoption, enterprise/infrastructure use cases, a visible role in China's inference stack, and substantial new capital. But the same public record also shows a business with negative public-cloud gross margins, supplier dependence, limited retention visibility, and meaningful regulatory / execution complexity. That mix argues for price sensitivity. The central valuation conclusion is therefore not that SiliconFlow is low quality; it is that the evidence quality lags the price. If the $1.2B valuation is buying access to a future high-margin infrastructure platform with defensible enterprise retention and improving supply economics, it may still work. If it is buying a scale story that remains subsidy-heavy and exposed to price wars, then the current entry point offers limited margin of safety. On public evidence alone, that is a Track recommendation rather than Buy.[CV001, CV002, CV003, CV004, CV005, CV006]

Recommendation summary table
RecommendationConfidenceRisk ratingValuation stanceDecision implication
TrackMedium-lowHighPrice-sensitive / not enough public support for Buy at $1.2BContinue diligence; participate only with stronger proof or better entry terms

Recommendation is based on public evidence only and is intentionally price-sensitive.

[CV001, CV002, CV003, CV040]
Thesis / anti-thesis table
ArgumentWhat would change the view
SiliconFlow is a strategically relevant China-based inference platform with broad model access, enterprise deployment options, and ecosystem momentum.Upgrade if 2026 revenue quality, retention, and margin trend are stronger than public evidence currently shows.
Recent financing and strategic investors could strengthen distribution and ecosystem leverage.Upgrade if investor/customer overlap demonstrably converts into durable enterprise revenue.
Public-cloud economics, compute dependence, and regulatory complexity make the current evidence base too weak for a buy call.Downgrade further if public-cloud margins stay deeply negative or supplier concentration worsens.
Customer breadth is real, but retention and concentration are opaque.Upgrade if cohort retention and spend concentration are comfortably better than implied by current gaps.

Each row is written as a falsifiable investment argument rather than a generic pro/con list.

[CV010, CV011, CV012, CV013, CV031]
FV001: Recommendation logic
[CV001, CV003, CV010, CV031, CV040]

8.2 Thesis, Anti-thesis, and Current Valuation Context

The thesis is straightforward: SiliconFlow is one of the few China-based inference platforms with enough breadth, speed, and ecosystem momentum to matter. It combines model access, API distribution, private deployment, and domestic chip adaptation in a market where inference is becoming the economic center of AI deployment. The recent financing also appears strategically rich, with industry investors that could matter for distribution and ecosystem pull. This is what gives the company optionality well beyond a narrow API arbitrage play. The anti-thesis is equally straightforward. Public-cloud economics remain poor, compute rental dominates costs, customer-retention quality is opaque, and the platform is exposed to both domestic AI governance and external semiconductor constraints. Compared with fast-growing Western inference peers, SiliconFlow's disclosed revenue scale is much smaller, and its public evidence on gross-margin quality is weaker. That means the company can be strategically important and still be a mediocre risk-adjusted entry at the current price. The disclosed $1.2B valuation is not obviously absurd in absolute terms, but it is not obviously supported either without a much stronger 2026 revenue and margin picture.[CV010, CV011, CV012, CV013, CV014, CV015]

Bull / base / bear scenario table
ScenarioAssumptionsValuation / return logicKey risksProbability signal
BullEnterprise retention proves strong; dedicated/private mix rises; public-cloud gross losses narrow sharply; strategic channels deepen.Exit value ~US$2.0B–3.0B over 3-4 years; current price can still work if margin inflection is real.Execution still depends on compute supply and regulatory control.Possible but not best-supported by current evidence.
BaseSiliconFlow remains important, grows, and avoids severe disruption, but revenue quality and margin path improve only gradually.Value range ~US$1.1B–1.6B; current entry offers limited upside for the risk.Price war, compliance cost, and mix quality keep returns muted.Best-supported by current public evidence.
BearScale does not cure economics; concentration or regulation bites; next funding occurs under weaker terms.Value range ~US$0.7B–1.0B; down-round or flat outcome plausible.Supplier shocks, losses, and weak retention compress the case.Must be taken seriously given public gaps.

Ranges are scenario-based public estimates, not management forecasts or a DCF.

[CV020, CV024, CV025, CV026, CV027, CV028]
Comparable valuation table
ComparableMetricMultiple / valuation / statusRelevanceLimitation
Together AI2026 valuation and bookingsUS$8.3B valuation; >US$1.15B annual bookingsDirect open-model inference and infrastructure peerBookings not the same as recognized revenue; different geography and scale
Fireworks AI2026 valuation and annualized revenueUS$17.5B valuation; >US$1B annualized revenueStrong direct inference-cloud peer with open-model focusMuch larger revenue scale and stronger disclosed monetization
Baseten2026 valuation and annualized revenueUS$13B valuation; ~US$600M annualized revenueInference-platform peer with enterprise deployment focusSacra estimates and US enterprise profile differ from China context
CoreWeave2025-2026 revenue and valuation contextUS$23B pre-IPO valuation reference; US$12B–13B 2026 revenue guideShows ceiling and risks of AI compute infrastructureMore capital-intensive GPU cloud than SiliconFlow; not a pure inference API peer

The table is intentionally partial because public China-specific pure-play inference comparables with direct valuation disclosure remain limited.

[CV014, CV021, CV022, CV023, CV029]
FV002: Valuation sensitivity
[CV020, CV024, CV025, CV026, CV028]

8.3 Bull / Base / Bear Scenarios, Comparable Set, and Return Logic

The peer set matters because it highlights a gap between strategic category value and current operating proof. Together AI, Fireworks AI, and Baseten all command multi-billion-dollar valuations, but public sources also show them at meaningfully higher revenue or bookings scale, with more explicit evidence around enterprise monetization. CoreWeave proves that infrastructure-adjacent AI platforms can become extremely valuable, but it also shows how capital intensity and customer concentration can remain severe even at much larger scale. SiliconFlow's lower headline valuation than these Western peers therefore should not be read automatically as cheap. Lower quality of disclosed economics can still justify a discount. This leads to three scenarios. In the bull case, SiliconFlow converts its scale and strategic relationships into better enterprise retention, stronger dedicated/private mix, and improving margins, allowing the business to outrun its current risk load. In the base case, the company remains important but operationally messy, with enough traction to defend the current round but not enough proof to produce compelling venture-style upside from that price. In the bear case, compute dependence, regulation, and price competition prevent margin inflection, producing either a flat valuation outcome or a future down-round. Public evidence today supports the base case more than the bull case.[CV020, CV021, CV022, CV023, CV024, CV025]

Thesis-break and kill triggers table
TriggerThresholdTransmission to thesisAction implication
Regulatory enforcementFormal action or remediation tied to AI labeling, privacy, or telecom complianceUndermines execution and enterprise trustPause / avoid until resolved
Margin failureNo meaningful improvement in public-cloud economics through next observable periodScale narrative fails to convert into business qualityMove from Track to Avoid at current price
Supplier / compute shockMaterial capacity disruption or worsening supplier concentrationThreatens reliability and gross margin simultaneouslyRe-cut downside case aggressively
Customer-quality disappointmentRetention or concentration is materially worse than hopedDestroys the case for premium infrastructure valuationDo not pay growth premium
Diligence miss on security / assuranceCannot produce enterprise-grade assurance packageWeakens high-value enterprise adoption thesisConstrain upside case and sales-quality assumptions

These triggers are observable and investment-actionable.

[CV032, CV033, CV034, CV035, CV039]
FV003: Valuation / return range
[CV020, CV024, CV025, CV026, CV027, CV028]

8.4 Final Diligence Asks and Thesis-break Triggers

The most important thing missing from the public file is not market demand but underwriting detail. Investors need current cash and burn, 2026 revenue run rate, gross-margin trend by product line, enterprise retention, top-customer concentration, supplier concentration, and conversion from serverless API usage into higher-quality dedicated or private deployments. Without those inputs, the valuation discussion becomes mostly narrative. That is why the final call is Track with high risk and medium-low confidence. The thesis can move up with evidence: improved margins, diversified compute supply, named enterprise references, audited security posture, and better revenue-quality metrics. It can move down with evidence too: regulatory action, supplier shocks, continued negative public-cloud economics, or proof that usage breadth still fails to convert into durable retained spend. A price cut alone would help, but only if the missing diligence items also stop pointing to structural weakness.[CV031, CV032, CV033, CV034, CV035, CV036]

Final diligence asks table
TopicMissing evidenceWhy it mattersOwner / diligence path
2026 revenue run-rate and mixCurrent run-rate by serverless, dedicated, and private deploymentDetermines whether the round is being underwritten on improving quality or on raw growthManagement / finance diligence
Retention and concentrationNRR, GRR, churn, top-customer share, spend decilesCore question for revenue durabilityManagement / customer diligence
Current cash and burnRun-date cash, monthly burn, financing proceeds net of restrictionsDetermines runway and next-round riskFinance diligence
Supplier concentration and contingencyTop suppliers, contract terms, domestic-chip fallback economicsDirectly affects reliability and margin riskOps / procurement diligence
Security and assuranceSOC 2 / ISO, uptime history, incident history, SLA packageRequired to underwrite enterprise moat qualitySecurity diligence
API-to-enterprise expansion motionConversion from pay-as-you-go usage into reserved/private contractsShows whether bottom-up adoption compounds into better revenue qualityGTM diligence

All asks are chosen because they directly move the investment call.

[CV036, CV037, CV038, CV040]
FV004: Investment KPIs
[CV001, CV004, CV005, CV036, CV040]

8.5 Exhibits

Disclaimer

This report is produced for diligence and informational purposes only. It is based on publicly available materials as of 2026-07-22 and does not constitute investment, legal, accounting, or tax advice. SiliconFlow is a private company; public disclosures remain incomplete on retention, concentration, current revenue run-rate, and enterprise assurance. Readers should independently verify all facts and obtain primary diligence materials before making investment decisions.

Evidence index

Claims
IDStatementConfidenceSources
CO001 Beijing SiliconFlow Technology Co., Ltd. was established in the PRC as a limited liability company on August 29, 2023. Medium SO011
CO002 SiliconFlow's registered office and head office are at Room 2301, 23/F, Tower D, Building 8, No. 1 Yard, Zhongguancun East Road, Haidian District, Beijing. High SO001, SO011
CO003 The company converted to a joint stock limited company on June 22, 2026. Medium SO011
CO004 SiliconFlow maintains a Hong Kong place of business while stating in the HKEX filing that its headquarters, senior management, and operations are primarily based outside Hong Kong. Medium SO011
CO005 The global .com service is governed by SiliconFlow Technology Pte. Ltd., and the global terms explicitly direct mainland-China users to use siliconflow.cn instead. High SO008, SO009
CO006 SiliconFlow publicly positions itself as a global AI infrastructure provider whose mission is to accelerate AGI for the benefit of all. High SO002, SO003
CO007 The China site presents SiliconFlow as a product suite spanning large-model APIs, reserved instances, inference acceleration services, and private deployment. High SO001, SO025
CO008 SiliconFlow's API documentation describes the platform as a one-stop cloud service for top-tier large-language-model APIs aimed at developers and enterprises. Medium SO007
CO009 The global homepage markets SiliconFlow as one platform for AI inference across text, image, video, audio, search, coding, and agent workloads. Medium SO002
CO010 SiliconFlow's GitHub organization says the company integrates hundreds of state-of-the-art models across language, speech, vision, and multimodal domains on top of a self-developed inference engine. Medium SO010
CO011 SiliconFlow's quickstart and cloud login surfaces show sign-in options via SMS, email, GitHub, and Google, alongside self-serve API key creation. High SO005, SO024
CO012 The chat-completions documentation exposes an OpenAI-style API surface with model selection, streaming, max_tokens, JSON-format support, and tool-related parameters. High SO006, SO023
CO013 As of April 30, 2026, SiliconFlow's platform had over 10 million registered users. Medium SO011
CO014 SiliconFlow recorded average daily token throughput of approximately 578.5 billion and peak daily throughput of approximately 1,071.4 billion in April 2026. Medium SO011
CO015 As of the latest practicable date in the HKEX filing, SiliconFlow had served over 13,000 enterprise customers. Medium SO011
CO016 As of the latest practicable date in the HKEX filing, SiliconFlow had supported a cumulative total of over 170 models. Medium SO011
CO017 By July 2026, SiliconFlow's public model library marketed more than 200 models, indicating platform expansion after the June 2026 filing snapshot. Medium SO004
CO018 The HKEX filing describes SiliconFlow as the largest independent ecosystem token supplier and one of the top five token suppliers overall in China. Medium SO011
CO019 The HKEX filing also describes SiliconFlow as a global leader in token throughput, registered users, and monthly active users, with globally leading overseas platform downloads and token throughput on authoritative platforms. Medium SO011
CO020 SiliconFlow's China product page advertises 10x-plus speed gains for language models, 66% image-model cost savings, 46% language-model cost savings, and BYOC plus isolation-based security controls. Medium SO025
CO021 The China homepage discloses a Beijing ICP filing and a Beijing value-added telecommunications business permit. Medium SO001
CO022 Dr. Yuan Jinhui is disclosed as SiliconFlow's founder, chairman, executive director, CEO, general manager, and financial controller. Medium SO011
CO023 Mr. Liu Juncheng is disclosed as executive director and chief technology officer. Medium SO011
CO024 Mr. Zeng Hua is disclosed as executive director and deputy general manager overseeing commercialization planning. Medium SO011
CO025 Mr. Chen Yingjie is disclosed as non-executive director and is also the managing director of Alibaba Group's strategic investment department. Medium SO011
CO026 Upon listing, SiliconFlow's board is expected to comprise seven directors: three executive, one non-executive, and three independent non-executive directors. High SO011, SO012
CO027 The HKEX history section says SiliconFlow developed under the leadership of co-founders Dr. Yuan, Liu Juncheng, Zeng Hua, Zhao Zhen, and Hu Jian. Medium SO011
CO028 Dr. Yuan previously served as a lead researcher at Microsoft China and later founded OneFlow, a deep-learning framework company. High SO011, SO018
CO029 Liu Juncheng previously worked in software development and AI system software, including as an R&D engineer at OneFlow from 2018 to 2023. Medium SO011
CO030 Zeng Hua previously held senior roles at Microsoft China, Baidu, and JD.com before joining SiliconFlow. Medium SO011
CO031 Independent profile databases also describe SiliconFlow as Beijing-based and founded in 2023. Medium SO020, SO021
CO032 The HKEX filing records angel, angel+, pre-A, and Series A capital injections of RMB47.20 million, RMB65.84 million, RMB71.48 million, and RMB285.99 million, respectively. Medium SO011
CO033 The HKEX filing records Series A+, Series B, and Series B+ capital injections of RMB220 million, RMB520 million, and RMB740 million, respectively, all in 2026. Medium SO011
CO034 The HKEX filing shows post-money valuation steps of RMB280.0 million, RMB565.8 million, RMB985.0 million, RMB2.286 billion, RMB3.120 billion, RMB5.020 billion, and RMB7.740 billion from angel through Series B+. Medium SO011
CO035 The HKEX filing says SiliconFlow completed the Series A+, Series B, and Series B+ financings after December 31, 2025 for aggregate cash consideration of approximately RMB1.48 billion. Medium SO011
CO036 External June 2026 financing coverage described SiliconFlow as completing over RMB2 billion of Series B financing backed by investors including Trip.com or Ctrip Strategic Investment, JinkoSolar, Kingdee, Unicom-linked capital, Biren, NIO Capital, SenseTime, GGV, and others. High SO013, SO014, SO015, SO016, SO017
CO037 Caixin Global said the June 2026 financing was SiliconFlow's fifth funding round since inception and that China Renaissance was the exclusive financial advisor. Medium SO013
CO038 Sina and 36Kr coverage said SiliconFlow served over 10 million users and 10,000 enterprise customers, grew revenue by over 10 times year-on-year, and reached millions of US dollars in overseas monthly revenue. Medium SO014, SO015, SO016, SO017
CO039 The HKEX filing says SiliconFlow's revenue rose from RMB7.3 million in 2024 to RMB55.3 million in 2025. Medium SO011
CO040 The HKEX filing says SiliconFlow's gross margin fell from 39.4% in 2024 to negative 24.0% in 2025 as cost of revenue rose faster than revenue. Medium SO011
CO041 The HKEX milestone section says SiliconFlow's overseas monthly revenue exceeded US$1 million in May 2026. Medium SO011
CO042 SiliconFlow said it started R&D on a large-model inference engine in August 2023. High SO017, SO011
CO043 SiliconFlow launched public-cloud MaaS in May 2024. High SO011, SO017
CO044 SiliconFlow launched DeepSeek inference services on Huawei Ascend in February 2025, which it described as an industry-first ultra-large-scale domestic-chip token-production service. High SO011, SO017
CO045 SiliconFlow launched private MaaS in September 2025 for customers with their own computing power and stricter data-compliance needs. High SO011, SO017
CO046 SiliconFlow launched the Elastic GPU heterogeneous-compute scheduling engine in April 2026. High SO011, SO017
CO047 The July 14, 2026 HKEX announcement added China Renaissance Securities (Hong Kong) as overall coordinator, signalling continued progress in the company's Hong Kong listing process. Medium SO012
CO048 KrASIA argued that the prospectus shows SiliconFlow is selling tokens at a loss in a compute-rental-heavy business exposed to pricing pressure. Medium SO018
CO049 Hello China Tech said SiliconFlow's public-cloud business carried a negative 119% gross margin in 2025 while on-premise deployment remained high-margin but harder to scale. Medium SO019
CO050 CB Insights lists SiliconFlow with $316.54 million total raised and a June 16, 2026 Series B of $295.84 million, a database presentation that does not fully reconcile to the HKEX B and B+ sequence. Low SO021
CO051 Tracxn lists Pan Yang and Jinhui Yuan as SiliconFlow's co-founders, which conflicts with the broader co-founder slate named in the HKEX filing. Low SO020
CO052 In July 2026, SiliconFlow announced Moonshot AI's Kimi K3 on its platform at $3 per million input tokens and $15 per million output tokens. High SO022, SO004
CO053 SiliconFlow's Kimi K3 API post says the serverless endpoint supports image input, tool calling, JSON Mode, streaming, and reasoning output. High SO023, SO006
CM001 SiliconFlow's relevant market sits in the inference middleware layer rather than in model training or end-user AI applications. High SM001, SM021
CM002 The included spend boundary covers serverless token APIs, dedicated instances, and private deployment of models through a unified interface. High SM001, SM021
CM003 Model training spend, semiconductor design, and end-user AI SaaS seats largely sit outside SiliconFlow's directly served market. Medium SM001, SM025
CM004 The main substitutes are hyperscaler model platforms, other independent inference APIs, direct model-lab APIs, and internal self-hosting. Medium SM001, SM008, SM011, SM016
CM005 The HKEX filing explicitly distinguishes independent ecosystem token supply platforms from closed ecosystems that bind users to proprietary compute and models. High SM001, SM024
CM006 SiliconFlow's own segmentation runs from developers and startups seeking cost-efficient on-demand access to large enterprises needing dedicated performance or supply assurance. High SM001, SM008
CM007 BYOK and preselected integrations reduce adoption friction by letting customers keep existing tools and workflows while switching token providers. High SM001, SM022
CM008 Multi-model inference platforms increasingly sell one-API access across text, image, video, code, and audio rather than single-modality access. High SM016, SM018, SM020, SM021
CM009 IDC says China enterprise MaaS token consumption rose from 114 trillion tokens in 2024 to 1,944 trillion tokens in 2025, roughly a 16x increase. Medium SM002
CM010 IDC projects China token consumption around 40,000 trillion in 2026, about 20x above 2025. High SM002, SM003
CM011 IDC sizes China public-cloud MaaS revenue at RMB3.07 billion in 2025. Medium SM002
CM012 IDC projects China public-cloud MaaS revenue reaching RMB18.6 billion in 2026. Medium SM002
CM013 Frost & Sullivan, as cited in the HKEX filing, says China's token supply market grew 1,602.6% from 2024 to 2025. Medium SM001
CM014 The same filing projects China's token supply market to reach approximately 53.2 quintillion tokens by 2030, implying a 638.3% CAGR from 2025 to 2030. Medium SM001
CM015 SiliconFlow held about 1.5% of China token-supply throughput in 2025, ranking fourth overall and first among independent ecosystem platforms. High SM001, SM024
CM016 MarketsandMarkets values the global AI inference market at USD106.15 billion in 2025 and USD254.98 billion in 2030, a 19.2% CAGR. Medium SM005
CM017 Grand View Research estimates the global AI inference market at USD97.24 billion in 2024 and USD253.75 billion in 2030, a 17.5% CAGR. Medium SM006
CM018 Fortune Business Insights sizes the global AI inference market at USD103.73 billion in 2025 and USD312.64 billion by 2034. Medium SM007
CM019 Independent analyst pages cluster the broad global AI inference market around roughly USD100 billion in the 2024-2025 base period. High SM005, SM006, SM007
CM020 Third-party global estimates agree that Asia Pacific is among the fastest-growing regions, even when they disagree on precise base-year values. Medium SM005, SM006, SM007
CM021 IDC's China MaaS lens and Frost & Sullivan's China token-supply lens agree on hypergrowth but do not use identical market definitions or baselines. Medium SM001, SM002
CM022 Buyer segments span individual developers, startups, enterprise platform teams, and regulated large organizations. High SM001, SM008, SM015
CM023 Dedicated performance, stable supply, latency, and private-environment deployment become important as buyers move from experiments to production-grade enterprise use. High SM001, SM011, SM015
CM024 AWS positions Bedrock as serving more than 100,000 organizations worldwide, from startups to global enterprises. Medium SM008
CM025 Microsoft Foundry packages models, agents, governance, RBAC, networking, and policy into one management plane, pointing to enterprise platform owners as the core buyer. High SM011, SM012
CM026 Alibaba's Token Plan shows inference demand is also being budgeted as team-productivity seats and shared credit pools rather than only as raw API consumption. Medium SM015
CM027 Together and Fireworks both emphasize OpenAI-compatible APIs and low-friction serverless onboarding, showing that portability is now a core expectation in this market. High SM016, SM018, SM020
CM028 SiliconFlow distributes through direct API integration, BYOK, and preselected tools such as LangChain, TRAE, Dify, and Cherry Studio. Medium SM001
CM029 Stanford's 2026 AI Index says frontier-model capability kept accelerating in 2025 and organizational AI adoption reached 88%. Medium SM004
CM030 Stanford also reports the U.S.-China frontier-model performance gap effectively closed by early 2026. Medium SM004
CM031 IDC says competition in China MaaS is shifting from pure price competition toward combined price, performance, and toolchain support. High SM002, SM003
CM032 IDC ranks performance, security and compliance, answer quality, platform availability, and cost effectiveness among the top enterprise selection factors, with cost only fifth for now. Medium SM002
CM033 Major platform vendors increasingly market routing, evaluation, observability, privacy, governance, and dedicated throughput as core product features, not extras. High SM008, SM011, SM012, SM020
CM034 AWS Bedrock and Microsoft Foundry both frame enterprise security, governance, and cost optimization as central to adoption. High SM008, SM011
CM035 Alibaba, Together, and Fireworks all offer packaging beyond simple pay-as-you-go, including seat subscriptions, dedicated endpoints, cached-token discounts, batch discounts, or hourly GPU deployment pricing. High SM015, SM017, SM019
CM036 The HKEX filing says SiliconFlow intentionally prioritized market share, user acquisition, and ecosystem building over immediate profitability in its public-cloud business. High SM001, SM024, SM025
CM037 Hello China Tech describes a severe 2023-2026 price war in model APIs, with mainstream prices down more than 90% since 2023 and further 2026 cuts from major vendors. Medium SM025
CM038 Compute rental, heterogeneous chip support, and supplier access remain structural constraints that matter for both service reliability and margins. Medium SM001, SM003, SM025
CM039 Global AI inference demand is broadening beyond text chat into edge, real-time, multimodal, and automation-heavy workloads. Medium SM005, SM007, SM018
CM040 Fortune Business Insights explicitly lists high hardware costs and integration challenges as adoption restraints for the AI inference market. Medium SM007
CM041 No retained public source isolates SiliconFlow's serviceable share by geography, customer segment, or category-specific retention. Low SM001, SM024, SM025
CM042 The most defensible market thesis is a stacked one: use global inference infrastructure for context, China MaaS growth for monetizable demand, and SiliconFlow throughput share only as a competitive signal. Medium SM001, SM002, SM005, SM006, SM007
CM043 The layered market-sizing pyramid is directional only because it mixes revenue, token-throughput, and market-share units rather than one additive denominator. Medium SM001, SM002, SM005
CM044 A common adoption path in this market is direct API experimentation first, then broader tool integration, then dedicated or private deployment once scale or governance requirements rise. Medium SM001, SM017, SM020
CP001 SiliconFlow competes across independent open inference platforms, incumbent cloud platforms, direct model-lab APIs, and internal build substitutes. High SP001, SP004, SP006, SP009, SP012, SP016, SP021
CP002 SiliconFlow ranks fourth in China token-supply throughput and first among independent ecosystem platforms in 2025. High SP001, SP023
CP003 Together AI positions itself around serverless and dedicated access to open models rather than a proprietary closed model stack. Medium SP012, SP013, SP014, SP015
CP004 Fireworks positions itself as a specialized training-and-inference platform for open models and says it processes 40T+ tokens per day. Medium SP016, SP017, SP018, SP020
CP005 OpenRouter positions itself as a unified API and broker that routes requests across hundreds of models and providers. Medium SP021, SP022
CP006 AWS Bedrock, Microsoft Foundry, and Alibaba Model Studio are incumbent substitutes with broader governance and control-plane depth than most independent peers. High SP004, SP006, SP007, SP008, SP009
CP007 Alibaba Model Studio combines official Qwen ownership with third-party model access and OpenAI-compatible APIs. High SP009, SP010, SP011
CP008 Retained sources here disclose far more about product scope and pricing than about current funding or revenue for Together, Fireworks, or OpenRouter. Low SP012, SP016, SP021
CP009 Internal build and direct model-lab APIs remain practical substitutes because many rivals expose portable, OpenAI-like integration paths. Medium SP013, SP020, SP021, SP025
CP010 OpenAI-compatible or drop-in API paths are common across SiliconFlow, OpenRouter, Fireworks, Together, and Alibaba. High SP009, SP015, SP020, SP021, SP025
CP011 Broad model-catalog competition is intense: SiliconFlow advertises 200+ models, Bedrock 100+, OpenRouter hundreds, and Foundry 1,900+ to 11,000+ access surfaces. Medium SP004, SP007, SP021, SP025
CP012 Together and Fireworks both offer a serverless-to-dedicated progression, while hyperscalers offer reserved or managed-compute paths for heavier production use. Medium SP005, SP008, SP013, SP014, SP018
CP013 Enterprise incumbents differentiate with guardrails, RBAC, policies, network isolation, observability, and managed-agent features. High SP004, SP006, SP007, SP008
CP014 OpenRouter differentiates on provider routing, fallbacks, price/throughput/latency sorting, and data-retention-aware controls. Medium SP021, SP022
CP015 Together differentiates on open-model access, a shared serverless API, and dedicated GPU deployments that reuse the same inference API. Medium SP012, SP013, SP014, SP015
CP016 Fireworks differentiates on serverless paths, prompt caching, on-demand infrastructure, and a combined training/inference posture. Medium SP016, SP018, SP019, SP020
CP017 SiliconFlow differentiates on heterogeneous compute, domestic-chip adaptation, and neutral multi-model positioning inside China. Medium SP001, SP003, SP025
CP018 Together lists Qwen 3.7 Max at $1.25 input and $3.75 output per 1M tokens, and DeepSeek V4 Pro at $1.74 input and $3.48 output. Medium SP015
CP019 Fireworks lists DeepSeek V4 Pro at $1.74 input / $0.145 cached input / $3.48 output and GPT OSS 20B at $0.07 input / $0.035 cached / $0.30 output. Medium SP019
CP020 AWS Bedrock pricing spans premium models such as Claude Opus 4.8 at $6 input / $30 output and lower-cost DeepSeek variants around $0.62 / $1.85 in listed regions. Medium SP005
CP021 Alibaba posts flagship Qwen list pricing with temporary regional discounts and also sells seat-based credit plans starting at $30 per seat per month. High SP010, SP011
CP022 OpenRouter makes pricing logic part of the product by letting customers sort providers by price, latency, or throughput and set max_price or performance thresholds. Medium SP022
CP023 Together and Fireworks both discount cached tokens and batch workloads, showing that heavy users are expected to demand lower effective pricing at scale. Medium SP013, SP019
CP024 API-surface switching costs are low because many rivals support OpenAI-compatible integration paths or drop-in migration. High SP006, SP009, SP020, SP021, SP025
CP025 OpenRouter is optimized for multi-homing rather than single-provider lock-in. Medium SP021, SP022
CP026 SiliconFlow's BYOK and preselected-tool distribution help adoption but also make customer multi-homing plausible. Medium SP001, SP025
CP027 Hyperscalers counter low API switching costs with identity, policy, networking, procurement, and broader platform integration. High SP004, SP006, SP007, SP009
CP028 AWS, Microsoft, and Alibaba can cross-sell inference from much broader cloud estates than independent startups can. Medium SP004, SP006, SP009
CP029 SiliconFlow's strongest retained moat candidate is China-local neutrality plus heterogeneous chip and deployment coverage. Medium SP001, SP003, SP025
CP030 Together and Fireworks compete more on speed, deployment control, and open-model operations than on exclusive model ownership. Medium SP013, SP014, SP016, SP018
CP031 OpenRouter's moat is routing intelligence and provider liquidity, but the model is inherently less lock-in-oriented than compute-owning platforms. Medium SP021, SP022
CP032 The top three China token suppliers above SiliconFlow are hyperscaler divisions, leaving SiliconFlow meaningfully smaller than the largest incumbents. High SP001, SP024
CP033 Hello China Tech argues mainstream model-API prices have fallen more than 90% since 2023, highlighting commoditization pressure. Medium SP024
CP034 SiliconFlow's filing shows that public-cloud market-share growth can coexist with negative gross margins and high compute-rental pressure. High SP001, SP023, SP024
CP035 Competitive risk is highest where buyers treat model access as commodity and can switch on price, latency, or uptime. Medium SP015, SP019, SP022
CP036 Competitive risk is lower where buyers need China-specific compute options, private deployment, or neutral multi-model support outside a single cloud. Medium SP001, SP009, SP025
CP037 No retained source proves customer lock-in for SiliconFlow comparable to hyperscaler IAM or network lock-in. Low SP004, SP006, SP009, SP025
CP038 No retained source here proves peer funding superiority or current margin superiority for SiliconFlow versus Together, Fireworks, or OpenRouter. Low
CP039 The competitive field is crowded partly because the same buyer job can be solved by a cloud platform, a neutral platform, a router, or internal build. Medium SP001, SP004, SP006, SP021
CP040 The best positioning map for this market places independent platforms high on openness and hyperscalers high on enterprise control and procurement strength. Medium SP004, SP006, SP009, SP012, SP016, SP021, SP025
CP041 Capability overlap is highest on basic API access and model breadth, while divergence is greatest on routing, governance, and deployment-control depth. Medium SP006, SP009, SP014, SP018, SP022, SP025
CP042 Moat readiness in this market depends more on supply access, governance, and operating efficiency than on simple model-count marketing. Medium SP001, SP002, SP004, SP006, SP024
CP043 Fireworks' vendor-authored 40T+ tokens/day claim indicates that some independent peers already operate at very large throughput scale. Medium SP016
CP044 A common adoption path is prototype on serverless or a broker, then shift toward dedicated capacity or deeper cloud control once scale and governance requirements rise. Medium SP005, SP013, SP014, SP018
CP045 SiliconFlow's competitive verdict is positive on relevance but still unproven on long-term durability versus hyperscaler bundling and category-wide price compression. Medium SP001, SP024, SP025
CP046 Because basic API access and model breadth overlap heavily, competitive selection often shifts to surrounding control surfaces, routing logic, and deployment guarantees. Medium SP006, SP009, SP015, SP020, SP025
CI001 SiliconFlow has two primary revenue lines: public cloud-based services and on-premise deployment solutions. Medium SI001, SI004
CI002 Public cloud-based services include serverless token services and dedicated instances. Medium SI001
CI003 On-premise deployment solutions install inference software in customer environments and currently carry much higher gross margins than public cloud. High SI001, SI004
CI004 In 2025 public cloud generated RMB29.261 million, or 52.9% of total revenue, while on-premise generated RMB26.069 million, or 47.1%. High SI001, SI004
CI005 SiliconFlow's public pricing surface is usage-based and model-level rather than seat-based. Medium SI002
CI006 List pricing in the inference market is shaped by cached-token discounts, batch discounts, and migrations toward dedicated capacity at scale. High SI006, SI007, SI010, SI012, SI014
CI007 Together explicitly frames serverless as cheaper for low or bursty traffic and dedicated replicas as cheaper when utilization stays high. Medium SI006
CI008 AWS, Fireworks, and Together all offer roughly 50% batch discounts in at least part of their pricing stack. High SI007, SI012, SI014, SI015
CI009 Alibaba monetizes the category through both pay-as-you-go model calls and seat-based subscription credits. High SI019, SI020, SI021
CI010 Official list pricing is useful for market context but is insufficient to infer SiliconFlow's realized pricing or margin. Medium SI002, SI019, SI024
CI011 SiliconFlow reported RMB55.33 million of revenue in 2025, up from RMB7.346 million in 2024. High SI001, SI004, SI005
CI012 SiliconFlow's 2025 blended gross margin was -24.0%. High SI001, SI004
CI013 The 2025 public-cloud gross loss margin was -119.0%, after -271.6% in 2024. High SI001, SI004
CI014 The 2025 on-premise deployment gross margin was 82.5%. Medium SI001
CI015 Compute rental fees accounted for 86.9% of 2025 cost of sales. High SI001, SI005
CI016 SiliconFlow explicitly prioritized market share, user acquisition, and ecosystem building over immediate profitability in public cloud. High SI001, SI004
CI017 Serverless paying accounts increased from 2,455 to 716,000 in 2025. Medium SI004, SI005
CI018 Using year-end paying-account count as a crude divisor, RMB14.3 million of serverless revenue implies very low annual spend density per paying account. Medium SI004
CI019 More than 64% of 2025 sales and marketing expense went to promotional compute credits. Medium SI004
CI020 Independent analyses argue that SiliconFlow is operating as a compute-renting middle layer inside a price-war environment. Medium SI004, SI005
CI021 IDC says cost effectiveness matters to buyers, but it trails performance, compliance, answer quality, and platform usability as a current selection factor. Medium SI022
CI022 Usage metrics such as registered users and paying-account growth prove demand but do not by themselves prove revenue quality. Medium SI001, SI004, SI005
CI023 At 2025 year-end SiliconFlow held RMB171 million of cash and cash equivalents plus RMB100 million of time deposits. High SI001, SI005
CI024 Net cash used in operating activities was RMB172 million in 2025. High SI001, SI005
CI025 On a simple backward-looking lens, SiliconFlow was not self-funding before the 2026 financings. Medium SI001, SI023, SI005
CI026 The filing discloses about RMB1.48 billion of post-2025 cash consideration across the A+, B, and B+ rounds. Medium SI001
CI027 SiliconFlow does not directly purchase chips; it mainly leases computing resources through partners. High SI005, SI001
CI028 Major purchases are concentrated among the top five suppliers, tying capital adequacy to supplier terms as well as to cash balances. High SI001, SI005
CI029 The retained public record does not support a precise current runway calculation as of 2026-07-22. Low SI001, SI005
CI030 The 2026 financings improved capital adequacy, but they did not by themselves solve revenue-quality or margin-path questions. Medium SI001, SI004, SI005
CI031 The public record does not reveal net revenue retention, gross retention, or customer concentration. Low SI001, SI002
CI032 The public record does not reveal realized effective pricing by model family or customer cohort. Low SI002, SI019, SI024
CI033 The public record does not reveal a current monthly cash-burn figure or a management-backed runway target as of run date. Low SI001, SI005
CI034 On-premise revenue appears higher quality on margin, but public sources are insufficient to show whether it can scale enough to change the whole-company profile. Medium SI001, SI004
CI035 Public-cloud token revenue remains the core growth engine but currently appears subsidy-heavy and compute-rental-heavy. High SI001, SI004, SI005
CI036 List pricing across the category increasingly pushes heavy users toward dedicated or reserved capacity once workloads stabilize. Medium SI006, SI013, SI014, SI018, SI024
CI037 Prompt caching is economically important because it can lower effective token costs or reduce compute wasted on repeated context. High SI008, SI010, SI016, SI025
CI038 Fine-tuning and dedicated hosting create additional monetization paths in the category, but retained sources do not show SiliconFlow currently monetizing those paths separately. Medium SI009, SI013, SI018
CI039 The clearest financial bridge is from usage into three revenue-quality tiers—serverless, dedicated, and on-premise—rather than into one uniform SaaS line. Medium SI001, SI002, SI004
CI040 The divergence between demand growth and financial quality is the central financial thesis of SiliconFlow as of 2026-07-22. Medium SI001, SI004, SI005
CE001 SiliconFlow is an API-first inference platform rather than a single-model application. High SE002, SE008
CE002 Its public product menu spans ready-to-use model APIs, reserved instances, inference acceleration services, and private deployment. High SE001, SE021, SE008
CE003 Serverless APIs are the core developer-facing surface. High SE001, SE002, SE004
CE004 Private deployment and BYOC are presented as enterprise options for privacy-sensitive workloads. Medium SE001, SE006
CE005 Users follow a simple path of model selection, API key creation, and endpoint integration. High SE002, SE004
CE006 The catalog spans text, speech, image, video, vector, reranking, and multimodal model classes. High SE002, SE013
CE007 Reserved instances are positioned for enterprise core inference scenarios needing dedicated capacity and cost optimization. High SE001, SE014
CE008 Inference acceleration is marketed both for open-source models and for self-developed models. High SE001, SE002
CE009 The product bundle reduces the need for customers to stitch together separate API, deployment, and optimization vendors. Medium SE001, SE002, SE019
CE010 SiliconFlow exposes model choice, streaming, context-window controls, tool calling, and request tracing through its chat API surface. Medium SE004
CE011 The platform uses a largely OpenAI-style integration pattern, lowering adoption friction for developers already familiar with that schema. Medium SE004, SE018
CE012 The docs show SiliconFlow regularly updates model availability and service capabilities over time. High SE003, SE004
CE013 Developers can stay on serverless or graduate toward reserved instances and private deployments as workloads mature. Medium SE001, SE014, SE016
CE014 SiliconFlow claims self-developed efficient operators, optimization frameworks, and a leading inference acceleration engine. High SE001, SE002
CE015 OneDiff is an acceleration library for diffusion models with optimized GPU kernels and compiler tooling. Medium SE010
CE016 OneDiff release history shows ongoing engineering maintenance rather than a one-off repository dump. Medium SE011
CE017 The API docs expose x-siliconcloud-trace-id response headers for request tracing and troubleshooting. Medium SE004
CE018 The strongest external engineering proof currently available is around inference-optimization tooling, not independently audited production reliability. Medium SE010, SE011, SE004
CE019 Public sources do not provide audited SLOs, public error-rate dashboards, or third-party performance benchmarks for the whole platform. Low SE001, SE002, SE004
CE020 SiliconFlow differentiates through a combination of catalog breadth, integration simplicity, deployment flexibility, and optimization claims. High SE001, SE002, SE005, SE013
CE021 The company continues to add current model launches, as shown by public launch messaging around recent models such as GLM-5.2 and Kimi K3. Medium SE001, SE012
CE022 The product moat is operational and integration-driven rather than based on ownership of a proprietary frontier model. Medium SE008, SE023, SE024
CE023 The platform depends on upstream model-provider availability and on continued access to external compute capacity. Medium SE003, SE008, SE023
CE024 The public product roadmap appears to be driven heavily by fast model onboarding and packaging cadence. Medium SE003, SE012, SE013
CE025 Private deployment increases appeal to regulated buyers but also increases implementation complexity and service-delivery burden. Medium SE001, SE006, SE014
CE026 The Kimi K3 launch post confirms a live developer surface with a standard API base URL. Medium SE012
CE027 Public maturity is strongest for the serverless API surface and weaker for enterprise assurance artifacts. Medium SE001, SE002, SE006, SE007
CE028 OneDiff release cadence is a useful roadmap proxy for SiliconFlow's optimization work, but it is not a substitute for a full platform changelog. Medium SE011, SE010
CE029 SiliconFlow publicly claims BYOC deployment and compute/network/storage isolation. Medium SE001
CE030 SiliconFlow maintains distinct China and international service surfaces with separate contractual language. High SE007, SE021
CE031 The privacy policy confirms the international surface is operated through siliconflow.com. Medium SE006
CE032 Trace identifiers in API responses are a real supportability control, though not a substitute for published reliability reporting. Medium SE004
CE033 The retained source set does not show a public SOC 2 report, ISO certificate list, or detailed public incident history. Low SE001, SE006, SE007
CE034 Large enterprise underwriting still requires third-party assurance artifacts that are not visible in the retained public set. Low SE006, SE007, SE001
CE035 The company's security and compliance messaging is plausible but currently stronger in claims than in public evidence depth. Medium SE001, SE006, SE007
CE036 SiliconFlow is built for supportability and enterprise packaging, but public evidence is insufficient to conclude best-in-class enterprise readiness. Medium SE001, SE004, SE006, SE007
CU001 SiliconFlow serves multiple customer types rather than a single homogeneous user base. High SU001, SU018, SU019
CU002 Public 2026 coverage cites more than 10 million users and more than 10,000 enterprise customers. Medium SU002, SU023
CU003 The company explicitly serves both developers and enterprises. High SU001, SU019
CU004 Enterprise demand includes reserved instances, private deployment, and infrastructure-style deployments rather than only lightweight API experimentation. High SU005, SU006, SU008
CU005 Community and developer-tool integrations prove API onboarding relevance across coding, search/RAG, translation, and agent workflows. High SU004, SU009, SU010, SU011
CU006 Telecom / compute partners are a distinct strategic customer surface for SiliconFlow. High SU007, SU003
CU007 The filing and public coverage imply that self-serve paying accounts and larger enterprise deployments coexist inside the customer base. Medium SU001, SU015
CU008 Broad user reach does not by itself prove high-quality revenue because developer and community adoption can be monetically light. Medium SU002, SU004, SU017
CU009 SiliconFlow appears especially strong in infra-heavy buyer scenarios where inference performance, private deployment, or国产化 adaptation matters. Medium SU005, SU006, SU007, SU018
CU010 Guizhou Mobile is the clearest named institutional counterparty in the retained customer corpus. High SU007, SU003
CU011 The Guizhou Mobile partnership covers inference framework deployment, compute coordination, token services, and joint operational systems. Medium SU007
CU012 The Guizhou Mobile relationship is described as a deepening of earlier 2025 cooperation rather than a first-touch pilot. Medium SU007
CU013 The reserved-instance case study shows at least one customer whose coding-agent workload reached 100 billion daily tokens in 2026. Medium SU008
CU014 MindSearch is a named public project with a documented SiliconFlow integration path. High SU009, SU004
CU015 Continue is a named public integration that positions SiliconFlow for coding-assistant workflows inside IDEs. High SU010, SU004
CU016 Cline is a named public integration that uses SiliconFlow through an OpenAI-compatible API pattern. High SU011, SU004
CU017 The named developer-tool proofs are useful evidence of adoption breadth but weak evidence of contract value or exclusivity. Medium SU009, SU010, SU011
CU018 Anonymous enterprise case studies add operational depth but limit concentration and logo-quality analysis because the customer names are withheld. High SU005, SU006, SU008
CU019 Named customer proof in this chapter is therefore stronger on breadth than on directly observable commercial weight. Medium SU007, SU009, SU010, SU011
CU020 SiliconFlow shows a land-and-expand pattern from pay-as-you-go usage toward reserved or deeper infrastructure deployment for some accounts. Medium SU008, SU025, SU001
CU021 The retained public set does not disclose NRR, GRR, or logo-retention rates. Medium SU001, SU015
CU022 The retained public set does not disclose average contract length, renewal rates, or top-customer concentration. Medium SU001, SU015, SU017
CU023 Public evidence for repeat usage is architectural and behavioral rather than cohort-based: heavy workloads, reserved-instance upgrades, and deeper partner cooperation. Medium SU007, SU008, SU025
CU024 Private deployment and国产化 case studies suggest deeper embedment than simple API testing, but the public record still lacks renewal proof. Medium SU005, SU006, SU021
CU025 Revenue concentration could still be high even with 10,000+ enterprise customers if a small number of token-heavy accounts dominate spend. Medium SU002, SU008, SU017
CU026 Strategic telecom or infrastructure partners may become concentrated routes to growth, creating bargaining-power and dependency risk. Medium SU007, SU021, SU017
CU027 Broad community integration can inflate awareness and usage without proving durable paid conversion. Medium SU004, SU009, SU010, SU011
CU028 The customer corpus supports diversification by use case, but not yet diversification by revenue contribution. Medium SU004, SU005, SU006, SU007
CU029 SiliconFlow has enough public evidence to show real adoption, not merely claimed logos. High SU002, SU007, SU009, SU010, SU011
CU030 Customer durability is materially less proven than customer breadth. Medium SU021, SU022, SU015
CU031 The strongest public customer signal is that some workloads are mission-critical enough to justify dedicated or private infrastructure decisions. High SU005, SU006, SU008
CU032 The cleanest public customer journey is discover the API, prove value in workflow, then deepen into reserved or private deployment as usage scales. Medium SU019, SU008, SU025
CU033 For diligence, the biggest remaining customer question is not acquisition but quality of retained spend. Medium SU002, SU015, SU017
CU034 A retention cohort figure can be shown only as a visibility proxy, not as a real renewal chart. Medium SU001, SU015
CU035 Developer adoption and enterprise adoption reinforce each other strategically, but they should not be valued as equivalent commercial proof. Medium SU004, SU007, SU010, SU011
CU036 The right customer underwriting request is a segmented cohort pack covering active customers, spend concentration, renewal, and upgrade from API to reserved/private deployment. Medium SU001, SU015, SU017
CR001 SiliconFlow's domestic terms expressly limit the service to developer-oriented content-generation scenarios and exclude automatic control, medical information, psychological counseling, and critical information infrastructure uses. Medium SR002
CR002 The domestic privacy policy says mainland-law requirements can trigger real-name verification and identity collection for service activation. Medium SR003
CR003 SiliconFlow's domestic terms explicitly reference the AI-generated-content labeling measures and prohibit malicious deletion or tampering with labels. Medium SR002
CR004 The domestic site publicly displays ICP and telecom-license identifiers, indicating regulated internet-service operations rather than a purely offshore surface. Medium SR004
CR005 SiliconFlow maintains separate China and international legal surfaces, reducing but not eliminating cross-surface governance complexity. High SR005, SR006
CR006 China's 2025 AI labeling measures create mandatory explicit and implicit labeling obligations for generated synthetic content. Medium SR007, SR032
CR007 Those labeling rules create operational risk because providers and platforms must implement metadata, logging, and downstream labeling workflows rather than merely publish a policy. Medium SR002, SR007, SR031
CR008 The public terms show compliance awareness, but they do not by themselves prove SiliconFlow's implementation quality for labeling, logging, or auditability. Medium SR002, SR007
CR009 SiliconFlow's domestic privacy policy says user information generated in operating the domestic site is stored within the PRC. Medium SR003
CR010 Mainland onboarding and identity checks can increase friction for some customers and enlarge the company's privacy-handling obligations. High SR002, SR003
CR011 The legal risk is a continuous compliance burden across content, privacy, telecom, and sector-specific use restrictions rather than a single one-time license hurdle. High SR002, SR003, SR004, SR007
CR012 SiliconFlow serves infrastructure-like workloads where outages, latency swings, or deployment failures can matter more than consumer feature gaps. High SR015, SR016, SR017
CR013 Public sources do not furnish a robust public uptime history, incident archive, or detailed SLO record. Low SR015, SR018
CR014 Public sources do not furnish SOC 2, ISO certification details, or equivalent third-party assurance artifacts in the retained corpus. Low SR004, SR005, SR006
CR015 The retained corpus suggests enterprise-grade operational aspirations, but not enough third-party evidence to conclude best-in-class operational maturity. Medium SR015, SR016, SR018
CR016 Privacy, labeling, and identity controls interact as a compound execution risk because each adds workflow and data-handling obligations to the platform. High SR002, SR003, SR007
CR017 The platform's own API reference warns that model availability and capability can change over time. Medium SR020
CR018 Rapid model churn creates support and migration risk for enterprise workloads even when the platform benefits from a broad model catalog. Medium SR019, SR020
CR019 Monitoring, fault tolerance, and traceability are publicly claimed mitigations, but retained sources do not independently validate their effectiveness at scale. Medium SR018, SR015
CR020 Buyer expectations in critical or regulated environments are rising toward explicit trustworthy-AI risk-management practices, even where frameworks such as NIST AI RMF are voluntary. High SR010, SR021
CR021 Operational-quality risk is therefore most acute where high-volume or regulated workloads demand both performance and documented assurance. Medium SR015, SR016, SR020
CR022 SiliconFlow depends on upstream model providers for catalog breadth and competitive relevance. High SR019, SR020, SR029
CR023 SiliconFlow depends heavily on leased compute and supplier relationships rather than on fully owned semiconductor supply. High SR001, SR011, SR012
CR024 U.S. export controls continue to target PRC access to advanced computing semiconductors, creating a real external dependency risk for Chinese AI-infrastructure players. High SR008, SR009
CR025 Domestic-chip adaptation and private-deployment narratives mitigate but do not eliminate supply-side risk. Medium SR016, SR017, SR024
CR026 Strategic telecom or compute partnerships can broaden distribution but also create concentration and execution dependencies. Medium SR014, SR023
CR027 Supplier / capacity planning is an execution-critical function because service quality and margin both depend on it. High SR001, SR011, SR015
CR028 Private deployment, reserved instances, and domestic-chip optimization increase implementation complexity and require strong enterprise delivery operations. High SR015, SR016, SR017
CR029 Independent analyses already frame SiliconFlow's economics as vulnerable to price war and compute-rental pressure. Medium SR011, SR012, SR013
CR030 Competitive pressure from hyperscalers and other inference platforms can intensify the margin impact of compliance and compute-supply costs. High SR025, SR026, SR027, SR028, SR029, SR030
CR031 Hidden concentration of token-heavy customers remains a real model risk because public breadth metrics do not reveal spend distribution. Medium SR001, SR013, SR011
CR032 SiliconFlow already has visible mitigations: separate legal surfaces, BYOC/private deployment, domestic adaptation, and strategic partner operations. Medium SR005, SR014, SR016, SR017
CR033 The public record does not yet prove whether SiliconFlow has sufficient second-line compliance, delivery, and supplier-management depth to scale safely. Low SR011, SR012, SR014
CR034 A formal regulatory event around labeling, privacy, or telecom compliance would be an immediate thesis-break trigger. Medium SR002, SR003, SR007
CR035 A material compute-capacity shock or supplier loss would directly threaten both reliability and gross margin. Medium SR001, SR008, SR023
CR036 Proof that enterprise retention or revenue concentration is worse than implied by public customer counts would materially weaken the investment case. Medium SR001, SR011, SR013
CR037 Failure of public-cloud margins to improve despite scale would indicate that growth is not translating into a durable business model. Medium SR001, SR011, SR029
CR038 The strongest risk transmission path is from regulation or supply shock into operating complexity, then into revenue quality, margin, financing need, and valuation. Medium SR007, SR023, SR029
CR039 The most likely early diligence flashpoint is weak evidence around enterprise assurance, retention, and supplier concentration rather than a near-term demand collapse. Medium SR011, SR012, SR015
CR040 SiliconFlow's risk profile is investable only if the company can keep converting scale into better margins faster than regulatory and supply complexity increase. Medium SR001, SR011, SR024
CV001 On public evidence alone, the right current recommendation is Track rather than Buy. Medium SV001, SV004, SV005, SV006, SV009
CV002 The risk rating for SiliconFlow at the current price is high. Medium SV001, SV025, SV026
CV003 Confidence in the recommendation is only medium-low because the price is public but the core underwriting metrics remain incomplete. Medium SV001, SV004, SV005
CV004 SiliconFlow is strategically relevant enough to merit continued diligence and tracking. High SV014, SV015, SV021, SV022
CV005 The company has real customer and usage proof rather than a purely narrative AI story. High SV002, SV023, SV024
CV006 Recent financing and strategic investors improve the company's optionality and time to execute. Medium SV002, SV003, SV022
CV007 Current public evidence is insufficient to justify a strong Buy at the disclosed June 2026 price. Medium SV001, SV004, SV005, SV025
CV008 The valuation stance is therefore price-sensitive rather than categorically negative. Medium SV002, SV006, SV009
CV009 A lower entry price or stronger proof package could move the recommendation upward. Medium SV001, SV004, SV006, SV009
CV010 The positive thesis begins with category positioning: SiliconFlow sits in a growing inference market with real product breadth and deployment flexibility. High SV014, SV015, SV021, SV028
CV011 The positive thesis also includes ecosystem and strategic-channel leverage, not only direct API usage. Medium SV002, SV003, SV022
CV012 The anti-thesis begins with disclosed weak economics, especially negative public-cloud gross margins. High SV001, SV004, SV005
CV013 Compute-rental dependence and supplier concentration weaken the claim that current scale already represents a mature infrastructure moat. High SV001, SV005, SV026
CV014 Compared with Western inference peers, SiliconFlow's disclosed revenue scale is much smaller. Medium SV001, SV006, SV009, SV012
CV015 SiliconFlow may still be strategically important even if it deserves a valuation discount to faster-scaling Western peers. Medium SV006, SV009, SV012, SV013
CV016 The current US$1.2B valuation is not obviously absurd in absolute terms, but it is not clearly supported by public underwriting data either. Medium SV002, SV004, SV005, SV006
CV017 The market-size thesis is real, but current evidence does not yet show that SiliconFlow has converted strategic relevance into high-quality economics. Medium SV001, SV014, SV028
CV018 Regulatory and supply-side risks deserve valuation weight because they can affect both growth and margin simultaneously. High SV025, SV026, SV029
CV019 The disclosed price already asks investors to underwrite future improvement rather than current financial quality. Medium SV001, SV002, SV004
CV020 The bull case requires better enterprise retention, richer dedicated/private-deployment mix, and meaningful margin improvement. Medium SV001, SV021, SV023
CV021 The base case assumes SiliconFlow remains strategically relevant and grows, but revenue quality improves only gradually. Medium SV001, SV004, SV028
CV022 The bear case assumes price competition, regulation, or supply constraints prevent margin inflection and weaken financing terms. Medium SV004, SV005, SV025, SV026
CV023 Together AI is a relevant direct comparable because it combines open-model access with inference infrastructure. High SV006, SV007
CV024 Together AI's 2026 $8.3B valuation comes with public evidence of >$1.15B annual bookings, a scale far above SiliconFlow's disclosed 2025 revenue base. Medium SV006
CV025 Fireworks AI is a relevant direct comparable because it is an inference-cloud platform rather than a generic hyperscaler. High SV009, SV010, SV030
CV026 Fireworks AI's 2026 $17.5B valuation is paired with public evidence of >$1B annualized revenue, making SiliconFlow look less obviously cheap despite its much lower headline price. High SV009, SV010
CV027 Baseten is a relevant inference-platform comparable because it combines developer entry with enterprise deployment options. Medium SV011, SV012
CV028 Baseten's 2026 ~$13B valuation and ~US$600M annualized revenue indicate that specialized inference platforms can justify high values—but with much stronger monetization evidence than SiliconFlow currently discloses. Medium SV012
CV029 CoreWeave is an infrastructure-adjacent valuation reference that shows both the upside and the concentration / capital-intensity risks of AI compute businesses. Medium SV013, SV026
CV030 Public comparables therefore support interest in the category, but not an automatic premium view on SiliconFlow at current evidence quality. Medium SV006, SV009, SV012, SV013
CV031 Retention, concentration, and supplier-risk visibility are the most important missing inputs preventing a stronger call. Medium SV001, SV004, SV005, SV026
CV032 A formal regulatory enforcement event would be a thesis-break trigger. Medium SV025, SV029
CV033 A material compute-supply shock or worsening supplier concentration would be a thesis-break trigger. Medium SV001, SV026
CV034 Failure of public-cloud economics to improve would be a thesis-break trigger. Medium SV001, SV004, SV005
CV035 Proof of weak enterprise retention or heavy whale concentration would be a thesis-break trigger. Medium SV001, SV004, SV023
CV036 The most important valuation diligence ask is current 2026 revenue run-rate and mix by product line. Medium SV001, SV004
CV037 The next most important asks are enterprise retention, concentration, and API-to-dedicated expansion metrics. Medium SV001, SV023, SV024
CV038 Security / assurance and supplier contingency are also valuation-critical because they affect enterprise quality and downside risk. Medium SV025, SV026, SV027
CV039 Price alone cannot rescue the investment case if the missing diligence items point to structural weakness rather than temporary opacity. Medium SV004, SV005, SV026
CV040 As of 2026-07-22, the best public-evidence call is Track: strategic relevance is clear, but upside from the current price is not. Medium SV001, SV002, SV006, SV009
Sources
IDPublisherTitleQuote
SO001 SiliconFlow 硅基流动 SiliconFlow - 致力于成为全球领先的 AI 能力提供商 地址:北京市海淀区中关村东路 1 号院 8 号楼 D 座 23 层 2301。
SO002 SiliconFlow SiliconFlow – AI Infrastructure for LLMs & Multimodal Models One Platform All Your AI Inference Needs.
SO003 SiliconFlow About Us - SiliconFlow | Global AI Infrastructure Provider We strive to become the world's most influential provider of AI infrastructure.
SO004 SiliconFlow SiliconFlow – AI Infrastructure for LLMs & Multimodal Models / Models One API to run inference on 200+ cutting-edge AI models, and deploy in seconds.
SO005 SiliconFlow Docs Quickstart - SiliconFlow Currently, the platform supports login via SMS, email, as well as OAuth login through GitHub and Google.
SO006 SiliconFlow Docs Chat - SiliconFlow API Reference We periodically update our models to enhance service quality. Changes may include model on/offlining or capability adjustments.
SO007 SiliconFlow Docs Product introduction - SiliconFlow SiliconFlow is committed to providing developers with faster, more comprehensive, and seamlessly integrated model APIs.
SO008 SILICONFLOW TECHNOLOGY PTE. LTD. Terms of Use - SiliconFlow The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn.
SO009 SILICONFLOW TECHNOLOGY PTE. LTD. Privacy Policy - SiliconFlow If you want to contact us... please contact us at contact@siliconflow.com.
SO010 GitHub SiliconFlow · GitHub SiliconFlow builds scalable, standardized, and high-performance AI infrastructure.
SO011 SiliconFlow / HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. As of April 30, 2026, our platform had over 10 million registered users.
SO012 HKEX Announcement on overall coordinator appointment for Beijing SiliconFlow Technology Co., Ltd. The Company has further appointed China Renaissance Securities (Hong Kong) Limited as its overall coordinator on July 14, 2026.
SO013 Caixin Global SiliconFlow Raises $294 Million as China's AI Inference Demand Surges Chinese AI model inference startup SiliconFlow has raised more than 2 billion yuan ($294 million) in a Series B funding round.
SO014 Sina Finance 硅基流动完成新一轮超 20 亿元融资 过去一年,公司在企业级市场实现爆发式增长。
SO015 Sina Finance 硅基流动完成新一轮超 20 亿元融资(转载) IDC预测,2026年中国市场的Token消费量将达到40,000万亿。
SO016 投资界 / InvestorsCN 硅基流动完成新一轮超20亿元融资,华兴资本担任独家财务顾问 近日,硅基流动已完成超20亿元B轮融资。
SO017 36Kr Europe Silicon Flow Completes New Round of Over 2-Billion-Yuan Financing By providing efficient MaaS through the Token factory model, the daily average Token call volume has reached trillions.
SO018 KrASIA Surging users, widening losses, and leased compute: Behind SiliconFlow's IPO filing But the prospectus also shows the cost of that growth. SiliconFlow is selling tokens at a loss.
SO019 Hello China Tech SiliconFlow IPO: What China's Token Boom Really Costs Its public cloud service ... generated 52.9% of 2025 revenue ... Its gross margin was negative 119%.
SO020 Tracxn SiliconFlow - 2026 Company Profile & Team - Tracxn SiliconFlow is a funded company based in Beijing (China), founded in 2023 by Pan YANG and Jinhui Yuan.
SO021 CB Insights SiliconFlow Stock Price, Funding, Valuation, Revenue & Financial Statements SiliconFlow has raised $316.54M over 8 rounds.
SO022 SiliconFlow Kimi K3 is now live on SiliconFlow SiliconFlow provides OpenAI- and Anthropic-compatible APIs for Kimi K3.
SO023 SiliconFlow Kimi K3 on SiliconFlow API Current support includes image input, tool calling, JSON Mode, streaming, and reasoning output.
SO024 SiliconFlow Cloud SiliconFlow 统一登录 / Welcome to SiliconFlow Blazing-fast, cost-effective Generative AI cloud services.
SO025 SiliconFlow 模型与产品页 - 硅基流动 支持 BYOC 部署,全面保护数据隐私与业务安全。
SM001 SiliconFlow / HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. In 2025, we were the fourth largest token supply platform in China in terms of annual token throughput, with a market share of 1.5% and ranked first among all independent ecosystem token supply platforms in China.
SM002 IDC China 中国MaaS市场进入高速增长期,Token经济从概念走向规模 2025年中国公有云MaaS市场的规模达到30.7亿元人民币。IDC预计2026年全年Token消耗量约为40,000万亿次。
SM003 36Kr Europe SiliconFlow completes funding round exceeding RMB 2 billion IDC predicts that the Token consumption in the Chinese market will reach 40,000 trillion in 2026, an increase of about 20 times compared to 2025.
SM004 Stanford HAI 2026 AI Index Report Industry produced over 90% of notable frontier models in 2025... Organizational adoption reached 88%.
SM005 MarketsandMarkets AI Inference Market Size, Share & Trends The AI inference market size was valued at USD 106.15 billion in 2025 and is projected to reach USD 254.98 billion by 2030, growing at a CAGR of 19.2%.
SM006 Grand View Research AI Inference Market Summary The global AI inference market size was estimated at USD 97.24 billion in 2024 and is projected to reach USD 253.75 billion by 2030, growing at a CAGR of 17.5%.
SM007 Fortune Business Insights AI Inference Market The global AI inference market size was valued at USD 103.73 billion in 2025 and is projected to grow from USD 117.80 billion in 2026 to USD 312.64 billion by 2034.
SM008 Amazon Web Services Amazon Bedrock Amazon Bedrock powers generative AI for more than 100,000 organizations worldwide—from startups to global enterprises across every industry.
SM009 Amazon Web Services Amazon Bedrock Pricing Amazon Bedrock offers select foundation models... for batch inference at a 50% lower price compared to on-demand inference pricing.
SM010 AWS Docs What is Amazon Bedrock? Bedrock supports 100+ foundation models from industry-leading providers, including Amazon, Anthropic, DeepSeek, Moonshot AI, MiniMax, and OpenAI.
SM011 Microsoft Learn What is Microsoft Foundry? Foundry unifies agents, models, and tools under a single management grouping with built-in enterprise-readiness capabilities including tracing, monitoring, evaluations, and customizable enterprise setup configurations.
SM012 Microsoft Azure Microsoft Foundry Pricing Foundry Models: Access more than 11,000 foundation, open, reasoning, multimodal and industry-specific models.
SM013 Microsoft Azure Foundry Models Pricing Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure... Choose from A100, H100, H200, and MI300 GPU families.
SM014 Alibaba Cloud Model Studio model pricing Model API calls are billed on a pay-as-you-go basis by default.
SM015 Alibaba Cloud Token Plan (Team Edition) overview Token Plan (Team Edition) is a monthly AI model subscription billed in Credits... and includes team management, data privacy, and dedicated throughput.
SM016 Together AI Serverless Inference Access all the top open-source models in one place.
SM017 Together AI Pricing Most teams start with serverless inference and move to dedicated endpoints at scale.
SM018 Together AI Docs Available models - Together AI docs Serverless models are the fastest way to run inference on Together... Pay only for the tokens you process.
SM019 Fireworks AI Pricing Additionally, please note that cached input tokens are by default priced at 50% for all text and vision language models... batch inference is priced at 50% of our serverless pricing.
SM020 Fireworks Docs Models & Inference Fireworks provides fast, cost-effective access to leading open-source text models through OpenAI-compatible APIs.
SM021 SiliconFlow Models - SiliconFlow One API to run inference on 200+ cutting-edge AI models, and deploy in seconds.
SM022 SiliconFlow Docs Quickstart - SiliconFlow Currently, the platform supports login via SMS, email, as well as OAuth login through GitHub and Google.
SM023 SiliconFlow Docs Chat Completions API Reference We periodically update our models to enhance service quality. Changes may include model on/offlining or capability adjustments.
SM024 KrASIA Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing According to third-party data cited in its prospectus, SiliconFlow was China's largest independent ecosystem token supplier by annual token throughput in 2025 and ranked among the top five token suppliers overall.
SM025 Hello China Tech SiliconFlow IPO: token economics As a middle-layer platform, SiliconFlow's main cost pressure comes from leasing computing power... In 2025, computing resource costs totaled RMB 59.627 million, accounting for 86.9% of cost of sales.
SP001 SiliconFlow / HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. In 2025, we were the fourth largest token supply platform in China in terms of annual token throughput, with a market share of 1.5% and ranked first among all independent ecosystem token supply platforms in China.
SP002 IDC China 中国MaaS市场进入高速增长期,Token经济从概念走向规模 IDC认为,MaaS厂商的竞争焦点正在从过去单纯的价格比拼,转向“价格、性能与工具链支持”的综合能力竞争。
SP003 36Kr Europe SiliconFlow completes funding round exceeding RMB 2 billion The latest "China AI Software Market Semi-Annual Tracker, 2025H2" released by IDC shows that SiliconFlow is the only startup company among the top four in the market share of China's public cloud MaaS.
SP004 Amazon Web Services Amazon Bedrock Amazon Bedrock powers generative AI for more than 100,000 organizations worldwide—from startups to global enterprises across every industry.
SP005 Amazon Web Services Amazon Bedrock Pricing Amazon Bedrock offers select foundation models ... for batch inference at a 50% lower price compared to on-demand inference pricing.
SP006 Microsoft Learn What is Microsoft Foundry? Foundry unifies agents, models, and tools under a single management grouping with built-in enterprise-readiness capabilities including tracing, monitoring, evaluations, and customizable enterprise setup configurations.
SP007 Microsoft Azure Microsoft Foundry Pricing Foundry Models: Access more than 11,000 foundation, open, reasoning, multimodal and industry-specific models.
SP008 Microsoft Azure Foundry Models Pricing Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure to deploy and run open-source and custom AI models on enterprise-grade GPUs.
SP009 Alibaba Cloud What is Model Studio? Alibaba Cloud Model Studio is a one-stop model service platform. It provides the full Qwen series and mainstream third-party LLMs through official Qwen APIs and OpenAI-compatible APIs.
SP010 Alibaba Cloud Model Studio model pricing Model API calls are billed on a pay-as-you-go basis by default.
SP011 Alibaba Cloud Token Plan (Team Edition) overview Token Plan (Team Edition) is a monthly AI model subscription billed in Credits ... and includes team management, data privacy, and dedicated throughput.
SP012 Together AI Pricing Most teams start with serverless inference and move to dedicated endpoints at scale.
SP013 Together AI Docs Serverless overview Run any supported model through a shared, per-token API with no provisioning and no minimums.
SP014 Together AI Docs Dedicated model inference overview Dedicated model inference bills per minute by hardware while a deployment runs, regardless of model or request volume.
SP015 Together AI Docs Available models - Together AI docs Serverless models are the fastest way to run inference on Together. You call any supported model through a shared per-token API, with no provisioning, no replicas to size, and no minimum cost.
SP016 Fireworks AI Fireworks home Fireworks processes 40T+ tokens per day.
SP017 Fireworks AI Pricing Pricing to seamlessly scale from idea to enterprise.
SP018 Fireworks Docs Serverless overview Serverless is multi-tenant inference for popular open models running on Fireworks-managed infrastructure.
SP019 Fireworks Docs Serverless pricing Per-token serverless pricing for text, vision, and embedding models, including Priority and Fast serving paths.
SP020 Fireworks Docs Models & Inference Fireworks provides fast, cost-effective access to leading open-source text models through OpenAI-compatible APIs.
SP021 OpenRouter Docs Quickstart OpenRouter provides a unified API that gives you access to hundreds of AI models through a single endpoint, while automatically handling fallbacks and selecting the most cost-effective options.
SP022 OpenRouter Docs Provider routing OpenRouter routes requests to the best available providers for your model. By default, requests are load balanced across the top providers to maximize uptime.
SP023 KrASIA Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing SiliconFlow was China's largest independent ecosystem token supplier by annual token throughput in 2025 and ranked among the top five token suppliers overall.
SP024 Hello China Tech SiliconFlow IPO: token economics The top three token suppliers by throughput, Volcengine, Alibaba Cloud, and Baidu AI Cloud, are all hyperscaler divisions.
SP025 SiliconFlow Models - SiliconFlow One API to run inference on 200+ cutting-edge AI models, and deploy in seconds.
SI001 SiliconFlow / HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. For our public cloud-based services, the gross loss margin was 271.6% in 2024 and 119.0% in 2025.
SI002 SiliconFlow Models - SiliconFlow One API to run inference on 200+ cutting-edge AI models, and deploy in seconds.
SI003 36Kr Europe SiliconFlow completes funding round exceeding RMB 2 billion By providing efficient MaaS through the Token factory model, the daily average Token call volume has reached trillions.
SI004 KrASIA Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing Its public cloud service... generated 52.9% of 2025 revenue. Its gross margin was negative 119%.
SI005 Hello China Tech SiliconFlow IPO: token economics In 2025, computing resource costs totaled RMB 59.627 million, accounting for 86.9% of cost of sales.
SI006 Together AI Docs Dedicated model inference pricing Dedicated model inference bills based on the hardware your deployments run on, regardless of model or request volume.
SI007 Together AI Docs Batch processing overview Run asynchronous batch workloads at up to 50% lower cost.
SI008 Together AI Docs Send requests on dedicated endpoints Prompt caching is enabled by default for dedicated model inference. No configuration is required.
SI009 Together AI Docs Fine-tuning pricing Fine-tuning is billed per token processed, scaled by model size, training method, and training type.
SI010 Fireworks Docs Prompt caching For serverless models, cached prompt tokens are discounted compared to regular prompt tokens. The default discount is 50%.
SI011 Fireworks Docs Serverless serving paths Priority tier is for workloads that require higher reliability during peak traffic periods, at a higher price point.
SI012 Fireworks Docs Serverless pricing Batch inference is billed at 50% of serverless pricing on both input and output.
SI013 Fireworks AI Pricing H100 80 GB GPU $7.00 per hour.
SI014 Amazon Web Services Amazon Bedrock Pricing Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing.
SI015 AWS Docs Batch inference in Amazon Bedrock With batch inference, you can submit multiple prompts and generate responses asynchronously.
SI016 AWS Docs Prompt caching in Amazon Bedrock Prompt caching can help when you have workloads with long and repeated contexts ... you're charged at a reduced rate for tokens read from cache.
SI017 Microsoft Azure Microsoft Foundry Pricing The Microsoft Agent pre-purchase plan allows you to save on Microsoft Foundry and Copilot Credit costs by purchasing Agent Commit Units up front.
SI018 Microsoft Azure Foundry Models Pricing Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure... Choose from A100, H100, H200, and MI300 GPU families.
SI019 Alibaba Cloud Model Studio model pricing Model API calls are billed on a pay-as-you-go basis by default.
SI020 Alibaba Cloud Token Plan overview Standard USD 30/seat/month ... Pro USD 100/seat/month ... Max USD 200/seat/month.
SI021 Alibaba Cloud What is Model Studio? Activating Model Studio is free. Costs apply only when you invoke models.
SI022 IDC China 中国MaaS市场进入高速增长期,Token经济从概念走向规模 影响大模型落地的Top5因素依次是:模型性能、安全合规要求、回答质量、在AI平台可用性以及成本效益。
SI023 Fireworks Docs Models & Inference Every response includes token usage information and performance metrics for debugging and observability.
SI024 Together AI Pricing Most teams start with serverless inference and move to dedicated endpoints at scale.
SI025 AWS Amazon Bedrock Features like Model Distillation, Prompt caching, and Intelligent Prompt Routing can reduce expenses while maintaining performance.
SE001 SiliconFlow SiliconFlow China homepage 覆盖语言、语音、图片、视频等场景,一站式提供大模型 API 服务,按量计费,助力应用快速上线。
SE002 SiliconFlow Docs Product introduction As a one-stop cloud service platform integrating top-tier large language models, SiliconFlow is committed to providing developers with faster, more comprehensive, and seamlessly integrated model APIs.
SE003 SiliconFlow Docs API reference home Corresponding Model Name. To better enhance service quality, we will make periodic changes to the models provided by this service.
SE004 SiliconFlow Docs Chat API reference The response header contains the x-siliconcloud-trace-id field, which serves as a unique identifier for tracing requests.
SE005 SiliconFlow Models catalog One API to run inference on 200+ cutting-edge AI models, and deploy in seconds.
SE006 SiliconFlow Docs Privacy Policy When you register, log in, and use the services provided through the https://siliconflow.com platform, we will collect and store the relevant information in accordance with this policy.
SE007 SiliconFlow Docs Terms of Use The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn.
SE008 HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. We offer API services for developers and enterprises ... and on-premise deployment solutions.
SE009 GitHub SiliconFlow organization High-performance inference for every SOTA model.
SE010 GitHub OneDiff repository onediff is an out-of-the-box acceleration library for diffusion models
SE011 GitHub OneDiff releases 1.2.0 ... support diffusers sd3 speedup ... add diffusers nexfort example
SE012 SiliconFlow Blog Kimi K3 API launch post Base URL: https://api.siliconflow.com/v1
SE013 36Kr Europe SiliconFlow completes funding round exceeding RMB 2 billion The platform brings together over 400 models spanning text, image, audio, and video generation.
SE014 KrASIA Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing Its public cloud service includes both a pay-as-you-go serverless model and dedicated computing instances.
SE015 IDC China 中国MaaS市场进入高速增长期,Token经济从概念走向规模 影响大模型落地的Top5因素 ... 在AI平台可用性以及成本效益。
SE016 Together AI Docs Serverless overview Use serverless inference to run models without managing infrastructure.
SE017 Fireworks AI Home The Generative AI Platform for production-ready workloads.
SE018 OpenRouter Docs Quickstart OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers.
SE019 Alibaba Cloud What is Model Studio? Model Studio provides model APIs and application development capabilities.
SE020 AWS Amazon Bedrock Features like Model Distillation, Prompt caching, and Intelligent Prompt Routing can reduce expenses while maintaining performance.
SE021 SiliconFlow SiliconFlow international site Get your Model API fast
SE022 SiliconFlow Cloud Cloud model catalog Model catalog and pricing surface for cloud users.
SE023 Hello China Tech SiliconFlow IPO: token economics It is, in essence, renting the hardware, bundling the software, and wrapping the models.
SE024 China Money Network SiliconFlow's role in China's AI ecosystem The platform supports over 400 models and serves thousands of enterprise customers.
SE025 GitHub OneFlow releases asset page reference in OneDiff README please install OneFlow by the links below.
SU001 HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. Our public cloud-based services primarily target developers and small- and medium-sized enterprises.
SU002 QQ News / Taimei 问AI · Token工厂模式为何吸引全产业链巨头联手投资? 过去一年,硅基流动日均Token调用量达数万亿,服务超1000万用户和1万家企业客户。
SU003 SiliconFlow News listing 贵州移动与硅基流动深化战略合作,加速“Token 工厂”建设
SU004 SiliconFlow Docs Scenarios and application cases Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases.
SU005 SiliconFlow Airline group AI infrastructure case 该企业引入硅基流动(SiliconFlow)推理加速框架,实现三层技术突破。
SU006 SiliconFlow Energy SOE AI infrastructure case 某能源央企正处于转型关键阶段。
SU007 SiliconFlow Guizhou Mobile strategic cooperation 贵州移动与硅基流动 ... 构建“推理框架 + 算力供给 + 模型服务 + 场景应用”全栈协同体系。
SU008 SiliconFlow Reserved instances support 100B-token daily workload 2026 年以来,其 Coding Agent 单日 Token 消耗冲上千亿量级。
SU009 SiliconFlow Docs Use SiliconCloud in MindSearch After adding this configuration, you can execute the relevant commands to start MindSearch.
SU010 SiliconFlow Docs Use Continue with SiliconFlow APIs By integrating SiliconFlow APIs into Continue, you can get access to 200+ open-source models.
SU011 SiliconFlow Docs Use Cline with SiliconFlow APIs we’ll show you how to integrate SiliconFlow’s APIs into Cline
SU012 SiliconFlow About us From open source to enterprise deployment, we accelerate what matters.
SU013 InforCapital SiliconFlow company profile SiliconFlow - AI Infrastructure
SU014 CB Insights SiliconFlow company profile SiliconFlow - Products, Competitors, Financials, Employees
SU015 KrASIA Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing
SU016 36Kr Europe SiliconFlow completes funding round exceeding RMB 2 billion The platform brings together over 400 models spanning text, image, audio, and video generation.
SU017 Hello China Tech SiliconFlow IPO: token economics It is, in essence, renting the hardware, bundling the software, and wrapping the models.
SU018 SiliconFlow China homepage 面向不同行业及需求场景,提供灵活的解决方案
SU019 SiliconFlow Docs Product introduction Our platform empowers developers and enterprises to focus on product innovation
SU020 SiliconFlow Models catalog One API to run inference on 200+ cutting-edge AI models
SU021 China Money Network SiliconFlow's role in China's AI ecosystem SiliconFlow's role in China's AI ecosystem
SU022 SiliconFlow International site Get your Model API fast
SU023 Tencent News Funding and ecosystem article 客户名单涵盖能源、金融、交通等核心行业的头部央企
SU024 GitHub SiliconFlow organization High-performance inference for every SOTA model.
SU025 SiliconFlow Reserved instances page 预留实例
SR001 HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. Our major purchases are concentrated with the top five suppliers.
SR002 SiliconFlow Docs CN 服务协议 服务适用于面向开发者的内容生成服务场合,不适用于自动控制、医疗信息服务、心理咨询和关键信息基础设施的场合。
SR003 SiliconFlow Docs CN 隐私政策 根据中华人民共和国大陆地区相关法律法规的规定,我们需要对您进行实名认证。
SR004 SiliconFlow CN China homepage 京ICP备2024051511号-1 增值电信业务经营许可证:京B2-20242084
SR005 SiliconFlow Docs EN Terms of Use The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn.
SR006 SiliconFlow Docs EN Privacy Policy When you register, log in, and use the services provided through the https://siliconflow.com platform, we will collect and store the relevant information.
SR007 Regulations.ai Measures for the Identification of AI-Generated (Synthetic) Content The Measures set a mandatory national baseline requiring that AI-generated or AI-synthesized content ... be clearly identified to users via explicit and implicit identification.
SR008 U.S. BIS Commerce strengthens restrictions on advanced computing semiconductors These rules reinforce and build on the October 7, 2022, October 17, 2023, and December 2, 2024, controls to restrict the PRC's ability to obtain certain high-end chips.
SR009 U.S. Department of Commerce Guidance on Advanced Computing Items (May 2026) a license is required to export advanced computing items to entities headquartered in Country Group D:5 or Macau
SR010 NIST AI Risk Management Framework On April 7, 2026, NIST released a concept note for an AI RMF Profile on Trustworthy AI in Critical Infrastructure.
SR011 KrASIA Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing
SR012 Hello China Tech SiliconFlow IPO: token economics It is, in essence, renting the hardware, bundling the software, and wrapping the models.
SR013 QQ News / Taimei 问AI · Token工厂模式为何吸引全产业链巨头联手投资? 当前市场在快速膨胀,但竞争也在加剧。
SR014 SiliconFlow Guizhou Mobile strategic cooperation 双方在 2025 年战略合作基础上再升级
SR015 SiliconFlow Reserved instances support 100B-token daily workload 配合企业级交付与运行保障、明确的 SLA 与完善的售后服务
SR016 SiliconFlow Energy SOE AI infrastructure case 实现完整的性能追踪与故障预警,保障复杂生产环境下的高可用与安全性。
SR017 SiliconFlow Airline group AI infrastructure case 模型更新周期从“数周”压缩至“3 天内”
SR018 SiliconFlow Docs Product introduction Provides comprehensive monitoring and fault tolerance mechanisms to guarantee service capabilities.
SR019 SiliconFlow Models catalog One API to run inference on 200+ cutting-edge AI models, and deploy in seconds.
SR020 SiliconFlow Docs API reference home we will make periodic changes to the models provided by this service
SR021 IDC China 中国MaaS市场进入高速增长期,Token经济从概念走向规模 影响大模型落地的Top5因素依次是:模型性能、安全合规要求、回答质量、在AI平台可用性以及成本效益。
SR022 36Kr Europe SiliconFlow completes funding round exceeding RMB 2 billion The daily average token call volume has reached trillions.
SR023 China Money Network SiliconFlow's role in China's AI ecosystem SiliconFlow's role in China's AI ecosystem
SR024 SiliconFlow Docs Scenarios and application cases Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases.
SR025 Together AI Pricing Most teams start with serverless inference and move to dedicated endpoints at scale.
SR026 Fireworks AI Pricing H100 80 GB GPU $7.00 per hour.
SR027 AWS Amazon Bedrock Pricing Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing.
SR028 Alibaba Cloud Model Studio model pricing Model API calls are billed on a pay-as-you-go basis by default.
SR029 OpenRouter Docs Quickstart OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers.
SR030 Together AI Docs Dedicated model inference pricing Dedicated model inference bills based on the hardware your deployments run on, regardless of model or request volume.
SR031 Regulations.ai Provisions on the Administration of Deep Synthesis of Internet Information Services Providers must implement real-identity verification before allowing publishing privileges, and all synthetic content must carry clear technical marks indicating its origin.
SR032 China Law Translate Measures for Labeling of AI-Generated Synthetic Content Service providers shall clearly explain the methods, styles, and other such specifications for labeling generated synthetic content in user service agreements, and notify users to carefully read and understand the corresponding labeling management requirements.
SV001 HKEX Application Proof of Beijing SiliconFlow Technology Co., Ltd. For our public cloud-based services, the gross loss margin was 119.0% in 2025.
SV002 QQ News / Taimei 问AI · Token工厂模式为何吸引全产业链巨头联手投资? 6月16日 ... 硅基流动宣布完成超20亿元B轮融资。
SV003 36Kr Europe SiliconFlow completes funding round exceeding RMB 2 billion The platform brings together over 400 models spanning text, image, audio, and video generation.
SV004 KrASIA Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing
SV005 Hello China Tech SiliconFlow IPO: token economics It is, in essence, renting the hardware, bundling the software, and wrapping the models.
SV006 Reuters / U.S. News Together AI raises $800 million at $8.3 billion valuation Together AI ... raised $800 million ... at $8.3 billion valuation ... annual bookings crossed $1.15 billion last quarter.
SV007 BusinessWire Together AI raises $800 million at $8.3 billion valuation Together AI Raises $800 Million at $8.3 Billion Valuation
SV008 BusinessWire Fireworks AI raises $250M Series C Fireworks AI Raises $250M Series C to Lead the AI Inference Market
SV009 CNBC Fireworks hits $17.5 billion valuation and $1B in annualized revenue The company said ... it has exceeded $1 billion in annualized revenue, and it has now raised a $1.5 billion round at a $17.5 billion valuation.
SV010 Sacra Fireworks AI revenue, valuation & funding At more than $1B in annualized revenue in July 2026, the latest valuation implies an approximate 17.5× revenue multiple.
SV011 BusinessWire Baseten raises $300M at a $5B valuation Baseten Raises $300M at a $5B Valuation to Power a Multi-Model Future
SV012 Sacra Baseten revenue, valuation & funding Baseten hit $600M in annualized revenue in March 2026 ... valued at $13B following its $1.5B Series F in June 2026.
SV013 Sacra CoreWeave revenue, valuation & funding CoreWeave guided full-year 2026 revenue of $12B–$13B.
SV014 SiliconFlow Models catalog One API to run inference on 200+ cutting-edge AI models, and deploy in seconds.
SV015 SiliconFlow Docs Product introduction Our platform empowers developers and enterprises to focus on product innovation while eliminating concerns about exorbitant computational costs.
SV016 Together AI Pricing Most teams start with serverless inference and move to dedicated endpoints at scale.
SV017 Fireworks AI Pricing H100 80 GB GPU $7.00 per hour.
SV018 AWS Amazon Bedrock Pricing Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing.
SV019 Alibaba Cloud Model Studio model pricing Model API calls are billed on a pay-as-you-go basis by default.
SV020 OpenRouter Docs Quickstart OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers.
SV021 SiliconFlow China homepage 大模型云服务 ... 预留实例 ... 私有化大模型服务平台
SV022 SiliconFlow Guizhou Mobile strategic cooperation 双方在 2025 年战略合作基础上再升级
SV023 SiliconFlow Reserved instances support 100B-token daily workload 其 Coding Agent 单日 Token 消耗冲上千亿量级。
SV024 SiliconFlow Docs Scenarios and application cases Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases.
SV025 Regulations.ai Measures for the Identification of AI-Generated (Synthetic) Content The Measures set a mandatory national baseline requiring that AI-generated or AI-synthesized content ... be clearly identified.
SV026 U.S. BIS Commerce strengthens restrictions on advanced computing semiconductors These rules ... restrict the PRC's ability to obtain certain high-end chips critical for military advantage.
SV027 NIST AI Risk Management Framework On April 7, 2026, NIST released a concept note for an AI RMF Profile on Trustworthy AI in Critical Infrastructure.
SV028 IDC China 中国MaaS市场进入高速增长期,Token经济从概念走向规模 影响大模型落地的Top5因素 ... 在AI平台可用性以及成本效益。
SV029 SiliconFlow Docs CN 服务协议 服务适用于面向开发者的内容生成服务场合,不适用于 ... 关键信息基础设施的场合。
SV030 CNBC Fireworks valuation and industry context By managing computing infrastructure for models, Fireworks does business in the inference cloud market, alongside startups such as Baseten and Together AI.