SiliconFlow
China-based AI inference infrastructure company with real scale and strategic relevance, but still negative public-cloud economics, thin public retention visibility, and a valuation that requires evidence-sensitive price discipline.
SiliconFlow looks like a real and increasingly important China-based inference platform, but the disclosed June 2026 price already asks investors to underwrite future margin improvement, enterprise durability, and supply / compliance execution that the public record still does not fully prove.
Cover facts
Company profile
SiliconFlow is a Beijing-founded AI inference infrastructure company established in August 2023. Public evidence describes a multi-surface business: pay-as-you-go model APIs, reserved or dedicated inference capacity, acceleration services, and private deployment for enterprise or compliance-sensitive workloads. By mid-2026 the company had reached meaningful adoption scale, with more than 10 million registered users, more than 13,000 enterprise customers in filing-era disclosure, and a public model library above 200 models. The company also disclosed seven funding rounds through June 2026, including a Series B/B+ sequence that lifted filing-disclosed post-money valuation to RMB7.74 billion while independent media described the round as roughly US$296 million at about a US$1.2 billion valuation. SiliconFlow is strategically relevant within the China inference stack, but the public record still shows weak public-cloud economics, supplier dependence, and incomplete disclosure on retention and concentration.
- Website
- siliconflow.cn
- Founded
- 2023-08-29
- Founders
- Yuan Jinhui
- Founding location
- Beijing, China
- Headquarters
- Beijing, China
- Product
- SiliconFlow delivers an API-first inference platform with usage-based access to large language, multimodal, image, audio, and video models, alongside reserved instances, inference-acceleration services, and private-deployment options for enterprise workloads.
- Customers
- Developers, AI builders, enterprise platform teams, regulated or state-linked institutions, and telecom / infrastructure partners that need model access, predictable inference performance, private deployment, or domestic-chip adaptation.
- Business model
- Usage-based serverless APIs, reserved or dedicated capacity, and private-deployment solutions. Public evidence suggests a two-line business mix between public cloud services and on-premise deployment, with the latter currently carrying materially better gross margins.
- Stage
- Late-stage private; June 2026 Series B / B+ financing and active HKEX Chapter 18C filing
- Funding status
- The public record supports a rapid funding climb from angel through Series B+ between December 2023 and June 2026. The HKEX filing shows post-money valuation steps up to RMB7.74 billion and aggregate post-2025 cash consideration of about RMB1.48 billion, while external media described the June 2026 financing as roughly US$296 million at a US$1.2 billion valuation.
Executive summary
Top strengths
- SiliconFlow has genuine strategic relevance in AI inference: broad model access, multiple deployment modes, and meaningful enterprise/infrastructure use cases rather than a simple demo API surface.
- Public adoption signals are substantial for a company founded in 2023, including more than 10 million registered users, more than 13,000 enterprise customers in filing-era disclosure, and token throughput large enough to support dedicated-capacity case studies.
- The company appears well positioned for China-specific inference demand because it combines model distribution, private deployment, and domestic-chip adaptation narratives.
- The June 2026 financing syndicate and HKEX filing process suggest that SiliconFlow has real capital-market relevance and time to continue executing.
Top risks
- Public-cloud economics are still weak: 2025 blended gross margin was negative and the public- cloud line remained deeply loss-making, indicating that scale has not yet solved quality of growth.
- Supplier and compute dependence remain material because the company rents or leases critical infrastructure rather than fully controlling its own semiconductor supply chain.
- Regulatory and compliance complexity is meaningful across AI labeling, privacy, identity, telecom licensing, and domestic/international surface management.
- Customer breadth is public, but retention, concentration, and revenue-quality metrics remain too thin to support a high-conviction underwriting case.
- The disclosed valuation can still prove expensive if current usage scale does not convert into higher-margin dedicated/private deployments and durable enterprise retention.
Open gaps
- Current 2026 revenue run-rate, product-line mix, and margin trend by serverless, dedicated, and private deployment.
- Enterprise retention, logo concentration, spend concentration, and API-to-dedicated expansion metrics.
- Current cash balance, monthly burn, and financing scenario analysis after the June 2026 round.
- Security / assurance package: uptime history, incident history, SOC 2 / ISO or equivalent evidence.
- Supplier concentration terms and contingency plans under export-control or capacity-shock scenarios.
Contents
01Company Overview
1.1 Identity, Legal Footprint, and Business Model
SiliconFlow's public identity is already more complex than a generic API startup. The China-facing domain presents Beijing SiliconFlow as an AI-capability provider with a concrete address in Haidian, a visible ICP filing, and a Beijing value-added telecom permit, while the global terms and privacy pages are issued by SiliconFlow Technology Pte. Ltd. and explicitly route mainland-China users back to siliconflow.cn. The HKEX filing adds a third layer: a PRC joint-stock company with a Beijing registered office and a Hong Kong place of business, but with headquarters, senior management, and operations still centered outside Hong Kong. In practical terms, the company is China-rooted, globally marketed, and already structuring itself for cross-border capital-market access. The underlying business model is clearer than the legal map. Across the China homepage, global homepage, GitHub profile, and API documentation, SiliconFlow consistently markets itself as an inference infrastructure layer rather than a model lab or end-application vendor. It sells standardized access to models through large-model APIs, dedicated or reserved instances, inference acceleration services, and private deployment. The docs emphasize rapid onboarding, API-key self-service, and an OpenAI-style surface; the marketing pages emphasize speed, cost control, multimodal coverage, and predictable usage-based delivery. That combination is important because it shows the company is trying to sit between model suppliers, compute suppliers, and downstream developers or enterprises as a neutral token-supply platform.[CO001, CO002, CO003, CO004, CO005, CO006]
| Metric | Value / status | Date / period | Confidence | Gap or caveat |
|---|---|---|---|---|
| Legal establishment | 2023-08-29 PRC LLC; converted to joint-stock company on 2026-06-22 | 2023-08 to 2026-06 | High | Conversion is clear; ultimate post-IPO structure still subject to listing completion |
| Registered HQ | Haidian District, Beijing | Current public record | High | Country-by-country office list beyond Beijing and HK is not fully public |
| Global legal entity | .com terms issued by SiliconFlow Technology Pte. Ltd. | Current public record | High | Public entity map between PRC and Singapore vehicles is still incomplete |
| Latest capital-market stage | HKEX Chapter 18C applicant / pre-commercial listing process active | 2026-06 to 2026-07 | High | Listing is not yet approved or completed |
| Latest disclosed valuation rung | RMB 7.74B post-money (Series B+) | 2026-06-09 / 2026-06-18 sequence | High | External headlines often refer instead to an over-RMB2B Series B |
| Post-2025 financing cash received | ~RMB 1.48B across A+, B, and B+ | After 2025-12-31 | High | Does not include earlier pre-2026 rounds |
| Registered users | 10.28M+ | 2026-04-30 | High | Point-in-time metric; current run-date level undisclosed |
| Enterprise customers | 13,000+ served | Latest practicable date before 2026-06-30 filing | High | No public split by paying status, cohort, or concentration |
| Model coverage | 170+ in filing; 200+ on July 2026 model library | 2026-06 to 2026-07 | High | Public website appears newer than filing snapshot |
| Token throughput | 578.5B average daily; 1,071.4B peak daily | 2026-04 | High | Usage scale does not directly imply healthy unit economics |
| Overseas commercialization | Monthly overseas revenue exceeded US$1M | 2026-05 | High | No public geographic or product mix disclosed |
| Current headcount | Not publicly supportable from retained sources | As of 2026-07-22 | High | Requires management confirmation |
Combines filing-backed operating metrics with public-website and financing context. Where filing and media snapshots differ, the table preserves the discrepancy rather than smoothing it away.
[CO001, CO002, CO003, CO004, CO005, CO013]Flow diagram linking SiliconFlow's legal footprint, platform architecture, customer base, investor set, and listing trajectory.
1.2 Leadership, Governance, and Control Signals
Public disclosure on leadership is better than a typical private startup because of the HKEX filing, but it still leaves important diligence gaps. Dr. Yuan Jinhui is clearly the central figure: founder, chairman, CEO, general manager, and financial controller. The filing also identifies Liu Juncheng as CTO and executive director, Zeng Hua as commercialization lead and executive director, and Chen Yingjie as the non-executive director linked to Alibaba's strategic-investment arm. Beyond the formal board, the history and employee-incentive sections point to a broader founding and core-team nucleus that includes Zhao Zhen and Hu Jian, both tied to the earlier OneFlow network. That matters because SiliconFlow appears to be less a lightly staffed API aggregator than a technically opinionated team with a pre-existing system-software lineage. The biographies reinforce that interpretation. Yuan came through Microsoft China and OneFlow, Liu also spent years at OneFlow in AI-system-software roles, and Zeng adds commercialization experience from Microsoft China, Baidu, and JD. Chen Yingjie adds a strategic-capital bridge to Alibaba. The board structure after listing is straightforward on paper—three executive directors, one non-executive, and three independent non-executives—but the public record still lacks details on control rights, founder dilution, secondary activity, and board-economics terms. One additional wrinkle is that third-party databases do not perfectly agree on the co-founder slate, which makes the filing the authoritative source but also shows how noisy startup databases remain for a fast-moving China AI company.[CO022, CO023, CO024, CO025, CO026, CO027]
| Person / group | Current role | Relevant prior background | Why it matters | Key-person or disclosure note |
|---|---|---|---|---|
| Dr. Yuan Jinhui | Founder, chairman, CEO, general manager, financial controller | Microsoft China lead researcher; founded OneFlow | Technical and strategic center of gravity for inference-platform thesis | Very high dependence on Yuan across product, finance, and capital markets |
| Liu Juncheng | Executive director, CTO | AI system-software and OneFlow R&D background | Owns core technology execution and architecture continuity | High dependence on core-system-software talent |
| Zeng Hua | Executive director, deputy general manager | Microsoft China, Baidu, JD commercialization background | Adds enterprise sales and go-to-market credibility | Important for translating technical scale into monetization |
| Chen Yingjie | Non-executive director | Alibaba strategic-investment managing director; ex-PwC; former XPeng and MiniMax board roles | Represents strategic-capital and ecosystem signal from a major platform investor | Board seat and rights are not fully disclosed in public materials |
| Zhao Zhen | Chief operating officer / co-founder per filing history | Former COO of OneFlow | Operational continuity from prior team network | Not a board member in public filing; external profile coverage is sparse |
| Hu Jian and broader core team | Deputy GM / core R&D and employee-incentive participants | Core team overlaps with OneFlow-era talent base | Shows deeper bench than a two-founder startup narrative | Public biographies for several core operators remain limited |
| Independent director slate | Wu Chuan, Li Dan, Hou Hong (proposed INEDs) | Academic, accounting, and strategy backgrounds | Suggests listing-readiness and added governance depth | Independent slate is filing-disclosed, but committee practice is not yet proven |
The HKEX filing materially improves leadership visibility, but several operational leaders and the economics of board control remain only partially disclosed.
[CO022, CO023, CO024, CO025, CO026, CO027]1.3 Funding History and Capital-Market Story
SiliconFlow's capital formation has been unusually compressed. The HKEX filing documents seven financing rounds from December 2023 through June 2026, with disclosed capital injections stepping from RMB47.2 million at angel to RMB65.84 million at angel+, RMB71.48 million at pre-A, RMB285.99 million at Series A, RMB220 million at Series A+, RMB520 million at Series B, and RMB740 million at Series B+. The corresponding post-money valuations rose from RMB280.0 million to RMB7.74 billion in roughly two and a half years. That is not just a fundraising story; it is a signal that investors rapidly repriced the company as the AI-token middle layer became investable in China. The June 2026 headlines need careful handling. External media consistently described SiliconFlow as completing an over-RMB2 billion Series B, backed by a strategic syndicate spanning Trip.com or Ctrip, JinkoSolar, Kingdee, Unicom-linked investors, Biren, NIO Capital, SenseTime, GGV, and others, with China Renaissance advising. But the HKEX filing then presents a more granular 2026 sequence of A+, B, and B+ financings, with aggregate post-2025 cash consideration of about RMB1.48 billion and valuation steps from RMB3.12 billion to RMB5.02 billion and then RMB7.74 billion. Public databases also show unreconciled totals, so the right conclusion is not that one source is necessarily wrong, but that the public capital history should be treated as directionally clear and transaction-label details as still needing data-room reconciliation.[CO032, CO033, CO034, CO035, CO036, CO037]
| Stakeholder / investor | Role or round context | Strategic relevance | What public record shows | Diligence ask |
|---|---|---|---|---|
| Alibaba / Hangzhou Duoxiang | Shareholder and board-linked strategic investor | Potential cloud, distribution, and ecosystem leverage | Alibaba-linked entity is a substantial shareholder; Chen Yingjie sits on the board | Confirm exact ownership, information rights, and commercial relationship scope |
| Trip.com / Ctrip Strategic Investment | Named June 2026 strategic investor | Travel, enterprise-demand, and capital-markets signal | Repeatedly named in June 2026 financing coverage | Verify whether Trip.com invested at B, B+, or both and whether commercial pilots exist |
| SenseTime strategic investment | Named June 2026 investor and existing AI-ecosystem peer | Signals model and AI-infra ecosystem alignment | Named in multiple June 2026 funding reports | Check whether the relationship is purely financial or includes workload, model, or channel cooperation |
| JinkoSolar Holding | Named June 2026 investor | Links AI infrastructure to power and datacenter economics | 36Kr frames Jinko as an energy-to-compute strategic partner | Determine whether this includes actual infrastructure contracts or only capital |
| Unicom-linked capital | Named June 2026 investor group | Telecom and computing-network integration may aid enterprise deployment | Funding coverage cites Unicom Xinwo and Unicom Capital vehicles | Clarify customer or network resource commitments |
| Biren Technology | Named June 2026 strategic investor | Domestic-chip adaptation is central to SiliconFlow's differentiation | Funding coverage frames Biren around domestic-chip inference collaboration | Check binding commercial terms, exclusivity, or co-marketing obligations |
| NIO Capital and other financial investors | Growth capital in late-stage rounds | Adds financial sponsorship beyond strategic corporates | Named across financing coverage and CB Insights investor table | Request full cap table and pro-rata or liquidation rights |
| China Renaissance | Exclusive adviser on June 2026 financing; later HKEX overall coordinator | Bridges private financing and public-listing preparation | Named in Caixin Global and HKEX coordinator announcement | Determine whether adviser economics or process dependencies create pressure for rapid listing |
Public information strongly supports a strategically dense investor base, but not the control rights or commercial obligations attached to those investors.
[CO025, CO026, CO033, CO036, CO037, CO047]1.4 Scale, Traction, and Geographic Reach
The filing and corroborating coverage show that SiliconFlow has already achieved real usage scale, even if its commercial density is still evolving. As of April 30, 2026, the company disclosed more than 10 million registered users, average daily throughput of 578.5 billion tokens, peak daily throughput above one trillion tokens, and more than 13,000 enterprise customers by the latest practicable date. The platform supported over 170 models in the filing and more than 200 on the public model library by July 2026. Those are unusually large infrastructure-side operating metrics for a company incorporated only in August 2023, and they help explain why investors were willing to fund it aggressively despite immature profitability. At the same time, the footprint evidence is broad but not exhaustive. Public pages firmly place the operating center in Beijing, and the filing confirms a Hong Kong business address as part of the IPO process. The official and third-party materials also suggest overseas commercialization is no longer hypothetical: the filing says overseas monthly revenue exceeded US$1 million in May 2026, while June coverage repeatedly described millions of dollars in overseas monthly revenue and global platform traction. What is still missing is a clean public breakdown by country, office, headcount, or overseas revenue mix. The company therefore looks scaled in usage, scaled enough to matter in enterprise adoption, and only partially disclosed in organizational footprint.[CO013, CO014, CO015, CO016, CO017, CO018]
Key maturity, scale, and caution indicators for SiliconFlow as of the 2026-07-22 run date.
1.5 Milestones, Compliance Signals, and Adverse Context
The operating chronology is coherent. SiliconFlow says it started inference-engine R&D in August 2023, launched public-cloud MaaS in May 2024, shipped DeepSeek inference on Huawei Ascend in February 2025, launched private MaaS in September 2025, released the Elastic GPU scheduler in April 2026, converted into a joint-stock company in June 2026, filed its HKEX application proof on June 30, and then added China Renaissance as an overall coordinator on July 14. That timeline shows a company moving in lockstep across productization, commercial expansion, financing, and listing preparation. The adverse frame is just as important. HKEX financial disclosure shows 2025 revenue scaling rapidly to RMB55.3 million, but gross margin swung negative at -24.0%. KrASIA and Hello China Tech sharpen the point: SiliconFlow appears to be renting compute, distributing third-party models, and competing in a price war where developer credits and below-cost token supply can accelerate adoption faster than profitability. Public-cloud economics look structurally tougher than on-premise deployment economics, which means the company's headline growth does not yet settle the question of durable monetization. In other words, the company overview supports a serious infrastructure platform with elite strategic backers, but not yet a fully de-risked business model.[CO039, CO040, CO042, CO043, CO044, CO045]
| Date | Event | Type | Amount / status | Participants | Implication |
|---|---|---|---|---|---|
| 2023-08-29 | Company established in Beijing | founding | PRC LLC formed | Yuan-led founding team | Creates the legal base for later product, financing, and listing work |
| 2023-08-01 | Inference-engine R&D begins | product | Engine development started | Founding engineering team | Shows the company began at the systems layer rather than at the application layer |
| 2024-05-01 | Public-cloud MaaS launches | product | Commercial public-cloud service live | SiliconFlow | Marks transition from R&D project to external platform |
| 2025-02-01 | DeepSeek services on Huawei Ascend launch | product | Domestic-chip token service milestone | SiliconFlow, DeepSeek, Huawei Ascend | Strengthens domestic-chip and heterogeneous-compute positioning |
| 2025-07-10 | Series A consideration fully paid | financing | ~RMB285.99M Series A | Private investors per filing | Moves valuation to RMB2.286B and funds broader commercialization |
| 2025-09-01 | Private MaaS launches | product | Private deployment offering live | SiliconFlow enterprise team | Adds higher-touch compliance-oriented revenue path |
| 2026-04-01 | Elastic GPU launches | product | Heterogeneous scheduling engine live | SiliconFlow | Supports the token-factory efficiency thesis |
| 2026-06-17 | Series B consideration fully paid | financing | ~RMB520M Series B | 2026 investor syndicate | Pushes disclosed post-money valuation to RMB5.02B |
| 2026-06-18 | Series B+ consideration paid | financing | ~RMB740M Series B+ | Late-2026 investors | Pushes disclosed post-money valuation to RMB7.74B just before listing |
| 2026-06-30 | HKEX application proof filed | governance | Chapter 18C pre-commercial filing | SiliconFlow, Huatai, Haitong | Creates the first deep public disclosure set |
| 2026-07-14 | China Renaissance added as overall coordinator | governance | Listing process expanded | SiliconFlow, China Renaissance | Signals active continuation of IPO preparation |
This is the single chronology of record for the overview chapter. Dates come from the filing or directly from the cited June-July 2026 announcements; month-only milestones use the first day of the disclosed month as an anchor.
[CO001, CO003, CO032, CO033, CO034, CO042]Chronology of SiliconFlow's formation, product launches, financing steps, and listing milestones from August 2023 through July 2026.
1.6 Exhibits
02Market Analysis
2.1 Market Boundary, Included Spend, and Substitutes
SiliconFlow is not selling a generic 'AI' product; it sits in the inference middleware layer that turns third-party models and rented or customer-owned compute into metered token output. The June 2026 HKEX filing, SiliconFlow's models catalog, and its onboarding docs all point to the same market boundary: public-cloud serverless token APIs, dedicated inference instances, and private deployments that let enterprises or developers consume many models through one interface. That means the included spend is not model training, semiconductor design, or end-user AI SaaS seats. It is the spend tied to model serving, token distribution, orchestration, and adjacent tooling that reduces the operational work of adopting AI in production. The most important substitutes are therefore not only other startups. Closed-ecosystem hyperscaler platforms such as Amazon Bedrock, Microsoft Foundry, and Alibaba Cloud Model Studio compete for the same model-access and agent-building budgets, while independent inference platforms such as Together and Fireworks compete on developer portability, speed, and price. In many buyer journeys the status quo is a mix of direct model-lab APIs, internal self-hosting, or an incumbent cloud already trusted for security and procurement. SiliconFlow's market matters because it sits between those options: more open than a single-cloud stack, but more productized than a do-it-yourself inference layer.[CM001, CM002, CM003, CM004, CM005, CM006]
| Segment / category | Included spend | Excluded spend | Buyer / payer | Why it matters |
|---|---|---|---|---|
| Public-cloud MaaS / token APIs | Serverless token calls, dedicated instances, model access and routing | Model training, chips, end-user AI SaaS seats | Developers, startups, product teams | This is SiliconFlow's core public-cloud revenue path |
| Private / on-prem MaaS | Deployment software and enterprise-specific token factories in customer environments | Generic IT outsourcing or unrelated cloud migration | Regulated enterprises, SOEs, finance, public sector | Captures buyers that need compliance and supply control |
| Developer distribution layer | BYOK usage inside apps, SDKs, orchestration tools, preintegrated marketplaces | Closed single-app subscriptions | Developers and tool operators | Reduces switching friction and broadens demand capture |
| Hyperscaler managed model platforms | Enterprise governance, agent tooling, model catalogs, routing, guardrails | Consumer AI assistants | Cloud platform owners, CIO budgets | Primary substitute set for enterprise-grade procurement |
| AI inference infrastructure context | Underlying inference hardware, cloud capacity, and serving software | Training-only infrastructure and non-AI compute | CSPs, infrastructure buyers | Useful for TAM context but too broad for a direct SiliconFlow SAM |
Boundary anchored on SiliconFlow's public-cloud and private-deployment positioning, then widened to the substitute set buyers actually compare during procurement.
[CM001, CM002, CM003, CM004, CM005, CM006]Inference-market value chain from model and compute supply through platforms, tools, and enterprise production deployment.
[CM004, CM006, CM007, CM022, CM027, CM028]2.2 Market Sizing Through China MaaS and Global Inference Lenses
No single public number cleanly captures SiliconFlow's opportunity, so the defensible approach is to preserve multiple lenses rather than force one TAM. On the China demand lens, IDC says enterprise MaaS token consumption jumped from 114 trillion tokens in 2024 to 1,944 trillion in 2025 and projects roughly 40,000 trillion in 2026, while public-cloud MaaS revenue reaches RMB3.07 billion in 2025 and RMB18.6 billion in 2026. On the filing lens, Frost & Sullivan's industry work inside the HKEX prospectus shows China's token-supply market growing 1,602.6% from 2024 to 2025 and reaching 53.2 quintillion tokens by 2030. Those two China-centric lenses broadly agree on explosive expansion, but they do not use identical definitions or baselines, which is exactly why they should be triangulated instead of averaged. On the broader global lens, third-party analyst pages from MarketsandMarkets, Grand View Research, and Fortune Business Insights all cluster the AI inference market around roughly USD100 billion in 2024-2025, with forecast endpoints ranging from roughly USD254 billion by 2030 to more than USD312 billion by 2034. That larger global figure is useful context, but it overstates SiliconFlow's directly reachable near-term market because much of it includes hardware and hyperscaler infrastructure spend far above an independent token-distribution layer. The better underwriting logic is stacked: global inference infrastructure for context, China public-cloud MaaS for monetizable demand, and SiliconFlow's 1.5% 2025 throughput share as a signal of present competitive relevance rather than a direct revenue-share bridge.[CM009, CM010, CM011, CM012, CM013, CM014]
| Lens / publisher | Geography | Value | CAGR / growth | Methodology cue | Confidence | Limitation |
|---|---|---|---|---|---|---|
| IDC public-cloud MaaS revenue | China | RMB3.07B (2025) to RMB18.6B (2026) | ~6.1x y/y | Enterprise MaaS revenue on public cloud | Medium | Revenue lens excludes private deployment |
| IDC token consumption | China | 1,944T tokens (2025); ~40,000T (2026) | ~16x in 2025; ~20x in 2026 | Token throughput / usage lens | High | Usage is not the same as monetized revenue |
| Frost & Sullivan via HKEX | China | 2,426.3T tokens total market in 2025; 53.2 quintillion by 2030 | 1,602.6% growth 2024-2025; 638.3% CAGR 2025-2030 | Token supply market throughput | Medium | Definition differs from IDC MaaS lens |
| SiliconFlow share signal via HKEX | China | 1.5% throughput share in 2025; rank #4 overall, #1 independent | n/a | Company ranking within token-supply platforms | High | Share is by throughput, not revenue |
| MarketsandMarkets AI inference | Global | USD106.15B (2025) to USD254.98B (2030) | 19.2% | Broad AI inference infrastructure market | Medium | Too broad for direct SiliconFlow SAM |
| Grand View AI inference | Global | USD97.24B (2024) to USD253.75B (2030) | 17.5% | Broad AI inference infrastructure market | Medium | One year earlier base than other reports |
| Fortune Business Insights AI inference | Global | USD103.73B (2025) to USD312.64B (2034) | 12.98% | Broad AI inference infrastructure market | Medium | Longer forecast window increases dispersion |
Use these lenses together, not as additive components. China MaaS demand, global AI-inference infrastructure, and SiliconFlow share signals answer different questions.
[CM009, CM010, CM011, CM012, CM013, CM014]A layered view from broad global AI-inference infrastructure down to SiliconFlow's current share signal in China token supply.
Layers are directional and not additive because the underlying units and market definitions differ.
[CM011, CM012, CM015, CM016, CM017, CM018]Low/base/high source-backed ranges for the broad global AI-inference market in consistent USD billions.
USD billions. Base-year draws from 2024-2025 report pages; forecast endpoint uses nearest published 2030 or 2034 figure from retained sources.
[CM016, CM017, CM018, CM019]2.3 Buyer Segmentation, Workflow Owners, and Budget Paths
The filing makes clear that SiliconFlow's market spans several buyer archetypes rather than one homogenous user. At the low-friction end are individual developers and startups that want commitment-free model access, rapid onboarding, prepaid usage, and the ability to keep existing tools through BYOK or compatible APIs. At the higher-control end are large enterprises and institutions that care about dedicated performance, stable supply, low latency, private deployment, or supply assurance. The same pattern shows up across comparable platforms: Bedrock markets to startups and global enterprises; Microsoft Foundry markets a unified resource plane for agents, governance, and fleet-wide controls; Alibaba's Token Plan converts inference into seat-based team productivity spend; and Together and Fireworks emphasize OpenAI-compatible APIs that reduce migration work. That mix implies multiple budget owners. For self-serve API usage the first budget can live with a developer lead, startup founder, or a team productivity manager. As deployments become larger or regulated, ownership migrates toward the CIO, platform engineering leader, procurement, or security-compliance gatekeepers. The adoption path usually runs from a quick serverless experiment to heavier usage, then to dedicated instances or private deployment once cost, stability, or compliance start to matter. SiliconFlow's advantage is that it can participate in more than one stage of that path; its challenge is that every stage brings stronger competition from incumbent clouds and lower tolerance for product or supply instability.[CM022, CM023, CM024, CM025, CM026, CM027]
| Segment | Buyer | User | Payer / budget owner | Workflow | Adoption trigger | Implication |
|---|---|---|---|---|---|---|
| Individual developers | Developer lead / founder | Engineer | Credit balance or small team budget | Prototype and direct API use | Need fast model access with no commitment | Self-serve onboarding matters more than contracts |
| AI-native startups and tools | CTO / product lead | Application team | Product or infra budget | Embed multi-model inference into apps | Need portability and speed to ship features | BYOK and pre-integrations reduce friction |
| Enterprise AI platform teams | CIO / platform engineering | Internal product teams | Cloud / platform budget | Standardize model access across business units | Need governance, routing, and observability | Compete directly with incumbent clouds |
| Regulated enterprises / institutions | CIO + security/compliance | Business units and knowledge workers | Enterprise procurement | Dedicated or private deployment | Need stability, latency, or data-control assurances | Higher-value but longer sales cycle |
| Team productivity / coding organizations | Engineering manager or team admin | Developers | Seat/subscription budget | Credits-based daily AI usage | Need predictable spend and team controls | Subscription packaging can expand budget pool |
Buyer map combines SiliconFlow's own segmentation with how AWS, Microsoft, Alibaba, Together, and Fireworks package inference for different budget owners.
[CM022, CM023, CM024, CM025, CM026, CM027]Segments mapped by budget owner, buying trigger, operational requirement, and migration path.
[CM022, CM023, CM024, CM025, CM026, CM027]2.4 Growth Drivers, Adoption Constraints, and Valuation Relevance
The most important tailwinds are visible in both macro and platform-level sources. Stanford's 2026 AI Index says frontier-model progress did not plateau in 2025 and that organizational adoption reached 88%, while IDC argues China's MaaS competition is no longer just about cheap tokens but about the combined package of price, performance, and toolchain support. The official product pages of Bedrock, Foundry, Alibaba, Together, and Fireworks all reinforce that shift: they sell routing, observability, governance, privacy controls, dedicated throughput, caching, and model evaluation as essential features. That matters for SiliconFlow because the market is maturing from raw API access into a more operationally demanding layer where orchestration and enterprise fit are increasingly monetizable. The constraints are equally material. IDC still ranks performance, security and compliance, answer quality, platform availability, and cost effectiveness among the top enterprise selection factors. Fortune's market page highlights hardware cost and integration complexity, while the filing and independent analyses show what this means economically for a neutral platform: public-cloud growth can be real even when the vendor is still renting compute, subsidizing developers, and operating at negative gross margins. Hello China Tech adds a sharper warning by describing a 2023-2026 price war in model APIs and compute-backed token supply. For valuation, that means the market can be enormous and still not translate cleanly into attractive independent-economics businesses unless the platform can protect pricing, improve utilization, and win higher-value enterprise workloads over time.[CM029, CM030, CM031, CM032, CM033, CM034]
| Driver / constraint | Direction | Timing | Evidence | Why it matters | Diligence ask |
|---|---|---|---|---|---|
| Frontier-model progress and broad AI adoption | Driver | Now | Stanford AI Index 2026 | More capable models create more inference demand | How much of new usage stays open/multi-model? |
| Multimodal and agent workloads | Driver | Now to medium term | IDC + hyperscaler product pages | Raises token intensity and expands use cases | Which workloads are highest value for SiliconFlow? |
| Price + performance + toolchain competition | Driver and filter | Now | IDC | Winners need more than low token prices | How differentiated is SiliconFlow tooling? |
| Governance, privacy, and enterprise controls | Driver | Now | AWS + Microsoft + Alibaba | Enterprise buyers increasingly require these features | Are SiliconFlow controls comparable enough to win large accounts? |
| Developer portability and OpenAI compatibility | Driver | Now | SiliconFlow + Together + Fireworks | Makes adoption easier and lowers migration cost | Does easy switching help SiliconFlow or hurt moat? |
| Hardware cost and integration complexity | Constraint | Now | Fortune Business Insights | Inference scale remains operationally expensive | What portion of cost can SiliconFlow structurally remove? |
| Price wars and subsidized user acquisition | Constraint | Now | HKEX + Hello China Tech + KrASIA | Market growth may outrun profit-pool growth | When do unit economics inflect positive? |
| Compute supply and heterogeneous orchestration | Constraint / differentiator | Now to medium term | HKEX + 36Kr | Access to supply can shape margins and reliability | How durable are supplier relationships and chip coverage? |
Several factors are two-sided: they enlarge the market while simultaneously raising the execution bar for independent platforms.
[CM029, CM030, CM031, CM032, CM033, CM034]2.5 Conflicting Estimates and Remaining Diligence Gaps
A disciplined market chapter should preserve what is still unknown. The first gap is definitional: IDC's China MaaS framing, Frost & Sullivan's token-supply framing, and the global AI-inference reports are all useful, but they are not measuring the same basket of spend or activity. The second gap is company-specific: no retained public source isolates SiliconFlow's serviceable revenue share by geography, by customer segment, or by model category, and none shows the retention or switching dynamics that would turn gross throughput into a durable share thesis. The practical conclusion is that market evidence is strongest at the level of demand direction and weakest at the level of precise independent-platform profit pools. SiliconFlow clearly participates in a fast-growing market; it is much harder to prove from public sources alone how much of that growth will accrue to neutral token platforms versus hyperscalers, model labs, or private deployments inside large enterprises. Investors should therefore treat market size as supportive context, but reserve underwriting conviction for evidence on unit economics, segment mix, and defensibility in later chapters.[CM021, CM041, CM042]
2.6 Exhibits
03Competitors
3.1 Competitive Landscape by Class
SiliconFlow's competitive set spans at least four classes. First are independent open inference platforms such as Together AI, Fireworks, and OpenRouter, which compete on model breadth, developer portability, speed, and pricing rather than on ownership of a proprietary frontier model. Second are incumbent cloud platforms such as Amazon Bedrock, Microsoft Foundry, and Alibaba Cloud Model Studio, which sell the same broad job—enterprise access to many models—but bundle it into larger governance, networking, identity, and procurement stacks. Third are direct model-lab APIs and single-vendor ecosystems, which can bypass neutral platforms entirely when a customer wants one preferred model rather than a marketplace or broker. Fourth is internal build: engineering teams can self-host open models or stitch together vendor APIs when they believe cost, control, or performance justify the complexity. SiliconFlow's own filing is clear that it wants to sit in the open, neutral middle. It explicitly contrasts itself with closed ecosystems, claims first place among independent ecosystem token-supply platforms in China, and positions itself as a connective layer across heterogeneous compute and multiple model families. That means the most relevant competitor question is not simply who has the biggest model catalog; it is which vendor class wins as buyers progress from experimentation to production. In early developer adoption, independent platforms and brokers can look interchangeable. In regulated or very large deployments, the balance often shifts toward the providers with stronger governance, procurement reach, and guaranteed supply.[CP001, CP002, CP003, CP004, CP005, CP006]
| Competitor | Class | Scale / funding signal | Target segment | Differentiation | Limitation / diligence note |
|---|---|---|---|---|---|
| SiliconFlow | Independent token-supply platform | HKEX applicant; 1.5% China throughput share in 2025; RMB7.74B latest disclosed valuation rung | Developers, startups, enterprises, private deployments | China-local neutrality, heterogeneous chips, 200+ models, domestic-chip adaptation | Much smaller than hyperscalers; public-cloud economics still negative in filing |
| Together AI | Independent open-model platform | Private company; retained sources here emphasize product rather than current funding or revenue | Developers and AI-native builders scaling from serverless to dedicated GPUs | Open-source focus, strong serverless catalog, dedicated endpoints, speed claims | No retained independent source here establishes current scale or margins |
| Fireworks AI | Independent open-model platform | Fireworks says it processes 40T+ tokens/day; retained sources here do not establish current funding or revenue | Developers, enterprises needing training/inference, on-demand GPU users | Specialized training + inference, serverless and on-demand, fine-tuning, performance positioning | Throughput claim is vendor-authored and not independently reconciled here |
| OpenRouter | Broker / router / aggregator | Unified API across hundreds of models; retained sources emphasize routing rather than disclosed scale | Developers and agent builders optimizing cost, latency, or fallback behavior | Provider routing, cost/latency sorting, multi-homing, BYOK-friendly design | Relies on external providers rather than owning core compute supply |
| Amazon Bedrock | Hyperscaler managed model platform | AWS platform serving 100,000+ organizations on Bedrock | Enterprise builders from startups to global enterprises | Governance, guardrails, model choice, batch/priority options, procurement strength | Less neutral than an independent broker and tied to AWS estate |
| Microsoft Foundry | Hyperscaler managed agent/model platform | Microsoft positions Foundry as unified control plane with 1,900+ to 11,000+ model access surfaces | Enterprise platform teams and agent builders | RBAC, policy, observability, governance, model catalog, managed compute | Pricing pages do not always expose token prices cleanly in retained text |
| Alibaba Cloud Model Studio | Regional hyperscaler / model platform | Alibaba combines Qwen ownership with third-party models and regional endpoints | Teams and enterprises across China and international regions | OpenAI compatibility, multimodal Qwen stack, regional deployments, team credit plans | Platform economics and realized discounting versus list pricing remain unclear |
Profile rows mix public-company incumbents, open independent peers, and a routing broker because buyers evaluate all of them against the same model-access job.
[CP001, CP002, CP003, CP004, CP005, CP006]Ordinal positioning by openness / portability and enterprise control / procurement strength.
Ordinal 1-10 scores synthesized from retained product, pricing, and governance evidence; not a third-party benchmark.
[CP001, CP013, CP014, CP015, CP016, CP017]3.2 Capability Breadth and Pricing Competition
Capability overlap is substantial. SiliconFlow, Together, Fireworks, OpenRouter, Alibaba, and other rivals all present some mix of OpenAI-compatible APIs, multi-model access, and low-friction onboarding. SiliconFlow advertises 200-plus models; OpenRouter says hundreds of models through a unified API; Bedrock says 100-plus foundation models; Microsoft Foundry markets 1,900-plus models in documentation and 11,000-plus models on its pricing surface; Alibaba packages Qwen plus DeepSeek, Kimi, GLM, and other third-party models; Together and Fireworks each combine serverless and higher-control deployment paths. This is a market where feature parity on basic access is becoming table stakes. Pricing and packaging are therefore increasingly strategic. Together and Fireworks both offer per-token serverless pricing with cached-token and batch discounts, then move heavier users toward dedicated hardware. AWS Bedrock layers token pricing, batch discounts, and provisioned or reserved capacity. Alibaba mixes pay-as-you-go rates with regional discounting and credit-based team plans. OpenRouter turns pricing itself into part of the product by routing across providers and sorting for price, latency, or throughput. The net result is that buyers can often find comparable base-model access across multiple vendors, but the all-in value proposition still differs meaningfully once deployment mode, rate limits, governance, routing, and performance controls enter the decision.[CP010, CP011, CP012, CP013, CP014, CP015]
| Buying criterion | SiliconFlow | Together | Fireworks | OpenRouter | AWS Bedrock | Microsoft Foundry | Alibaba Model Studio |
|---|---|---|---|---|---|---|---|
| OpenAI-compatible API | Yes | Yes | Yes | Drop-in OpenAI SDK path | n/a in retained source | Yes via OpenAI()/project endpoint | Yes |
| Broad multi-model catalog | 200+ models | Yes | 100+ open text + vision + more | Hundreds of models | 100+ models | 1,900+ to 11,000+ models | Qwen + DeepSeek/Kimi/GLM + multimodal |
| Dedicated / reserved capacity | Dedicated instances + private deployment | Dedicated model inference | On-demand deployments | Routes to providers rather than own dedicated fleet | Provisioned / reserved tiers | Managed compute / provisioned throughput | Regional deployment scopes and team plans |
| Routing / fallback controls | Model selection + BYOK channels | Model choice; serverless to dedicated | Serving paths and tiering | Explicit provider routing and fallbacks | Prompt routing | Model router and unified control plane | Regional endpoints and billing controls |
| Enterprise governance / controls | Partially public from docs and filing | Some docs; limited retained trust evidence | Usage metrics and dashboards | Routing/data controls but lighter enterprise stack | Guardrails, privacy, compliance | RBAC, network, policy, observability | Data privacy statement and monitoring |
Unsupported cells are phrased conservatively and limited to retained-source evidence only.
[CP010, CP011, CP012, CP013, CP014, CP015]| Vendor | Example unit / packaging | Illustrative retained pricing | Discount / control lever | Implication |
|---|---|---|---|---|
| SiliconFlow | Per-model inference API | Public model library exposes per-model prices; exact realized rates not equal to margin | BYOK and tool integrations | Competes on breadth and convenience, but realized economics are undisclosed |
| Together | Serverless per token; dedicated per GPU-minute | Qwen 3.7 Max $1.25 input / $3.75 output per 1M tokens; DeepSeek V4 Pro $1.74 / $3.48 | Cached-input discounts and batch discounts; dedicated cheaper at high utilization | Strong direct comparable for open-model token pricing |
| Fireworks | Serverless per token; on-demand GPU-hour; fine-tuning | DeepSeek V4 Pro $1.74 / $0.145 cached / $3.48; H100 on-demand $7.00/hr | Priority/Fast tiers, cached-token discount, batch 50% of standard | Competes on both token pricing and higher-control infrastructure |
| OpenRouter | Provider-routed per-token API | Rates vary by underlying provider; router can sort by price, throughput, or latency | Fallbacks, max_price, preferred latency/throughput, ZDR | Turns routing logic into part of the commercial offer |
| AWS Bedrock | Per-token plus batch / provisioned options | Claude Opus 4.8 $6 input / $30 output per 1M; DeepSeek v3.2 $0.62 / $1.85 in listed regions | 50% batch discount; standard / priority / reserved tiers | Enterprise incumbent can span premium and low-cost models |
| Microsoft Foundry | Serverless / managed compute / provisioned | Managed GPU families listed; many token prices obscured as $- in retained surface | ACU pre-purchase plans and resource-wide governance | Commercial model is broader platform contract, not just one API rate |
| Alibaba Model Studio | Pay-as-you-go plus credits subscription | qwen3.7-max list price $2.5 input / $7.5 output with temporary discounts; team seats from $30 to $200 per month | Night/day discounts, free quota, shared credit packs | Aggressive regional discounting and seat packaging widen the competitive set |
Illustrative prices are list or posted rates captured on 2026-07-22 and should not be read as realized net pricing.
[CP018, CP019, CP020, CP021, CP022, CP023]Capability comparison emphasizing where overlap is high and where meaningful divergence remains.
Evidence-backed qualitative labels derived from retained official docs and pricing pages only.
[CP010, CP011, CP012, CP013, CP014, CP015]3.3 Switching Costs, Lock-in, and Distribution Power
For independent inference platforms, switching costs look structurally low at the API surface and much higher at the surrounding workflow layer. The common use of OpenAI-compatible endpoints across SiliconFlow, Together, Fireworks, OpenRouter, and Alibaba means that many developers can test or swap providers with limited code change. OpenRouter's routing controls and BYOK logic go even further by encouraging multi-homing instead of hard lock-in. SiliconFlow's own BYOK and preselected-tool distribution serve a similar goal: they make adoption easier, but they also make competitive displacement easier if another provider offers better cost, throughput, or reliability. Incumbents answer that weakness with distribution power. Bedrock and Foundry do not just sell tokens; they sell inference inside larger enterprise control planes with IAM, policies, guardrails, observability, and existing procurement relationships. Alibaba can similarly lean on regional cloud infrastructure, Qwen ownership, and packaged subscriptions. These surrounding advantages can create more durable attachment than model access alone. SiliconFlow therefore needs to win where openness, neutral routing across models, local chip heterogeneity, or China-specific deployment fit outweighs the convenience of staying inside a hyperscaler's broader estate.[CP024, CP025, CP026, CP027, CP028, CP029]
3.4 Moat Durability and Adverse Competitive Risks
The adverse evidence is unusually important here because the entire sector is moving toward commoditization on core API access. Hello China Tech argues mainstream model-API prices have fallen more than 90% since 2023, while the filing shows that SiliconFlow's public-cloud business prioritized share and user acquisition over near-term profitability. In China, the top three token suppliers by throughput are still hyperscaler divisions, and SiliconFlow's 1.5% share leaves it materially smaller than the largest incumbents. That does not invalidate the business, but it does mean investors should be careful about treating usage growth as equivalent to durable competitive advantage. The strongest moat candidate in retained sources is not proprietary model ownership or hard customer lock-in. It is SiliconFlow's neutral position in the China stack: broad open-model access, heterogeneous chip support including domestic compute, and the ability to serve both self-serve and private-deployment workloads. The problem is that similar neutrality and portability are also selling points for Together, Fireworks, and OpenRouter, while hyperscalers can mimic many API features and underwrite them with deeper balance sheets. Competitive durability therefore depends less on surface feature checklists and more on supply access, operational efficiency, trust, and the ability to convert price-sensitive developer traffic into harder-to-displace enterprise accounts.[CP030, CP031, CP032, CP033, CP034, CP035]
| Moat claim | Threat | Severity | Why it matters | Mitigation / diligence ask |
|---|---|---|---|---|
| China-local neutral platform | Hyperscalers replicate API surface and underprice traffic | High | Low switching costs can erase feature-only advantages | Test enterprise win reasons beyond raw price |
| Heterogeneous chip orchestration | Competitors improve multi-chip support or secure preferred supply | High | Supply access and unit economics are central to reliability and margin | Verify supplier contracts and domestic-chip depth |
| Broad model access | Catalog breadth commoditizes quickly across routers and clouds | Medium | Hundreds of models alone do not create durable lock-in | Measure actual usage concentration by model family |
| Developer distribution through tools | BYOK and compatibility also enable multi-homing away from SiliconFlow | Medium | Channels boost top-of-funnel but may weaken retention | Request cohort retention by acquisition channel |
| Enterprise/private deployment path | Incumbent clouds can bundle governance, network, and procurement | High | Bundling can overpower neutral-platform advantages in big accounts | Probe security posture and procurement wins against clouds |
| Independent positioning | Price wars and negative-margin public cloud compress the whole category | High | Market growth may not translate into profit-pool growth | Benchmark gross margins and promotional-credit discipline versus peers |
The risk register focuses on durability, not feature checklists.
[CP029, CP030, CP031, CP032, CP033, CP034]Compact view of what matters most in competitive durability for SiliconFlow.
[CP002, CP024, CP029, CP032, CP034, CP042]3.5 Exhibits
04Financials
4.1 Revenue Model and Pricing Architecture
SiliconFlow's filing shows two economically different businesses under one brand. Public cloud-based services include serverless token services and dedicated instances; on-premise deployment solutions install inference software in customer environments. Serverless services are prepaid, low-priced, consumption-driven products aimed at developers and smaller customers, while dedicated instances and private deployments serve larger buyers that need stability, reserved capacity, or compliance control. The public models page and API docs make the commercial surface visible: list pricing is usage-based at the model level, access is API-key driven, and the company relies on rapid onboarding and broad model choice to drive adoption. The problem is that list pricing is only the top of the revenue story. Peer pricing pages from Together, Fireworks, AWS, and Alibaba show the same market logic: cached tokens, batch jobs, dedicated capacity, and fine-tuning or deployment charges all change realized economics relative to simple per-token list rates. SiliconFlow therefore operates in a market where customers are trained to expect pay-as-you-go flexibility and frequent discounts, while heavy users can often migrate toward more cost-efficient dedicated capacity. That makes revenue quality highly sensitive to customer mix, effective discounts, and whether higher-value enterprise workloads eventually outweigh low-ARPU developer traffic.[CI001, CI002, CI003, CI004, CI005, CI006]
| Stream | Mechanism | Unit | Current value / status | Revenue quality | Diligence ask |
|---|---|---|---|---|---|
| Serverless token services | Prepaid, pay-as-you-go token consumption | Tokens / API calls | RMB14.3M revenue in 2025 | Low current quality: high volume, low spend density | Need cohort retention and effective realized price per token |
| Dedicated instances | Reserved computing capacity on public cloud | Instance / reserved capacity | RMB15.0M revenue in 2025 | Better than pure serverless but still compute-rental exposed | Need utilization and contract-length disclosure |
| Public cloud total | Serverless + dedicated instances | RMB revenue | RMB29.261M, 52.9% of 2025 revenue | Growth engine but negative margin | Need line-level margin bridge and enterprise mix |
| On-premise deployment | Software / deployment inside customer environment | Project revenue | RMB26.069M, 47.1% of 2025 revenue | Higher quality on margin, lower scalability | Need sales-cycle length and repeatability data |
| Model-level list pricing | API price menu on public model library | Per-token / per-unit list price | Publicly visible list prices on platform | Not equal to realized pricing or gross margin | Need discount and promotion policy by cohort |
Revenue mix is filing-backed; pricing posture is supplemented by public product surfaces and peer list-pricing evidence.
[CI001, CI002, CI003, CI004, CI005, CI009]| Vendor / surface | Price / unit / contract | List vs realized | Discount / control lever | Source / implication |
|---|---|---|---|---|
| SiliconFlow models page | Usage-based model pricing | List only | Model choice and credits affect realized spend | Confirms pay-as-you-go positioning, not realized revenue |
| Together serverless | Per-token pricing | List only | Cached-input and batch discounts; migrate to dedicated | Illustrates market pressure toward lower effective unit cost |
| Together dedicated | Per-GPU-minute / hour | List only | Autoscaling and reserved capacity | Dedicated can be cheaper at high utilization |
| Fireworks serverless | Per-token pricing | List only | Priority / Fast tiers, cached-input discount, batch 50% of standard | Operational attributes alter effective economics |
| AWS Bedrock | Per-token plus batch / provisioned options | List only | 50% batch discount; provisioned throughput | Large incumbents can span price segments |
| Alibaba Model Studio | Per-token plus seat / credit plan | List only | Regional discounting, free quota, subscription seats | Packaging broadens addressable budgets but muddies realized pricing |
Official pricing pages are list pricing only; none of them reveal realized contract economics or contribution margins.
[CI006, CI007, CI008, CI010]How token consumption and deployment choices convert into different revenue-quality profiles.
[CI001, CI002, CI003, CI004, CI039]4.2 Cost Structure and Unit Economics
The filing is unusually clear about what hurts current economics. In 2025 SiliconFlow generated RMB55.33 million of revenue but recorded cost of sales of RMB68.632 million, implying a negative gross profit and a blended gross margin of -24.0%. The public-cloud line was much worse: its gross loss margin was -119.0% in 2025 after -271.6% in 2024, while on-premise deployment remained high-margin at 82.5%. Compute rental dominated the cost structure, accounting for 86.9% of cost of sales, and the filing explicitly says the company prioritized market share, user acquisition, and ecosystem building over immediate profitability. Independent commentary sharpens the unit-economics picture. Hello China Tech and KrASIA both note that SiliconFlow is effectively a middle layer renting compute and packaging third-party models into a price-war environment. The best public proxy for monetization quality is not total registered users but spend density. Serverless paying accounts rose from 2,455 to 716,000 in one year, yet serverless token revenue was only RMB14.3 million, implying extremely small average annual spend per account. More than 64% of 2025 sales and marketing expense went to promotional compute credits. That mix can create large usage numbers quickly, but it does not yet prove attractive CAC payback or durable gross-margin recovery.[CI011, CI012, CI013, CI014, CI015, CI016]
| Metric | Value / status | Confidence | Why it matters | Diligence ask |
|---|---|---|---|---|
| 2025 revenue | RMB55.33M | High | Top-line scale anchor | Confirm 2026 run-rate and revenue recognition cadence |
| 2025 blended gross margin | -24.0% | High | Shows scale has not yet fixed economics | Bridge public-cloud vs on-prem mix effects |
| 2025 public-cloud gross loss margin | -119.0% | High | Core growth engine still loss-making | Need per-product margin path and utilization targets |
| 2025 on-prem gross margin | 82.5% | High | Higher-quality but lower-scale revenue stream | Assess repeatability and ceiling of this line |
| Compute rental share of COGS | 86.9% | High | Supply cost dominates economics | Need supplier contracts and pricing roadmap |
| Serverless paying accounts | 716,000 at 2025 year-end | High | Demonstrates acquisition scale | Need active spend distribution by cohort |
| Approx. serverless revenue per paying account | ~RMB20/year using year-end accounts as rough divisor | Medium | Suggests very low spend density | Need average monthly active payer and revenue by decile |
| Promotional credits share of S&M | >64% in 2025 | High | Customer acquisition may be subsidy-heavy | Need CAC payback and promo-credit conversion |
Uses public figures and simple cautionary proxies; the RMB20/account figure is explicitly approximate and not a management KPI.
[CI011, CI012, CI013, CI014, CI015, CI016]Why scale has not yet translated into attractive public-cloud unit economics.
[CI011, CI013, CI015, CI018, CI020, CI040]Source-backed range view for key financial inputs and rough adequacy signals.
Range mixes business-line and balance-sheet anchors to show spread in economics; not a single management scenario model.
[CI012, CI014, CI023, CI026]4.3 Capital Adequacy and Financing Dependency
Capital adequacy improved meaningfully after the 2026 financing sequence, but the business still looks financing-dependent rather than self-funding. At the end of 2025 SiliconFlow held RMB171 million of cash and cash equivalents plus RMB100 million of time deposits, while net cash used in operating activities was RMB172 million for the year. On that simple backward-looking lens, the company would not have had comfortable standalone runway without fresh capital. The post-2025 A+, B, and B+ financings disclosed in the filing changed that picture by injecting roughly RMB1.48 billion in cash consideration after year-end, which is why the question shifts from immediate survival to how efficiently that new capital can be converted into better unit economics. The use-of-proceeds and supplier sections still point to meaningful dependency on external compute supply and continuing commercialization spend. The company does not buy chips directly; it mainly leases computing resources through partners, and purchases are concentrated among a handful of major suppliers. That means liquidity needs are tied not just to R&D burn but to compute-availability strategy, promotional usage credits, and the mix between serverless traffic and higher-quality enterprise work. The retained public record does not support a clean current runway calculation as of 2026-07-22, so investors should treat capital adequacy as improved but still linked to operating execution rather than solved by the June 2026 rounds alone.[CI023, CI024, CI025, CI026, CI027, CI028]
| Item | Value / status | Why it matters | Confidence | Diligence ask |
|---|---|---|---|---|
| Cash and cash equivalents | RMB171M at 2025 year-end | Immediate liquidity base before 2026 financings | High | Confirm current unrestricted cash |
| Time deposits | RMB100M at 2025 year-end | Adds liquidity cushion but may not equal instant operating cash | High | Clarify maturity and restrictions |
| Operating cash burn | RMB172M net cash used in operations in 2025 | Shows pre-2026 business was not self-funding | High | Provide monthly 2026 burn trend |
| Post-2025 financing cash consideration | ~RMB1.48B across A+, B, and B+ | Materially extends runway relative to year-end cash | High | Confirm cash received net of fees and restricted uses |
| Compute-supply dependency | Leases compute rather than buying chips; supplier concentration is high | Runway depends on supplier terms as much as cash | Medium | Need payables, prepayment, and contract structure |
| Current runway as of 2026-07-22 | Not publicly supportable from retained sources | Prevents precise underwriting | High | Management should provide cash, burn, and runway at run date |
Historical funding chronology lives in Chapter 1; this table focuses on forward adequacy and dependency.
[CI023, CI024, CI025, CI026, CI027, CI028]Cash requirements are shaped by compute procurement, commercialization subsidies, and financing support.
[CI024, CI025, CI026, CI027, CI028, CI029]4.4 Public Gaps and Financial Verdict
The public record supports a strong financial caution signal but not a full underwriting model. We have revenue, gross margin, some mix data, customer-count proxies, supplier concentration, and a post-2025 financing bridge. We do not have clean net revenue retention, cohort margins, realized discount rates, customer concentration, monthly burn as of run date, or a management-backed timetable for public-cloud breakeven. Because official pricing pages across the market are list prices, not realized rates, they are useful for competitive context but not enough to infer actual SiliconFlow contribution margins. The financial verdict is therefore mixed. Revenue growth and user acquisition show demand is real. On-premise deployments appear economically healthier than public cloud. But the current core growth engine—public-cloud token supply—still looks subsidy-heavy, compute-rental-heavy, and margin-negative. For investors, the most important diligence blocker is not whether there is a large market, but whether SiliconFlow can translate its token-factory scale into materially better revenue quality and structurally improved gross margins before the next capital-reliance cycle begins.[CI031, CI032, CI033, CI034, CI035, CI036]
| Missing private metric | Impact | Why missing matters | Exact diligence path |
|---|---|---|---|
| Net revenue retention / gross retention | Material | Without retention data, user growth may mask churn or weak monetization | Request retention by developer, enterprise, and private-deployment cohorts |
| Realized effective price per token | Material | List prices do not reveal discounting or promotion dependence | Ask for realized net pricing by top model families |
| Customer concentration | Material | Large enterprise concentration can distort revenue quality | Request top-customer revenue share and contract terms |
| Current monthly burn | Material | Post-2026 financing adequacy cannot be underwritten from 2025 numbers alone | Ask for latest monthly cash burn and cash balance |
| Public-cloud breakeven timeline | Important | Core growth engine valuation depends on margin inflection timing | Request management operating plan and utilization milestones |
| Sales efficiency / CAC payback | Important | Promotional credits may be masking expensive acquisition | Request payback by acquisition channel and customer type |
Every missing field is directly tied to underwriting, not curiosity.
[CI031, CI032, CI033, CI034, CI035, CI036]4.5 Exhibits
05Product & Technology
5.1 Product Definition and Module Map
SiliconFlow's product is best understood as a multilayer inference delivery platform. The Chinese homepage, English docs introduction, models catalog, and API references all present a workflow where developers or enterprise teams pick a model family, obtain an API key, and then call hosted inference endpoints using usage-based pricing. That core serverless surface is surrounded by adjacent offerings—reserved instances for stable enterprise workloads, private deployment for regulated or data-sensitive buyers, and acceleration services for customers that want faster inference on self-developed or open-source models. In other words, the company is selling access, orchestration, optimization, and deployment modes around model inference, not a single proprietary frontier model. This module mix matters strategically. It lets SiliconFlow serve different buyer jobs with one control plane: fast experimentation for developers, higher-SLA reserved capacity for enterprises, and private or hybrid deployment for customers that cannot stay on a shared public cloud. The breadth is a strength because it reduces the number of separate vendors a customer must test. It is also a complexity risk because the company must keep model inventory current, maintain heterogeneous serving paths, and explain which delivery mode actually fits each workload.[CE001, CE002, CE003, CE004, CE005, CE006]
| Module / asset | Primary user | Status / maturity | Differentiation | Diligence gap |
|---|---|---|---|---|
| Serverless model APIs | Developers, startups, enterprise builders | GA / publicly documented | Broad model catalog with pay-as-you-go access | Need realized uptime / latency / error-rate metrics |
| Reserved instances | Enterprise platform teams | GA / homepage-promoted | Dedicated capacity and cost optimization for core inference workloads | Need SLA terms and utilization economics |
| Inference acceleration service | Model builders, infra teams | GA / homepage-promoted | Performance optimization for self-developed or open-source models | Need independent benchmarks outside company claims |
| Private deployment | Regulated or data-sensitive enterprises | GA / homepage-promoted | BYOC, isolation, and private deployment options | Need deployment references and security attestations |
| Models catalog / control plane | All users | GA / publicly visible | Single surface across text, speech, image, video, and multimodal models | Need lifecycle / deprecation policy and migration tooling detail |
Status is public-surface verified, not independently audited production maturity.
[CE001, CE002, CE003, CE004, CE006]| User job | Current workflow | SiliconFlow solution | Measurable benefit | Limitation |
|---|---|---|---|---|
| Prototype an AI app quickly | Select model, create API key, call endpoint | Serverless API catalog | Fast experimentation without GPU ops burden | Cost / performance may change as model roster changes |
| Run stable enterprise inference | Move from bursty demand to predictable load | Reserved instances | Capacity control and stronger performance predictability | Need public SLA and support-package details |
| Deploy in private environment | Keep data or workloads off shared public cloud | Private deployment / BYOC | More control over privacy and compliance posture | Implementation complexity and reference depth unclear |
| Optimize custom or open-source model serving | Tune runtime for speed and cost | Acceleration service | Lower latency and better hardware utilization | Independent benchmark coverage is limited |
| Switch across model families for specific tasks | Compare model capabilities by task | Broad catalog across modalities | Reduced integration switching cost | Platform still depends on upstream model availability |
Use-case mapping focuses on customer workflow, not company marketing categories.
[CE005, CE007, CE008, CE009]SiliconFlow combines customer-facing APIs, deployment modes, optimization software, and underlying compute into one inference-delivery stack.
Layer values are qualitative weights representing relative functional scope within the public architecture, not resource allocation or revenue mix.
[CE001, CE003, CE010, CE022]5.2 Architecture and Developer Workflow
The public technical surface is unmistakably API-first. The chat-completions reference exposes model selection, streaming, context-window controls, tool calling, and tracing headers. The broader docs introduction describes pay-as-you-go API access across text, image, speech, video, vector, reranking, and multimodal models. The workflow implied by these sources is simple for users: register, create an API key, choose a model from the catalog, integrate against a largely OpenAI-style request pattern, and then decide later whether to remain on shared serverless, migrate to reserved instances, or move into private deployment. That low-friction path is central to why SiliconFlow can aggregate so many models under one platform. Under the hood, the platform is not only an API broker. SiliconFlow repeatedly emphasizes self-developed efficient operators, optimization frameworks, acceleration engines, dynamic scaling, monitoring, and fault tolerance. The open-source OneDiff project provides the clearest external engineering proof that the company actually builds inference-optimization software rather than only packaging other vendors' models. OneDiff focuses on acceleration libraries, compiler backends, and optimized kernels for diffusion models, which does not prove the entire platform architecture but does support the broader claim that inference performance engineering is an in-house competency. The technical story is therefore more credible than pure marketing copy, though still stronger on acceleration tooling than on independently audited reliability metrics.[CE010, CE011, CE012, CE013, CE014, CE015]
| Layer / component | Role | Dependency | Risk |
|---|---|---|---|
| API gateway and auth | Entry point for application calls | Developer credentials and request management | Customer-facing outages or auth failures hit all modules |
| Model catalog / routing layer | Maps requests to available hosted models | Upstream model providers and version changes | Frequent model churn can force migration work |
| Inference acceleration layer | Improves latency / throughput / cost | Internal optimization software and GPU-specific tuning | Performance claims may not generalize across workloads |
| Compute orchestration and scaling | Allocates shared or dedicated capacity | Leased GPU resources and autoscaling logic | Capacity shortages or cost spikes affect margins and reliability |
| Monitoring / traceability | Operational debugging and service assurance | Logging, trace ids, observability stack | Lack of public SLO reporting reduces outside verification |
Architecture is public-source grounded and avoids unsupported hidden-layer speculation.
[CE010, CE011, CE012, CE014, CE023]Typical path from evaluation to scaled production use on SiliconFlow.
[CE005, CE011, CE013, CE017]The product relies on upstream model, compute, and open-source optimization dependencies.
[CE014, CE018, CE022, CE023, CE025]5.3 Differentiation, Roadmap, and Dependencies
SiliconFlow's main product differentiation is combinational: breadth of model access, fast integration, multiple deployment modes, and a performance/cost optimization narrative. The homepages cite 10x+ speed improvement for language models, 1-second image generation, and double-digit cost savings for several scenarios; the docs emphasize large model coverage; and the product menu spans public cloud, reserved instances, acceleration services, and private deployment. That bundle is more differentiated than a simple model marketplace because it gives the company room to compete on service quality and deployment flexibility, not only on catalog breadth. At the same time, the dependency map is substantial. SiliconFlow depends on upstream model providers staying available, on continued access to leased GPU capacity, on developer trust in the API abstraction, and on its own ability to keep pace with a rapidly changing model roster. The API docs explicitly warn that model availability and capabilities will change over time and that some services are still being updated. The Kimi K3 launch post shows the platform can add new models quickly, but it also underlines that roadmap execution is partly a model-onboarding race rather than only a deep proprietary R&D race. That creates a product moat that is real in operations and integration quality, but potentially more fragile than the moats of model creators or hyperscalers with first-party compute and distribution.[CE020, CE021, CE022, CE023, CE024, CE025]
| Control / metric | Status | Scope | Gap |
|---|---|---|---|
| BYOC deployment support | Publicly claimed | Enterprise / private deployments | Need architecture review and customer reference |
| Compute / network / storage isolation | Publicly claimed | Data security controls | Need third-party attestations |
| Privacy policy | Publicly available | International platform data handling | Need DPA / subprocessors / retention detail |
| Separate terms for China vs international surface | Publicly available | Jurisdiction and contracting structure | Need legal-entity-by-customer mapping |
| Trace identifiers in API responses | Publicly documented | Support and troubleshooting | Need public incident / status history |
| Industry-standard compliance claim | Marketing-level claim | Enterprise positioning | Need named certifications and audit reports |
Public trust signals exist, but assurance evidence is incomplete.
[CE029, CE030, CE031, CE032, CE033]| Date / stage | Feature / milestone | Status | Implication | Source |
|---|---|---|---|---|
| 2026-07-22 current | Extensive model catalog with multimodal coverage | Live | Breadth supports one-platform positioning | Docs / models surfaces |
| 2026-07-22 current | Reasoning and thinking-mode parameters for supported models | Live | API surface keeps pace with newer model behaviors | Chat API reference |
| 2026-07-22 current | Reserved instance and private deployment offerings | Live | Signals enterprise packaging beyond developer APIs | Chinese homepage |
| 2026-07-22 current | Rapid onboarding of Kimi K3 / GLM-5.2 and similar launches | Live / recent | Roadmap execution partly depends on fast partner/model integration | Official blog / homepage |
| 2024 open-source release cadence | OneDiff acceleration library releases | Historical but relevant | Shows continuing engineering work in inference optimization | GitHub releases |
Roadmap is inferred from public release behavior; no full changelog or committed forward roadmap was found.
[CE020, CE021, CE024, CE026, CE028]Publicly visible maturity and verification quality across SiliconFlow capability areas.
[CE020, CE027, CE034, CE035]5.4 Trust, Safety, Security, and Quality Controls
SiliconFlow does publish meaningful trust and compliance signals, but they are not yet the same as comprehensive enterprise assurance. The Chinese homepage cites data isolation across compute, network, and storage, support for BYOC deployment, and compliance with industry standards and regulatory requirements. The privacy policy and terms confirm that the company has distinct China and international service surfaces, which matters for data handling and jurisdictional separation. The public API surface also includes trace identifiers that aid issue troubleshooting. Together these are real operational signals that the platform is built for supportability rather than only demo usage. However, the retained source set does not provide a public SOC 2 report, ISO certification list, public uptime history, security incident history, detailed DPA package, or benchmarked error-rate / latency SLOs. The gap is important because SiliconFlow is selling critical inference infrastructure to enterprise workloads. For diligence, the relevant question is not whether the company knows security matters—the website clearly says it does—but whether it can furnish the third-party attestations, reliability reports, and governance artifacts that large enterprise buyers will expect before standardizing on the platform.[CE029, CE030, CE031, CE032, CE033, CE034]
5.5 Exhibits
06Customers
6.1 Customer Segmentation and Adoption Surfaces
SiliconFlow's customer base is not one homogeneous SaaS list. The public sources point to at least four practical segments: developers building directly on the API, enterprise platform teams using reserved or private deployments, telecom / compute partners that embed the inference stack into regional infrastructure, and community projects that route external user demand through SiliconFlow-hosted models. The QQ financing coverage and independent reports give the broadest top-line adoption markers—more than 10 million users and 10,000 enterprise customers—while the filing and case studies show that enterprise use is concentrated in compute-intensive, infrastructure-like scenarios rather than lightweight chatbot experiments alone. The adoption surfaces also matter because they imply different monetization and durability profiles. Community and developer integrations such as MindSearch, Continue, and Cline prove API relevance and onboarding ease, but they are weak proxies for high-ARPU contracts. In contrast, Guizhou Mobile cooperation, reserved-instance case studies, and private-deployment stories show deeper operational embedding, but often without public contract values or renewal terms. That split means SiliconFlow clearly has demand breadth, yet investors still need to separate proof of usage from proof of durable, high-quality revenue.[CU001, CU002, CU003, CU004, CU005, CU006]
| Segment | Buyer / user / payer | Use case | Scale signal | Revenue / strategic value | Gap |
|---|---|---|---|---|---|
| Developers and AI builders | User=developer; payer=individual/team | Prototype and launch AI apps via API | Named integrations and community guides are abundant | Broad top-of-funnel demand and ecosystem relevance | Need active-paid-developer count and revenue share |
| Enterprise platform teams | Buyer=IT/AI platform lead; user=internal app teams | Reserved instances, private deployment, coding or knowledge workloads | Anonymous case studies plus reserved-instance example | Likely higher ARPU and stronger embedding | Need named logo roster and contract duration |
| Telecom / compute partners | Buyer=regional infra operator; user=industry customers | Token factory / inference infrastructure co-build | Guizhou Mobile partnership is explicit | Strategic distribution and supply leverage | Need commercial terms and scale of live workloads |
| Regulated / state-linked institutions | Buyer=央企 / public-sector style institutions | 国产化 private deployment and cost/performance optimization | Airline and energy央企 cases show pattern | Important for trust and domestic moat narrative | Named customer list mostly withheld |
| Community tools / open-source apps | User=end developers or end users of third-party tools | Search/RAG, coding, translation, agent workflows | MindSearch, Continue, Cline, and broader usercase hub | High awareness and repeated usage surface | Not enough evidence on monetization or exclusivity |
Segments distinguish buyer, user, and payer roles rather than treating all “customers” as equal.
[CU001, CU003, CU004, CU005, CU006]| Metric | Value | Date | Source | Confidence | Implication | Missing denominator |
|---|---|---|---|---|---|---|
| Users served | 10M+ users | 2026-06-16 | QQ financing coverage / official disclosure | Medium | Very broad adoption surface | Unknown monthly active share |
| Enterprise customers | 10,000+ enterprise customers | 2026-06-16 | QQ financing coverage / official disclosure | Medium | Enterprise reach is materially larger than a small pilot roster | Unknown production-grade share |
| Serverless paying accounts | 716,000 year-end accounts | 2025-12-31 | IPO filing / coverage | Medium | Self-serve monetization exists at scale | Unknown active monthly payer rate |
| Reserved-instance example demand | Single customer at 100B daily tokens | 2026 | Official case study | Medium | Some workloads are large enough to justify dedicated capacity | Anonymous customer limits segmentation |
| Guizhou Mobile strategic upgrade | Deepened 2026 agreement after 2025 cooperation | 2026-06-10 | Official partnership announcement | High | Institutional relationship appears multi-phase rather than one-off | Commercial value undisclosed |
| Community usercase breadth | Multiple named integration guides across search, coding, translation, and RAG | 2026-07-22 | Official docs usercase hub | High | API relevance spans many practitioner workflows | Unknown how many convert into paid production use |
This table mixes broad customer-reach markers with deeper deployment signals to show funnel layers, not a single conversion chain.
[CU002, CU007, CU010, CU013, CU020, CU029]SiliconFlow customer motion starts with easy API access and can expand into dedicated or joint-infrastructure relationships.
Stages are supported by retained sources but public conversion rates between stages are unavailable.
[CU003, CU013, CU020, CU032]Public proof narrows from broad user and enterprise counts to the much smaller subset with deep deployment evidence.
The funnel mixes counts from different layers of proof. The final zero means no public retention cohort disclosure, not zero retention.
[CU002, CU007, CU011, CU022]6.2 Named Customer and User Proof
The strongest public named proof comes from two channels. First are enterprise and ecosystem partnerships, especially Guizhou Mobile, where SiliconFlow publicly describes a deepened strategic collaboration around inference frameworks, token services, and joint operating systems. Second are developer-facing user cases where named tools or open-source projects publish integration guides around SiliconFlow APIs. MindSearch, Continue, and Cline are especially relevant because they show the platform being used in search/RAG and coding-agent workflows that are naturally token-intensive and repeat usage-heavy. These proofs are not all equal. Guizhou Mobile is a named institutional counterparty with executive quotes and a concrete operating scope, which is much stronger evidence than a how-to guide. The MindSearch, Continue, and Cline pages prove awareness and practical integration, but they do not disclose paid conversion, scale, or exclusive commitment. Anonymous enterprise case studies in aviation, energy, and coding agents add depth on deployment patterns and outcomes, yet their anonymity limits concentration analysis. The correct reading is therefore that named proof exists and spans both developer and enterprise surfaces, but only part of it qualifies as strong commercial evidence.[CU010, CU011, CU012, CU013, CU014, CU015]
| Customer / project | Segment | Deployment / use case | Production vs pilot | Outcome / proof | Limitation |
|---|---|---|---|---|---|
| Guizhou Mobile | Telecom / regional AI infrastructure | Inference framework deployment, token services, joint operations | Production / strategic cooperation evidence stronger than pilot | Official signed agreement with executive quotes and multi-workstream scope | Revenue, deployment count, and renewal terms undisclosed |
| MindSearch | Developer / search-RAG project | Integration with SiliconFlow API and deployment guide | Practical integration proof | Official how-to guide with config and deployment steps, including HuggingFace Space path | Does not prove paid production scale or exclusivity |
| Continue | Developer / IDE coding assistant | VS Code / JetBrains integration with SiliconFlow-hosted models | Practical integration proof | Official guide maps real model selection, caching, and verification workflow | No disclosed conversion or spend data |
| Cline | Developer / coding agent | OpenAI-compatible API integration for coding-agent workflows | Practical integration proof | Official guide details base URL, model IDs, and multi-mode setup | No customer-count or revenue disclosure |
Named proof spans one institutional partner and multiple developer-tool integrations; this is useful evidence of adoption breadth, but it is not a substitute for a named enterprise revenue roster.
[CU011, CU014, CU015, CU016, CU017, CU018]Institutional partnership proof is stronger than classic retention visibility; developer-tool proof is broad but lighter on commercial weight.
Labels reflect the combination of naming, deployment detail, quantified outcomes, and revenue-proximity.
[CU014, CU018, CU019, CU023, CU030, CU033]6.3 Durability, Expansion, and Concentration
Public sources support an expansion story more convincingly than a retention story. The reserved-instance case shows one lifecycle from pay-as-you-go usage toward locked dedicated capacity as token demand grows. The Guizhou Mobile partnership points to platform embedding deeper into regional compute and service infrastructure. The enterprise case studies show SiliconFlow moving beyond public-cloud API calls into private deployment,国产 chip adaptation, and broader operational integration. Those patterns are consistent with land-and-expand behavior: start with easy API access, then deepen into reserved instances, private deployment, or joint operations as usage becomes mission-critical. What is missing is classic SaaS durability evidence. The retained public set does not disclose NRR, GRR, churn, average contract length, renewal rates, top-customer concentration, or segment-level revenue contribution. The official and independent record also does not separate how many of the 10,000 enterprise customers are truly production-grade, how many are small-budget experimentation accounts, or what percentage of demand is mediated through a handful of strategic partners. That means customer breadth is real, but customer quality and concentration remain unresolved diligence questions.[CU020, CU021, CU022, CU023, CU024, CU025]
| Metric | Value / null | Segment | Confidence | Diligence ask |
|---|---|---|---|---|
| Public NRR | null | All paying customers | High | Request NRR by developer, enterprise, and private-deployment cohorts |
| Public GRR / logo retention | null | All paying customers | High | Request logo retention and renewal rates by segment |
| Contract duration | null | Enterprise / telecom | High | Request average term and renewal options for reserved and private deals |
| Repeat usage proof | Qualitative only: some workloads move from pay-as-you-go to reserved capacity | Enterprise heavy users | Medium | Provide conversion rate from API spend to reserved-instance deployment |
| Satisfaction / reference depth | Mixed: strong workflow detail, weak public formal testimonials | Developers and enterprise buyers | Medium | Request NPS / CSAT and customer references |
| Public adverse churn evidence | No robust churn dataset in retained corpus | All segments | Medium | Provide churn reasons and lost-logo analysis |
Null means no public disclosure in retained sources, not that the metric is irrelevant.
[CU021, CU022, CU023, CU024, CU030]| Expansion driver | Concentration risk | Impact | Diligence path |
|---|---|---|---|
| API to reserved-instance upgrade path | Unknown whether expansion depends on a small set of whales | Large revenue concentration could hide beneath broad user counts | Request spend deciles and expansion cohorts |
| Private deployment for regulated enterprises | Named proof is thin because key case studies are anonymous | Hard to assess renewal risk and sales-cycle durability | Request named references under NDA |
| Telecom / infra partnerships | Strategic partners may become concentrated routes to volume growth | Could create bargaining-power and dependency issues | Request partner contribution to bookings and compute supply |
| Community-tool integrations | High awareness may not equal durable monetization | Can inflate usage without proving high-LTV customers | Request paid conversion from community integrations |
| Cross-industry央企 adoption narrative | Official story is strong, but exact vertical mix is vague | Could overstate diversification if revenue is concentrated | Request revenue by vertical and top ten accounts |
The main uncertainty is not whether demand exists, but where revenue concentration hides inside the demand base.
[CU025, CU026, CU027, CU028, CU031]Public sources show customer formation and some expansion clues, but not true retention percentages.
This is a visibility proxy. “1” marks one clear public expansion pattern rather than a rate; zeros mean no public percentage or cohort disclosure.
[CU021, CU022, CU024, CU034]6.4 Customer Verdict and Gaps
The overall customer verdict is positive on relevance and mixed on durability. SiliconFlow has enough public proof to show that developers actively integrate the API, that institutional buyers are willing to use the company in significant inference contexts, and that at least some workloads become large enough to justify reserved capacity or private deployment. This is more than superficial logo collection. The customer narrative is supported by case-study detail, integration depth, and specific workflow framing. However, the public evidence is still not strong enough to underwrite concentration risk, expansion efficiency, or retention quality. Community user cases are useful but not equivalent to paid production logos. Anonymous enterprise stories prove pattern, not precise account value. Even Guizhou Mobile is a strategic partnership rather than a disclosed revenue contract. The diligence priority is therefore straightforward: obtain segmented active-customer counts, retained cohorts, top-customer exposure, and expansion conversion from serverless into reserved or private deployments.[CU029, CU030, CU031, CU032, CU033, CU034]
6.5 Exhibits
07Risks
7.1 Regulatory and Legal Risk
SiliconFlow sits inside multiple overlapping compliance regimes. Its China terms and privacy policy require real-name verification for many users, restrict certain application domains, commit the company to AI-generated-content labeling obligations, and place user-information storage within mainland China for domestic operations. These are not generic website boilerplate points: they directly affect onboarding friction, enterprise contracting, product design, and logging obligations. The terms explicitly state the service is not suitable for automatic control, medical information services, psychological counseling, and critical information infrastructure scenarios, which narrows some of the highest-consequence deployment categories unless customized controls exist outside the public surface. The broader Chinese policy stack is becoming denser rather than lighter. The AI labeling measures effective in 2025 require explicit and implicit labels for generated synthetic content, with logging, metadata, and platform-distribution responsibilities. SiliconFlow's domestic terms already acknowledge those rules and prohibit users from deleting or tampering with labels. That alignment is positive, but it also means compliance risk is operational: if SiliconFlow's tools or enterprise customers mis-handle labeling, metadata, or downstream distribution, the company may face administrative scrutiny even when model output generation itself is legal. The legal risk is therefore less about a single banned activity and more about the burden of continuously implementing a moving regulatory stack across domestic and international surfaces.[CR001, CR002, CR003, CR004, CR005, CR006]
| Rule / case | Jurisdiction | Status | Likelihood | Severity | Mitigation | Residual exposure | Diligence path |
|---|---|---|---|---|---|---|---|
| AI-generated content labeling obligations | China | Effective / active | Medium-high | High | Terms acknowledge labeling duties; provider can add labels and retain logs | Operational mis-implementation risk remains | Request labeling architecture and compliance audit |
| Real-name verification and identity checks | China | Active contractual / legal requirement | Medium | Medium-high | Domestic onboarding and verification procedures in privacy/terms | Friction and privacy risk remain for some users | Request KYC flow and exception handling |
| Telecom / internet service licensing | China | Active | Medium | High | ICP / telecom filings visible on domestic site | Scope of license fit for evolving services is not publicly tested here | Obtain counsel memo mapping licenses to current services |
| Sensitive-domain exclusions (medical, psych, auto-control, CII) | China / customer contract | Active contractual limitation | Medium | High | Public terms narrow scope unless custom arrangements exist | Enterprise misuse could create legal or reputational issues | Request sector-specific deployment controls and approval process |
| Cross-border legal-surface split | China vs international | Active | Medium | Medium | Separate China and international terms/privacy surfaces | Operational confusion or inconsistent controls can still occur | Review entity, data-flow, and contracting map |
Rows are ordered by investment relevance rather than by comprehensive legal taxonomy.
[CR001, CR002, CR003, CR004, CR005, CR006]Highest residual risk clusters around regulation, compute dependency, and low-visibility customer economics.
[CR006, CR015, CR024, CR029, CR037]7.2 Operational, Security, and Quality Risk
Operationally, SiliconFlow looks more like infrastructure than like a simple application layer. That raises the cost of outages, latency swings, bad deployments, and compliance failures. Case studies and product pages emphasize private deployment, heterogeneous chip adaptation, reserved instances, monitoring, fault tolerance, and cost optimization—all signals that customers are placing meaningful workloads on the platform. But the public assurance surface remains incomplete. The retained corpus does not show public uptime reports, detailed SLOs, SOC 2 or ISO attestations, or a public incident archive. That makes it hard for investors to judge whether operational maturity matches the company's market position. Security and content-governance exposure also interact. The domestic terms say the service should not be used for certain sensitive domains and require users to comply with wide content restrictions; the privacy policy explains that SiliconFlow collects logs and identity information for compliance and security purposes; and labeling-related rules preserve logs for six months in some cases. These controls can help reduce misuse, but they also expand the surface area for privacy, content, and access-control failures. For an inference platform, the relevant risk is not only whether the model hallucinates, but whether the platform can reliably enforce account controls, data handling, labeling, and operational isolation at scale while still shipping fast enough for the market.[CR012, CR013, CR014, CR015, CR016, CR017]
| Failure mode | Likelihood | Severity | Mitigation maturity | Residual exposure | Unresolved gap |
|---|---|---|---|---|---|
| Platform outage / latency instability on critical workloads | Medium | High | Medium | High | No public uptime/SLO record in retained corpus |
| Security-control gap versus enterprise expectations | Medium | High | Low-Medium | High | No public SOC 2 / ISO / incident archive identified |
| Privacy / log-handling failure | Medium | High | Medium | Medium-high | Domestic privacy policy shows significant data-handling duties |
| Labeling / metadata enforcement failure | Medium | Medium-high | Medium | Medium-high | Need proof of pipeline-level explicit and implicit labeling controls |
| Quality or model-lifecycle disruption from rapid model churn | Medium-high | Medium | Medium | Medium | API docs explicitly warn model availability can change |
This register focuses on infrastructure-like failure modes rather than abstract model-risk discourse.
[CR012, CR013, CR014, CR015, CR016, CR017]Shows how regulation and operations transmit into revenue quality, margin, financing, and valuation.
[CR007, CR016, CR030, CR037]7.3 Dependency, Financial, and Competitive Risk
The most acute non-regulatory risk is external dependency. SiliconFlow does not own a first-party frontier model franchise or a first-party global hyperscale cloud. Its model breadth depends on upstream model partners staying relevant and available; its gross margin depends on access to external compute; and its supply resilience depends on both domestic chip adaptation and continued ability to source or lease advanced semiconductor capacity. The company's own filing and independent analyses show compute rental dominating cost of sales and major supplier concentration remaining high. That makes the business vulnerable to both cost spikes and strategic bargaining from suppliers or strategic channels. Geopolitics amplifies the problem. U.S. export-control guidance continues to tighten around advanced computing semiconductors and diversion to PRC-linked entities, while China simultaneously increases domestic governance over AI-generated content and internet services. SiliconFlow may adapt with domestic chips, private deployment, and token-factory operating models, but those are mitigations, not immunity. Competitive pressure compounds the exposure: the same market that validates SiliconFlow also trains buyers to compare it against hyperscalers and other inference providers on price, reliability, and integration convenience. If price war, labeling compliance costs, and compute scarcity all intensify together, SiliconFlow's weakest link becomes margin and capital adequacy rather than simple demand generation.[CR022, CR023, CR024, CR025, CR026, CR027]
| Dependency | Counterparty / class | Role | Concentration | Failure scenario | Severity | Mitigation | Residual exposure |
|---|---|---|---|---|---|---|---|
| Upstream model providers | Model labs / OSS ecosystems | Catalog breadth and user demand | Medium-high | Model access or competitiveness weakens quickly | High | Multi-model routing and fast onboarding | Still lacks first-party model control |
| Leased compute suppliers | GPU / cloud / infrastructure partners | Core service delivery | High | Cost spikes or capacity constraints compress margin and reliability | High | Domestic-chip adaptation and partner network | Supply bargaining power remains external |
| Strategic telecom / infra partners | Guizhou Mobile and similar channels | Distribution plus compute coordination | Medium | Partner-led growth becomes concentrated or politically exposed | Medium-high | Joint operations and regional ecosystem strategy | Economic contribution mix is not public |
| Developer ecosystem surfaces | Continue / Cline / MindSearch and community tools | Demand generation and usage growth | Low-Medium | Awareness fails to convert into paid durable accounts | Medium | Low-friction API and many models | Commercial conversion data absent |
| Regulatory dependencies | CAC / MIIT / local compliance regimes | Operating permission and content governance | High | Rule changes force workflow redesign or stricter audits | High | Terms, privacy controls, labeling alignment | Rules continue to evolve |
Dependencies mix commercial, technical, and regulatory counterparties because all can transmit into growth and margin.
[CR022, CR023, CR024, CR025, CR032]| Role / function | Dependency or gap | Likelihood | Severity | Mitigation | Diligence path |
|---|---|---|---|---|---|
| Founder / senior technical leadership | Inference infrastructure and regulatory navigation remain founder-heavy | Medium | Medium-high | Recent financing broadens resources | Assess second-line operating bench |
| Compliance / legal operations | Must map fast-moving AI, privacy, telecom, and labeling rules into products | Medium-high | High | Public terms show awareness | Request compliance org chart and escalation process |
| Enterprise delivery / support | Private deployment and reserved instances require strong implementation discipline | Medium | High | Case studies imply experience | Request deployment timelines, staffing, and support SLAs |
| Supplier / capacity planning | Margin and service quality depend on accurate capacity and supplier management | High | High | Partner relationships and domestic adaptation | Request procurement governance and contingency plans |
Execution risk is tied more to operating depth than to raw headcount size.
[CR026, CR027, CR028, CR033]SiliconFlow depends on regulators, compute suppliers, model providers, and enterprise channels at the same time.
[CR023, CR024, CR025, CR027, CR032]7.4 Mitigations, Monitoring, and Kill Triggers
The retained evidence does show mitigation paths. SiliconFlow has distinct domestic and international legal surfaces, real-name and content controls, BYOC / private-deployment options, domestic-chip optimization narratives, reserved instances, and joint operations with compute partners like Guizhou Mobile. Those mechanisms can reduce some of the most obvious risks: private deployment can help customers with data localization concerns, labeling obligations are now reflected in terms, and dedicated capacity can stabilize heavy workloads. The question is whether execution quality will keep pace with regulatory complexity and commercial scaling. From an investment perspective, the right kill criteria are observable. A negative regulatory event around AI-content labeling or telecom compliance, evidence of export-control-induced compute bottlenecks, a failure to improve public-cloud margin, or proof that enterprise customer concentration is much higher than implied by broad user counts would each materially damage the thesis. In contrast, the thesis strengthens if SiliconFlow can show audited security controls, stable supplier diversification, clear enterprise retention, and continued migration of high-volume accounts into higher-quality dedicated or private-deployment contracts.[CR032, CR033, CR034, CR035, CR036, CR037]
| Risk | Monitorable trigger | Threshold / event | Action implication |
|---|---|---|---|
| China regulatory compliance drift | Regulator notice, enforcement, or required remediation | Formal action tied to labeling, privacy, or telecom compliance | Pause / re-underwrite compliance readiness |
| Compute-supply shock | Capacity shortage, material supplier loss, or cost spike | Service degradation or gross-margin re-deterioration | Re-cut downside case and runway assumptions |
| Customer-quality shortfall | Weak enterprise retention or whale concentration | Retention materially below expectation or top-customer share unexpectedly high | Reduce conviction and pressure-test valuation |
| Operational maturity gap | Security incident or inability to furnish enterprise assurance artifacts | Major incident or failed enterprise audit | Delay or avoid underwriting enterprise moat |
| Price-war / competition escalation | Persistent margin pressure with no mix improvement | Public-cloud losses fail to narrow despite scale | Treat scale as low-quality growth |
Kill criteria are designed to be observable in diligence or early post-investment monitoring.
[CR034, CR035, CR036, CR037, CR038, CR039]7.5 Exhibits
08Valuation
8.1 Recommendation, Risk Rating, and Price Discipline
SiliconFlow clears the threshold for strategic relevance but not for conviction buying at the publicly described June 2026 round price. The company has several features investors want: a real product, a large market, developer adoption, enterprise/infrastructure use cases, a visible role in China's inference stack, and substantial new capital. But the same public record also shows a business with negative public-cloud gross margins, supplier dependence, limited retention visibility, and meaningful regulatory / execution complexity. That mix argues for price sensitivity. The central valuation conclusion is therefore not that SiliconFlow is low quality; it is that the evidence quality lags the price. If the $1.2B valuation is buying access to a future high-margin infrastructure platform with defensible enterprise retention and improving supply economics, it may still work. If it is buying a scale story that remains subsidy-heavy and exposed to price wars, then the current entry point offers limited margin of safety. On public evidence alone, that is a Track recommendation rather than Buy.[CV001, CV002, CV003, CV004, CV005, CV006]
| Recommendation | Confidence | Risk rating | Valuation stance | Decision implication |
|---|---|---|---|---|
| Track | Medium-low | High | Price-sensitive / not enough public support for Buy at $1.2B | Continue diligence; participate only with stronger proof or better entry terms |
Recommendation is based on public evidence only and is intentionally price-sensitive.
[CV001, CV002, CV003, CV040]| Argument | What would change the view |
|---|---|
| SiliconFlow is a strategically relevant China-based inference platform with broad model access, enterprise deployment options, and ecosystem momentum. | Upgrade if 2026 revenue quality, retention, and margin trend are stronger than public evidence currently shows. |
| Recent financing and strategic investors could strengthen distribution and ecosystem leverage. | Upgrade if investor/customer overlap demonstrably converts into durable enterprise revenue. |
| Public-cloud economics, compute dependence, and regulatory complexity make the current evidence base too weak for a buy call. | Downgrade further if public-cloud margins stay deeply negative or supplier concentration worsens. |
| Customer breadth is real, but retention and concentration are opaque. | Upgrade if cohort retention and spend concentration are comfortably better than implied by current gaps. |
Each row is written as a falsifiable investment argument rather than a generic pro/con list.
[CV010, CV011, CV012, CV013, CV031]8.2 Thesis, Anti-thesis, and Current Valuation Context
The thesis is straightforward: SiliconFlow is one of the few China-based inference platforms with enough breadth, speed, and ecosystem momentum to matter. It combines model access, API distribution, private deployment, and domestic chip adaptation in a market where inference is becoming the economic center of AI deployment. The recent financing also appears strategically rich, with industry investors that could matter for distribution and ecosystem pull. This is what gives the company optionality well beyond a narrow API arbitrage play. The anti-thesis is equally straightforward. Public-cloud economics remain poor, compute rental dominates costs, customer-retention quality is opaque, and the platform is exposed to both domestic AI governance and external semiconductor constraints. Compared with fast-growing Western inference peers, SiliconFlow's disclosed revenue scale is much smaller, and its public evidence on gross-margin quality is weaker. That means the company can be strategically important and still be a mediocre risk-adjusted entry at the current price. The disclosed $1.2B valuation is not obviously absurd in absolute terms, but it is not obviously supported either without a much stronger 2026 revenue and margin picture.[CV010, CV011, CV012, CV013, CV014, CV015]
| Scenario | Assumptions | Valuation / return logic | Key risks | Probability signal |
|---|---|---|---|---|
| Bull | Enterprise retention proves strong; dedicated/private mix rises; public-cloud gross losses narrow sharply; strategic channels deepen. | Exit value ~US$2.0B–3.0B over 3-4 years; current price can still work if margin inflection is real. | Execution still depends on compute supply and regulatory control. | Possible but not best-supported by current evidence. |
| Base | SiliconFlow remains important, grows, and avoids severe disruption, but revenue quality and margin path improve only gradually. | Value range ~US$1.1B–1.6B; current entry offers limited upside for the risk. | Price war, compliance cost, and mix quality keep returns muted. | Best-supported by current public evidence. |
| Bear | Scale does not cure economics; concentration or regulation bites; next funding occurs under weaker terms. | Value range ~US$0.7B–1.0B; down-round or flat outcome plausible. | Supplier shocks, losses, and weak retention compress the case. | Must be taken seriously given public gaps. |
Ranges are scenario-based public estimates, not management forecasts or a DCF.
[CV020, CV024, CV025, CV026, CV027, CV028]| Comparable | Metric | Multiple / valuation / status | Relevance | Limitation |
|---|---|---|---|---|
| Together AI | 2026 valuation and bookings | US$8.3B valuation; >US$1.15B annual bookings | Direct open-model inference and infrastructure peer | Bookings not the same as recognized revenue; different geography and scale |
| Fireworks AI | 2026 valuation and annualized revenue | US$17.5B valuation; >US$1B annualized revenue | Strong direct inference-cloud peer with open-model focus | Much larger revenue scale and stronger disclosed monetization |
| Baseten | 2026 valuation and annualized revenue | US$13B valuation; ~US$600M annualized revenue | Inference-platform peer with enterprise deployment focus | Sacra estimates and US enterprise profile differ from China context |
| CoreWeave | 2025-2026 revenue and valuation context | US$23B pre-IPO valuation reference; US$12B–13B 2026 revenue guide | Shows ceiling and risks of AI compute infrastructure | More capital-intensive GPU cloud than SiliconFlow; not a pure inference API peer |
The table is intentionally partial because public China-specific pure-play inference comparables with direct valuation disclosure remain limited.
[CV014, CV021, CV022, CV023, CV029]8.3 Bull / Base / Bear Scenarios, Comparable Set, and Return Logic
The peer set matters because it highlights a gap between strategic category value and current operating proof. Together AI, Fireworks AI, and Baseten all command multi-billion-dollar valuations, but public sources also show them at meaningfully higher revenue or bookings scale, with more explicit evidence around enterprise monetization. CoreWeave proves that infrastructure-adjacent AI platforms can become extremely valuable, but it also shows how capital intensity and customer concentration can remain severe even at much larger scale. SiliconFlow's lower headline valuation than these Western peers therefore should not be read automatically as cheap. Lower quality of disclosed economics can still justify a discount. This leads to three scenarios. In the bull case, SiliconFlow converts its scale and strategic relationships into better enterprise retention, stronger dedicated/private mix, and improving margins, allowing the business to outrun its current risk load. In the base case, the company remains important but operationally messy, with enough traction to defend the current round but not enough proof to produce compelling venture-style upside from that price. In the bear case, compute dependence, regulation, and price competition prevent margin inflection, producing either a flat valuation outcome or a future down-round. Public evidence today supports the base case more than the bull case.[CV020, CV021, CV022, CV023, CV024, CV025]
| Trigger | Threshold | Transmission to thesis | Action implication |
|---|---|---|---|
| Regulatory enforcement | Formal action or remediation tied to AI labeling, privacy, or telecom compliance | Undermines execution and enterprise trust | Pause / avoid until resolved |
| Margin failure | No meaningful improvement in public-cloud economics through next observable period | Scale narrative fails to convert into business quality | Move from Track to Avoid at current price |
| Supplier / compute shock | Material capacity disruption or worsening supplier concentration | Threatens reliability and gross margin simultaneously | Re-cut downside case aggressively |
| Customer-quality disappointment | Retention or concentration is materially worse than hoped | Destroys the case for premium infrastructure valuation | Do not pay growth premium |
| Diligence miss on security / assurance | Cannot produce enterprise-grade assurance package | Weakens high-value enterprise adoption thesis | Constrain upside case and sales-quality assumptions |
These triggers are observable and investment-actionable.
[CV032, CV033, CV034, CV035, CV039]8.4 Final Diligence Asks and Thesis-break Triggers
The most important thing missing from the public file is not market demand but underwriting detail. Investors need current cash and burn, 2026 revenue run rate, gross-margin trend by product line, enterprise retention, top-customer concentration, supplier concentration, and conversion from serverless API usage into higher-quality dedicated or private deployments. Without those inputs, the valuation discussion becomes mostly narrative. That is why the final call is Track with high risk and medium-low confidence. The thesis can move up with evidence: improved margins, diversified compute supply, named enterprise references, audited security posture, and better revenue-quality metrics. It can move down with evidence too: regulatory action, supplier shocks, continued negative public-cloud economics, or proof that usage breadth still fails to convert into durable retained spend. A price cut alone would help, but only if the missing diligence items also stop pointing to structural weakness.[CV031, CV032, CV033, CV034, CV035, CV036]
| Topic | Missing evidence | Why it matters | Owner / diligence path |
|---|---|---|---|
| 2026 revenue run-rate and mix | Current run-rate by serverless, dedicated, and private deployment | Determines whether the round is being underwritten on improving quality or on raw growth | Management / finance diligence |
| Retention and concentration | NRR, GRR, churn, top-customer share, spend deciles | Core question for revenue durability | Management / customer diligence |
| Current cash and burn | Run-date cash, monthly burn, financing proceeds net of restrictions | Determines runway and next-round risk | Finance diligence |
| Supplier concentration and contingency | Top suppliers, contract terms, domestic-chip fallback economics | Directly affects reliability and margin risk | Ops / procurement diligence |
| Security and assurance | SOC 2 / ISO, uptime history, incident history, SLA package | Required to underwrite enterprise moat quality | Security diligence |
| API-to-enterprise expansion motion | Conversion from pay-as-you-go usage into reserved/private contracts | Shows whether bottom-up adoption compounds into better revenue quality | GTM diligence |
All asks are chosen because they directly move the investment call.
[CV036, CV037, CV038, CV040]8.5 Exhibits
Disclaimer
This report is produced for diligence and informational purposes only. It is based on publicly available materials as of 2026-07-22 and does not constitute investment, legal, accounting, or tax advice. SiliconFlow is a private company; public disclosures remain incomplete on retention, concentration, current revenue run-rate, and enterprise assurance. Readers should independently verify all facts and obtain primary diligence materials before making investment decisions.
Evidence index
| ID | Statement | Confidence | Sources |
|---|---|---|---|
| CO001 | Beijing SiliconFlow Technology Co., Ltd. was established in the PRC as a limited liability company on August 29, 2023. | Medium | SO011 |
| CO002 | SiliconFlow's registered office and head office are at Room 2301, 23/F, Tower D, Building 8, No. 1 Yard, Zhongguancun East Road, Haidian District, Beijing. | High | SO001, SO011 |
| CO003 | The company converted to a joint stock limited company on June 22, 2026. | Medium | SO011 |
| CO004 | SiliconFlow maintains a Hong Kong place of business while stating in the HKEX filing that its headquarters, senior management, and operations are primarily based outside Hong Kong. | Medium | SO011 |
| CO005 | The global .com service is governed by SiliconFlow Technology Pte. Ltd., and the global terms explicitly direct mainland-China users to use siliconflow.cn instead. | High | SO008, SO009 |
| CO006 | SiliconFlow publicly positions itself as a global AI infrastructure provider whose mission is to accelerate AGI for the benefit of all. | High | SO002, SO003 |
| CO007 | The China site presents SiliconFlow as a product suite spanning large-model APIs, reserved instances, inference acceleration services, and private deployment. | High | SO001, SO025 |
| CO008 | SiliconFlow's API documentation describes the platform as a one-stop cloud service for top-tier large-language-model APIs aimed at developers and enterprises. | Medium | SO007 |
| CO009 | The global homepage markets SiliconFlow as one platform for AI inference across text, image, video, audio, search, coding, and agent workloads. | Medium | SO002 |
| CO010 | SiliconFlow's GitHub organization says the company integrates hundreds of state-of-the-art models across language, speech, vision, and multimodal domains on top of a self-developed inference engine. | Medium | SO010 |
| CO011 | SiliconFlow's quickstart and cloud login surfaces show sign-in options via SMS, email, GitHub, and Google, alongside self-serve API key creation. | High | SO005, SO024 |
| CO012 | The chat-completions documentation exposes an OpenAI-style API surface with model selection, streaming, max_tokens, JSON-format support, and tool-related parameters. | High | SO006, SO023 |
| CO013 | As of April 30, 2026, SiliconFlow's platform had over 10 million registered users. | Medium | SO011 |
| CO014 | SiliconFlow recorded average daily token throughput of approximately 578.5 billion and peak daily throughput of approximately 1,071.4 billion in April 2026. | Medium | SO011 |
| CO015 | As of the latest practicable date in the HKEX filing, SiliconFlow had served over 13,000 enterprise customers. | Medium | SO011 |
| CO016 | As of the latest practicable date in the HKEX filing, SiliconFlow had supported a cumulative total of over 170 models. | Medium | SO011 |
| CO017 | By July 2026, SiliconFlow's public model library marketed more than 200 models, indicating platform expansion after the June 2026 filing snapshot. | Medium | SO004 |
| CO018 | The HKEX filing describes SiliconFlow as the largest independent ecosystem token supplier and one of the top five token suppliers overall in China. | Medium | SO011 |
| CO019 | The HKEX filing also describes SiliconFlow as a global leader in token throughput, registered users, and monthly active users, with globally leading overseas platform downloads and token throughput on authoritative platforms. | Medium | SO011 |
| CO020 | SiliconFlow's China product page advertises 10x-plus speed gains for language models, 66% image-model cost savings, 46% language-model cost savings, and BYOC plus isolation-based security controls. | Medium | SO025 |
| CO021 | The China homepage discloses a Beijing ICP filing and a Beijing value-added telecommunications business permit. | Medium | SO001 |
| CO022 | Dr. Yuan Jinhui is disclosed as SiliconFlow's founder, chairman, executive director, CEO, general manager, and financial controller. | Medium | SO011 |
| CO023 | Mr. Liu Juncheng is disclosed as executive director and chief technology officer. | Medium | SO011 |
| CO024 | Mr. Zeng Hua is disclosed as executive director and deputy general manager overseeing commercialization planning. | Medium | SO011 |
| CO025 | Mr. Chen Yingjie is disclosed as non-executive director and is also the managing director of Alibaba Group's strategic investment department. | Medium | SO011 |
| CO026 | Upon listing, SiliconFlow's board is expected to comprise seven directors: three executive, one non-executive, and three independent non-executive directors. | High | SO011, SO012 |
| CO027 | The HKEX history section says SiliconFlow developed under the leadership of co-founders Dr. Yuan, Liu Juncheng, Zeng Hua, Zhao Zhen, and Hu Jian. | Medium | SO011 |
| CO028 | Dr. Yuan previously served as a lead researcher at Microsoft China and later founded OneFlow, a deep-learning framework company. | High | SO011, SO018 |
| CO029 | Liu Juncheng previously worked in software development and AI system software, including as an R&D engineer at OneFlow from 2018 to 2023. | Medium | SO011 |
| CO030 | Zeng Hua previously held senior roles at Microsoft China, Baidu, and JD.com before joining SiliconFlow. | Medium | SO011 |
| CO031 | Independent profile databases also describe SiliconFlow as Beijing-based and founded in 2023. | Medium | SO020, SO021 |
| CO032 | The HKEX filing records angel, angel+, pre-A, and Series A capital injections of RMB47.20 million, RMB65.84 million, RMB71.48 million, and RMB285.99 million, respectively. | Medium | SO011 |
| CO033 | The HKEX filing records Series A+, Series B, and Series B+ capital injections of RMB220 million, RMB520 million, and RMB740 million, respectively, all in 2026. | Medium | SO011 |
| CO034 | The HKEX filing shows post-money valuation steps of RMB280.0 million, RMB565.8 million, RMB985.0 million, RMB2.286 billion, RMB3.120 billion, RMB5.020 billion, and RMB7.740 billion from angel through Series B+. | Medium | SO011 |
| CO035 | The HKEX filing says SiliconFlow completed the Series A+, Series B, and Series B+ financings after December 31, 2025 for aggregate cash consideration of approximately RMB1.48 billion. | Medium | SO011 |
| CO036 | External June 2026 financing coverage described SiliconFlow as completing over RMB2 billion of Series B financing backed by investors including Trip.com or Ctrip Strategic Investment, JinkoSolar, Kingdee, Unicom-linked capital, Biren, NIO Capital, SenseTime, GGV, and others. | High | SO013, SO014, SO015, SO016, SO017 |
| CO037 | Caixin Global said the June 2026 financing was SiliconFlow's fifth funding round since inception and that China Renaissance was the exclusive financial advisor. | Medium | SO013 |
| CO038 | Sina and 36Kr coverage said SiliconFlow served over 10 million users and 10,000 enterprise customers, grew revenue by over 10 times year-on-year, and reached millions of US dollars in overseas monthly revenue. | Medium | SO014, SO015, SO016, SO017 |
| CO039 | The HKEX filing says SiliconFlow's revenue rose from RMB7.3 million in 2024 to RMB55.3 million in 2025. | Medium | SO011 |
| CO040 | The HKEX filing says SiliconFlow's gross margin fell from 39.4% in 2024 to negative 24.0% in 2025 as cost of revenue rose faster than revenue. | Medium | SO011 |
| CO041 | The HKEX milestone section says SiliconFlow's overseas monthly revenue exceeded US$1 million in May 2026. | Medium | SO011 |
| CO042 | SiliconFlow said it started R&D on a large-model inference engine in August 2023. | High | SO017, SO011 |
| CO043 | SiliconFlow launched public-cloud MaaS in May 2024. | High | SO011, SO017 |
| CO044 | SiliconFlow launched DeepSeek inference services on Huawei Ascend in February 2025, which it described as an industry-first ultra-large-scale domestic-chip token-production service. | High | SO011, SO017 |
| CO045 | SiliconFlow launched private MaaS in September 2025 for customers with their own computing power and stricter data-compliance needs. | High | SO011, SO017 |
| CO046 | SiliconFlow launched the Elastic GPU heterogeneous-compute scheduling engine in April 2026. | High | SO011, SO017 |
| CO047 | The July 14, 2026 HKEX announcement added China Renaissance Securities (Hong Kong) as overall coordinator, signalling continued progress in the company's Hong Kong listing process. | Medium | SO012 |
| CO048 | KrASIA argued that the prospectus shows SiliconFlow is selling tokens at a loss in a compute-rental-heavy business exposed to pricing pressure. | Medium | SO018 |
| CO049 | Hello China Tech said SiliconFlow's public-cloud business carried a negative 119% gross margin in 2025 while on-premise deployment remained high-margin but harder to scale. | Medium | SO019 |
| CO050 | CB Insights lists SiliconFlow with $316.54 million total raised and a June 16, 2026 Series B of $295.84 million, a database presentation that does not fully reconcile to the HKEX B and B+ sequence. | Low | SO021 |
| CO051 | Tracxn lists Pan Yang and Jinhui Yuan as SiliconFlow's co-founders, which conflicts with the broader co-founder slate named in the HKEX filing. | Low | SO020 |
| CO052 | In July 2026, SiliconFlow announced Moonshot AI's Kimi K3 on its platform at $3 per million input tokens and $15 per million output tokens. | High | SO022, SO004 |
| CO053 | SiliconFlow's Kimi K3 API post says the serverless endpoint supports image input, tool calling, JSON Mode, streaming, and reasoning output. | High | SO023, SO006 |
| CM001 | SiliconFlow's relevant market sits in the inference middleware layer rather than in model training or end-user AI applications. | High | SM001, SM021 |
| CM002 | The included spend boundary covers serverless token APIs, dedicated instances, and private deployment of models through a unified interface. | High | SM001, SM021 |
| CM003 | Model training spend, semiconductor design, and end-user AI SaaS seats largely sit outside SiliconFlow's directly served market. | Medium | SM001, SM025 |
| CM004 | The main substitutes are hyperscaler model platforms, other independent inference APIs, direct model-lab APIs, and internal self-hosting. | Medium | SM001, SM008, SM011, SM016 |
| CM005 | The HKEX filing explicitly distinguishes independent ecosystem token supply platforms from closed ecosystems that bind users to proprietary compute and models. | High | SM001, SM024 |
| CM006 | SiliconFlow's own segmentation runs from developers and startups seeking cost-efficient on-demand access to large enterprises needing dedicated performance or supply assurance. | High | SM001, SM008 |
| CM007 | BYOK and preselected integrations reduce adoption friction by letting customers keep existing tools and workflows while switching token providers. | High | SM001, SM022 |
| CM008 | Multi-model inference platforms increasingly sell one-API access across text, image, video, code, and audio rather than single-modality access. | High | SM016, SM018, SM020, SM021 |
| CM009 | IDC says China enterprise MaaS token consumption rose from 114 trillion tokens in 2024 to 1,944 trillion tokens in 2025, roughly a 16x increase. | Medium | SM002 |
| CM010 | IDC projects China token consumption around 40,000 trillion in 2026, about 20x above 2025. | High | SM002, SM003 |
| CM011 | IDC sizes China public-cloud MaaS revenue at RMB3.07 billion in 2025. | Medium | SM002 |
| CM012 | IDC projects China public-cloud MaaS revenue reaching RMB18.6 billion in 2026. | Medium | SM002 |
| CM013 | Frost & Sullivan, as cited in the HKEX filing, says China's token supply market grew 1,602.6% from 2024 to 2025. | Medium | SM001 |
| CM014 | The same filing projects China's token supply market to reach approximately 53.2 quintillion tokens by 2030, implying a 638.3% CAGR from 2025 to 2030. | Medium | SM001 |
| CM015 | SiliconFlow held about 1.5% of China token-supply throughput in 2025, ranking fourth overall and first among independent ecosystem platforms. | High | SM001, SM024 |
| CM016 | MarketsandMarkets values the global AI inference market at USD106.15 billion in 2025 and USD254.98 billion in 2030, a 19.2% CAGR. | Medium | SM005 |
| CM017 | Grand View Research estimates the global AI inference market at USD97.24 billion in 2024 and USD253.75 billion in 2030, a 17.5% CAGR. | Medium | SM006 |
| CM018 | Fortune Business Insights sizes the global AI inference market at USD103.73 billion in 2025 and USD312.64 billion by 2034. | Medium | SM007 |
| CM019 | Independent analyst pages cluster the broad global AI inference market around roughly USD100 billion in the 2024-2025 base period. | High | SM005, SM006, SM007 |
| CM020 | Third-party global estimates agree that Asia Pacific is among the fastest-growing regions, even when they disagree on precise base-year values. | Medium | SM005, SM006, SM007 |
| CM021 | IDC's China MaaS lens and Frost & Sullivan's China token-supply lens agree on hypergrowth but do not use identical market definitions or baselines. | Medium | SM001, SM002 |
| CM022 | Buyer segments span individual developers, startups, enterprise platform teams, and regulated large organizations. | High | SM001, SM008, SM015 |
| CM023 | Dedicated performance, stable supply, latency, and private-environment deployment become important as buyers move from experiments to production-grade enterprise use. | High | SM001, SM011, SM015 |
| CM024 | AWS positions Bedrock as serving more than 100,000 organizations worldwide, from startups to global enterprises. | Medium | SM008 |
| CM025 | Microsoft Foundry packages models, agents, governance, RBAC, networking, and policy into one management plane, pointing to enterprise platform owners as the core buyer. | High | SM011, SM012 |
| CM026 | Alibaba's Token Plan shows inference demand is also being budgeted as team-productivity seats and shared credit pools rather than only as raw API consumption. | Medium | SM015 |
| CM027 | Together and Fireworks both emphasize OpenAI-compatible APIs and low-friction serverless onboarding, showing that portability is now a core expectation in this market. | High | SM016, SM018, SM020 |
| CM028 | SiliconFlow distributes through direct API integration, BYOK, and preselected tools such as LangChain, TRAE, Dify, and Cherry Studio. | Medium | SM001 |
| CM029 | Stanford's 2026 AI Index says frontier-model capability kept accelerating in 2025 and organizational AI adoption reached 88%. | Medium | SM004 |
| CM030 | Stanford also reports the U.S.-China frontier-model performance gap effectively closed by early 2026. | Medium | SM004 |
| CM031 | IDC says competition in China MaaS is shifting from pure price competition toward combined price, performance, and toolchain support. | High | SM002, SM003 |
| CM032 | IDC ranks performance, security and compliance, answer quality, platform availability, and cost effectiveness among the top enterprise selection factors, with cost only fifth for now. | Medium | SM002 |
| CM033 | Major platform vendors increasingly market routing, evaluation, observability, privacy, governance, and dedicated throughput as core product features, not extras. | High | SM008, SM011, SM012, SM020 |
| CM034 | AWS Bedrock and Microsoft Foundry both frame enterprise security, governance, and cost optimization as central to adoption. | High | SM008, SM011 |
| CM035 | Alibaba, Together, and Fireworks all offer packaging beyond simple pay-as-you-go, including seat subscriptions, dedicated endpoints, cached-token discounts, batch discounts, or hourly GPU deployment pricing. | High | SM015, SM017, SM019 |
| CM036 | The HKEX filing says SiliconFlow intentionally prioritized market share, user acquisition, and ecosystem building over immediate profitability in its public-cloud business. | High | SM001, SM024, SM025 |
| CM037 | Hello China Tech describes a severe 2023-2026 price war in model APIs, with mainstream prices down more than 90% since 2023 and further 2026 cuts from major vendors. | Medium | SM025 |
| CM038 | Compute rental, heterogeneous chip support, and supplier access remain structural constraints that matter for both service reliability and margins. | Medium | SM001, SM003, SM025 |
| CM039 | Global AI inference demand is broadening beyond text chat into edge, real-time, multimodal, and automation-heavy workloads. | Medium | SM005, SM007, SM018 |
| CM040 | Fortune Business Insights explicitly lists high hardware costs and integration challenges as adoption restraints for the AI inference market. | Medium | SM007 |
| CM041 | No retained public source isolates SiliconFlow's serviceable share by geography, customer segment, or category-specific retention. | Low | SM001, SM024, SM025 |
| CM042 | The most defensible market thesis is a stacked one: use global inference infrastructure for context, China MaaS growth for monetizable demand, and SiliconFlow throughput share only as a competitive signal. | Medium | SM001, SM002, SM005, SM006, SM007 |
| CM043 | The layered market-sizing pyramid is directional only because it mixes revenue, token-throughput, and market-share units rather than one additive denominator. | Medium | SM001, SM002, SM005 |
| CM044 | A common adoption path in this market is direct API experimentation first, then broader tool integration, then dedicated or private deployment once scale or governance requirements rise. | Medium | SM001, SM017, SM020 |
| CP001 | SiliconFlow competes across independent open inference platforms, incumbent cloud platforms, direct model-lab APIs, and internal build substitutes. | High | SP001, SP004, SP006, SP009, SP012, SP016, SP021 |
| CP002 | SiliconFlow ranks fourth in China token-supply throughput and first among independent ecosystem platforms in 2025. | High | SP001, SP023 |
| CP003 | Together AI positions itself around serverless and dedicated access to open models rather than a proprietary closed model stack. | Medium | SP012, SP013, SP014, SP015 |
| CP004 | Fireworks positions itself as a specialized training-and-inference platform for open models and says it processes 40T+ tokens per day. | Medium | SP016, SP017, SP018, SP020 |
| CP005 | OpenRouter positions itself as a unified API and broker that routes requests across hundreds of models and providers. | Medium | SP021, SP022 |
| CP006 | AWS Bedrock, Microsoft Foundry, and Alibaba Model Studio are incumbent substitutes with broader governance and control-plane depth than most independent peers. | High | SP004, SP006, SP007, SP008, SP009 |
| CP007 | Alibaba Model Studio combines official Qwen ownership with third-party model access and OpenAI-compatible APIs. | High | SP009, SP010, SP011 |
| CP008 | Retained sources here disclose far more about product scope and pricing than about current funding or revenue for Together, Fireworks, or OpenRouter. | Low | SP012, SP016, SP021 |
| CP009 | Internal build and direct model-lab APIs remain practical substitutes because many rivals expose portable, OpenAI-like integration paths. | Medium | SP013, SP020, SP021, SP025 |
| CP010 | OpenAI-compatible or drop-in API paths are common across SiliconFlow, OpenRouter, Fireworks, Together, and Alibaba. | High | SP009, SP015, SP020, SP021, SP025 |
| CP011 | Broad model-catalog competition is intense: SiliconFlow advertises 200+ models, Bedrock 100+, OpenRouter hundreds, and Foundry 1,900+ to 11,000+ access surfaces. | Medium | SP004, SP007, SP021, SP025 |
| CP012 | Together and Fireworks both offer a serverless-to-dedicated progression, while hyperscalers offer reserved or managed-compute paths for heavier production use. | Medium | SP005, SP008, SP013, SP014, SP018 |
| CP013 | Enterprise incumbents differentiate with guardrails, RBAC, policies, network isolation, observability, and managed-agent features. | High | SP004, SP006, SP007, SP008 |
| CP014 | OpenRouter differentiates on provider routing, fallbacks, price/throughput/latency sorting, and data-retention-aware controls. | Medium | SP021, SP022 |
| CP015 | Together differentiates on open-model access, a shared serverless API, and dedicated GPU deployments that reuse the same inference API. | Medium | SP012, SP013, SP014, SP015 |
| CP016 | Fireworks differentiates on serverless paths, prompt caching, on-demand infrastructure, and a combined training/inference posture. | Medium | SP016, SP018, SP019, SP020 |
| CP017 | SiliconFlow differentiates on heterogeneous compute, domestic-chip adaptation, and neutral multi-model positioning inside China. | Medium | SP001, SP003, SP025 |
| CP018 | Together lists Qwen 3.7 Max at $1.25 input and $3.75 output per 1M tokens, and DeepSeek V4 Pro at $1.74 input and $3.48 output. | Medium | SP015 |
| CP019 | Fireworks lists DeepSeek V4 Pro at $1.74 input / $0.145 cached input / $3.48 output and GPT OSS 20B at $0.07 input / $0.035 cached / $0.30 output. | Medium | SP019 |
| CP020 | AWS Bedrock pricing spans premium models such as Claude Opus 4.8 at $6 input / $30 output and lower-cost DeepSeek variants around $0.62 / $1.85 in listed regions. | Medium | SP005 |
| CP021 | Alibaba posts flagship Qwen list pricing with temporary regional discounts and also sells seat-based credit plans starting at $30 per seat per month. | High | SP010, SP011 |
| CP022 | OpenRouter makes pricing logic part of the product by letting customers sort providers by price, latency, or throughput and set max_price or performance thresholds. | Medium | SP022 |
| CP023 | Together and Fireworks both discount cached tokens and batch workloads, showing that heavy users are expected to demand lower effective pricing at scale. | Medium | SP013, SP019 |
| CP024 | API-surface switching costs are low because many rivals support OpenAI-compatible integration paths or drop-in migration. | High | SP006, SP009, SP020, SP021, SP025 |
| CP025 | OpenRouter is optimized for multi-homing rather than single-provider lock-in. | Medium | SP021, SP022 |
| CP026 | SiliconFlow's BYOK and preselected-tool distribution help adoption but also make customer multi-homing plausible. | Medium | SP001, SP025 |
| CP027 | Hyperscalers counter low API switching costs with identity, policy, networking, procurement, and broader platform integration. | High | SP004, SP006, SP007, SP009 |
| CP028 | AWS, Microsoft, and Alibaba can cross-sell inference from much broader cloud estates than independent startups can. | Medium | SP004, SP006, SP009 |
| CP029 | SiliconFlow's strongest retained moat candidate is China-local neutrality plus heterogeneous chip and deployment coverage. | Medium | SP001, SP003, SP025 |
| CP030 | Together and Fireworks compete more on speed, deployment control, and open-model operations than on exclusive model ownership. | Medium | SP013, SP014, SP016, SP018 |
| CP031 | OpenRouter's moat is routing intelligence and provider liquidity, but the model is inherently less lock-in-oriented than compute-owning platforms. | Medium | SP021, SP022 |
| CP032 | The top three China token suppliers above SiliconFlow are hyperscaler divisions, leaving SiliconFlow meaningfully smaller than the largest incumbents. | High | SP001, SP024 |
| CP033 | Hello China Tech argues mainstream model-API prices have fallen more than 90% since 2023, highlighting commoditization pressure. | Medium | SP024 |
| CP034 | SiliconFlow's filing shows that public-cloud market-share growth can coexist with negative gross margins and high compute-rental pressure. | High | SP001, SP023, SP024 |
| CP035 | Competitive risk is highest where buyers treat model access as commodity and can switch on price, latency, or uptime. | Medium | SP015, SP019, SP022 |
| CP036 | Competitive risk is lower where buyers need China-specific compute options, private deployment, or neutral multi-model support outside a single cloud. | Medium | SP001, SP009, SP025 |
| CP037 | No retained source proves customer lock-in for SiliconFlow comparable to hyperscaler IAM or network lock-in. | Low | SP004, SP006, SP009, SP025 |
| CP038 | No retained source here proves peer funding superiority or current margin superiority for SiliconFlow versus Together, Fireworks, or OpenRouter. | Low | |
| CP039 | The competitive field is crowded partly because the same buyer job can be solved by a cloud platform, a neutral platform, a router, or internal build. | Medium | SP001, SP004, SP006, SP021 |
| CP040 | The best positioning map for this market places independent platforms high on openness and hyperscalers high on enterprise control and procurement strength. | Medium | SP004, SP006, SP009, SP012, SP016, SP021, SP025 |
| CP041 | Capability overlap is highest on basic API access and model breadth, while divergence is greatest on routing, governance, and deployment-control depth. | Medium | SP006, SP009, SP014, SP018, SP022, SP025 |
| CP042 | Moat readiness in this market depends more on supply access, governance, and operating efficiency than on simple model-count marketing. | Medium | SP001, SP002, SP004, SP006, SP024 |
| CP043 | Fireworks' vendor-authored 40T+ tokens/day claim indicates that some independent peers already operate at very large throughput scale. | Medium | SP016 |
| CP044 | A common adoption path is prototype on serverless or a broker, then shift toward dedicated capacity or deeper cloud control once scale and governance requirements rise. | Medium | SP005, SP013, SP014, SP018 |
| CP045 | SiliconFlow's competitive verdict is positive on relevance but still unproven on long-term durability versus hyperscaler bundling and category-wide price compression. | Medium | SP001, SP024, SP025 |
| CP046 | Because basic API access and model breadth overlap heavily, competitive selection often shifts to surrounding control surfaces, routing logic, and deployment guarantees. | Medium | SP006, SP009, SP015, SP020, SP025 |
| CI001 | SiliconFlow has two primary revenue lines: public cloud-based services and on-premise deployment solutions. | Medium | SI001, SI004 |
| CI002 | Public cloud-based services include serverless token services and dedicated instances. | Medium | SI001 |
| CI003 | On-premise deployment solutions install inference software in customer environments and currently carry much higher gross margins than public cloud. | High | SI001, SI004 |
| CI004 | In 2025 public cloud generated RMB29.261 million, or 52.9% of total revenue, while on-premise generated RMB26.069 million, or 47.1%. | High | SI001, SI004 |
| CI005 | SiliconFlow's public pricing surface is usage-based and model-level rather than seat-based. | Medium | SI002 |
| CI006 | List pricing in the inference market is shaped by cached-token discounts, batch discounts, and migrations toward dedicated capacity at scale. | High | SI006, SI007, SI010, SI012, SI014 |
| CI007 | Together explicitly frames serverless as cheaper for low or bursty traffic and dedicated replicas as cheaper when utilization stays high. | Medium | SI006 |
| CI008 | AWS, Fireworks, and Together all offer roughly 50% batch discounts in at least part of their pricing stack. | High | SI007, SI012, SI014, SI015 |
| CI009 | Alibaba monetizes the category through both pay-as-you-go model calls and seat-based subscription credits. | High | SI019, SI020, SI021 |
| CI010 | Official list pricing is useful for market context but is insufficient to infer SiliconFlow's realized pricing or margin. | Medium | SI002, SI019, SI024 |
| CI011 | SiliconFlow reported RMB55.33 million of revenue in 2025, up from RMB7.346 million in 2024. | High | SI001, SI004, SI005 |
| CI012 | SiliconFlow's 2025 blended gross margin was -24.0%. | High | SI001, SI004 |
| CI013 | The 2025 public-cloud gross loss margin was -119.0%, after -271.6% in 2024. | High | SI001, SI004 |
| CI014 | The 2025 on-premise deployment gross margin was 82.5%. | Medium | SI001 |
| CI015 | Compute rental fees accounted for 86.9% of 2025 cost of sales. | High | SI001, SI005 |
| CI016 | SiliconFlow explicitly prioritized market share, user acquisition, and ecosystem building over immediate profitability in public cloud. | High | SI001, SI004 |
| CI017 | Serverless paying accounts increased from 2,455 to 716,000 in 2025. | Medium | SI004, SI005 |
| CI018 | Using year-end paying-account count as a crude divisor, RMB14.3 million of serverless revenue implies very low annual spend density per paying account. | Medium | SI004 |
| CI019 | More than 64% of 2025 sales and marketing expense went to promotional compute credits. | Medium | SI004 |
| CI020 | Independent analyses argue that SiliconFlow is operating as a compute-renting middle layer inside a price-war environment. | Medium | SI004, SI005 |
| CI021 | IDC says cost effectiveness matters to buyers, but it trails performance, compliance, answer quality, and platform usability as a current selection factor. | Medium | SI022 |
| CI022 | Usage metrics such as registered users and paying-account growth prove demand but do not by themselves prove revenue quality. | Medium | SI001, SI004, SI005 |
| CI023 | At 2025 year-end SiliconFlow held RMB171 million of cash and cash equivalents plus RMB100 million of time deposits. | High | SI001, SI005 |
| CI024 | Net cash used in operating activities was RMB172 million in 2025. | High | SI001, SI005 |
| CI025 | On a simple backward-looking lens, SiliconFlow was not self-funding before the 2026 financings. | Medium | SI001, SI023, SI005 |
| CI026 | The filing discloses about RMB1.48 billion of post-2025 cash consideration across the A+, B, and B+ rounds. | Medium | SI001 |
| CI027 | SiliconFlow does not directly purchase chips; it mainly leases computing resources through partners. | High | SI005, SI001 |
| CI028 | Major purchases are concentrated among the top five suppliers, tying capital adequacy to supplier terms as well as to cash balances. | High | SI001, SI005 |
| CI029 | The retained public record does not support a precise current runway calculation as of 2026-07-22. | Low | SI001, SI005 |
| CI030 | The 2026 financings improved capital adequacy, but they did not by themselves solve revenue-quality or margin-path questions. | Medium | SI001, SI004, SI005 |
| CI031 | The public record does not reveal net revenue retention, gross retention, or customer concentration. | Low | SI001, SI002 |
| CI032 | The public record does not reveal realized effective pricing by model family or customer cohort. | Low | SI002, SI019, SI024 |
| CI033 | The public record does not reveal a current monthly cash-burn figure or a management-backed runway target as of run date. | Low | SI001, SI005 |
| CI034 | On-premise revenue appears higher quality on margin, but public sources are insufficient to show whether it can scale enough to change the whole-company profile. | Medium | SI001, SI004 |
| CI035 | Public-cloud token revenue remains the core growth engine but currently appears subsidy-heavy and compute-rental-heavy. | High | SI001, SI004, SI005 |
| CI036 | List pricing across the category increasingly pushes heavy users toward dedicated or reserved capacity once workloads stabilize. | Medium | SI006, SI013, SI014, SI018, SI024 |
| CI037 | Prompt caching is economically important because it can lower effective token costs or reduce compute wasted on repeated context. | High | SI008, SI010, SI016, SI025 |
| CI038 | Fine-tuning and dedicated hosting create additional monetization paths in the category, but retained sources do not show SiliconFlow currently monetizing those paths separately. | Medium | SI009, SI013, SI018 |
| CI039 | The clearest financial bridge is from usage into three revenue-quality tiers—serverless, dedicated, and on-premise—rather than into one uniform SaaS line. | Medium | SI001, SI002, SI004 |
| CI040 | The divergence between demand growth and financial quality is the central financial thesis of SiliconFlow as of 2026-07-22. | Medium | SI001, SI004, SI005 |
| CE001 | SiliconFlow is an API-first inference platform rather than a single-model application. | High | SE002, SE008 |
| CE002 | Its public product menu spans ready-to-use model APIs, reserved instances, inference acceleration services, and private deployment. | High | SE001, SE021, SE008 |
| CE003 | Serverless APIs are the core developer-facing surface. | High | SE001, SE002, SE004 |
| CE004 | Private deployment and BYOC are presented as enterprise options for privacy-sensitive workloads. | Medium | SE001, SE006 |
| CE005 | Users follow a simple path of model selection, API key creation, and endpoint integration. | High | SE002, SE004 |
| CE006 | The catalog spans text, speech, image, video, vector, reranking, and multimodal model classes. | High | SE002, SE013 |
| CE007 | Reserved instances are positioned for enterprise core inference scenarios needing dedicated capacity and cost optimization. | High | SE001, SE014 |
| CE008 | Inference acceleration is marketed both for open-source models and for self-developed models. | High | SE001, SE002 |
| CE009 | The product bundle reduces the need for customers to stitch together separate API, deployment, and optimization vendors. | Medium | SE001, SE002, SE019 |
| CE010 | SiliconFlow exposes model choice, streaming, context-window controls, tool calling, and request tracing through its chat API surface. | Medium | SE004 |
| CE011 | The platform uses a largely OpenAI-style integration pattern, lowering adoption friction for developers already familiar with that schema. | Medium | SE004, SE018 |
| CE012 | The docs show SiliconFlow regularly updates model availability and service capabilities over time. | High | SE003, SE004 |
| CE013 | Developers can stay on serverless or graduate toward reserved instances and private deployments as workloads mature. | Medium | SE001, SE014, SE016 |
| CE014 | SiliconFlow claims self-developed efficient operators, optimization frameworks, and a leading inference acceleration engine. | High | SE001, SE002 |
| CE015 | OneDiff is an acceleration library for diffusion models with optimized GPU kernels and compiler tooling. | Medium | SE010 |
| CE016 | OneDiff release history shows ongoing engineering maintenance rather than a one-off repository dump. | Medium | SE011 |
| CE017 | The API docs expose x-siliconcloud-trace-id response headers for request tracing and troubleshooting. | Medium | SE004 |
| CE018 | The strongest external engineering proof currently available is around inference-optimization tooling, not independently audited production reliability. | Medium | SE010, SE011, SE004 |
| CE019 | Public sources do not provide audited SLOs, public error-rate dashboards, or third-party performance benchmarks for the whole platform. | Low | SE001, SE002, SE004 |
| CE020 | SiliconFlow differentiates through a combination of catalog breadth, integration simplicity, deployment flexibility, and optimization claims. | High | SE001, SE002, SE005, SE013 |
| CE021 | The company continues to add current model launches, as shown by public launch messaging around recent models such as GLM-5.2 and Kimi K3. | Medium | SE001, SE012 |
| CE022 | The product moat is operational and integration-driven rather than based on ownership of a proprietary frontier model. | Medium | SE008, SE023, SE024 |
| CE023 | The platform depends on upstream model-provider availability and on continued access to external compute capacity. | Medium | SE003, SE008, SE023 |
| CE024 | The public product roadmap appears to be driven heavily by fast model onboarding and packaging cadence. | Medium | SE003, SE012, SE013 |
| CE025 | Private deployment increases appeal to regulated buyers but also increases implementation complexity and service-delivery burden. | Medium | SE001, SE006, SE014 |
| CE026 | The Kimi K3 launch post confirms a live developer surface with a standard API base URL. | Medium | SE012 |
| CE027 | Public maturity is strongest for the serverless API surface and weaker for enterprise assurance artifacts. | Medium | SE001, SE002, SE006, SE007 |
| CE028 | OneDiff release cadence is a useful roadmap proxy for SiliconFlow's optimization work, but it is not a substitute for a full platform changelog. | Medium | SE011, SE010 |
| CE029 | SiliconFlow publicly claims BYOC deployment and compute/network/storage isolation. | Medium | SE001 |
| CE030 | SiliconFlow maintains distinct China and international service surfaces with separate contractual language. | High | SE007, SE021 |
| CE031 | The privacy policy confirms the international surface is operated through siliconflow.com. | Medium | SE006 |
| CE032 | Trace identifiers in API responses are a real supportability control, though not a substitute for published reliability reporting. | Medium | SE004 |
| CE033 | The retained source set does not show a public SOC 2 report, ISO certificate list, or detailed public incident history. | Low | SE001, SE006, SE007 |
| CE034 | Large enterprise underwriting still requires third-party assurance artifacts that are not visible in the retained public set. | Low | SE006, SE007, SE001 |
| CE035 | The company's security and compliance messaging is plausible but currently stronger in claims than in public evidence depth. | Medium | SE001, SE006, SE007 |
| CE036 | SiliconFlow is built for supportability and enterprise packaging, but public evidence is insufficient to conclude best-in-class enterprise readiness. | Medium | SE001, SE004, SE006, SE007 |
| CU001 | SiliconFlow serves multiple customer types rather than a single homogeneous user base. | High | SU001, SU018, SU019 |
| CU002 | Public 2026 coverage cites more than 10 million users and more than 10,000 enterprise customers. | Medium | SU002, SU023 |
| CU003 | The company explicitly serves both developers and enterprises. | High | SU001, SU019 |
| CU004 | Enterprise demand includes reserved instances, private deployment, and infrastructure-style deployments rather than only lightweight API experimentation. | High | SU005, SU006, SU008 |
| CU005 | Community and developer-tool integrations prove API onboarding relevance across coding, search/RAG, translation, and agent workflows. | High | SU004, SU009, SU010, SU011 |
| CU006 | Telecom / compute partners are a distinct strategic customer surface for SiliconFlow. | High | SU007, SU003 |
| CU007 | The filing and public coverage imply that self-serve paying accounts and larger enterprise deployments coexist inside the customer base. | Medium | SU001, SU015 |
| CU008 | Broad user reach does not by itself prove high-quality revenue because developer and community adoption can be monetically light. | Medium | SU002, SU004, SU017 |
| CU009 | SiliconFlow appears especially strong in infra-heavy buyer scenarios where inference performance, private deployment, or国产化 adaptation matters. | Medium | SU005, SU006, SU007, SU018 |
| CU010 | Guizhou Mobile is the clearest named institutional counterparty in the retained customer corpus. | High | SU007, SU003 |
| CU011 | The Guizhou Mobile partnership covers inference framework deployment, compute coordination, token services, and joint operational systems. | Medium | SU007 |
| CU012 | The Guizhou Mobile relationship is described as a deepening of earlier 2025 cooperation rather than a first-touch pilot. | Medium | SU007 |
| CU013 | The reserved-instance case study shows at least one customer whose coding-agent workload reached 100 billion daily tokens in 2026. | Medium | SU008 |
| CU014 | MindSearch is a named public project with a documented SiliconFlow integration path. | High | SU009, SU004 |
| CU015 | Continue is a named public integration that positions SiliconFlow for coding-assistant workflows inside IDEs. | High | SU010, SU004 |
| CU016 | Cline is a named public integration that uses SiliconFlow through an OpenAI-compatible API pattern. | High | SU011, SU004 |
| CU017 | The named developer-tool proofs are useful evidence of adoption breadth but weak evidence of contract value or exclusivity. | Medium | SU009, SU010, SU011 |
| CU018 | Anonymous enterprise case studies add operational depth but limit concentration and logo-quality analysis because the customer names are withheld. | High | SU005, SU006, SU008 |
| CU019 | Named customer proof in this chapter is therefore stronger on breadth than on directly observable commercial weight. | Medium | SU007, SU009, SU010, SU011 |
| CU020 | SiliconFlow shows a land-and-expand pattern from pay-as-you-go usage toward reserved or deeper infrastructure deployment for some accounts. | Medium | SU008, SU025, SU001 |
| CU021 | The retained public set does not disclose NRR, GRR, or logo-retention rates. | Medium | SU001, SU015 |
| CU022 | The retained public set does not disclose average contract length, renewal rates, or top-customer concentration. | Medium | SU001, SU015, SU017 |
| CU023 | Public evidence for repeat usage is architectural and behavioral rather than cohort-based: heavy workloads, reserved-instance upgrades, and deeper partner cooperation. | Medium | SU007, SU008, SU025 |
| CU024 | Private deployment and国产化 case studies suggest deeper embedment than simple API testing, but the public record still lacks renewal proof. | Medium | SU005, SU006, SU021 |
| CU025 | Revenue concentration could still be high even with 10,000+ enterprise customers if a small number of token-heavy accounts dominate spend. | Medium | SU002, SU008, SU017 |
| CU026 | Strategic telecom or infrastructure partners may become concentrated routes to growth, creating bargaining-power and dependency risk. | Medium | SU007, SU021, SU017 |
| CU027 | Broad community integration can inflate awareness and usage without proving durable paid conversion. | Medium | SU004, SU009, SU010, SU011 |
| CU028 | The customer corpus supports diversification by use case, but not yet diversification by revenue contribution. | Medium | SU004, SU005, SU006, SU007 |
| CU029 | SiliconFlow has enough public evidence to show real adoption, not merely claimed logos. | High | SU002, SU007, SU009, SU010, SU011 |
| CU030 | Customer durability is materially less proven than customer breadth. | Medium | SU021, SU022, SU015 |
| CU031 | The strongest public customer signal is that some workloads are mission-critical enough to justify dedicated or private infrastructure decisions. | High | SU005, SU006, SU008 |
| CU032 | The cleanest public customer journey is discover the API, prove value in workflow, then deepen into reserved or private deployment as usage scales. | Medium | SU019, SU008, SU025 |
| CU033 | For diligence, the biggest remaining customer question is not acquisition but quality of retained spend. | Medium | SU002, SU015, SU017 |
| CU034 | A retention cohort figure can be shown only as a visibility proxy, not as a real renewal chart. | Medium | SU001, SU015 |
| CU035 | Developer adoption and enterprise adoption reinforce each other strategically, but they should not be valued as equivalent commercial proof. | Medium | SU004, SU007, SU010, SU011 |
| CU036 | The right customer underwriting request is a segmented cohort pack covering active customers, spend concentration, renewal, and upgrade from API to reserved/private deployment. | Medium | SU001, SU015, SU017 |
| CR001 | SiliconFlow's domestic terms expressly limit the service to developer-oriented content-generation scenarios and exclude automatic control, medical information, psychological counseling, and critical information infrastructure uses. | Medium | SR002 |
| CR002 | The domestic privacy policy says mainland-law requirements can trigger real-name verification and identity collection for service activation. | Medium | SR003 |
| CR003 | SiliconFlow's domestic terms explicitly reference the AI-generated-content labeling measures and prohibit malicious deletion or tampering with labels. | Medium | SR002 |
| CR004 | The domestic site publicly displays ICP and telecom-license identifiers, indicating regulated internet-service operations rather than a purely offshore surface. | Medium | SR004 |
| CR005 | SiliconFlow maintains separate China and international legal surfaces, reducing but not eliminating cross-surface governance complexity. | High | SR005, SR006 |
| CR006 | China's 2025 AI labeling measures create mandatory explicit and implicit labeling obligations for generated synthetic content. | Medium | SR007, SR032 |
| CR007 | Those labeling rules create operational risk because providers and platforms must implement metadata, logging, and downstream labeling workflows rather than merely publish a policy. | Medium | SR002, SR007, SR031 |
| CR008 | The public terms show compliance awareness, but they do not by themselves prove SiliconFlow's implementation quality for labeling, logging, or auditability. | Medium | SR002, SR007 |
| CR009 | SiliconFlow's domestic privacy policy says user information generated in operating the domestic site is stored within the PRC. | Medium | SR003 |
| CR010 | Mainland onboarding and identity checks can increase friction for some customers and enlarge the company's privacy-handling obligations. | High | SR002, SR003 |
| CR011 | The legal risk is a continuous compliance burden across content, privacy, telecom, and sector-specific use restrictions rather than a single one-time license hurdle. | High | SR002, SR003, SR004, SR007 |
| CR012 | SiliconFlow serves infrastructure-like workloads where outages, latency swings, or deployment failures can matter more than consumer feature gaps. | High | SR015, SR016, SR017 |
| CR013 | Public sources do not furnish a robust public uptime history, incident archive, or detailed SLO record. | Low | SR015, SR018 |
| CR014 | Public sources do not furnish SOC 2, ISO certification details, or equivalent third-party assurance artifacts in the retained corpus. | Low | SR004, SR005, SR006 |
| CR015 | The retained corpus suggests enterprise-grade operational aspirations, but not enough third-party evidence to conclude best-in-class operational maturity. | Medium | SR015, SR016, SR018 |
| CR016 | Privacy, labeling, and identity controls interact as a compound execution risk because each adds workflow and data-handling obligations to the platform. | High | SR002, SR003, SR007 |
| CR017 | The platform's own API reference warns that model availability and capability can change over time. | Medium | SR020 |
| CR018 | Rapid model churn creates support and migration risk for enterprise workloads even when the platform benefits from a broad model catalog. | Medium | SR019, SR020 |
| CR019 | Monitoring, fault tolerance, and traceability are publicly claimed mitigations, but retained sources do not independently validate their effectiveness at scale. | Medium | SR018, SR015 |
| CR020 | Buyer expectations in critical or regulated environments are rising toward explicit trustworthy-AI risk-management practices, even where frameworks such as NIST AI RMF are voluntary. | High | SR010, SR021 |
| CR021 | Operational-quality risk is therefore most acute where high-volume or regulated workloads demand both performance and documented assurance. | Medium | SR015, SR016, SR020 |
| CR022 | SiliconFlow depends on upstream model providers for catalog breadth and competitive relevance. | High | SR019, SR020, SR029 |
| CR023 | SiliconFlow depends heavily on leased compute and supplier relationships rather than on fully owned semiconductor supply. | High | SR001, SR011, SR012 |
| CR024 | U.S. export controls continue to target PRC access to advanced computing semiconductors, creating a real external dependency risk for Chinese AI-infrastructure players. | High | SR008, SR009 |
| CR025 | Domestic-chip adaptation and private-deployment narratives mitigate but do not eliminate supply-side risk. | Medium | SR016, SR017, SR024 |
| CR026 | Strategic telecom or compute partnerships can broaden distribution but also create concentration and execution dependencies. | Medium | SR014, SR023 |
| CR027 | Supplier / capacity planning is an execution-critical function because service quality and margin both depend on it. | High | SR001, SR011, SR015 |
| CR028 | Private deployment, reserved instances, and domestic-chip optimization increase implementation complexity and require strong enterprise delivery operations. | High | SR015, SR016, SR017 |
| CR029 | Independent analyses already frame SiliconFlow's economics as vulnerable to price war and compute-rental pressure. | Medium | SR011, SR012, SR013 |
| CR030 | Competitive pressure from hyperscalers and other inference platforms can intensify the margin impact of compliance and compute-supply costs. | High | SR025, SR026, SR027, SR028, SR029, SR030 |
| CR031 | Hidden concentration of token-heavy customers remains a real model risk because public breadth metrics do not reveal spend distribution. | Medium | SR001, SR013, SR011 |
| CR032 | SiliconFlow already has visible mitigations: separate legal surfaces, BYOC/private deployment, domestic adaptation, and strategic partner operations. | Medium | SR005, SR014, SR016, SR017 |
| CR033 | The public record does not yet prove whether SiliconFlow has sufficient second-line compliance, delivery, and supplier-management depth to scale safely. | Low | SR011, SR012, SR014 |
| CR034 | A formal regulatory event around labeling, privacy, or telecom compliance would be an immediate thesis-break trigger. | Medium | SR002, SR003, SR007 |
| CR035 | A material compute-capacity shock or supplier loss would directly threaten both reliability and gross margin. | Medium | SR001, SR008, SR023 |
| CR036 | Proof that enterprise retention or revenue concentration is worse than implied by public customer counts would materially weaken the investment case. | Medium | SR001, SR011, SR013 |
| CR037 | Failure of public-cloud margins to improve despite scale would indicate that growth is not translating into a durable business model. | Medium | SR001, SR011, SR029 |
| CR038 | The strongest risk transmission path is from regulation or supply shock into operating complexity, then into revenue quality, margin, financing need, and valuation. | Medium | SR007, SR023, SR029 |
| CR039 | The most likely early diligence flashpoint is weak evidence around enterprise assurance, retention, and supplier concentration rather than a near-term demand collapse. | Medium | SR011, SR012, SR015 |
| CR040 | SiliconFlow's risk profile is investable only if the company can keep converting scale into better margins faster than regulatory and supply complexity increase. | Medium | SR001, SR011, SR024 |
| CV001 | On public evidence alone, the right current recommendation is Track rather than Buy. | Medium | SV001, SV004, SV005, SV006, SV009 |
| CV002 | The risk rating for SiliconFlow at the current price is high. | Medium | SV001, SV025, SV026 |
| CV003 | Confidence in the recommendation is only medium-low because the price is public but the core underwriting metrics remain incomplete. | Medium | SV001, SV004, SV005 |
| CV004 | SiliconFlow is strategically relevant enough to merit continued diligence and tracking. | High | SV014, SV015, SV021, SV022 |
| CV005 | The company has real customer and usage proof rather than a purely narrative AI story. | High | SV002, SV023, SV024 |
| CV006 | Recent financing and strategic investors improve the company's optionality and time to execute. | Medium | SV002, SV003, SV022 |
| CV007 | Current public evidence is insufficient to justify a strong Buy at the disclosed June 2026 price. | Medium | SV001, SV004, SV005, SV025 |
| CV008 | The valuation stance is therefore price-sensitive rather than categorically negative. | Medium | SV002, SV006, SV009 |
| CV009 | A lower entry price or stronger proof package could move the recommendation upward. | Medium | SV001, SV004, SV006, SV009 |
| CV010 | The positive thesis begins with category positioning: SiliconFlow sits in a growing inference market with real product breadth and deployment flexibility. | High | SV014, SV015, SV021, SV028 |
| CV011 | The positive thesis also includes ecosystem and strategic-channel leverage, not only direct API usage. | Medium | SV002, SV003, SV022 |
| CV012 | The anti-thesis begins with disclosed weak economics, especially negative public-cloud gross margins. | High | SV001, SV004, SV005 |
| CV013 | Compute-rental dependence and supplier concentration weaken the claim that current scale already represents a mature infrastructure moat. | High | SV001, SV005, SV026 |
| CV014 | Compared with Western inference peers, SiliconFlow's disclosed revenue scale is much smaller. | Medium | SV001, SV006, SV009, SV012 |
| CV015 | SiliconFlow may still be strategically important even if it deserves a valuation discount to faster-scaling Western peers. | Medium | SV006, SV009, SV012, SV013 |
| CV016 | The current US$1.2B valuation is not obviously absurd in absolute terms, but it is not clearly supported by public underwriting data either. | Medium | SV002, SV004, SV005, SV006 |
| CV017 | The market-size thesis is real, but current evidence does not yet show that SiliconFlow has converted strategic relevance into high-quality economics. | Medium | SV001, SV014, SV028 |
| CV018 | Regulatory and supply-side risks deserve valuation weight because they can affect both growth and margin simultaneously. | High | SV025, SV026, SV029 |
| CV019 | The disclosed price already asks investors to underwrite future improvement rather than current financial quality. | Medium | SV001, SV002, SV004 |
| CV020 | The bull case requires better enterprise retention, richer dedicated/private-deployment mix, and meaningful margin improvement. | Medium | SV001, SV021, SV023 |
| CV021 | The base case assumes SiliconFlow remains strategically relevant and grows, but revenue quality improves only gradually. | Medium | SV001, SV004, SV028 |
| CV022 | The bear case assumes price competition, regulation, or supply constraints prevent margin inflection and weaken financing terms. | Medium | SV004, SV005, SV025, SV026 |
| CV023 | Together AI is a relevant direct comparable because it combines open-model access with inference infrastructure. | High | SV006, SV007 |
| CV024 | Together AI's 2026 $8.3B valuation comes with public evidence of >$1.15B annual bookings, a scale far above SiliconFlow's disclosed 2025 revenue base. | Medium | SV006 |
| CV025 | Fireworks AI is a relevant direct comparable because it is an inference-cloud platform rather than a generic hyperscaler. | High | SV009, SV010, SV030 |
| CV026 | Fireworks AI's 2026 $17.5B valuation is paired with public evidence of >$1B annualized revenue, making SiliconFlow look less obviously cheap despite its much lower headline price. | High | SV009, SV010 |
| CV027 | Baseten is a relevant inference-platform comparable because it combines developer entry with enterprise deployment options. | Medium | SV011, SV012 |
| CV028 | Baseten's 2026 ~$13B valuation and ~US$600M annualized revenue indicate that specialized inference platforms can justify high values—but with much stronger monetization evidence than SiliconFlow currently discloses. | Medium | SV012 |
| CV029 | CoreWeave is an infrastructure-adjacent valuation reference that shows both the upside and the concentration / capital-intensity risks of AI compute businesses. | Medium | SV013, SV026 |
| CV030 | Public comparables therefore support interest in the category, but not an automatic premium view on SiliconFlow at current evidence quality. | Medium | SV006, SV009, SV012, SV013 |
| CV031 | Retention, concentration, and supplier-risk visibility are the most important missing inputs preventing a stronger call. | Medium | SV001, SV004, SV005, SV026 |
| CV032 | A formal regulatory enforcement event would be a thesis-break trigger. | Medium | SV025, SV029 |
| CV033 | A material compute-supply shock or worsening supplier concentration would be a thesis-break trigger. | Medium | SV001, SV026 |
| CV034 | Failure of public-cloud economics to improve would be a thesis-break trigger. | Medium | SV001, SV004, SV005 |
| CV035 | Proof of weak enterprise retention or heavy whale concentration would be a thesis-break trigger. | Medium | SV001, SV004, SV023 |
| CV036 | The most important valuation diligence ask is current 2026 revenue run-rate and mix by product line. | Medium | SV001, SV004 |
| CV037 | The next most important asks are enterprise retention, concentration, and API-to-dedicated expansion metrics. | Medium | SV001, SV023, SV024 |
| CV038 | Security / assurance and supplier contingency are also valuation-critical because they affect enterprise quality and downside risk. | Medium | SV025, SV026, SV027 |
| CV039 | Price alone cannot rescue the investment case if the missing diligence items point to structural weakness rather than temporary opacity. | Medium | SV004, SV005, SV026 |
| CV040 | As of 2026-07-22, the best public-evidence call is Track: strategic relevance is clear, but upside from the current price is not. | Medium | SV001, SV002, SV006, SV009 |
| ID | Publisher | Title | Quote |
|---|---|---|---|
| SO001 | SiliconFlow | 硅基流动 SiliconFlow - 致力于成为全球领先的 AI 能力提供商 | 地址:北京市海淀区中关村东路 1 号院 8 号楼 D 座 23 层 2301。 |
| SO002 | SiliconFlow | SiliconFlow – AI Infrastructure for LLMs & Multimodal Models | One Platform All Your AI Inference Needs. |
| SO003 | SiliconFlow | About Us - SiliconFlow | Global AI Infrastructure Provider | We strive to become the world's most influential provider of AI infrastructure. |
| SO004 | SiliconFlow | SiliconFlow – AI Infrastructure for LLMs & Multimodal Models / Models | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SO005 | SiliconFlow Docs | Quickstart - SiliconFlow | Currently, the platform supports login via SMS, email, as well as OAuth login through GitHub and Google. |
| SO006 | SiliconFlow Docs | Chat - SiliconFlow API Reference | We periodically update our models to enhance service quality. Changes may include model on/offlining or capability adjustments. |
| SO007 | SiliconFlow Docs | Product introduction - SiliconFlow | SiliconFlow is committed to providing developers with faster, more comprehensive, and seamlessly integrated model APIs. |
| SO008 | SILICONFLOW TECHNOLOGY PTE. LTD. | Terms of Use - SiliconFlow | The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn. |
| SO009 | SILICONFLOW TECHNOLOGY PTE. LTD. | Privacy Policy - SiliconFlow | If you want to contact us... please contact us at contact@siliconflow.com. |
| SO010 | GitHub | SiliconFlow · GitHub | SiliconFlow builds scalable, standardized, and high-performance AI infrastructure. |
| SO011 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | As of April 30, 2026, our platform had over 10 million registered users. |
| SO012 | HKEX | Announcement on overall coordinator appointment for Beijing SiliconFlow Technology Co., Ltd. | The Company has further appointed China Renaissance Securities (Hong Kong) Limited as its overall coordinator on July 14, 2026. |
| SO013 | Caixin Global | SiliconFlow Raises $294 Million as China's AI Inference Demand Surges | Chinese AI model inference startup SiliconFlow has raised more than 2 billion yuan ($294 million) in a Series B funding round. |
| SO014 | Sina Finance | 硅基流动完成新一轮超 20 亿元融资 | 过去一年,公司在企业级市场实现爆发式增长。 |
| SO015 | Sina Finance | 硅基流动完成新一轮超 20 亿元融资(转载) | IDC预测,2026年中国市场的Token消费量将达到40,000万亿。 |
| SO016 | 投资界 / InvestorsCN | 硅基流动完成新一轮超20亿元融资,华兴资本担任独家财务顾问 | 近日,硅基流动已完成超20亿元B轮融资。 |
| SO017 | 36Kr Europe | Silicon Flow Completes New Round of Over 2-Billion-Yuan Financing | By providing efficient MaaS through the Token factory model, the daily average Token call volume has reached trillions. |
| SO018 | KrASIA | Surging users, widening losses, and leased compute: Behind SiliconFlow's IPO filing | But the prospectus also shows the cost of that growth. SiliconFlow is selling tokens at a loss. |
| SO019 | Hello China Tech | SiliconFlow IPO: What China's Token Boom Really Costs | Its public cloud service ... generated 52.9% of 2025 revenue ... Its gross margin was negative 119%. |
| SO020 | Tracxn | SiliconFlow - 2026 Company Profile & Team - Tracxn | SiliconFlow is a funded company based in Beijing (China), founded in 2023 by Pan YANG and Jinhui Yuan. |
| SO021 | CB Insights | SiliconFlow Stock Price, Funding, Valuation, Revenue & Financial Statements | SiliconFlow has raised $316.54M over 8 rounds. |
| SO022 | SiliconFlow | Kimi K3 is now live on SiliconFlow | SiliconFlow provides OpenAI- and Anthropic-compatible APIs for Kimi K3. |
| SO023 | SiliconFlow | Kimi K3 on SiliconFlow API | Current support includes image input, tool calling, JSON Mode, streaming, and reasoning output. |
| SO024 | SiliconFlow Cloud | SiliconFlow 统一登录 / Welcome to SiliconFlow | Blazing-fast, cost-effective Generative AI cloud services. |
| SO025 | SiliconFlow | 模型与产品页 - 硅基流动 | 支持 BYOC 部署,全面保护数据隐私与业务安全。 |
| SM001 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | In 2025, we were the fourth largest token supply platform in China in terms of annual token throughput, with a market share of 1.5% and ranked first among all independent ecosystem token supply platforms in China. |
| SM002 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 2025年中国公有云MaaS市场的规模达到30.7亿元人民币。IDC预计2026年全年Token消耗量约为40,000万亿次。 |
| SM003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | IDC predicts that the Token consumption in the Chinese market will reach 40,000 trillion in 2026, an increase of about 20 times compared to 2025. |
| SM004 | Stanford HAI | 2026 AI Index Report | Industry produced over 90% of notable frontier models in 2025... Organizational adoption reached 88%. |
| SM005 | MarketsandMarkets | AI Inference Market Size, Share & Trends | The AI inference market size was valued at USD 106.15 billion in 2025 and is projected to reach USD 254.98 billion by 2030, growing at a CAGR of 19.2%. |
| SM006 | Grand View Research | AI Inference Market Summary | The global AI inference market size was estimated at USD 97.24 billion in 2024 and is projected to reach USD 253.75 billion by 2030, growing at a CAGR of 17.5%. |
| SM007 | Fortune Business Insights | AI Inference Market | The global AI inference market size was valued at USD 103.73 billion in 2025 and is projected to grow from USD 117.80 billion in 2026 to USD 312.64 billion by 2034. |
| SM008 | Amazon Web Services | Amazon Bedrock | Amazon Bedrock powers generative AI for more than 100,000 organizations worldwide—from startups to global enterprises across every industry. |
| SM009 | Amazon Web Services | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models... for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SM010 | AWS Docs | What is Amazon Bedrock? | Bedrock supports 100+ foundation models from industry-leading providers, including Amazon, Anthropic, DeepSeek, Moonshot AI, MiniMax, and OpenAI. |
| SM011 | Microsoft Learn | What is Microsoft Foundry? | Foundry unifies agents, models, and tools under a single management grouping with built-in enterprise-readiness capabilities including tracing, monitoring, evaluations, and customizable enterprise setup configurations. |
| SM012 | Microsoft Azure | Microsoft Foundry Pricing | Foundry Models: Access more than 11,000 foundation, open, reasoning, multimodal and industry-specific models. |
| SM013 | Microsoft Azure | Foundry Models Pricing | Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure... Choose from A100, H100, H200, and MI300 GPU families. |
| SM014 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SM015 | Alibaba Cloud | Token Plan (Team Edition) overview | Token Plan (Team Edition) is a monthly AI model subscription billed in Credits... and includes team management, data privacy, and dedicated throughput. |
| SM016 | Together AI | Serverless Inference | Access all the top open-source models in one place. |
| SM017 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SM018 | Together AI Docs | Available models - Together AI docs | Serverless models are the fastest way to run inference on Together... Pay only for the tokens you process. |
| SM019 | Fireworks AI | Pricing | Additionally, please note that cached input tokens are by default priced at 50% for all text and vision language models... batch inference is priced at 50% of our serverless pricing. |
| SM020 | Fireworks Docs | Models & Inference | Fireworks provides fast, cost-effective access to leading open-source text models through OpenAI-compatible APIs. |
| SM021 | SiliconFlow | Models - SiliconFlow | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SM022 | SiliconFlow Docs | Quickstart - SiliconFlow | Currently, the platform supports login via SMS, email, as well as OAuth login through GitHub and Google. |
| SM023 | SiliconFlow Docs | Chat Completions API Reference | We periodically update our models to enhance service quality. Changes may include model on/offlining or capability adjustments. |
| SM024 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | According to third-party data cited in its prospectus, SiliconFlow was China's largest independent ecosystem token supplier by annual token throughput in 2025 and ranked among the top five token suppliers overall. |
| SM025 | Hello China Tech | SiliconFlow IPO: token economics | As a middle-layer platform, SiliconFlow's main cost pressure comes from leasing computing power... In 2025, computing resource costs totaled RMB 59.627 million, accounting for 86.9% of cost of sales. |
| SP001 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | In 2025, we were the fourth largest token supply platform in China in terms of annual token throughput, with a market share of 1.5% and ranked first among all independent ecosystem token supply platforms in China. |
| SP002 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | IDC认为,MaaS厂商的竞争焦点正在从过去单纯的价格比拼,转向“价格、性能与工具链支持”的综合能力竞争。 |
| SP003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The latest "China AI Software Market Semi-Annual Tracker, 2025H2" released by IDC shows that SiliconFlow is the only startup company among the top four in the market share of China's public cloud MaaS. |
| SP004 | Amazon Web Services | Amazon Bedrock | Amazon Bedrock powers generative AI for more than 100,000 organizations worldwide—from startups to global enterprises across every industry. |
| SP005 | Amazon Web Services | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models ... for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SP006 | Microsoft Learn | What is Microsoft Foundry? | Foundry unifies agents, models, and tools under a single management grouping with built-in enterprise-readiness capabilities including tracing, monitoring, evaluations, and customizable enterprise setup configurations. |
| SP007 | Microsoft Azure | Microsoft Foundry Pricing | Foundry Models: Access more than 11,000 foundation, open, reasoning, multimodal and industry-specific models. |
| SP008 | Microsoft Azure | Foundry Models Pricing | Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure to deploy and run open-source and custom AI models on enterprise-grade GPUs. |
| SP009 | Alibaba Cloud | What is Model Studio? | Alibaba Cloud Model Studio is a one-stop model service platform. It provides the full Qwen series and mainstream third-party LLMs through official Qwen APIs and OpenAI-compatible APIs. |
| SP010 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SP011 | Alibaba Cloud | Token Plan (Team Edition) overview | Token Plan (Team Edition) is a monthly AI model subscription billed in Credits ... and includes team management, data privacy, and dedicated throughput. |
| SP012 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SP013 | Together AI Docs | Serverless overview | Run any supported model through a shared, per-token API with no provisioning and no minimums. |
| SP014 | Together AI Docs | Dedicated model inference overview | Dedicated model inference bills per minute by hardware while a deployment runs, regardless of model or request volume. |
| SP015 | Together AI Docs | Available models - Together AI docs | Serverless models are the fastest way to run inference on Together. You call any supported model through a shared per-token API, with no provisioning, no replicas to size, and no minimum cost. |
| SP016 | Fireworks AI | Fireworks home | Fireworks processes 40T+ tokens per day. |
| SP017 | Fireworks AI | Pricing | Pricing to seamlessly scale from idea to enterprise. |
| SP018 | Fireworks Docs | Serverless overview | Serverless is multi-tenant inference for popular open models running on Fireworks-managed infrastructure. |
| SP019 | Fireworks Docs | Serverless pricing | Per-token serverless pricing for text, vision, and embedding models, including Priority and Fast serving paths. |
| SP020 | Fireworks Docs | Models & Inference | Fireworks provides fast, cost-effective access to leading open-source text models through OpenAI-compatible APIs. |
| SP021 | OpenRouter Docs | Quickstart | OpenRouter provides a unified API that gives you access to hundreds of AI models through a single endpoint, while automatically handling fallbacks and selecting the most cost-effective options. |
| SP022 | OpenRouter Docs | Provider routing | OpenRouter routes requests to the best available providers for your model. By default, requests are load balanced across the top providers to maximize uptime. |
| SP023 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | SiliconFlow was China's largest independent ecosystem token supplier by annual token throughput in 2025 and ranked among the top five token suppliers overall. |
| SP024 | Hello China Tech | SiliconFlow IPO: token economics | The top three token suppliers by throughput, Volcengine, Alibaba Cloud, and Baidu AI Cloud, are all hyperscaler divisions. |
| SP025 | SiliconFlow | Models - SiliconFlow | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SI001 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | For our public cloud-based services, the gross loss margin was 271.6% in 2024 and 119.0% in 2025. |
| SI002 | SiliconFlow | Models - SiliconFlow | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SI003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | By providing efficient MaaS through the Token factory model, the daily average Token call volume has reached trillions. |
| SI004 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Its public cloud service... generated 52.9% of 2025 revenue. Its gross margin was negative 119%. |
| SI005 | Hello China Tech | SiliconFlow IPO: token economics | In 2025, computing resource costs totaled RMB 59.627 million, accounting for 86.9% of cost of sales. |
| SI006 | Together AI Docs | Dedicated model inference pricing | Dedicated model inference bills based on the hardware your deployments run on, regardless of model or request volume. |
| SI007 | Together AI Docs | Batch processing overview | Run asynchronous batch workloads at up to 50% lower cost. |
| SI008 | Together AI Docs | Send requests on dedicated endpoints | Prompt caching is enabled by default for dedicated model inference. No configuration is required. |
| SI009 | Together AI Docs | Fine-tuning pricing | Fine-tuning is billed per token processed, scaled by model size, training method, and training type. |
| SI010 | Fireworks Docs | Prompt caching | For serverless models, cached prompt tokens are discounted compared to regular prompt tokens. The default discount is 50%. |
| SI011 | Fireworks Docs | Serverless serving paths | Priority tier is for workloads that require higher reliability during peak traffic periods, at a higher price point. |
| SI012 | Fireworks Docs | Serverless pricing | Batch inference is billed at 50% of serverless pricing on both input and output. |
| SI013 | Fireworks AI | Pricing | H100 80 GB GPU $7.00 per hour. |
| SI014 | Amazon Web Services | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SI015 | AWS Docs | Batch inference in Amazon Bedrock | With batch inference, you can submit multiple prompts and generate responses asynchronously. |
| SI016 | AWS Docs | Prompt caching in Amazon Bedrock | Prompt caching can help when you have workloads with long and repeated contexts ... you're charged at a reduced rate for tokens read from cache. |
| SI017 | Microsoft Azure | Microsoft Foundry Pricing | The Microsoft Agent pre-purchase plan allows you to save on Microsoft Foundry and Copilot Credit costs by purchasing Agent Commit Units up front. |
| SI018 | Microsoft Azure | Foundry Models Pricing | Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure... Choose from A100, H100, H200, and MI300 GPU families. |
| SI019 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SI020 | Alibaba Cloud | Token Plan overview | Standard USD 30/seat/month ... Pro USD 100/seat/month ... Max USD 200/seat/month. |
| SI021 | Alibaba Cloud | What is Model Studio? | Activating Model Studio is free. Costs apply only when you invoke models. |
| SI022 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素依次是:模型性能、安全合规要求、回答质量、在AI平台可用性以及成本效益。 |
| SI023 | Fireworks Docs | Models & Inference | Every response includes token usage information and performance metrics for debugging and observability. |
| SI024 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SI025 | AWS | Amazon Bedrock | Features like Model Distillation, Prompt caching, and Intelligent Prompt Routing can reduce expenses while maintaining performance. |
| SE001 | SiliconFlow | SiliconFlow China homepage | 覆盖语言、语音、图片、视频等场景,一站式提供大模型 API 服务,按量计费,助力应用快速上线。 |
| SE002 | SiliconFlow Docs | Product introduction | As a one-stop cloud service platform integrating top-tier large language models, SiliconFlow is committed to providing developers with faster, more comprehensive, and seamlessly integrated model APIs. |
| SE003 | SiliconFlow Docs | API reference home | Corresponding Model Name. To better enhance service quality, we will make periodic changes to the models provided by this service. |
| SE004 | SiliconFlow Docs | Chat API reference | The response header contains the x-siliconcloud-trace-id field, which serves as a unique identifier for tracing requests. |
| SE005 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SE006 | SiliconFlow Docs | Privacy Policy | When you register, log in, and use the services provided through the https://siliconflow.com platform, we will collect and store the relevant information in accordance with this policy. |
| SE007 | SiliconFlow Docs | Terms of Use | The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn. |
| SE008 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | We offer API services for developers and enterprises ... and on-premise deployment solutions. |
| SE009 | GitHub | SiliconFlow organization | High-performance inference for every SOTA model. |
| SE010 | GitHub | OneDiff repository | onediff is an out-of-the-box acceleration library for diffusion models |
| SE011 | GitHub | OneDiff releases | 1.2.0 ... support diffusers sd3 speedup ... add diffusers nexfort example |
| SE012 | SiliconFlow Blog | Kimi K3 API launch post | Base URL: https://api.siliconflow.com/v1 |
| SE013 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The platform brings together over 400 models spanning text, image, audio, and video generation. |
| SE014 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Its public cloud service includes both a pay-as-you-go serverless model and dedicated computing instances. |
| SE015 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素 ... 在AI平台可用性以及成本效益。 |
| SE016 | Together AI Docs | Serverless overview | Use serverless inference to run models without managing infrastructure. |
| SE017 | Fireworks AI | Home | The Generative AI Platform for production-ready workloads. |
| SE018 | OpenRouter Docs | Quickstart | OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers. |
| SE019 | Alibaba Cloud | What is Model Studio? | Model Studio provides model APIs and application development capabilities. |
| SE020 | AWS | Amazon Bedrock | Features like Model Distillation, Prompt caching, and Intelligent Prompt Routing can reduce expenses while maintaining performance. |
| SE021 | SiliconFlow | SiliconFlow international site | Get your Model API fast |
| SE022 | SiliconFlow Cloud | Cloud model catalog | Model catalog and pricing surface for cloud users. |
| SE023 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SE024 | China Money Network | SiliconFlow's role in China's AI ecosystem | The platform supports over 400 models and serves thousands of enterprise customers. |
| SE025 | GitHub | OneFlow releases asset page reference in OneDiff README | please install OneFlow by the links below. |
| SU001 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | Our public cloud-based services primarily target developers and small- and medium-sized enterprises. |
| SU002 | QQ News / Taimei | 问AI · Token工厂模式为何吸引全产业链巨头联手投资? | 过去一年,硅基流动日均Token调用量达数万亿,服务超1000万用户和1万家企业客户。 |
| SU003 | SiliconFlow | News listing | 贵州移动与硅基流动深化战略合作,加速“Token 工厂”建设 |
| SU004 | SiliconFlow Docs | Scenarios and application cases | Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases. |
| SU005 | SiliconFlow | Airline group AI infrastructure case | 该企业引入硅基流动(SiliconFlow)推理加速框架,实现三层技术突破。 |
| SU006 | SiliconFlow | Energy SOE AI infrastructure case | 某能源央企正处于转型关键阶段。 |
| SU007 | SiliconFlow | Guizhou Mobile strategic cooperation | 贵州移动与硅基流动 ... 构建“推理框架 + 算力供给 + 模型服务 + 场景应用”全栈协同体系。 |
| SU008 | SiliconFlow | Reserved instances support 100B-token daily workload | 2026 年以来,其 Coding Agent 单日 Token 消耗冲上千亿量级。 |
| SU009 | SiliconFlow Docs | Use SiliconCloud in MindSearch | After adding this configuration, you can execute the relevant commands to start MindSearch. |
| SU010 | SiliconFlow Docs | Use Continue with SiliconFlow APIs | By integrating SiliconFlow APIs into Continue, you can get access to 200+ open-source models. |
| SU011 | SiliconFlow Docs | Use Cline with SiliconFlow APIs | we’ll show you how to integrate SiliconFlow’s APIs into Cline |
| SU012 | SiliconFlow | About us | From open source to enterprise deployment, we accelerate what matters. |
| SU013 | InforCapital | SiliconFlow company profile | SiliconFlow - AI Infrastructure |
| SU014 | CB Insights | SiliconFlow company profile | SiliconFlow - Products, Competitors, Financials, Employees |
| SU015 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing |
| SU016 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The platform brings together over 400 models spanning text, image, audio, and video generation. |
| SU017 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SU018 | SiliconFlow | China homepage | 面向不同行业及需求场景,提供灵活的解决方案 |
| SU019 | SiliconFlow Docs | Product introduction | Our platform empowers developers and enterprises to focus on product innovation |
| SU020 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models |
| SU021 | China Money Network | SiliconFlow's role in China's AI ecosystem | SiliconFlow's role in China's AI ecosystem |
| SU022 | SiliconFlow | International site | Get your Model API fast |
| SU023 | Tencent News | Funding and ecosystem article | 客户名单涵盖能源、金融、交通等核心行业的头部央企 |
| SU024 | GitHub | SiliconFlow organization | High-performance inference for every SOTA model. |
| SU025 | SiliconFlow | Reserved instances page | 预留实例 |
| SR001 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | Our major purchases are concentrated with the top five suppliers. |
| SR002 | SiliconFlow Docs CN | 服务协议 | 服务适用于面向开发者的内容生成服务场合,不适用于自动控制、医疗信息服务、心理咨询和关键信息基础设施的场合。 |
| SR003 | SiliconFlow Docs CN | 隐私政策 | 根据中华人民共和国大陆地区相关法律法规的规定,我们需要对您进行实名认证。 |
| SR004 | SiliconFlow CN | China homepage | 京ICP备2024051511号-1 增值电信业务经营许可证:京B2-20242084 |
| SR005 | SiliconFlow Docs EN | Terms of Use | The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn. |
| SR006 | SiliconFlow Docs EN | Privacy Policy | When you register, log in, and use the services provided through the https://siliconflow.com platform, we will collect and store the relevant information. |
| SR007 | Regulations.ai | Measures for the Identification of AI-Generated (Synthetic) Content | The Measures set a mandatory national baseline requiring that AI-generated or AI-synthesized content ... be clearly identified to users via explicit and implicit identification. |
| SR008 | U.S. BIS | Commerce strengthens restrictions on advanced computing semiconductors | These rules reinforce and build on the October 7, 2022, October 17, 2023, and December 2, 2024, controls to restrict the PRC's ability to obtain certain high-end chips. |
| SR009 | U.S. Department of Commerce | Guidance on Advanced Computing Items (May 2026) | a license is required to export advanced computing items to entities headquartered in Country Group D:5 or Macau |
| SR010 | NIST | AI Risk Management Framework | On April 7, 2026, NIST released a concept note for an AI RMF Profile on Trustworthy AI in Critical Infrastructure. |
| SR011 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing |
| SR012 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SR013 | QQ News / Taimei | 问AI · Token工厂模式为何吸引全产业链巨头联手投资? | 当前市场在快速膨胀,但竞争也在加剧。 |
| SR014 | SiliconFlow | Guizhou Mobile strategic cooperation | 双方在 2025 年战略合作基础上再升级 |
| SR015 | SiliconFlow | Reserved instances support 100B-token daily workload | 配合企业级交付与运行保障、明确的 SLA 与完善的售后服务 |
| SR016 | SiliconFlow | Energy SOE AI infrastructure case | 实现完整的性能追踪与故障预警,保障复杂生产环境下的高可用与安全性。 |
| SR017 | SiliconFlow | Airline group AI infrastructure case | 模型更新周期从“数周”压缩至“3 天内” |
| SR018 | SiliconFlow Docs | Product introduction | Provides comprehensive monitoring and fault tolerance mechanisms to guarantee service capabilities. |
| SR019 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SR020 | SiliconFlow Docs | API reference home | we will make periodic changes to the models provided by this service |
| SR021 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素依次是:模型性能、安全合规要求、回答质量、在AI平台可用性以及成本效益。 |
| SR022 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The daily average token call volume has reached trillions. |
| SR023 | China Money Network | SiliconFlow's role in China's AI ecosystem | SiliconFlow's role in China's AI ecosystem |
| SR024 | SiliconFlow Docs | Scenarios and application cases | Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases. |
| SR025 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SR026 | Fireworks AI | Pricing | H100 80 GB GPU $7.00 per hour. |
| SR027 | AWS | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SR028 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SR029 | OpenRouter Docs | Quickstart | OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers. |
| SR030 | Together AI Docs | Dedicated model inference pricing | Dedicated model inference bills based on the hardware your deployments run on, regardless of model or request volume. |
| SR031 | Regulations.ai | Provisions on the Administration of Deep Synthesis of Internet Information Services | Providers must implement real-identity verification before allowing publishing privileges, and all synthetic content must carry clear technical marks indicating its origin. |
| SR032 | China Law Translate | Measures for Labeling of AI-Generated Synthetic Content | Service providers shall clearly explain the methods, styles, and other such specifications for labeling generated synthetic content in user service agreements, and notify users to carefully read and understand the corresponding labeling management requirements. |
| SV001 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | For our public cloud-based services, the gross loss margin was 119.0% in 2025. |
| SV002 | QQ News / Taimei | 问AI · Token工厂模式为何吸引全产业链巨头联手投资? | 6月16日 ... 硅基流动宣布完成超20亿元B轮融资。 |
| SV003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The platform brings together over 400 models spanning text, image, audio, and video generation. |
| SV004 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing |
| SV005 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SV006 | Reuters / U.S. News | Together AI raises $800 million at $8.3 billion valuation | Together AI ... raised $800 million ... at $8.3 billion valuation ... annual bookings crossed $1.15 billion last quarter. |
| SV007 | BusinessWire | Together AI raises $800 million at $8.3 billion valuation | Together AI Raises $800 Million at $8.3 Billion Valuation |
| SV008 | BusinessWire | Fireworks AI raises $250M Series C | Fireworks AI Raises $250M Series C to Lead the AI Inference Market |
| SV009 | CNBC | Fireworks hits $17.5 billion valuation and $1B in annualized revenue | The company said ... it has exceeded $1 billion in annualized revenue, and it has now raised a $1.5 billion round at a $17.5 billion valuation. |
| SV010 | Sacra | Fireworks AI revenue, valuation & funding | At more than $1B in annualized revenue in July 2026, the latest valuation implies an approximate 17.5× revenue multiple. |
| SV011 | BusinessWire | Baseten raises $300M at a $5B valuation | Baseten Raises $300M at a $5B Valuation to Power a Multi-Model Future |
| SV012 | Sacra | Baseten revenue, valuation & funding | Baseten hit $600M in annualized revenue in March 2026 ... valued at $13B following its $1.5B Series F in June 2026. |
| SV013 | Sacra | CoreWeave revenue, valuation & funding | CoreWeave guided full-year 2026 revenue of $12B–$13B. |
| SV014 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SV015 | SiliconFlow Docs | Product introduction | Our platform empowers developers and enterprises to focus on product innovation while eliminating concerns about exorbitant computational costs. |
| SV016 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SV017 | Fireworks AI | Pricing | H100 80 GB GPU $7.00 per hour. |
| SV018 | AWS | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SV019 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SV020 | OpenRouter Docs | Quickstart | OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers. |
| SV021 | SiliconFlow | China homepage | 大模型云服务 ... 预留实例 ... 私有化大模型服务平台 |
| SV022 | SiliconFlow | Guizhou Mobile strategic cooperation | 双方在 2025 年战略合作基础上再升级 |
| SV023 | SiliconFlow | Reserved instances support 100B-token daily workload | 其 Coding Agent 单日 Token 消耗冲上千亿量级。 |
| SV024 | SiliconFlow Docs | Scenarios and application cases | Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases. |
| SV025 | Regulations.ai | Measures for the Identification of AI-Generated (Synthetic) Content | The Measures set a mandatory national baseline requiring that AI-generated or AI-synthesized content ... be clearly identified. |
| SV026 | U.S. BIS | Commerce strengthens restrictions on advanced computing semiconductors | These rules ... restrict the PRC's ability to obtain certain high-end chips critical for military advantage. |
| SV027 | NIST | AI Risk Management Framework | On April 7, 2026, NIST released a concept note for an AI RMF Profile on Trustworthy AI in Critical Infrastructure. |
| SV028 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素 ... 在AI平台可用性以及成本效益。 |
| SV029 | SiliconFlow Docs CN | 服务协议 | 服务适用于面向开发者的内容生成服务场合,不适用于 ... 关键信息基础设施的场合。 |
| SV030 | CNBC | Fireworks valuation and industry context | By managing computing infrastructure for models, Fireworks does business in the inference cloud market, alongside startups such as Baseten and Together AI. |