SiliconFlow
中国 AI 推理基础设施公司,已有真实规模和战略相关性;但公有云经济性仍为负,公开留存数据偏薄,当前估值也要求投资人把价格纪律牢牢系在证据上。
SiliconFlow 像是一家真实且越来越重要的中国推理平台,但 2026 年 6 月披露价格已经要求投资人承销未来利润率改善、企业客户耐久性,以及供应 / 合规执行;公开记录尚未完全证明这些。
封面要素
公司概况
SiliconFlow 是一家 2023 年 8 月在北京创立的 AI 推理基础设施公司。公开证据显示,它的业务触点多元:按量付费模型 API、预留或专用推理容量、加速服务,以及面向企业或合规敏感工作负载的私有化部署。到 2026 年中,公司已经跑出有意义的采用规模:超过 1000 万注册用户、上市申请期披露的超过 13,000 家企业客户,以及超过 200 个模型的公开模型库。公司还披露了截至 2026 年 6 月的七轮融资,其中 Series B/B+ 序列把申请文件披露的投后估值推至 RMB7.74 billion;独立媒体则称该轮约 US$296 million、估值约 US$1.2 billion。SiliconFlow 在中国推理栈中具有战略相关性,但公开记录仍显示公有云经济性偏弱、供应商依赖较重,留存和集中度披露也不完整。
- 成立时间
- 2023-08-29
- 创始人
- Yuan Jinhui
- 创立地点
- Beijing, China
- 总部
- Beijing, China
- 产品
- SiliconFlow 提供 API 优先的推理平台,按量接入大语言、多模态、图像、音频和视频模型,同时提供预留实例、推理加速服务和面向企业工作负载的私有化部署选项。
- 客户
- 开发者、AI 构建者、企业平台团队、受监管或有国资关联的机构,以及需要模型接入、可预测推理性能、私有化部署或国产芯片适配的电信 / 基础设施伙伴。
- 商业模式
- 基于使用量的 serverless API、预留或专用容量、私有化部署方案。公开证据显示,公司业务大致分为公有云服务和本地部署两条线,后者目前毛利率明显更好。
- 阶段
- Late-stage private; June 2026 Series B / B+ financing and active HKEX Chapter 18C filing
- 融资情况
- 公开记录支持公司在 2023 年 12 月至 2026 年 6 月间从天使轮到 Series B+ 快速融资爬升。HKEX 申请文件显示投后估值一路升至 RMB7.74 billion,2025 年后累计现金对价约 RMB1.48 billion;外部媒体则将 2026 年 6 月融资描述为约 US$296 million、估值 US$1.2 billion。
执行摘要
主要优势
- SiliconFlow 在 AI 推理中具备真实战略相关性:模型访问广、部署模式多,并有实质企业 / 基础设施用例,不只是 demo API。
- 对一家 2023 年成立的公司,公开采用信号很强:注册用户超过 10 million,申报期披露企业客户超过 13,000 家,token 吞吐足以支撑专属容量案例。
- 公司把模型分发、私有部署和国产芯片适配叙事合在一起,较适合中国特定的推理需求。
- 2026 年 6 月融资财团和 HKEX 申报流程显示,SiliconFlow 具备真实资本市场相关性,也还有时间继续执行。
主要风险
- 公有云经济性仍弱:2025 年综合毛利率为负,公有云业务线仍深度亏损,说明规模尚未解决增长质量。
- 供应商和算力依赖仍重要,因为公司租用或租赁关键基础设施,并未完全控制自己的半导体供应链。
- AI 标识、隐私、身份、电信许可和境内外表面管理,都让监管与合规复杂度很高。
- 客户广度已经公开,但留存、集中度和收入质量指标仍太薄,不足以支撑高确信承销。
- 如果现有使用规模不能转成更高毛利的专属 / 私有部署和耐久企业留存,披露估值仍可能显贵。
未决问题
- 当前 2026 年收入运行率,以及 serverless、专属、私有部署分产品线组合和毛利率趋势。
- 企业留存、logo 集中度、支出集中度,以及 API 向专属容量扩张指标。
- 2026 年 6 月融资后的当前现金余额、月度烧钱速度和融资情景分析。
- 安全 / 保障包:正常运行时间历史、事故历史、SOC 2 / ISO 或同等证据。
- 出口管制或容量冲击情景下的供应商集中条款和应急计划。
目录
01公司概览
1.1 身份、法律足迹与商业模式
SiliconFlow 的公开身份已经比普通 API 创业公司复杂。面向中国的域名展示的是北京 SiliconFlow:一家 AI 能力提供商,在海淀有具体地址,ICP 备案和北京增值电信业务许可证都可见;全球条款和隐私页面则由 SiliconFlow Technology Pte. Ltd. 发布,并明确把中国大陆用户导回 siliconflow.cn。HKEX 申请文件又加了第三层:一家中国股份有限公司,在北京有注册办公室,在香港有营业地点,但总部、高管和运营仍集中在香港以外。落到实操上,公司扎根中国、面向全球营销,也已经开始为跨境资本市场准入搭架构。 底层商业模式比法律地图更清楚。中国官网、全球官网、GitHub profile 和 API 文档都持续把 SiliconFlow 包装为推理基础设施层,而不是模型实验室或终端应用厂商。它卖的是一组标准化模型接入能力:大模型 API、专用或预留实例、推理加速服务和私有化部署。文档强调快速上手、API key 自助和 OpenAI 风格接口;营销页面强调速度、成本控制、多模态覆盖和可预测的按量交付。这个组合很重要,因为它显示公司想卡在模型供应商、算力供应商和下游开发者或企业之间,成为中立的 token 供给平台。[CO001, CO002, CO003, CO004, CO005, CO006]
| 指标 | 数值 / 状态 | 日期 / 期间 | 置信度 | 缺口或注意事项 |
|---|---|---|---|---|
| 法律设立 | 2023-08-29 中国有限责任公司;2026-06-22 改制为股份有限公司 | 2023-08 至 2026-06 | 高 | 改制清楚;IPO 后最终架构仍取决于上市完成 |
| 注册总部 | 北京海淀区 | 当前公开记录 | 高 | 北京和香港之外的分国办公室清单尚未完全公开 |
| 全球法律主体 | .com 条款由 SiliconFlow Technology Pte. Ltd. 发布 | 当前公开记录 | 高 | 中国大陆与新加坡主体之间的公开实体图谱仍不完整 |
| 最新资本市场阶段 | HKEX 第 18C 章申请人 / 商业化前上市流程进行中 | 2026-06 至 2026-07 | 高 | 上市尚未获批或完成 |
| 最新披露估值档位 | RMB 7.74B 投后(Series B+) | 2026-06-09 / 2026-06-18 序列 | 高 | 外部标题常改称超 RMB2B 的 Series B |
| 2025 年后收到的融资现金 | ~RMB 1.48B,覆盖 A+、B 和 B+ | 2025-12-31 之后 | 高 | 不包括 2026 年前更早轮次 |
| 注册用户 | 10.28M+ | 2026-04-30 | 高 | 时点指标;当前运行日期水平未披露 |
| 企业客户 | 已服务 13,000+ | 2026-06-30 申报前最后实际可行日期 | 高 | 未按付费状态、队列或集中度公开拆分 |
| 模型覆盖 | 申报文件中 170+;2026 年 7 月模型库为 200+ | 2026-06 至 2026-07 | 高 | 公开网站看起来比申报快照更新 |
| Token 吞吐量 | 日均 578.5B;单日峰值 1,071.4B | 2026-04 | 高 | 使用规模不直接意味着单位经济健康 |
| 海外商业化 | 月度海外收入超过 US$1M | 2026-05 | 高 | 未公开披露地理或产品组合 |
| 当前员工数 | 留存来源无法公开支撑 | 截至 2026-07-22 | 高 | 需要管理层确认 |
结合申报文件支撑的运营指标、公开网站和融资背景。申报文件与媒体快照不一致时,本表保留差异,而不是抹平。
[CO001, CO002, CO003, CO004, CO005, CO013]该流程图连接 SiliconFlow 的法律足迹、平台架构、客户基础、投资方阵容和上市轨迹。
1.2 领导力、治理与控制信号
由于 HKEX 申请文件,领导层公开披露比典型私营创业公司更好,但关键尽调缺口仍在。袁进辉博士显然是核心人物:创始人、董事长、CEO、总经理兼财务负责人。申请文件还列出 CTO 兼执行董事刘俊诚、商业化负责人兼执行董事曾华,以及与 Alibaba 战略投资部门相关的非执行董事陈英杰。正式董事会之外,历史和员工激励章节指向更宽的创始及核心团队内核,包括赵真和胡健,两人都与早期 OneFlow 网络相关。这一点重要,因为 SiliconFlow 看起来并不是人手很轻的 API 聚合器,而是一支技术主张很强、带有既有系统软件谱系的团队。 履历进一步强化了这个判断。袁进辉来自 Microsoft China 和 OneFlow;刘俊诚也曾在 OneFlow 从事多年 AI 系统软件工作;曾华补上来自 Microsoft China、Baidu 和 JD 的商业化经验;陈英杰则把 Alibaba 的战略资本桥接进来。上市后的董事会结构纸面上很直接——三名执行董事、一名非执行董事、三名独立非执行董事——但公开记录仍缺少控制权、创始人稀释、二级交易和董事会经济条款细节。还有一个小褶皱:第三方数据库对联合创始人名单并不完全一致,这让申请文件成为权威来源,也说明对一家高速变化的中国 AI 公司,创业公司数据库仍然噪声很大。[CO022, CO023, CO024, CO025, CO026, CO027]
| 人物 / 群体 | 当前角色 | 相关过往背景 | 重要性 | 关键人物或披露备注 |
|---|---|---|---|---|
| Yuan Jinhui 博士 | 创始人、董事长、CEO、总经理、财务负责人 | Microsoft China 首席研究员;创立 OneFlow | 推理平台论点的技术与战略重心 | 产品、财务和资本市场对 Yuan 的依赖极高 |
| Liu Juncheng | 执行董事、CTO | AI 系统软件与 OneFlow 研发背景 | 负责核心技术执行和架构延续 | 高度依赖核心系统软件人才 |
| Zeng Hua | 执行董事、副总经理 | Microsoft China、Baidu、JD 商业化背景 | 增加企业销售和 GTM 可信度 | 对把技术规模转化为变现很重要 |
| Chen Yingjie | 非执行董事 | Alibaba 战略投资董事总经理;前 PwC;曾任 XPeng 和 MiniMax 董事 | 代表主要平台投资者的战略资本与生态信号 | 董事席位和权利未在公开材料中完整披露 |
| Zhao Zhen | 首席运营官 / 联合创始人(据申报历史) | OneFlow 前 COO | 延续此前团队网络的运营连续性 | 公开申报中不是董事会成员;外部资料覆盖稀疏 |
| Hu Jian 和更广核心团队 | 副总经理 / 核心研发及员工激励参与者 | 核心团队与 OneFlow 时代人才基础重叠 | 说明团队厚度超过双创始人叙事 | 多名核心运营者的公开履历仍有限 |
| 独立董事名单 | Wu Chuan、Li Dan、Hou Hong(拟任 INED) | 学术、会计和战略背景 | 指向上市准备度和更深治理厚度 | 独立董事名单已在申报中披露,但委员会实践尚未验证 |
HKEX 申报文件显著提高了领导层可见度,但几名运营负责人和董事会控制权的经济安排仍只披露了一部分。
[CO022, CO023, CO024, CO025, CO026, CO027]1.3 融资历史与资本市场叙事
SiliconFlow 的融资节奏异常压缩。HKEX 申请文件记录了 2023 年 12 月至 2026 年 6 月的七轮融资,披露出资额从天使轮 RMB47.2 million,升至天使+ 轮 RMB65.84 million、pre-A 轮 RMB71.48 million、Series A 轮 RMB285.99 million、Series A+ 轮 RMB220 million、Series B 轮 RMB520 million、Series B+ 轮 RMB740 million。对应投后估值在约两年半内从 RMB280.0 million 升至 RMB7.74 billion。这不只是融资故事;它释放的信号是,当 AI token 中间层在中国变得可投资,投资人迅速重估了这家公司。 2026 年 6 月的新闻标题需要谨慎处理。外部媒体一致称 SiliconFlow 完成了超 RMB2 billion 的 Series B 轮,由覆盖 Trip.com 或 Ctrip、JinkoSolar、Kingdee、Unicom 关联投资方、Biren、NIO Capital、SenseTime、GGV 等的战略财团支持,China Renaissance 担任顾问。但 HKEX 申请文件随后给出更细的 2026 年 A+、B、B+ 融资序列,2025 年后累计现金对价约 RMB1.48 billion,估值从 RMB3.12 billion 升至 RMB5.02 billion,再到 RMB7.74 billion。公开数据库也有尚未对齐的总额,因此正确结论不是某个来源必然错了,而是公开资本历史方向清楚,交易标签细节仍需要在资料室里对账。[CO032, CO033, CO034, CO035, CO036, CO037]
| 利益相关方 / 投资者 | 角色或轮次背景 | 战略相关性 | 公开记录显示 | 尽调要求 |
|---|---|---|---|---|
| Alibaba / Hangzhou Duoxiang | 股东且与董事会关联的战略投资者 | 潜在云、分发和生态杠杆 | Alibaba 关联实体是重要股东;Chen Yingjie 任董事 | 确认确切持股、信息权和商业关系范围 |
| Trip.com / Ctrip 战略投资 | 2026 年 6 月具名战略投资者 | 旅游、企业需求和资本市场信号 | 2026 年 6 月融资报道反复具名 | 核验 Trip.com 投在 B、B+ 还是两者,并确认是否存在商业试点 |
| SenseTime 战略投资 | 2026 年 6 月具名投资者,也是现有 AI 生态同业 | 指向模型和 AI 基础设施生态协同 | 多篇 2026 年 6 月融资报告具名 | 检查关系是纯财务投资,还是包含工作负载、模型或渠道合作 |
| JinkoSolar Holding | 2026 年 6 月具名投资者 | 连接 AI 基础设施与电力、数据中心经济性 | 36Kr 将 Jinko 定位为能源到算力战略伙伴 | 判断是否包含实际基础设施合同,还是只有资本 |
| Unicom 关联资本 | 2026 年 6 月具名投资者群体 | 电信与算网融合可能帮助企业部署 | 融资报道提到 Unicom Xinwo 和 Unicom Capital 载体 | 澄清客户或网络资源承诺 |
| Biren Technology | 2026 年 6 月具名战略投资者 | 国产芯片适配是 SiliconFlow 差异化的核心 | 融资报道围绕国产芯片推理协作描述 Biren | 检查有约束力的商业条款、排他性或联合营销义务 |
| NIO Capital 和其他财务投资者 | 后期轮次增长资本 | 在战略企业之外增加财务支持 | 融资报道和 CB Insights 投资者表均有具名 | 要求完整股权表,以及跟投权或清算权 |
| China Renaissance | 2026 年 6 月融资独家顾问;后任 HKEX 整体协调人 | 衔接私募融资与上市准备 | Caixin Global 和 HKEX 协调人公告均有具名 | 判断顾问经济利益或流程依赖是否造成快速上市压力 |
公开信息强力支持投资者基础具有高战略密度,但不能说明这些投资者附带的控制权或商业义务。
[CO025, CO026, CO033, CO036, CO037, CO047]1.4 规模、牵引力与地域覆盖
申请文件和相互印证的报道显示,SiliconFlow 已经取得真实使用规模,尽管商业密度仍在演进。截至 2026-04-30,公司披露超过 1000 万注册用户,平均日吞吐量 578.5 billion tokens,峰值日吞吐量超过 1 trillion tokens,并在最新可行日期拥有超过 13,000 家企业客户。申请文件显示平台支持超过 170 个模型;到 2026 年 7 月,公开模型库已超过 200 个。对一家 2023 年 8 月才注册成立的公司来说,这些是异常大的基础设施侧运营指标,也解释了为什么即使盈利能力尚不成熟,投资人仍愿意激进出资。 与此同时,足迹证据很宽,但并不完整。公开页面明确把运营中心放在北京,申请文件也确认香港营业地址是 IPO 流程的一部分。官方和第三方材料还显示,海外商业化已不再是假设:申请文件称 2026 年 5 月海外月收入超过 US$1 million,6 月报道也反复提到数百万美元海外月收入和全球平台牵引力。仍缺的是按国家、办公室、员工人数或海外收入组合拆开的干净公开口径。因此,公司在使用规模上已经成规模,在企业采用上也足以重要,但组织足迹只披露了一部分。[CO013, CO014, CO015, CO016, CO017, CO018]
截至 2026-07-22 运行日,SiliconFlow 的关键成熟度、规模和警示指标。
1.5 里程碑、合规信号与不利背景
运营时间线是连贯的。SiliconFlow 称公司 2023 年 8 月启动推理引擎研发,2024 年 5 月推出公有云 MaaS,2025 年 2 月在 Huawei Ascend 上交付 DeepSeek 推理,2025 年 9 月推出私有 MaaS,2026 年 4 月发布 Elastic GPU scheduler,2026 年 6 月改制为股份有限公司,6 月 30 日递交 HKEX 申请版本,7 月 14 日又加入 China Renaissance 作为整体协调人。这条时间线显示,公司在产品化、商业扩张、融资和上市准备上同步推进。 不利框架同样重要。HKEX 财务披露显示,2025 年收入迅速增至 RMB55.3 million,但毛利率转负至 -24.0%。KrASIA 和 Hello China Tech 把问题说得更尖锐:SiliconFlow 似乎在租用算力、分发第三方模型,并卷入一场价格战;开发者补贴和低于成本的 token 供给能比盈利更快拉动采用。公有云经济性看起来比本地部署经济性结构性更难,这意味着公司的亮眼增长尚未回答可持续变现问题。换句话说,公司概览支持的是一家严肃的基础设施平台,背后有顶级战略投资人,但商业模式还没有完全风险出清。[CO039, CO040, CO042, CO043, CO044, CO045]
| 日期 | 事件 | 类型 | 金额 / 状态 | 参与方 | 含义 |
|---|---|---|---|---|---|
| 2023-08-29 | 公司在北京成立 | 创立 | 中国有限责任公司成立 | Yuan 牵头的创始团队 | 为后续产品、融资和上市工作搭建法律基础 |
| 2023-08-01 | 推理引擎研发启动 | 产品 | 引擎开发启动 | 创始工程团队 | 说明公司从系统层而非应用层起步 |
| 2024-05-01 | 公有云 MaaS 上线 | 产品 | 商用公有云服务上线 | SiliconFlow | 标志着从研发项目转向外部平台 |
| 2025-02-01 | 基于 Huawei Ascend 的 DeepSeek 服务上线 | 产品 | 国产芯片 Token 服务里程碑 | SiliconFlow、DeepSeek、Huawei Ascend | 强化国产芯片和异构算力定位 |
| 2025-07-10 | Series A 对价全额支付 | 融资 | ~RMB285.99M Series A | 申报文件所列私人投资者 | 推动估值至 RMB2.286B,并支持更广商业化 |
| 2025-09-01 | 私有 MaaS 上线 | 产品 | 私有化部署产品上线 | SiliconFlow 企业团队 | 新增更重交付、面向合规的收入路径 |
| 2026-04-01 | Elastic GPU 上线 | 产品 | 异构调度引擎上线 | SiliconFlow | 支撑 Token 工厂效率论点 |
| 2026-06-17 | Series B 对价全额支付 | 融资 | ~RMB520M Series B | 2026 年投资者银团 | 将披露投后估值推至 RMB5.02B |
| 2026-06-18 | Series B+ 对价支付 | 融资 | ~RMB740M Series B+ | 2026 年后期投资者 | 上市前夕,将披露投后估值推至 RMB7.74B |
| 2026-06-30 | HKEX 申请版本递交 | 治理 | 第 18C 章商业化前申报 | SiliconFlow、Huatai、Haitong | 形成第一套深度公开披露材料 |
| 2026-07-14 | China Renaissance 被增聘为整体协调人 | 治理 | 上市流程扩展 | SiliconFlow、China Renaissance | 表明 IPO 准备继续积极推进 |
这是概览章节唯一的记录年表。日期来自申报文件,或直接来自 2026 年 6–7 月公告;仅披露月份的里程碑以该月第一天作为锚点。
[CO001, CO003, CO032, CO033, CO034, CO042]SiliconFlow 从 2023 年 8 月到 2026 年 7 月的成立、产品发布、融资步骤和上市里程碑年表。
1.6 图表
02市场分析
2.1 市场边界、纳入支出与替代方案
SiliconFlow 卖的不是泛泛的「AI」产品;它坐在推理中间件层,把第三方模型和租用或客户自有算力转成按量计费的 token 输出。2026 年 6 月 HKEX 申请文件、SiliconFlow 模型目录和上手文档都指向同一个市场边界:公有云 serverless token API、专用推理实例,以及让企业或开发者通过一个接口消费多模型的私有化部署。这意味着纳入的支出不是模型训练、半导体设计或终端 AI SaaS 席位,而是与模型服务、token 分发、编排及邻近工具相关的支出;这些工具减少 AI 在生产环境落地时的运营工作。 因此,最重要的替代方案不只是其他创业公司。Amazon Bedrock、Microsoft Foundry、Alibaba Cloud Model Studio 等封闭生态的超大规模云平台争夺同样的模型接入和 agent 构建预算;Together、Fireworks 等独立推理平台则在开发者可迁移性、速度和价格上竞争。在很多买方路径里,现状是直接使用模型实验室 API、内部自托管,或沿用一家已在安全和采购上被信任的现有云。SiliconFlow 的市场之所以重要,是因为它夹在这些选项之间:比单一云栈更开放,又比自建推理层更产品化。[CM001, CM002, CM003, CM004, CM005, CM006]
| 细分 / 类别 | 纳入支出 | 排除支出 | 买方 / 付款方 | 重要性 |
|---|---|---|---|---|
| 公有云 MaaS / Token API | Serverless Token 调用、专用实例、模型访问与路由 | 模型训练、芯片、终端用户 AI SaaS 席位 | 开发者、初创公司、产品团队 | 这是 SiliconFlow 核心公有云收入路径 |
| 私有 / 本地 MaaS | 客户环境中的部署软件和企业专属 Token 工厂 | 通用 IT 外包或无关云迁移 | 受监管企业、国企、金融、公共部门 | 覆盖需要合规和供应控制的买方 |
| 开发者分发层 | 应用、SDK、编排工具、预集成市场中的 BYOK 使用 | 封闭单应用订阅 | 开发者和工具运营方 | 降低切换摩擦,扩大需求捕获 |
| 超大规模云托管模型平台 | 企业治理、Agent 工具、模型目录、路由、护栏 | 消费者 AI 助手 | 云平台所有者、CIO 预算 | 企业级采购的主要替代集合 |
| AI 推理基础设施背景 | 底层推理硬件、云容量和服务软件 | 仅训练基础设施和非 AI 计算 | CSP、基础设施买家 | 适合做 TAM 背景,但对 SiliconFlow 直接 SAM 来说太宽 |
边界锚定在 SiliconFlow 的公有云和私有化部署定位上,再扩展到买方采购时真正比较的替代集合。
[CM001, CM002, CM003, CM004, CM005, CM006]推理市场价值链:从模型和算力供给,到平台、工具,再到企业生产部署。
[CM004, CM006, CM007, CM022, CM027, CM028]2.2 通过中国 MaaS 与全球推理视角测算市场
没有一个公开数字能干净捕捉 SiliconFlow 的机会,因此更稳的做法是保留多个视角,而不是硬凑一个 TAM。从中国需求视角看,IDC 称企业 MaaS token 消费量从 2024 年的 114 trillion tokens 跳至 2025 年的 1,944 trillion,并预计 2026 年约 40,000 trillion;公有云 MaaS 收入在 2025 年达到 RMB3.07 billion,2026 年达到 RMB18.6 billion。从申请文件视角看,HKEX 招股书中的 Frost & Sullivan 行业研究显示,中国 token 供给市场 2024 至 2025 年增长 1,602.6%,到 2030 年达到 53.2 quintillion tokens。两个以中国为中心的视角大体同意爆发式扩张,但定义和基线并不相同,这正是它们应该被交叉校验、而不是简单平均的原因。 从更宽的全球视角看,MarketsandMarkets、Grand View Research 和 Fortune Business Insights 的第三方分析页面都把 AI 推理市场聚在 2024-2025 年约 USD100 billion 的量级,预测终点从 2030 年约 USD254 billion 到 2034 年超过 USD312 billion 不等。这个更大的全球数字有上下文价值,但会高估 SiliconFlow 近期可直接触达的市场,因为其中很大一部分包含硬件和超大规模云基础设施支出,层级远高于独立 token 分发层。更好的承销逻辑是分层:全球推理基础设施用于背景,中国公有云 MaaS 用于可变现需求,SiliconFlow 2025 年 1.5% 吞吐份额则是当前竞争相关性的信号,而不是直接收入份额桥。[CM009, CM010, CM011, CM012, CM013, CM014]
| 口径 / 发布方 | 地域 | 数值 | CAGR / 增长 | 口径提示 | 置信度 | 局限 |
|---|---|---|---|---|---|---|
| IDC 公有云 MaaS 收入 | 中国 | RMB3.07B(2025)至 RMB18.6B(2026) | 同比 ~6.1x | 公有云上的企业 MaaS 收入口径 | 中 | 收入口径不含私有化部署 |
| IDC token 消耗 | 中国 | 1,944T tokens(2025);约 40,000T(2026) | 2025 年 ~16x;2026 年 ~20x | token 吞吐量 / 使用量口径 | 高 | 使用量不等于货币化收入 |
| Frost & Sullivan(经 HKEX) | 中国 | 2025 年总市场 2,426.3T tokens;2030 年 53.2 quintillion | 2024-2025 增长 1,602.6%;2025-2030 CAGR 638.3% | token 供给市场吞吐量 | 中 | 定义不同于 IDC MaaS 口径 |
| SiliconFlow 份额信号(经 HKEX) | 中国 | 2025 年吞吐量份额 1.5%;整体排名 #4,独立平台 #1 | n/a | 公司在 token 供给平台中的排名 | 高 | 份额按吞吐量计算,不按收入计算 |
| MarketsandMarkets AI 推理 | 全球 | USD106.15B(2025)至 USD254.98B(2030) | 19.2% | 广义 AI 推理基础设施市场 | 中 | 口径过宽,不能直接作为 SiliconFlow SAM |
| Grand View AI 推理 | 全球 | USD97.24B(2024)至 USD253.75B(2030) | 17.5% | 广义 AI 推理基础设施市场 | 中 | 基期比其他报告早一年 |
| Fortune Business Insights AI 推理 | 全球 | USD103.73B(2025)至 USD312.64B(2034) | 12.98% | 广义 AI 推理基础设施市场 | 中 | 预测窗口更长,离散度更高 |
几类口径要一起看,不能相加。中国 MaaS 需求、全球 AI 推理基础设施、SiliconFlow 份额信号回答的是不同问题。
[CM009, CM010, CM011, CM012, CM013, CM014]从广义全球 AI 推理基础设施,到 SiliconFlow 当前在中国 token 供给中的份额信号,分层查看。
各层只作方向性参考,不能相加,因为底层单位和市场定义不同。
[CM011, CM012, CM015, CM016, CM017, CM018]广义全球 AI 推理市场的低 / 基准 / 高三档来源支撑区间,统一为 USD 十亿美元。
单位为 USD 十亿美元。基准年取自 2024-2025 年报告页;预测终值使用保留来源中最接近的已发布 2030 或 2034 年数字。
[CM016, CM017, CM018, CM019]2.3 买方分层、工作流负责人和预算路径
申请文件清楚显示,SiliconFlow 的市场覆盖多个买方原型,而不是单一同质用户。低摩擦端是个人开发者和创业公司,他们要的是无需承诺的模型接入、快速上手、预付费使用,以及通过 BYOK 或兼容 API 保留既有工具的能力。高控制端是大型企业和机构,他们更看重专用性能、稳定供给、低延迟、私有化部署或供应保障。类似模式也出现在可比平台:Bedrock 面向创业公司和全球企业营销;Microsoft Foundry 主打 agent、治理和全资源池控制的统一资源平面;Alibaba 的 Token Plan 把推理转成按席位的团队生产力支出;Together 和 Fireworks 强调 OpenAI 兼容 API,以降低迁移工作。 这种组合意味着预算负责人不止一种。自助 API 使用的第一笔预算可能在开发负责人、创业公司创始人或团队生产力经理手里。部署规模变大或受到监管后,所有权会迁移到 CIO、平台工程负责人、采购或安全合规把关人。采用路径通常从一次快速 serverless 实验开始,走向更重的使用;当成本、稳定性或合规开始重要,再转向专用实例或私有化部署。SiliconFlow 的优势是能参与这条路径的多个阶段;挑战是每个阶段都会带来更强的现有云竞争,也更不能容忍产品或供应不稳定。[CM022, CM023, CM024, CM025, CM026, CM027]
| 客群 | 买方 | 用户 | 付款方 / 预算负责人 | 工作流 | 采用触发点 | 含义 |
|---|---|---|---|---|---|---|
| 个人开发者 | 开发负责人 / 创始人 | 工程师 | 积分余额或小团队预算 | 原型开发和直接 API 调用 | 需要无需长期承诺的快速模型访问 | 自助开通比合同更关键 |
| AI 原生创业公司与工具 | CTO / 产品负责人 | 应用团队 | 产品或基础设施预算 | 把多模型推理嵌入应用 | 需要可迁移性,也要更快上线功能 | BYOK 和预集成降低摩擦 |
| 企业 AI 平台团队 | CIO / 平台工程 | 内部产品团队 | 云 / 平台预算 | 在各业务部门统一模型访问 | 需要治理、路由和可观测性 | 直接与既有云厂商竞争 |
| 受监管企业 / 机构 | CIO + 安全 / 合规 | 业务部门和知识工作者 | 企业采购 | 专属或私有化部署 | 需要稳定性、低延迟或数据控制承诺 | 价值更高,但销售周期更长 |
| 团队生产力 / 研发组织 | 工程经理或团队管理员 | 开发者 | 席位 / 订阅预算 | 基于积分的日常 AI 使用 | 需要可预测支出和团队管控 | 订阅打包可扩大预算池 |
买方图谱把 SiliconFlow 自身分层,与 AWS、Microsoft、Alibaba、Together、Fireworks 面向不同预算负责人的推理打包方式放在一起看。
[CM022, CM023, CM024, CM025, CM026, CM027]按预算负责人、购买触发、运营要求和迁移路径映射细分市场。
[CM022, CM023, CM024, CM025, CM026, CM027]2.4 增长驱动、采用约束与估值相关性
最重要的顺风同时出现在宏观和平台层来源中。Stanford 2026 AI Index 称,2025 年前沿模型进步没有停滞,组织采用率达到 88%;IDC 则认为,中国 MaaS 竞争已不再只是便宜 token,而是价格、性能和工具链支持的组合包。Bedrock、Foundry、Alibaba、Together 和 Fireworks 的官方产品页也强化了这种转向:它们把路由、可观测性、治理、隐私控制、专用吞吐、缓存和模型评估作为核心功能出售。这对 SiliconFlow 很重要,因为市场正从原始 API 接入成熟为运营要求更高的一层;编排能力和企业适配越来越能变现。 约束同样实质。IDC 仍把性能、安全与合规、答案质量、平台可用性和成本效益列为企业选型的顶级因素。Fortune 的市场页面强调硬件成本和集成复杂度;申请文件和独立分析则显示,对中立平台来说,这些约束会如何落到经济性上:即便供应商仍在租用算力、补贴开发者、毛利率为负,公有云增长也可能是真的。Hello China Tech 的警告更锋利,称 2023-2026 年模型 API 和算力支撑的 token 供给经历了价格战。放到估值上,这意味着市场可以巨大,却未必会顺畅转化为有吸引力的独立经济性业务,除非平台能守住定价、提高利用率,并随时间赢下更高价值的企业工作负载。[CM029, CM030, CM031, CM032, CM033, CM034]
| 驱动因素 / 约束 | 方向 | 时间 | 证据 | 重要性 | 尽调问题 |
|---|---|---|---|---|---|
| 前沿模型进步和 AI 广泛采用 | 驱动 | 当前 | Stanford AI Index 2026 | 模型能力越强,推理需求越大 | 新增用量中,多少会留在开放 / 多模型形态? |
| 多模态和 Agent 工作负载 | 驱动 | 当前至中期 | IDC + 超大云厂商产品页 | 提高 token 强度,拓宽使用场景 | 哪些工作负载对 SiliconFlow 价值最高? |
| 价格 + 性能 + 工具链竞争 | 驱动兼筛选 | 当前 | IDC | 胜出者不能只靠低 token 价格 | SiliconFlow 工具链差异化有多强? |
| 治理、隐私和企业控制 | 驱动 | 当前 | AWS + Microsoft + Alibaba | 企业买家越来越把这些功能视为必选项 | SiliconFlow 的控制能力是否足以比肩对手、拿下大客户? |
| 开发者可迁移性和 OpenAI 兼容性 | 驱动 | 当前 | SiliconFlow + Together + Fireworks | 降低采用门槛和迁移成本 | 易切换是在帮 SiliconFlow 获客,还是削弱护城河? |
| 硬件成本和集成复杂度 | 约束 | 当前 | Fortune Business Insights | 推理规模化仍然吃运营成本 | SiliconFlow 能结构性削掉多少成本? |
| 价格战和补贴式获客 | 约束 | 当前 | HKEX、Hello China Tech 与 KrASIA | 市场规模可能跑得比利润池更快 | 单位经济何时转正? |
| 算力供给和异构编排 | 约束 / 差异化点 | 当前至中期 | HKEX + 36Kr | 供给获取能力会影响利润率和可靠性 | 供应商关系和芯片覆盖能撑多久? |
多个因素都有两面性:一边放大市场,一边抬高独立平台的执行门槛。
[CM029, CM030, CM031, CM032, CM033, CM034]2.5 冲突估算与剩余尽调缺口
有纪律的市场章节应该保留未知项。第一个缺口是定义:IDC 的中国 MaaS 框架、Frost & Sullivan 的 token 供给框架和全球 AI 推理报告都有用,但它们衡量的支出或活动篮子不同。第二个缺口是公司特定的:已留存公开来源没有按地域、客户细分或模型类别拆出 SiliconFlow 的可服务收入份额,也没有展示能把总吞吐变成可持续份额论点的留存或切换动态。 实际结论是,市场证据在需求方向层面最强,在独立平台精准利润池层面最弱。SiliconFlow 显然参与了一个高速增长的市场;仅凭公开来源,很难证明这部分增长会有多少归属于中立 token 平台,而不是超大规模云厂商、模型实验室或大型企业内部的私有化部署。因此,投资人应把市场规模当作支撑性背景,但把承销确信度留给后续章节中的单位经济性、细分组合和防御性证据。[CM021, CM041, CM042]
2.6 图表
03竞争对手
3.1 竞争格局分层
SiliconFlow 的竞争集合至少横跨四类。第一类是 Together AI、Fireworks、OpenRouter 等独立开放推理平台,它们比拼模型广度、开发者可迁移性、速度和定价,而不是自有专有前沿模型。第二类是 Amazon Bedrock、Microsoft Foundry、Alibaba Cloud Model Studio 等现有云平台,卖的是同一类大工作——让企业接入多模型——但把它打包进更大的治理、网络、身份和采购栈。第三类是直接的模型实验室 API 和单一厂商生态;当客户只想要一个偏好的模型,而不是模型市场或代理层,它们可以完全绕过中立平台。第四类是内部自建:如果工程团队认为成本、控制或性能足以抵消复杂度,他们可以自托管开放模型,或把多个供应商 API 拼在一起。 SiliconFlow 自己的申请文件很清楚:公司想坐在开放、中立的中间位置。它明确对比封闭生态,声称在中国独立生态 token 供给平台中排名第一,并把自己定位为连接异构算力和多模型家族的一层。这意味着最相关的竞争问题不只是「谁的模型目录最大」,而是买方从实验走向生产后,哪一类供应商会胜出。早期开发者采用阶段,独立平台和代理层可能看起来可以互换。到了受监管或超大规模部署,天平往往转向治理、采购触达和供给保障更强的供应商。[CP001, CP002, CP003, CP004, CP005, CP006]
| 竞争对手 | 类型 | 规模 / 融资信号 | 目标客群 | 差异化 | 局限 / 尽调备注 |
|---|---|---|---|---|---|
| SiliconFlow | 独立 token 供给平台 | HKEX 递表公司;2025 年中国吞吐量份额 1.5%;最新披露估值档位 RMB7.74B | 开发者、创业公司、企业、私有化部署 | 中国本土中立性、异构芯片、200+ 模型、国产芯片适配 | 规模远小于超大云厂商;申报文件显示公有云经济性仍为负 |
| Together AI | 独立开放模型平台 | 私营公司;本次留存来源更强调产品,未锁定当前融资或收入 | 从无服务器扩到专属 GPU 的开发者和 AI 原生构建者 | 聚焦开源,无服务器模型目录强,提供专属端点,也主打速度 | 本次留存的独立来源未证实当前规模或利润率 |
| Fireworks AI | 独立开放模型平台 | Fireworks 称每日处理 40T+ tokens;本次留存来源未证实当前融资或收入 | 开发者、有训练 / 推理需求的企业、按需 GPU 用户 | 专项训练 + 推理、无服务器和按需部署、微调、性能定位 | 吞吐量说法来自厂商,本报告未独立校验 |
| OpenRouter | 中介 / 路由 / 聚合器 | 跨数百个模型的统一 API;留存来源更强调路由,未披露规模 | 优化成本、延迟或回退行为的开发者和 Agent 构建者 | 供应商路由、按成本 / 延迟排序、多宿主、BYOK 友好设计 | 依赖外部供应商,不掌握核心算力供给 |
| Amazon Bedrock | 超大云厂商托管模型平台 | AWS 称 Bedrock 服务 100,000+ 家组织 | 从创业公司到全球企业的企业构建者 | 治理、护栏、模型选择、批处理 / 优先级选项、采购优势 | 中立性弱于独立中介,且绑定 AWS 资产栈 |
| Microsoft Foundry | 超大云厂商托管 Agent / 模型平台 | Microsoft 将 Foundry 定位为统一控制平面,提供 1,900+ 至 11,000+ 个模型访问入口 | 企业平台团队和 Agent 构建者 | RBAC、策略、可观测性、治理、模型目录、托管算力 | 留存文本中的定价页并不总能清楚展示 token 价格 |
| Alibaba Cloud Model Studio | 区域超大云厂商 / 模型平台 | Alibaba 拥有 Qwen,同时整合第三方模型和区域端点 | 中国及海外区域的团队和企业 | OpenAI 兼容性、多模态 Qwen 栈、区域部署、团队积分计划 | 平台经济性以及相对标价的实际折扣仍不清楚 |
画像行混合了上市公司既有厂商、开放独立同行和路由中介,因为买家会把它们放在同一个模型访问任务下评估。
[CP001, CP002, CP003, CP004, CP005, CP006]按开放性 / 可移植性,以及企业控制 / 采购强度做序数定位。
1-10 的序数分数由保留的产品、定价和治理证据综合得出;不是第三方基准。
[CP001, CP013, CP014, CP015, CP016, CP017]3.2 能力广度与价格竞争
能力重叠很大。SiliconFlow、Together、Fireworks、OpenRouter、Alibaba 及其他对手都呈现出 OpenAI 兼容 API、多模型接入和低摩擦上手的某种组合。SiliconFlow 宣传 200+ 模型;OpenRouter 称通过统一 API 提供数百个模型;Bedrock 称有 100+ foundation models;Microsoft Foundry 在文档中营销 1,900+ 模型,在定价界面上展示 11,000+ 模型;Alibaba 打包 Qwen 以及 DeepSeek、Kimi、GLM 等第三方模型;Together 和 Fireworks 各自结合 serverless 和更高控制度的部署路径。在这个市场,基础接入的功能平价正在变成入场券。 因此,定价和包装越来越有战略意义。Together 和 Fireworks 都提供按 token 的 serverless 定价,配套 cached-token 和 batch 折扣,再把重度用户推向专用硬件。AWS Bedrock 叠加 token 定价、batch 折扣以及 provisioned 或 reserved capacity。Alibaba 混合按量费率、区域折扣和 credit-based 团队计划。OpenRouter 则把定价本身变成产品的一部分,通过跨供应商路由并按价格、延迟或吞吐排序。最终结果是,买方通常能在多个供应商处找到可比的基础模型接入;但当部署模式、速率限制、治理、路由和性能控制进入决策,总体价值主张仍会明显不同。[CP010, CP011, CP012, CP013, CP014, CP015]
| 采购标准 | SiliconFlow | Together | Fireworks | OpenRouter | AWS Bedrock | Microsoft Foundry | Alibaba Model Studio |
|---|---|---|---|---|---|---|---|
| OpenAI 兼容 API | 是 | 是 | 是 | 可直接替换的 OpenAI SDK 路径 | 留存来源未覆盖 | 通过 OpenAI()/project 端点支持 | 是 |
| 广泛多模型目录 | 200+ 模型 | 是 | 100+ 开放文本 + 视觉等模型 | 数百个模型 | 100+ 模型 | 1,900+ 至 11,000+ 模型 | Qwen + DeepSeek/Kimi/GLM + 多模态 |
| 专属 / 预留容量 | 专属实例 + 私有化部署 | 专属模型推理 | 按需部署 | 路由至供应商,而非自有专属集群 | 预配置 / 预留档位 | 托管算力 / 预配置吞吐量 | 区域部署范围和团队计划 |
| 路由 / 回退控制 | 模型选择 + BYOK 通道 | 模型选择;从无服务器到专属 | 服务路径和分层 | 明确供应商路由和回退 | 提示词路由 | 模型路由器和统一控制平面 | 区域端点和计费控制 |
| 企业治理 / 控制 | 文档和申报文件部分公开 | 有部分文档;留存的信任证据有限 | 用量指标和仪表盘 | 有路由 / 数据控制,但企业栈较轻 | 护栏、隐私、合规 | RBAC、网络、策略、可观测性 | 数据隐私声明和监控 |
未获支持的单元格采用保守表述,只限于留存来源证据。
[CP010, CP011, CP012, CP013, CP014, CP015]| 供应商 | 示例计费单位 / 打包 | 留存示例价格 | 折扣 / 控制杠杆 | 含义 |
|---|---|---|---|---|
| SiliconFlow | 按模型计费的推理 API | 公开模型库展示按模型价格;实际成交费率不等于利润率 | BYOK 和工具集成 | 靠广度和便利性竞争,但实际经济性未披露 |
| Together | 无服务器按 token;专属按 GPU-minute | Qwen 3.7 Max 每 1M tokens 输入 $1.25 / 输出 $3.75;DeepSeek V4 Pro $1.74 / $3.48 | 缓存输入折扣和批处理折扣;高利用率下专属更便宜 | 开放模型 token 定价的强直接可比对象 |
| Fireworks | 无服务器按 token;按需 GPU-hour;微调 | DeepSeek V4 Pro $1.74 / 缓存 $0.145 / $3.48;H100 按需 $7.00/hr | Priority/Fast 档位、缓存 token 折扣、批处理为标准价 50% | 同时在 token 定价和更高控制度基础设施上竞争 |
| OpenRouter | 供应商路由的按 token API | 费率随底层供应商变化;路由器可按价格、吞吐量或延迟排序 | 回退、max_price、优先延迟 / 吞吐量、ZDR | 把路由逻辑变成商业方案的一部分 |
| AWS Bedrock | 按 token,加批处理 / 预配置选项 | Claude Opus 4.8 每 1M 输入 $6 / 输出 $30;DeepSeek v3.2 在列示区域 $0.62 / $1.85 | 批处理 50% 折扣;标准 / 优先 / 预留档位 | 企业既有厂商可同时覆盖高端和低成本模型 |
| Microsoft Foundry | 无服务器 / 托管算力 / 预配置 | 列出托管 GPU 系列;留存页面中许多 token 价格显示为 $-,并不清楚 | ACU 预购计划和资源级治理 | 商业模式是更宽的平台合同,不只是一条 API 费率 |
| Alibaba Model Studio | 按量付费加积分订阅 | qwen3.7-max 标价为输入 $2.5 / 输出 $7.5,并有临时折扣;团队席位每月 $30 至 $200 | 夜间 / 白天折扣、免费额度、共享积分包 | 激进的区域折扣和席位打包扩大竞争范围 |
示例价格为 2026-07-22 抓取的标价或公开价格,不应视为实际净价。
[CP018, CP019, CP020, CP021, CP022, CP023]能力对比,突出重叠度高的领域,以及仍有实质差异的地方。
定性标签仅来自保留的官方文档和价格页证据。
[CP010, CP011, CP012, CP013, CP014, CP015]3.3 切换成本、锁定与分发力
对独立推理平台来说,API 表面的切换成本结构性偏低,周边工作流层则高得多。SiliconFlow、Together、Fireworks、OpenRouter 和 Alibaba 普遍使用 OpenAI 兼容端点,许多开发者因此可以用有限代码改动测试或替换供应商。OpenRouter 的路由控制和 BYOK 逻辑更进一步,鼓励 multi-homing,而不是硬锁定。SiliconFlow 自己的 BYOK 和预选工具分发也服务同一目标:降低采用门槛,但如果另一家供应商提供更好的成本、吞吐或可靠性,也会降低被替换的门槛。 现有厂商用分发力回应这个弱点。Bedrock 和 Foundry 卖的不只是 tokens;它们把推理卖进更大的企业控制平面,里面有 IAM、策略、guardrails、可观测性和既有采购关系。Alibaba 也可以借助区域云基础设施、Qwen 所有权和打包订阅。这些周边优势能比模型接入本身创造更持久的粘性。因此,SiliconFlow 需要在开放性、跨模型中立路由、本地芯片异构性或中国特定部署适配胜过留在超大规模云厂商更大资产内的便利时取胜。[CP024, CP025, CP026, CP027, CP028, CP029]
3.4 护城河耐久性与不利竞争风险
不利证据在这里异常重要,因为整个行业的核心 API 接入正走向商品化。Hello China Tech 认为,主流模型 API 价格自 2023 年以来已下跌超过 90%;申请文件则显示,SiliconFlow 的公有云业务把份额和用户获取放在短期盈利之前。在中国,按吞吐量计算的前三大 token 供应商仍是超大规模云厂商部门,而 SiliconFlow 的 1.5% 份额明显小于最大现有厂商。这并不否定业务,但确实意味着投资人不应轻易把使用增长等同于可持续竞争优势。 已留存来源中最强的护城河候选,不是专有模型所有权,也不是强客户锁定,而是 SiliconFlow 在中国栈中的中立位置:广泛的开放模型接入、包括国产算力在内的异构芯片支持,以及同时服务自助和私有化部署工作负载的能力。问题在于,类似的中立性和可迁移性也是 Together、Fireworks 和 OpenRouter 的卖点;超大规模云厂商也能模仿许多 API 功能,并用更深的资产负债表补贴。竞争耐久性因此更少取决于表面功能清单,更取决于供应获取、运营效率、信任,以及把价格敏感的开发者流量转化为更难替换的企业账户的能力。[CP030, CP031, CP032, CP033, CP034, CP035]
| 护城河主张 | 威胁 | 严重性 | 重要性 | 缓解措施 / 尽调问题 |
|---|---|---|---|---|
| 中国本土中立平台 | 超大云厂商复制 API 界面,并以更低价格抢流量 | 高 | 切换成本低,单靠功能的优势可能被抹平 | 验证企业客户选择 SiliconFlow 是否不只是因为价格 |
| 异构芯片编排 | 竞争对手补强多芯片支持,或锁定优先供应 | 高 | 供应可得性和单位经济性决定可靠性与利润率 | 核验供应商合同和国产芯片适配深度 |
| 广泛模型接入 | 模型目录广度在路由平台和云平台间很快商品化 | 中 | 仅有数百个模型,撑不起持久锁定 | 按模型家族衡量实际用量集中度 |
| 借助工具触达开发者 | BYOK 和兼容性也让客户更容易在 SiliconFlow 之外多平台并用 | 中 | 渠道能拉高漏斗入口,但可能削弱留存 | 按获客渠道索取 cohort 留存数据 |
| 企业 / 私有化部署路径 | 既有云厂商可以把治理、网络和采购打包出售 | 高 | 大客户里,打包方案可能压过中立平台优势 | 追问安全能力,以及与云厂商竞争时的采购胜率 |
| 独立定位 | 价格战和负毛利公有云会挤压整个品类 | 高 | 市场增长未必转化为利润池增长 | 对标同行毛利率和促销额度纪律 |
风险登记表关注优势能否持久,而不是功能清单。
[CP029, CP030, CP031, CP032, CP033, CP034]压缩呈现 SiliconFlow 竞争耐久性最关键的指标。
[CP002, CP024, CP029, CP032, CP034, CP042]3.5 图表
04财务
4.1 收入模式与定价架构
SiliconFlow 的申请文件显示,一个品牌下面有两块经济性不同的业务。公有云服务包括 serverless token 服务和专用实例;本地部署解决方案则把推理软件安装到客户环境中。Serverless 服务是预付费、低价、消费驱动的产品,面向开发者和较小客户;专用实例和私有化部署服务于需要稳定性、预留容量或合规控制的大买方。公开模型页和 API 文档把商业界面展示得很清楚:标价按模型层级的使用量计费,接入由 API key 驱动,公司靠快速上手和广泛模型选择拉动采用。 问题在于,标价只是收入故事的最上层。Together、Fireworks、AWS 和 Alibaba 的同业定价页显示了同样的市场逻辑:cached tokens、batch jobs、专用容量、fine-tuning 或部署收费,都会让实际经济性偏离简单的按 token 标价。因此,SiliconFlow 所在市场训练客户预期按量灵活性和频繁折扣;重度用户又往往能迁移到成本效率更高的专用容量。这让收入质量高度取决于客户组合、有效折扣,以及更高价值企业工作负载最终能否压过低 ARPU 的开发者流量。[CI001, CI002, CI003, CI004, CI005, CI006]
| 收入来源 | 机制 | 计量单位 | 当前数值 / 状态 | 收入质量 | 尽调问题 |
|---|---|---|---|---|---|
| Serverless token 服务 | 预付费,按用量消耗 token | token / API 调用 | 2025 年收入 RMB14.3M | 当前质量低:量大,但消费密度低 | 需要 cohort 留存和每 token 实际有效成交价 |
| 专属实例 | 公有云上的预留算力 | 实例 / 预留容量 | 2025 年收入 RMB15.0M | 优于纯 Serverless,但仍暴露在算力租赁成本下 | 需要披露利用率和合同期限 |
| 公有云合计 | Serverless + 专属实例 | RMB 收入 | RMB29.261M,2025 年收入的 52.9% | 增长引擎,但毛利为负 | 需要业务线级毛利拆解和企业客户占比 |
| 本地部署 | 客户环境内的软件 / 部署 | 项目收入 | RMB26.069M,2025 年收入的 47.1% | 毛利质量更高,但可扩展性更弱 | 需要销售周期长度和可复制性数据 |
| 模型级标价 | 公共模型库上的 API 价格表 | 按 token / 按单位标价 | 平台公开可见标价 | 不等于实际成交价或毛利率 | 需要按 cohort 拆分的折扣和促销政策 |
收入结构有申报文件支撑;定价姿态由公开产品界面和同行标价证据补充。
[CI001, CI002, CI003, CI004, CI005, CI009]| 供应商 / 界面 | 价格 / 单位 / 合同 | 标价 vs 实际成交 | 折扣 / 控制杠杆 | 来源 / 含义 |
|---|---|---|---|---|
| SiliconFlow 模型页面 | 按用量计费的模型定价 | 仅标价 | 模型选择和抵扣额度会影响实际支出 | 证实按用量付费定位,但不能证实实际收入 |
| Together Serverless 服务 | 按 token 定价 | 仅标价 | 缓存输入和批处理折扣;可迁移至专属方案 | 显示市场在把有效单位成本往下压 |
| Together Dedicated 服务 | 按 GPU 分钟 / 小时 | 仅标价 | 自动扩缩容和预留容量 | 高利用率下,专属方案可能更便宜 |
| Fireworks Serverless 服务 | 按 token 定价 | 仅标价 | Priority / Fast 层级、缓存输入折扣、批处理为标准价 50% | 运营属性会改变实际经济性 |
| AWS Bedrock | 按 token 计费,另有批处理 / 预置选项 | 仅标价 | 批处理 50% 折扣;预置吞吐 | 大型既有厂商能覆盖多个价格带 |
| Alibaba Model Studio | 按 token 计费,另有席位 / 额度方案 | 仅标价 | 区域折扣、免费额度、订阅席位 | 打包方案扩大可覆盖预算,但也让实际成交价更模糊 |
官方定价页只给标价;没有任何页面披露实际合同经济性或贡献毛利。
[CI006, CI007, CI008, CI010]token 消耗和部署选择如何转化成不同收入质量画像。
[CI001, CI002, CI003, CI004, CI039]4.2 成本结构与单位经济性
申请文件少见地清楚说明了当前经济性痛点。2025 年,SiliconFlow 录得收入 RMB55.33 million,但销售成本 RMB68.632 million,意味着毛利为负,综合毛利率为 -24.0%。公有云线更差:2025 年毛损率为 -119.0%,2024 年为 -271.6%;本地部署仍保持 82.5% 的高毛利。算力租赁主导成本结构,占销售成本 86.9%;申请文件也明确称,公司把市场份额、用户获取和生态建设置于即时盈利之上。 独立评论把单位经济性画得更清楚。Hello China Tech 和 KrASIA 都指出,SiliconFlow 实际上是一个中间层:租用算力,把第三方模型打包进价格战环境。公开材料中衡量变现质量的最佳代理指标不是注册用户总数,而是消费密度。Serverless 付费账户一年内从 2,455 增至 716,000,但 serverless token 收入只有 RMB14.3 million,意味着每个账户年均消费极小。2025 年销售和营销费用中,超过 64% 用于促销算力额度。这种组合能快速做大使用量,但尚未证明有吸引力的 CAC 回收或可持续毛利率修复。[CI011, CI012, CI013, CI014, CI015, CI016]
| 指标 | 数值 / 状态 | 置信度 | 为什么重要 | 尽调问题 |
|---|---|---|---|---|
| 2025 年收入 | RMB55.33M | 高 | 收入规模锚点 | 确认 2026 年 run-rate 和收入确认节奏 |
| 2025 年综合毛利率 | -24.0% | 高 | 显示规模尚未修复经济性 | 拆解公有云与本地部署收入结构的影响 |
| 2025 年公有云毛亏损率 | -119.0% | 高 | 核心增长引擎仍在亏损 | 需要各产品毛利路径和利用率目标 |
| 2025 年本地部署毛利率 | 82.5% | 高 | 收入质量更高,但规模更小 | 评估该业务线的可复制性和天花板 |
| 算力租赁占 COGS 比重 | 86.9% | 高 | 供应成本主导经济性 | 需要供应商合同和定价路线图 |
| Serverless 付费账户 | 2025 年末 716,000 个 | 高 | 显示获客规模 | 需要按 cohort 拆分的活跃支出分布 |
| 约每个 Serverless 付费账户收入 | 以年末账户数粗略作分母,约 ~RMB20 / 年 | 中 | 指向很低的消费密度 | 需要月均活跃付费方和按十分位拆分的收入 |
| 促销额度占 S&M 比重 | 2025 年 >64% | 高 | 获客可能高度依赖补贴 | 需要 CAC 回收期和促销额度转化率 |
使用公开数据和简单的审慎代理指标;RMB20 / 账户这个数字明确只是近似值,不是管理层 KPI。
[CI011, CI012, CI013, CI014, CI015, CI016]为什么规模尚未转化成有吸引力的公有云单位经济。
[CI011, CI013, CI015, CI018, CI020, CI040]关键财务输入和粗略充足性信号的来源支撑区间。
区间混合业务线和资产负债表锚点,以展示经济性分布;不是单一管理层情景模型。
[CI012, CI014, CI023, CI026]4.3 资本充足性与融资依赖
2026 年融资序列后,资本充足性明显改善,但业务看起来仍依赖融资,而不是自我造血。2025 年末,SiliconFlow 持有 RMB171 million 现金及现金等价物,加上 RMB100 million 定期存款;全年经营活动现金净流出为 RMB172 million。用这个简单的后视镜口径看,如果没有新资本,公司独立现金跑道并不宽裕。申请文件披露的 2025 年后 A+、B、B+ 融资改变了这幅图,年末之后注入约 RMB1.48 billion 现金对价。因此,问题从即时生存转为:新资本能多高效地转化为更好的单位经济性。 募资用途和供应商章节仍指向外部算力供给和持续商业化支出的实质依赖。公司不直接买芯片,主要通过伙伴租赁计算资源,采购也集中在少数主要供应商。这意味着流动性需求不只绑定 R&D 消耗,还绑定算力可用性策略、促销使用额度,以及 serverless 流量和更高质量企业工作之间的组合。已留存公开记录不支持截至 2026-07-22 的干净当前现金跑道计算,所以投资人应把资本充足性看作已经改善,但仍与运营执行相连,而不是被 2026 年 6 月轮次一次性解决。[CI023, CI024, CI025, CI026, CI027, CI028]
| 项目 | 数值 / 状态 | 为什么重要 | 置信度 | 尽调问题 |
|---|---|---|---|---|
| 现金及现金等价物 | 2025 年末 RMB171M | 2026 年融资前的即时流动性基础 | 高 | 确认当前不受限现金 |
| 定期存款 | 2025 年末 RMB100M | 增加流动性缓冲,但未必等同于可立即动用的经营现金 | 高 | 明确期限和限制 |
| 经营现金消耗 | 2025 年经营活动净现金流出 RMB172M | 显示 2026 年前业务无法自我供血 | 高 | 提供 2026 年月度 burn 趋势 |
| 2025 年后融资现金对价 | A+、B 和 B+ 轮合计 ~RMB1.48B | 相比年末现金,显著拉长 runway | 高 | 确认扣除费用后的到账现金及受限用途 |
| 算力供应依赖 | 租赁算力而非购买芯片;供应商集中度高 | runway 不只取决于现金,也取决于供应商条款 | 中 | 需要应付款、预付款和合同结构 |
| 截至 2026-07-22 的当前 runway | 现有来源无法公开支撑 | 无法精确投资判断 | 高 | 管理层应提供报告生成日的现金、burn 和 runway |
历史融资时间线见第 1 章;本表聚焦未来充足性和依赖关系。
[CI023, CI024, CI025, CI026, CI027, CI028]现金需求受算力采购、商业化补贴和融资支持塑造。
[CI024, CI025, CI026, CI027, CI028, CI029]4.4 公开缺口与财务判断
公开记录支持一个强烈的财务警示信号,但不足以搭出完整承销模型。我们有收入、毛利率、部分组合数据、客户数量代理指标、供应商集中度和 2025 年后融资桥。我们没有干净的净收入留存、cohort 毛利、实际折扣率、客户集中度、截至报告日的月度消耗,或管理层背书的公有云盈亏平衡时间表。由于整个市场的官方定价页都是标价而非实际成交价,它们可用于竞争背景,却不足以推断 SiliconFlow 的实际贡献毛利。 因此,财务判断是混合的。收入增长和用户获取显示需求真实。本地部署经济性看起来比公有云更健康。但当前核心增长引擎——公有云 token 供给——仍显得补贴重、算力租赁重、毛利为负。对投资人来说,最重要的尽调阻碍不是市场是否足够大,而是 SiliconFlow 能否在下一轮资本依赖周期开始前,把 token 工厂规模转化为明显更好的收入质量和结构性改善的毛利率。[CI031, CI032, CI033, CI034, CI035, CI036]
| 缺失的私有指标 | 影响 | 缺失为何重要 | 具体尽调路径 |
|---|---|---|---|
| 净收入留存 / 总留存 | 重大 | 没有留存数据,用户增长可能掩盖流失或变现弱 | 索取开发者、企业和私有化部署 cohort 的留存 |
| 每 token 实际有效成交价 | 重大 | 标价无法揭示折扣或促销依赖 | 要求按头部模型家族提供实际净价 |
| 客户集中度 | 重大 | 大企业客户集中可能扭曲收入质量 | 索取头部客户收入占比和合同条款 |
| 当前月度 burn | 重大 | 仅靠 2025 年数据,无法判断 2026 年融资后的资金是否充足 | 索取最新月度现金消耗和现金余额 |
| 公有云盈亏平衡时间表 | 重要 | 核心增长引擎的估值取决于毛利拐点时间 | 索取管理层经营计划和利用率里程碑 |
| 销售效率 / CAC 回收期 | 重要 | 促销额度可能掩盖高成本获客 | 按获客渠道和客户类型索取回收期 |
每个缺失字段都直接关系投资判断,不是出于好奇。
[CI031, CI032, CI033, CI034, CI035, CI036]4.5 图表
05产品与技术
5.1 产品定义与模块图谱
理解 SiliconFlow 产品,最好把它看成多层推理交付平台。中文官网、英文文档介绍、模型目录和 API 参考文档都呈现出同一条工作流:开发者或企业团队选择一个模型家族,获取 API key,然后按使用量计费调用托管推理端点。这个核心 serverless 表面周围,还有几类相邻供给——面向稳定企业工作负载的预留实例,面向受监管或数据敏感买方的私有化部署,以及面向希望在自研或开源模型上获得更快推理的客户的加速服务。换句话说,公司卖的是围绕模型推理的接入、编排、优化和部署模式,而不是单一专有前沿模型。 这个模块组合有战略意义。它让 SiliconFlow 用一个控制平面服务不同买方任务:开发者的快速实验、企业的更高 SLA 预留容量,以及无法停留在共享公有云上的客户需要的私有或混合部署。广度是优势,因为它减少了客户必须分别测试的供应商数量;广度也是复杂性风险,因为公司必须保持模型库存更新,维护异构 serving 路径,并解释哪种交付模式真正适合每类工作负载。[CE001, CE002, CE003, CE004, CE005, CE006]
| 模块 / 资产 | 主要用户 | 状态 / 成熟度 | 差异化 | 尽调缺口 |
|---|---|---|---|---|
| Serverless 模型 API | 开发者、初创公司、企业构建者 | GA / 公开文档可查 | 模型目录广,支持按用量付费 | 需要实际可用性 / 延迟 / 错误率指标 |
| 预留实例 | 企业平台团队 | GA / 官网首页推广 | 为核心推理工作负载提供专属容量和成本优化 | 需要 SLA 条款和利用率经济性 |
| 推理加速服务 | 模型构建者、基础设施团队 | GA / 官网首页推广 | 面向自研或开源模型的性能优化 | 需要公司口径之外的独立 benchmark |
| 私有化部署 | 受监管或数据敏感型企业 | GA / 官网首页推广 | BYOC、隔离和私有化部署选项 | 需要部署案例和安全证明 |
| 模型目录 / 控制平面 | 所有用户 | GA / 公开可见 | 统一界面覆盖文本、语音、图像、视频和多模态模型 | 需要生命周期 / 下线政策和迁移工具细节 |
状态来自公开界面核验,不等于独立审计过的生产成熟度。
[CE001, CE002, CE003, CE004, CE006]| 用户任务 | 当前工作流 | SiliconFlow 方案 | 可衡量收益 | 限制 |
|---|---|---|---|---|
| 快速原型化 AI 应用 | 选择模型、创建 API key、调用 endpoint | Serverless API 目录 | 不背 GPU 运维负担,也能快速试验 | 模型阵容变化时,成本 / 性能可能随之变化 |
| 运行稳定的企业推理 | 将突发需求转为可预测负载 | 预留实例 | 控制容量,性能可预测性更强 | 需要公开 SLA 和支持包细节 |
| 在私有环境部署 | 让数据或工作负载不进入共享公有云 | 私有化部署 / BYOC | 对隐私和合规态势掌控更多 | 实施复杂度和案例深度不清楚 |
| 优化自定义或开源模型服务 | 为速度和成本调优运行时 | 加速服务 | 降低延迟,提高硬件利用率 | 独立 benchmark 覆盖有限 |
| 针对特定任务切换模型家族 | 按任务比较模型能力 | 跨模态的广泛目录 | 降低集成切换成本 | 平台仍依赖上游模型可用性 |
用例映射聚焦客户工作流,不按公司营销分类。
[CE005, CE007, CE008, CE009]SiliconFlow 把面向客户的 API、部署模式、优化软件和底层算力拼成一套推理交付栈。
层级数值是定性权重,代表公开架构中的相对功能范围,不代表资源分配或收入组合。
[CE001, CE003, CE010, CE022]5.2 架构与开发者工作流
公开技术界面明显是 API 优先。chat-completions 参考文档暴露了模型选择、streaming、context-window 控制、tool calling 和 tracing headers。更宽的文档介绍描述了跨文本、图像、语音、视频、向量、reranking 和多模态模型的按量 API 接入。这些来源暗示的用户工作流很简单:注册、创建 API key、从目录选择模型、按大体 OpenAI 风格的请求模式集成,然后再决定留在共享 serverless、迁移到预留实例,还是进入私有化部署。低摩擦路径是 SiliconFlow 能把这么多模型聚合到一个平台下的核心原因。 在底层,平台不只是 API 代理层。SiliconFlow 反复强调自研高效算子、优化框架、加速引擎、动态扩缩容、监控和容错。开源 OneDiff 项目提供了最清楚的外部工程证据,说明公司确实在构建推理优化软件,而不只是包装其他供应商的模型。OneDiff 聚焦 diffusion models 的加速库、compiler backends 和优化 kernels;这不能证明整个平台架构,但支持一个更宽的主张:推理性能工程是内部能力。因此,技术故事比纯营销文案更可信,虽然在加速工具上强于在独立审计的可靠性指标上。[CE010, CE011, CE012, CE013, CE014, CE015]
| 层 / 组件 | 作用 | 依赖 | 风险 |
|---|---|---|---|
| API 网关和认证 | 应用调用入口 | 开发者凭证和请求管理 | 客户侧服务中断或认证故障会波及所有模块 |
| 模型目录 / 路由层 | 将请求映射到可用托管模型 | 上游模型供应商和版本变更 | 模型频繁更替可能迫使客户迁移 |
| 推理加速层 | 改善延迟 / 吞吐 / 成本 | 内部优化软件和面向特定 GPU 的调优 | 性能主张未必能泛化到不同工作负载 |
| 算力编排和扩缩容 | 分配共享或专属容量 | 租赁 GPU 资源和自动扩缩容逻辑 | 容量短缺或成本飙升会挤压利润率和可靠性 |
| 监控 / 可追溯性 | 运营调试和服务保障 | 日志、trace id、可观测性栈 | 公开 SLO 报告不足,外部验证打折 |
架构基于公开来源,不做缺少依据的隐藏层猜测。
[CE010, CE011, CE012, CE014, CE023]SiliconFlow 上从评估到规模化生产使用的典型路径。
[CE005, CE011, CE013, CE017]产品靠上游模型、算力和开源优化组件支撑。
[CE014, CE018, CE022, CE023, CE025]5.3 差异化、路线图与依赖
SiliconFlow 的主要产品差异化是组合式的:模型接入广度、快速集成、多种部署模式,以及性能 / 成本优化叙事。官网称语言模型速度提升 10x+、图像生成 1 秒、多个场景有两位数成本节省;文档强调大模型覆盖;产品菜单横跨公有云、预留实例、加速服务和私有化部署。这个组合包比简单模型市场更有差异化,因为它让公司不只在目录广度上竞争,还能在服务质量和部署灵活性上竞争。 与此同时,依赖图谱很重。SiliconFlow 依赖上游模型提供方持续可用,依赖租赁 GPU 容量的持续获取,依赖开发者信任 API 抽象层,也依赖自己跟上快速变化的模型清单。API 文档明确警告,模型可用性和能力会随时间变化,一些服务仍在更新。Kimi K3 发布帖显示平台能快速加入新模型,但也说明路线图执行部分是一场模型接入竞赛,而不只是深层专有 R&D 竞赛。这造就了一条在运营和集成质量上真实存在的产品护城河,但它可能比模型创造者或拥有一方算力与分发的超大规模云厂商护城河更脆弱。[CE020, CE021, CE022, CE023, CE024, CE025]
| 控制项 / 指标 | 状态 | 范围 | 缺口 |
|---|---|---|---|
| BYOC 部署支持 | 公开声称 | 企业 / 私有化部署 | 需要架构评审和客户背书 |
| 计算 / 网络 / 存储隔离 | 公开声称 | 数据安全控制 | 需要第三方鉴证 |
| 隐私政策 | 公开可得 | 国际平台数据处理 | 需要 DPA、子处理方和留存细节 |
| 中国站与国际站条款分开 | 公开可得 | 司法辖区和签约结构 | 需要按客户映射签约法律实体 |
| API 响应里的 trace id | 公开文档可查 | 支持和故障排查 | 需要公开事故 / 状态历史 |
| 行业标准合规声明 | 营销层面声明 | 企业级定位 | 需要列明认证和审计报告 |
公开信任信号已经存在,但保障证据仍不完整。
[CE029, CE030, CE031, CE032, CE033]| 日期 / 阶段 | 功能 / 里程碑 | 状态 | 影响 | 来源 |
|---|---|---|---|---|
| 2026-07-22 当前 | 覆盖多模态的丰富模型目录 | 已上线 | 覆盖面支撑一站式平台定位 | 文档 / 模型页面 |
| 2026-07-22 当前 | 支持模型可用 reasoning 和 thinking-mode 参数 | 已上线 | API 能跟上更新的模型行为 | Chat API 参考 |
| 2026-07-22 当前 | 预留实例和私有化部署产品 | 已上线 | 说明公司不只卖开发者 API,也在打包企业级方案 | 中文官网 |
| 2026-07-22 当前 | 快速接入 Kimi K3 / GLM-5.2 等新发布模型 | 已上线 / 近期 | 路线图执行部分取决于伙伴 / 模型能否快速接入 | 官方博客 / 首页 |
| 2024 开源发布节奏 | OneDiff 加速库发布 | 历史信息,但仍相关 | 说明推理优化仍在持续投入工程 | GitHub 发布 |
路线图根据公开发布行为推断;未找到完整变更日志或已承诺的前瞻路线图。
[CE020, CE021, CE024, CE026, CE028]SiliconFlow 各能力领域的公开可见成熟度和验证质量。
[CE020, CE027, CE034, CE035]5.4 信任、安全、安保与质量控制
SiliconFlow 确实发布了有意义的信任与合规信号,但还不等于完整的企业级保障。中文官网提到算力、网络和存储层面的数据隔离,支持 BYOC 部署,并符合行业标准和监管要求。隐私政策和条款确认,公司有不同的中国和国际服务界面;这对数据处理和司法辖区隔离很重要。公开 API 界面还包括 trace 标识符,便于问题排障。合在一起,这些都是真实运营信号,说明平台面向可支持性搭建,而不只是演示使用。 不过,已留存来源集没有提供公开 SOC 2 报告、ISO 认证清单、公开可用性历史、安全事件历史、详细 DPA 包,或经过基准测试的错误率 / 延迟 SLO。这个缺口重要,因为 SiliconFlow 正把关键推理基础设施卖给企业工作负载。尽调中相关问题不是公司是否知道安全重要——网站显然说知道——而是它能否提供大型企业买方在平台标准化前会要求的第三方证明、可靠性报告和治理材料。[CE029, CE030, CE031, CE032, CE033, CE034]
5.5 图表
06客户
6.1 客户分层与采用入口
SiliconFlow 的客户基础不是一张同质化 SaaS 名单。公开来源至少指向四类实际客群:直接接入 API 的开发者、使用预留或私有化部署的企业平台团队、把推理栈嵌入区域基础设施的电信 / 算力伙伴,以及通过 SiliconFlow 托管模型承接外部用户需求的社区项目。QQ 融资报道和独立报告给出了最宽口径的采用指标——超过 1,000 万用户、10,000 家企业客户;申报文件和案例研究则显示,企业使用更多集中在算力密集、接近基础设施的场景,而不只是轻量聊天机器人试验。 采用入口同样重要,因为它们对应不同的变现和持续性。MindSearch、Continue、Cline 等社区和开发者集成证明 API 有相关性、上手容易,但它们不是高 ARPU 合同的强代理。相比之下,贵州移动合作、预留实例案例和私有化部署故事说明运营嵌入更深,但公开合同金额或续约条款往往缺失。因此,SiliconFlow 显然有需求广度;投资人仍需把使用证据与持久、高质量收入证据拆开看。[CU001, CU002, CU003, CU004, CU005, CU006]
| 客群 | 买方 / 用户 / 付款方 | 用例 | 规模信号 | 收入 / 战略价值 | 缺口 |
|---|---|---|---|---|---|
| 开发者与 AI 构建者 | 用户=开发者;付款方=个人 / 团队 | 通过 API 原型开发并上线 AI 应用 | 具名集成和社区指南很多 | 漏斗顶端需求广,也能切进生态 | 需要活跃付费开发者数和收入占比 |
| 企业平台团队 | 买方=IT / AI 平台负责人;用户=内部应用团队 | 预留实例、私有化部署、编码或知识工作负载 | 匿名案例研究加预留实例示例 | ARPU 可能更高,嵌入也更深 | 需要具名客户名单和合同期限 |
| 电信 / 算力伙伴 | 买方=区域基础设施运营商;用户=行业客户 | Token 工厂 / 推理基础设施共建 | Guizhou Mobile 合作明确 | 战略分发和供给杠杆 | 需要商业条款和在线工作负载规模 |
| 受监管 / 国资相关机构 | 买方=央企 / 公共部门类机构 | 国产化私有化部署和性价比优化 | 航空和能源央企案例呈现同一模式 | 能支撑信任和国内护城河叙事 | 具名客户名单大多未披露 |
| 社区工具 / 开源应用 | 用户=终端开发者或第三方工具终端用户 | 搜索 / RAG、编码、翻译、Agent 工作流 | MindSearch、Continue、Cline,以及更广的用例中心 | 认知度高,重复使用触点多 | 变现或排他性证据不足 |
客群划分区分买方、用户和付款方角色,而不是把所有“客户”一概而论。
[CU001, CU003, CU004, CU005, CU006]| 指标 | 值 | 日期 | 来源 | 置信度 | 含义 | 缺失分母 |
|---|---|---|---|---|---|---|
| 服务用户数 | 10M+ 用户 | 2026-06-16 | QQ 融资报道 / 官方披露 | 中 | 采用面很广 | 月活占比未知 |
| 企业客户 | 10,000+ 家企业客户 | 2026-06-16 | QQ 融资报道 / 官方披露 | 中 | 企业触达显著大于小规模试点名单 | 生产级占比未知 |
| Serverless 付费账户 | 年末 716,000 个账户 | 2025-12-31 | IPO 文件 / 报道 | 中 | 自助式变现已有规模 | 月活付费账户率未知 |
| 预留实例示例需求 | 单一客户日用量 100B token | 2026 | 官方案例研究 | 中 | 部分工作负载大到足以支撑专用容量 | 匿名客户限制分层判断 |
| Guizhou Mobile 战略升级 | 2025 年合作后,2026 年协议加深 | 2026-06-10 | 官方合作公告 | 高 | 机构关系看起来是多阶段,而非一次性合作 | 商业价值未披露 |
| 社区用例广度 | 搜索、编码、翻译和 RAG 中有多个具名集成指南 | 2026-07-22 | 官方文档用例中心 | 高 | API 在多类从业者工作流里都有用 | 多少转化为付费生产使用未知 |
这张表把广义客户触达标记与更深部署信号放在一起,用来展示漏斗层次,而非单一转化链。
[CU002, CU007, CU010, CU013, CU020, CU029]SiliconFlow 客户路径从轻量 API 访问开始,可扩展为专属容量或联合基础设施关系。
报告保留来源能支撑这些阶段,但公开资料没有披露阶段间转化率。
[CU003, CU013, CU020, CU032]公开证据从宽口径用户和企业客户数量,收窄到有深度部署证据的少数样本。
漏斗混用了不同证据层级的计数;最后的零表示没有公开留存队列披露,不是留存为零。
[CU002, CU007, CU011, CU022]6.2 具名客户与用户证据
最强的公开具名证据来自两条渠道。第一是企业和生态合作,尤其是贵州移动;SiliconFlow 公开描述了双方围绕推理框架、Token 服务和联合操作系统深化战略协作。第二是面向开发者的用户案例,具名工具或开源项目围绕 SiliconFlow API 发布集成指南。MindSearch、Continue、Cline 尤其相关,因为它们说明平台进入了搜索 / RAG 和编码 Agent 工作流;这些场景天然消耗大量 Token,也需要高频重复使用。 这些证据强弱不一。贵州移动是具名机构交易对手,有高管引述和具体运营范围,证据强度远高于操作指南。MindSearch、Continue、Cline 页面证明认知度和实际集成,但没有披露付费转化、规模或排他承诺。航空、能源和编码 Agent 的匿名企业案例进一步说明部署模式和结果,但匿名性限制了集中度分析。因此,正确读法是:具名证据存在,且覆盖开发者和企业两个入口;但只有一部分能算强商业证据。[CU010, CU011, CU012, CU013, CU014, CU015]
| 客户 / 项目 | 客群 | 部署 / 用例 | 生产 / 试点 | 结果 / 证据 | 限制 |
|---|---|---|---|---|---|
| Guizhou Mobile | 电信 / 区域 AI 基础设施 | 推理框架部署、token 服务、联合运营 | 生产 / 战略合作证据强于试点 | 官方签署协议,含高管引述和多条工作流范围 | 收入、部署数量和续约条款未披露 |
| MindSearch | 开发者 / 搜索-RAG 项目 | 接入 SiliconFlow API 并提供部署指南 | 实际集成证据 | 官方操作指南包含配置和部署步骤,包括 HuggingFace Space 路径 | 不能证明付费生产规模或排他性 |
| Continue | 开发者 / IDE 编码助手 | VS Code / JetBrains 与 SiliconFlow 托管模型集成 | 实际集成证据 | 官方指南覆盖实际模型选择、缓存和验证流程 | 未披露转化或支出数据 |
| Cline | 开发者 / 编码 Agent | 面向编码 Agent 工作流的 OpenAI 兼容 API 集成 | 实际集成证据 | 官方指南详述 base URL、model ID 和多模式设置 | 未披露客户数或收入 |
具名证据覆盖一个机构伙伴和多个开发者工具集成;它能证明采用广度,但不能替代具名企业收入客户名单。
[CU011, CU014, CU015, CU016, CU017, CU018]机构合作证据强于传统留存可见性;开发者工具证据覆盖面宽,但商业分量更轻。
标签综合了是否具名、部署细节、量化结果,以及离收入有多近。
[CU014, CU018, CU019, CU023, CU030, CU033]6.3 持续性、扩张与集中度
公开来源支撑扩张故事,力度强于留存故事。预留实例案例展示了一条生命周期:Token 需求增长后,客户从按量付费转向锁定专用容量。贵州移动合作指向平台更深嵌入区域算力和服务基础设施。企业案例研究显示,SiliconFlow 从公有云 API 调用延伸到私有化部署、国产芯片适配和更广的运营集成。这些模式符合先落地再扩张:先用易接入的 API 起步,使用变成关键业务后,再加深到预留实例、私有化部署或联合运营。 缺口在经典 SaaS 持续性证据。留存的公开资料没有披露 NRR、GRR、流失率、平均合同期限、续约率、头部客户集中度或分客群收入贡献。官方和独立记录也没有拆出 10,000 家企业客户中有多少真正进入生产级使用、多少只是小预算试验账户、多少需求由少数战略伙伴中介。因此,客户广度真实存在,但客户质量和集中度仍是未解的尽调问题。[CU020, CU021, CU022, CU023, CU024, CU025]
| 指标 | 值 / null | 客群 | 置信度 | 尽调问题 |
|---|---|---|---|---|
| 公开 NRR | null | 所有付费客户 | 高 | 要求按开发者、企业和私有化部署分组提供 NRR |
| 公开 GRR / 客户 logo 留存 | null | 所有付费客户 | 高 | 要求按客群提供 logo 留存和续约率 |
| 合同期限 | null | 企业 / 电信 | 高 | 要求提供预留和私有化交易的平均期限及续约选项 |
| 重复使用证据 | 只有定性证据:部分工作负载从按量付费转向预留容量 | 企业重度用户 | 中 | 提供从 API 消费转向预留实例部署的转化率 |
| 满意度 / 客户背书深度 | 证据混合:工作流细节强,公开正式证言弱 | 开发者和企业买方 | 中 | 要求提供 NPS / CSAT 和客户背书 |
| 公开负面流失证据 | 留存语料中没有可靠流失数据集 | 全部客群 | 中 | 提供流失原因和丢失客户 logo 分析 |
null 表示保留来源中没有公开披露,不代表指标无关。
[CU021, CU022, CU023, CU024, CU030]| 扩张驱动 | 集中度风险 | 影响 | 尽调路径 |
|---|---|---|---|
| 从 API 升级到预留实例的路径 | 扩张是否依赖少数大客户未知 | 广泛用户数下可能藏着很高收入集中度 | 要求提供消费十分位和扩张分组 |
| 面向受监管企业的私有化部署 | 关键案例匿名,具名证据偏薄 | 难以评估续约风险和销售周期能否持续 | 在 NDA 下要求具名客户背书 |
| 电信 / 基础设施伙伴关系 | 战略伙伴可能成为用量增长的集中通道 | 可能带来议价权和依赖问题 | 要求提供伙伴对订单额和算力供给的贡献 |
| 社区工具集成 | 高认知度不等于可持续变现 | 可能推高使用量,却不能证明高 LTV 客户 | 要求提供社区集成带来的付费转化 |
| 跨行业央企采用叙事 | 官方叙事强,但具体垂直行业组合模糊 | 如果收入集中,可能夸大多元化 | 要求按垂直行业和前十大账户拆分收入 |
主要不确定性不是需求是否存在,而是需求基础中收入集中度藏在哪里。
[CU025, CU026, CU027, CU028, CU031]公开来源能看到客户形成和少量扩展线索,但看不到真实留存率。
这是可见性代理指标。“1”表示一条清晰的公开扩展路径,不是比率;0 表示没有公开百分比或队列披露。
[CU021, CU022, CU024, CU034]6.4 客户判断与缺口
整体客户判断是:相关性为正,持续性喜忧参半。SiliconFlow 有足够公开证据说明,开发者在主动集成 API,机构买方愿意在重要推理场景中使用公司产品,且至少部分工作负载大到足以支撑预留容量或私有化部署。这不只是表层 Logo 收集;客户叙事有案例细节、集成深度和具体工作流框架支撑。 但公开证据还不足以支撑集中度风险、扩张效率或留存质量的核保。社区用户案例有用,却不等同于付费生产 Logo。匿名企业故事证明模式,不证明精确账户价值。即便贵州移动也是战略合作,而非披露收入合同。因此,尽调优先级很直接:拿到分客群活跃客户数、留存队列、头部客户敞口,以及从 Serverless 转向预留或私有化部署的扩张转化。[CU029, CU030, CU031, CU032, CU033, CU034]
6.5 图表证据
07风险
7.1 监管与法律风险
SiliconFlow 处在多套重叠合规制度之内。其中国条款和隐私政策要求许多用户完成实名验证,限制某些应用领域,承诺 AI 生成内容标识义务,并要求境内运营中的用户信息存储在中国大陆。这些不是普通网站样板条款:它们直接影响入驻摩擦、企业签约、产品设计和日志义务。条款明确写明,服务不适用于自动控制、医疗信息服务、心理咨询和关键信息基础设施场景;除非公开界面之外存在定制化控制措施,否则部分后果最重的部署类别会被收窄。 更广的中国政策栈正在变密,而不是变轻。2025 年生效的 AI 标识办法要求对生成式合成内容设置显式和隐式标识,并承担日志、元数据和平台分发责任。SiliconFlow 的境内条款已承认这些规则,并禁止用户删除或篡改标识。这种对齐是正面信号,但也意味着合规风险落在运营层面:如果 SiliconFlow 的工具或企业客户处理标识、元数据或下游分发不当,即便模型输出生成本身合法,公司也可能受到行政审查。因此,法律风险不只是某个被禁活动,而是要在境内外业务面上持续执行一套不断变化的监管栈。[CR001, CR002, CR003, CR004, CR005, CR006]
| 规则 / 案例 | 司法辖区 | 状态 | 可能性 | 严重性 | 缓释措施 | 剩余风险敞口 | 尽调路径 |
|---|---|---|---|---|---|---|---|
| AI 生成内容标识义务 | 中国 | 已生效 / 执行中 | 中高 | 高 | 条款写明标识义务;服务商可添加标识并留存日志 | 仍有运营执行不到位风险 | 要求提供标识架构和合规审计 |
| 实名认证和身份核验 | 中国 | 有效的合同 / 法律要求 | 中 | 中高 | 国内隐私政策 / 条款写明开户和核验流程 | 部分用户仍面临摩擦和隐私风险 | 要求提供 KYC 流程和例外处理 |
| 电信 / 互联网服务许可 | 中国 | 有效 | 中 | 高 | 国内站点可见 ICP / 电信备案 | 许可证范围能否匹配不断演进的服务,公开材料没有检验 | 获取律师备忘录,把许可证逐项对应到当前服务 |
| 敏感领域排除(医疗、心理、汽车控制、CII) | 中国 / 客户合同 | 有效合同限制 | 中 | 高 | 除非有定制安排,公开条款会收窄适用范围 | 企业误用可能引发法律或声誉问题 | 要求提供分行业部署控制和审批流程 |
| 跨境法律界面拆分 | 中国 / 国际 | 有效 | 中 | 中 | 中国与国际条款 / 隐私界面分开 | 仍可能出现运营混淆或控制不一致 | 审查实体、数据流和签约地图 |
行按投资相关性排序,而不是按完整法律分类排序。
[CR001, CR002, CR003, CR004, CR005, CR006]最高残余风险集中在监管、算力依赖和客户经济性不透明三处。
[CR006, CR015, CR024, CR029, CR037]7.2 运营、安全与质量风险
从运营上看,SiliconFlow 更像基础设施,而不是简单应用层。宕机、时延波动、错误部署和合规失败的代价因此更高。案例研究和产品页强调私有化部署、异构芯片适配、预留实例、监控、容错和成本优化——这些都说明客户把有分量的工作负载放在平台上。但公开保证材料仍不完整。留存语料没有展示公开可用性报告、详细 SLO、SOC 2 或 ISO 证明,也没有公开事故档案。投资人因此难以判断运营成熟度是否匹配公司的市场位置。 安全和内容治理敞口也会互相牵连。境内条款称服务不应用于某些敏感领域,并要求用户遵守广泛内容限制;隐私政策说明 SiliconFlow 会出于合规和安全目的收集日志和身份信息;标识相关规则在部分情形下要求日志保存六个月。这些控制能帮助降低滥用,但也扩大了隐私、内容和访问控制失败的攻击面。对推理平台来说,相关风险不只是模型是否幻觉,而是平台能否在规模化时可靠执行账户控制、数据处理、标识和运营隔离,同时保持足够快的发版速度。[CR012, CR013, CR014, CR015, CR016, CR017]
| 失效模式 | 可能性 | 严重性 | 缓释成熟度 | 剩余风险敞口 | 未解决缺口 |
|---|---|---|---|---|---|
| 关键工作负载的平台宕机 / 延迟不稳 | 中 | 高 | 中 | 高 | 留存语料中没有公开正常运行时间 / SLO 记录 |
| 安全控制与企业预期存在差距 | 中 | 高 | 低-中 | 高 | 未识别到公开 SOC 2 / ISO / 事故档案 |
| 隐私 / 日志处理失效 | 中 | 高 | 中 | 中-高 | 国内隐私政策显示数据处理责任很重 |
| 标识 / 元数据执行失效 | 中 | 中-高 | 中 | 中-高 | 需要证明流水线层面有显式和隐式标识控制 |
| 模型快速更替扰乱质量或模型生命周期 | 中-高 | 中 | 中 | 中 | API 文档明确提醒模型可用性可能变化 |
这份登记表聚焦基础设施类失效模式,而不是抽象的模型风险讨论。
[CR012, CR013, CR014, CR015, CR016, CR017]展示监管和运营如何传导到收入质量、毛利率、融资和估值。
[CR007, CR016, CR030, CR037]7.3 依赖、财务与竞争风险
最尖锐的非监管风险是外部依赖。SiliconFlow 没有自有前沿模型阵地,也没有自有全球超大规模云。模型广度取决于上游模型伙伴能否持续有竞争力且可用;毛利率取决于能否拿到外部算力;供应韧性则取决于国产芯片适配,以及持续采购或租用先进半导体能力的能力。公司自己的申报文件和独立分析显示,算力租赁主导销售成本,主要供应商集中度仍高。这让业务同时暴露在成本飙升和供应商或战略渠道议价之下。 地缘政治会放大这个问题。美国出口管制指引继续围绕先进计算半导体及流向中国关联实体的转移收紧;与此同时,中国也在加强对 AI 生成内容和互联网服务的境内治理。SiliconFlow 可以借国产芯片、私有化部署和 Token 工厂运营模式适配,但这些是缓释措施,不是免疫。竞争压力进一步叠加风险:同一个验证 SiliconFlow 的市场,也在训练买方按价格、可靠性和集成便利性,把它与超大规模云和其他推理提供商比较。如果价格战、标识合规成本和算力稀缺同时加剧,SiliconFlow 最薄弱的环节就会从需求生成变成利润率和资本充足性。[CR022, CR023, CR024, CR025, CR026, CR027]
| 依赖项 | 交易对手 / 类别 | 角色 | 集中度 | 失效情景 | 严重性 | 缓释措施 | 剩余风险暴露 |
|---|---|---|---|---|---|---|---|
| 上游模型提供方 | 模型实验室 / 开源生态 | 模型目录广度和用户需求 | 中-高 | 模型访问或竞争力快速走弱 | 高 | 多模型路由和快速接入 | 仍缺少第一方模型控制权 |
| 租用算力供应商 | GPU / 云 / 基础设施合作伙伴 | 核心服务交付 | 高 | 成本飙升或容量约束挤压毛利率和可靠性 | 高 | 国产芯片适配和合作伙伴网络 | 供应议价权仍在外部 |
| 战略电信 / 基础设施合作伙伴 | 贵州移动及类似渠道 | 渠道分发与算力协同 | 中 | 合作伙伴驱动的增长变得集中,或暴露在政策风险下 | 中-高 | 联合运营和区域生态策略 | 经济贡献结构未公开 |
| 开发者生态触点 | Continue / Cline / MindSearch 及社区工具 | 拉动需求和用量增长 | 低-中 | 认知度未能转化为可持续付费账户 | 中 | 低摩擦 API 和丰富模型 | 缺少商业转化数据 |
| 监管依赖 | CAC / MIIT / 地方合规制度 | 运营许可和内容治理 | 高 | 规则变化迫使工作流重做或审计收紧 | 高 | 条款、隐私控制、标识对齐 | 规则仍在演进 |
依赖项混合了商业、技术和监管对手方,因为三者都会传导到增长和利润率。
[CR022, CR023, CR024, CR025, CR032]| 角色 / 职能 | 依赖或缺口 | 发生概率 | 严重性 | 缓释措施 | 尽调路径 |
|---|---|---|---|---|---|
| 创始人 / 高级技术领导层 | 推理基础设施和监管应对仍高度依赖创始人 | 中 | 中-高 | 近期融资拓宽资源 | 评估二线运营班底 |
| 合规 / 法务运营 | 必须把快速变化的 AI、隐私、电信和标识规则落到产品里 | 中-高 | 高 | 公开条款显示公司有意识 | 索取合规组织架构和升级处理流程 |
| 企业交付 / 支持 | 私有部署和预留实例需要很强的实施纪律 | 中 | 高 | 案例研究显示已有经验 | 索取部署周期、人员配置和支持 SLA |
| 供应商 / 容量规划 | 毛利率和服务质量取决于容量预测和供应商管理是否准确 | 高 | 高 | 合作伙伴关系和国产化适配 | 索取采购治理和应急预案 |
执行风险更多取决于运营纵深,而不是单纯员工规模。
[CR026, CR027, CR028, CR033]SiliconFlow 同时依赖监管机构、算力供应商、模型提供商和企业渠道。
[CR023, CR024, CR025, CR027, CR032]7.4 缓释措施、监控指标与否决触发条件
留存证据确实显示了缓释路径。SiliconFlow 有区分境内和国际的法律界面、实名和内容控制、BYOC / 私有化部署选项、国产芯片优化叙事、预留实例,以及与贵州移动等算力伙伴的联合运营。这些机制能降低部分最明显的风险:私有化部署能帮助有数据本地化顾虑的客户,标识义务已写入条款,专用容量能稳定重负载。问题在于,执行质量能否跟上监管复杂度和商业规模化。 从投资角度看,合适的否决标准是可观察的。围绕 AI 内容标识或电信合规的负面监管事件、出口管制引发算力瓶颈的证据、公有云毛利率改善失败,或企业客户集中度远高于广泛用户数所暗示水平的证据,都会实质性损害投资论点。相反,如果 SiliconFlow 能展示经审计的安全控制、稳定的供应商多元化、清晰的企业留存,以及高容量账户持续迁移到质量更高的专用或私有化部署合同,投资论点会增强。[CR032, CR033, CR034, CR035, CR036, CR037]
| 风险 | 可监测触发项 | 阈值 / 事件 | 行动含义 |
|---|---|---|---|
| 中国监管合规偏离 | 监管通知、执法或要求整改 | 与标识、隐私或电信合规相关的正式行动 | 暂停 / 重新评估合规就绪度 |
| 算力供应冲击 | 容量短缺、重要供应商流失或成本飙升 | 服务降级或毛利率再次恶化 | 重切下行情景和现金跑道假设 |
| 客户质量不达标 | 企业留存偏弱或大客户集中 | 留存显著低于预期,或头部客户占比意外偏高 | 降低确信度,并压力测试估值 |
| 运营成熟度缺口 | 安全事件,或无法提供企业级保障材料 | 重大事件或企业审计失败 | 暂缓或避免给企业护城河背书 |
| 价格战 / 竞争升级 | 利润率持续承压,结构没有改善 | 尽管规模扩大,公有云亏损仍未收窄 | 将规模视为低质量增长 |
终止标准设计成可在尽调或投资后早期监测中观察。
[CR034, CR035, CR036, CR037, CR038, CR039]7.5 图表证据
08估值
8.1 建议、风险评级与价格纪律
SiliconFlow 达到了战略相关性的门槛,但不足以在公开描述的 2026 年 6 月轮次价格上坚定买入。公司有几项投资人想要的特征:真实产品、大市场、开发者采用、企业 / 基础设施用例、在中国推理栈中的可见角色,以及大量新资本。但同一组公开记录也显示,这家公司有公有云毛利为负、供应商依赖、留存可见度有限,以及显著的监管 / 执行复杂度。这种组合要求投资人保持价格敏感。 因此,核心估值结论不是 SiliconFlow 质量低,而是证据质量落后于价格。如果 $1.2B 估值买的是一个未来高毛利基础设施平台,且具备可防守的企业留存和改善中的供应经济性,这笔投资仍可能成立。如果买的是一个仍高度依赖补贴、暴露在价格战中的规模故事,那么当前进入点的安全边际有限。仅凭公开证据,结论应是跟踪,而不是买入。[CV001, CV002, CV003, CV004, CV005, CV006]
| 建议 | 信心 | 风险评级 | 估值立场 | 决策含义 |
|---|---|---|---|---|
| 观察 | 中-低 | 高 | 价格敏感 / 公开证据不足以支撑以 $1.2B 买入 | 继续尽调;只有证据更强或入场条款更好才参与 |
该建议仅基于公开证据,并刻意保持价格敏感。
[CV001, CV002, CV003, CV040]| 论点 | 哪些情况会改变判断 |
|---|---|
| SiliconFlow 是中国具有战略相关性的推理平台,具备广泛模型接入、企业部署选项和生态势能。 | 若 2026 年收入质量、留存和利润率趋势强于当前公开证据,再上调判断。 |
| 近期融资和战略投资方可能强化分发和生态杠杆。 | 若投资方与客户重叠能明确转化为可持续企业收入,再上调判断。 |
| 公有云经济性、算力依赖和监管复杂度,让当前证据基础不足以支持买入判断。 | 若公有云毛利率仍深度为负,或供应商集中度恶化,再进一步下调判断。 |
| 客户广度真实存在,但留存和集中度不透明。 | 若队列留存和支出集中度明显好于当前缺口所暗示的水平,再上调判断。 |
每一行都写成可证伪的投资论点,而不是泛泛的利弊清单。
[CV010, CV011, CV012, CV013, CV031]8.2 投资论点、反论点与当前估值背景
投资论点很直接:SiliconFlow 是少数有足够广度、速度和生态动量的中国推理平台之一,已经重要到不可忽视。它把模型访问、API 分发、私有化部署和国产芯片适配组合在一起,而推理正在成为 AI 部署的经济中心。近期融资看起来也有较高战略含量,产业投资人可能带来分发和生态拉力。这给了公司远超狭窄 API 套利的可选性。 反论点同样直接。公有云经济性仍差,算力租赁主导成本,客户留存质量不透明,平台同时暴露在境内 AI 治理和外部半导体约束之下。与快速增长的西方推理同业相比,SiliconFlow 披露的收入规模小得多,关于毛利质量的公开证据也更弱。这意味着公司可以有战略重要性,但以当前价格进入,风险调整后回报仍可能平庸。披露的 $1.2B 估值从绝对值看并不明显荒谬,但如果没有更强的 2026 年收入和利润率图景,也并不明显有支撑。[CV010, CV011, CV012, CV013, CV014, CV015]
| 情景 | 假设 | 估值 / 回报逻辑 | 关键风险 | 概率信号 |
|---|---|---|---|---|
| 乐观 | 企业留存被证明强劲;专用 / 私有部署占比上升;公有云毛亏损大幅收窄;战略渠道加深。 | 3-4 年退出价值约 US$2.0B–3.0B;若利润率拐点真实,当前价格仍可能跑通。 | 执行仍取决于算力供应和监管控制。 | 有可能,但当前证据并不最支持这一情景。 |
| 基准 | SiliconFlow 仍具重要性、持续增长,并避开严重冲击,但收入质量和利润率路径只会逐步改善。 | 价值区间约 US$1.1B–1.6B;按当前入场价,风险对应的上行空间有限。 | 价格战、合规成本和收入结构质量压住回报。 | 当前公开证据最支持这一情景。 |
| 悲观 | 规模未能修复经济性;集中度或监管开始反噬;下一轮融资条款走弱。 | 价值区间约 US$0.7B–1.0B;下轮降估值或平轮都有可能。 | 供应商冲击、亏损和留存偏弱挤压投资逻辑。 | 鉴于公开缺口,必须认真看待。 |
上述区间是基于情景的公开估计,不是管理层预测或 DCF。
[CV020, CV024, CV025, CV026, CV027, CV028]| 可比公司 | 指标 | 倍数 / 估值 / 状态 | 参考意义 | 局限 |
|---|---|---|---|---|
| Together AI | 2026 年估值和预订额 | 估值 US$8.3B;年度预订额 >US$1.15B | 直接的开放模型推理和基础设施可比公司 | 预订额不等于确认收入;地域和规模不同 |
| Fireworks AI | 2026 年估值和年化收入 | 估值 US$17.5B;年化收入 >US$1B | 聚焦开放模型的强直接推理云可比公司 | 收入规模大得多,披露的变现能力也更强 |
| Baseten | 2026 年估值和年化收入 | 估值 US$13B;年化收入 ~US$600M | 聚焦企业部署的推理平台可比公司 | Sacra 估计和美国企业画像不同于中国语境 |
| CoreWeave | 2025-2026 年收入和估值参照 | IPO 前估值参考 US$23B;2026 年收入指引 US$12B–13B | 显示 AI 算力基础设施的上限和风险 | 比 SiliconFlow 更偏重资本密集型 GPU 云,不是纯推理 API 可比公司 |
该表刻意不完整,因为公开披露直接估值的中国本土纯推理可比公司仍然有限。
[CV014, CV021, CV022, CV023, CV029]8.3 乐观 / 基准 / 悲观情景、可比组与回报逻辑
可比组很重要,因为它凸显了战略品类价值与当前运营证据之间的缺口。Together AI、Fireworks AI 和 Baseten 都拿到数十亿美元估值,但公开来源也显示,它们的收入或订单规模明显更高,企业变现证据更明确。CoreWeave 证明,贴近基础设施的 AI 平台可以变得极有价值;但它也说明,即使规模大得多,资本强度和客户集中度仍可能很严峻。因此,SiliconFlow 名义估值低于这些西方同业,不应被自动解读为便宜。披露经济性的质量较低,仍可以合理支持折价。 这引出三种情景。乐观情景下,SiliconFlow 把规模和战略关系转化为更好的企业留存、更强的专用 / 私有化组合和改善的利润率,业务跑赢当前风险负荷。基准情景下,公司仍然重要但运营杂乱,牵引力足以支撑当前轮次,却不足以从这个价格打出有吸引力的风险投资式上行。悲观情景下,算力依赖、监管和价格竞争阻止利润率拐点,结果要么估值走平,要么未来下轮融资。今天的公开证据更支持基准情景,而不是乐观情景。[CV020, CV021, CV022, CV023, CV024, CV025]
| 触发项 | 阈值 | 如何传导到投资逻辑 | 行动含义 |
|---|---|---|---|
| 监管执法 | 与 AI 标识、隐私或电信合规相关的正式行动或整改 | 削弱执行力和企业信任 | 在解决前暂停 / 回避 |
| 利润率失效 | 到下一个可观察期,公有云经济性仍无实质改善 | 规模叙事未能转化为业务质量 | 按当前价格从观察转为回避 |
| 供应商 / 算力冲击 | 重大容量中断或供应商集中度恶化 | 同时威胁可靠性和毛利率 | 激进重切下行情景 |
| 客户质量低于预期 | 留存或集中度显著差于预期 | 摧毁高溢价基础设施估值的依据 | 不要支付增长溢价 |
| 安全 / 保障尽调不达标 | 无法拿出企业级保障包 | 削弱高价值企业采用逻辑 | 收窄上行情景和销售质量假设 |
这些触发项可观察,也能直接转化为投资行动。
[CV032, CV033, CV034, CV035, CV039]8.4 最终尽调要求与论点破裂触发条件
公开材料中最重要的缺口不是市场需求,而是核保细节。投资人需要当前现金和消耗、2026 年收入运行率、分产品线毛利率趋势、企业留存、头部客户集中度、供应商集中度,以及从 Serverless API 使用转向质量更高的专用或私有化部署的转化。没有这些输入,估值讨论就大多只是叙事。 因此,最终判断是跟踪,高风险、中低置信度。证据可以把论点上调:利润率改善、算力供应多元化、具名企业参考、经审计的安全姿态,以及更好的收入质量指标。证据也可以把论点下调:监管行动、供应商冲击、公有云经济性持续为负,或使用广度仍无法转化为持久留存支出的证明。单纯降价会有帮助,但前提是缺失的尽调项也不再指向结构性弱点。[CV031, CV032, CV033, CV034, CV035, CV036]
| 主题 | 缺失证据 | 重要性 | 负责人 / 尽调路径 |
|---|---|---|---|
| 2026 年收入运行率和结构 | 按无服务器、专用和私有部署拆分的当前运行率 | 决定本轮承销依据是质量改善,还是单纯增长 | 管理层 / 财务尽调 |
| 留存和集中度 | NRR、GRR、流失率、头部客户占比、支出十分位 | 收入耐久性的核心问题 | 管理层 / 客户尽调 |
| 当前现金和烧钱速度 | 报告日现金、月度烧钱、扣除限制后的融资到账 | 决定现金跑道和下一轮风险 | 财务尽调 |
| 供应商集中度和应急方案 | 头部供应商、合同条款、国产芯片备份方案经济性 | 直接影响可靠性和毛利率风险 | 运营 / 采购尽调 |
| 安全和保障 | SOC 2 / ISO、正常运行时间记录、事故历史、SLA 包 | 为企业护城河质量背书所必需 | 安全尽调 |
| API 到企业合约的扩张动作 | 从按量付费用量转化为预留 / 私有合约 | 显示自下而上的采用能否叠加成更好的收入质量 | GTM 尽调 |
所有问题都被选入,是因为它们会直接移动投资判断。
[CV036, CV037, CV038, CV040]8.5 图表证据
免责声明
本报告仅用于尽调和信息参考。报告基于截至 2026-07-22 的公开材料,不构成投资、法律、会计或税务建议。SiliconFlow 是一家私营公司;公开披露仍未完整覆盖留存、客户集中度、当前年化收入水平和企业级保障。读者在做出投资决策前,应自行核验所有事实,并获取一手尽调材料。
证据索引
| 编号 | 陈述 | 可信度 | 来源 |
|---|---|---|---|
| CO001 | Beijing SiliconFlow Technology Co., Ltd. was established in the PRC as a limited liability company on August 29, 2023. | 中 | SO011 |
| CO002 | SiliconFlow's registered office and head office are at Room 2301, 23/F, Tower D, Building 8, No. 1 Yard, Zhongguancun East Road, Haidian District, Beijing. | 高 | SO001, SO011 |
| CO003 | The company converted to a joint stock limited company on June 22, 2026. | 中 | SO011 |
| CO004 | SiliconFlow maintains a Hong Kong place of business while stating in the HKEX filing that its headquarters, senior management, and operations are primarily based outside Hong Kong. | 中 | SO011 |
| CO005 | The global .com service is governed by SiliconFlow Technology Pte. Ltd., and the global terms explicitly direct mainland-China users to use siliconflow.cn instead. | 高 | SO008, SO009 |
| CO006 | SiliconFlow publicly positions itself as a global AI infrastructure provider whose mission is to accelerate AGI for the benefit of all. | 高 | SO002, SO003 |
| CO007 | The China site presents SiliconFlow as a product suite spanning large-model APIs, reserved instances, inference acceleration services, and private deployment. | 高 | SO001, SO025 |
| CO008 | SiliconFlow's API documentation describes the platform as a one-stop cloud service for top-tier large-language-model APIs aimed at developers and enterprises. | 中 | SO007 |
| CO009 | The global homepage markets SiliconFlow as one platform for AI inference across text, image, video, audio, search, coding, and agent workloads. | 中 | SO002 |
| CO010 | SiliconFlow's GitHub organization says the company integrates hundreds of state-of-the-art models across language, speech, vision, and multimodal domains on top of a self-developed inference engine. | 中 | SO010 |
| CO011 | SiliconFlow's quickstart and cloud login surfaces show sign-in options via SMS, email, GitHub, and Google, alongside self-serve API key creation. | 高 | SO005, SO024 |
| CO012 | The chat-completions documentation exposes an OpenAI-style API surface with model selection, streaming, max_tokens, JSON-format support, and tool-related parameters. | 高 | SO006, SO023 |
| CO013 | As of April 30, 2026, SiliconFlow's platform had over 10 million registered users. | 中 | SO011 |
| CO014 | SiliconFlow recorded average daily token throughput of approximately 578.5 billion and peak daily throughput of approximately 1,071.4 billion in April 2026. | 中 | SO011 |
| CO015 | As of the latest practicable date in the HKEX filing, SiliconFlow had served over 13,000 enterprise customers. | 中 | SO011 |
| CO016 | As of the latest practicable date in the HKEX filing, SiliconFlow had supported a cumulative total of over 170 models. | 中 | SO011 |
| CO017 | By July 2026, SiliconFlow's public model library marketed more than 200 models, indicating platform expansion after the June 2026 filing snapshot. | 中 | SO004 |
| CO018 | The HKEX filing describes SiliconFlow as the largest independent ecosystem token supplier and one of the top five token suppliers overall in China. | 中 | SO011 |
| CO019 | The HKEX filing also describes SiliconFlow as a global leader in token throughput, registered users, and monthly active users, with globally leading overseas platform downloads and token throughput on authoritative platforms. | 中 | SO011 |
| CO020 | SiliconFlow's China product page advertises 10x-plus speed gains for language models, 66% image-model cost savings, 46% language-model cost savings, and BYOC plus isolation-based security controls. | 中 | SO025 |
| CO021 | The China homepage discloses a Beijing ICP filing and a Beijing value-added telecommunications business permit. | 中 | SO001 |
| CO022 | Dr. Yuan Jinhui is disclosed as SiliconFlow's founder, chairman, executive director, CEO, general manager, and financial controller. | 中 | SO011 |
| CO023 | Mr. Liu Juncheng is disclosed as executive director and chief technology officer. | 中 | SO011 |
| CO024 | Mr. Zeng Hua is disclosed as executive director and deputy general manager overseeing commercialization planning. | 中 | SO011 |
| CO025 | Mr. Chen Yingjie is disclosed as non-executive director and is also the managing director of Alibaba Group's strategic investment department. | 中 | SO011 |
| CO026 | Upon listing, SiliconFlow's board is expected to comprise seven directors: three executive, one non-executive, and three independent non-executive directors. | 高 | SO011, SO012 |
| CO027 | The HKEX history section says SiliconFlow developed under the leadership of co-founders Dr. Yuan, Liu Juncheng, Zeng Hua, Zhao Zhen, and Hu Jian. | 中 | SO011 |
| CO028 | Dr. Yuan previously served as a lead researcher at Microsoft China and later founded OneFlow, a deep-learning framework company. | 高 | SO011, SO018 |
| CO029 | Liu Juncheng previously worked in software development and AI system software, including as an R&D engineer at OneFlow from 2018 to 2023. | 中 | SO011 |
| CO030 | Zeng Hua previously held senior roles at Microsoft China, Baidu, and JD.com before joining SiliconFlow. | 中 | SO011 |
| CO031 | Independent profile databases also describe SiliconFlow as Beijing-based and founded in 2023. | 中 | SO020, SO021 |
| CO032 | The HKEX filing records angel, angel+, pre-A, and Series A capital injections of RMB47.20 million, RMB65.84 million, RMB71.48 million, and RMB285.99 million, respectively. | 中 | SO011 |
| CO033 | The HKEX filing records Series A+, Series B, and Series B+ capital injections of RMB220 million, RMB520 million, and RMB740 million, respectively, all in 2026. | 中 | SO011 |
| CO034 | The HKEX filing shows post-money valuation steps of RMB280.0 million, RMB565.8 million, RMB985.0 million, RMB2.286 billion, RMB3.120 billion, RMB5.020 billion, and RMB7.740 billion from angel through Series B+. | 中 | SO011 |
| CO035 | The HKEX filing says SiliconFlow completed the Series A+, Series B, and Series B+ financings after December 31, 2025 for aggregate cash consideration of approximately RMB1.48 billion. | 中 | SO011 |
| CO036 | External June 2026 financing coverage described SiliconFlow as completing over RMB2 billion of Series B financing backed by investors including Trip.com or Ctrip Strategic Investment, JinkoSolar, Kingdee, Unicom-linked capital, Biren, NIO Capital, SenseTime, GGV, and others. | 高 | SO013, SO014, SO015, SO016, SO017 |
| CO037 | Caixin Global said the June 2026 financing was SiliconFlow's fifth funding round since inception and that China Renaissance was the exclusive financial advisor. | 中 | SO013 |
| CO038 | Sina and 36Kr coverage said SiliconFlow served over 10 million users and 10,000 enterprise customers, grew revenue by over 10 times year-on-year, and reached millions of US dollars in overseas monthly revenue. | 中 | SO014, SO015, SO016, SO017 |
| CO039 | The HKEX filing says SiliconFlow's revenue rose from RMB7.3 million in 2024 to RMB55.3 million in 2025. | 中 | SO011 |
| CO040 | The HKEX filing says SiliconFlow's gross margin fell from 39.4% in 2024 to negative 24.0% in 2025 as cost of revenue rose faster than revenue. | 中 | SO011 |
| CO041 | The HKEX milestone section says SiliconFlow's overseas monthly revenue exceeded US$1 million in May 2026. | 中 | SO011 |
| CO042 | SiliconFlow said it started R&D on a large-model inference engine in August 2023. | 高 | SO017, SO011 |
| CO043 | SiliconFlow launched public-cloud MaaS in May 2024. | 高 | SO011, SO017 |
| CO044 | SiliconFlow launched DeepSeek inference services on Huawei Ascend in February 2025, which it described as an industry-first ultra-large-scale domestic-chip token-production service. | 高 | SO011, SO017 |
| CO045 | SiliconFlow launched private MaaS in September 2025 for customers with their own computing power and stricter data-compliance needs. | 高 | SO011, SO017 |
| CO046 | SiliconFlow launched the Elastic GPU heterogeneous-compute scheduling engine in April 2026. | 高 | SO011, SO017 |
| CO047 | The July 14, 2026 HKEX announcement added China Renaissance Securities (Hong Kong) as overall coordinator, signalling continued progress in the company's Hong Kong listing process. | 中 | SO012 |
| CO048 | KrASIA argued that the prospectus shows SiliconFlow is selling tokens at a loss in a compute-rental-heavy business exposed to pricing pressure. | 中 | SO018 |
| CO049 | Hello China Tech said SiliconFlow's public-cloud business carried a negative 119% gross margin in 2025 while on-premise deployment remained high-margin but harder to scale. | 中 | SO019 |
| CO050 | CB Insights lists SiliconFlow with $316.54 million total raised and a June 16, 2026 Series B of $295.84 million, a database presentation that does not fully reconcile to the HKEX B and B+ sequence. | 低 | SO021 |
| CO051 | Tracxn lists Pan Yang and Jinhui Yuan as SiliconFlow's co-founders, which conflicts with the broader co-founder slate named in the HKEX filing. | 低 | SO020 |
| CO052 | In July 2026, SiliconFlow announced Moonshot AI's Kimi K3 on its platform at $3 per million input tokens and $15 per million output tokens. | 高 | SO022, SO004 |
| CO053 | SiliconFlow's Kimi K3 API post says the serverless endpoint supports image input, tool calling, JSON Mode, streaming, and reasoning output. | 高 | SO023, SO006 |
| CM001 | SiliconFlow's relevant market sits in the inference middleware layer rather than in model training or end-user AI applications. | 高 | SM001, SM021 |
| CM002 | The included spend boundary covers serverless token APIs, dedicated instances, and private deployment of models through a unified interface. | 高 | SM001, SM021 |
| CM003 | Model training spend, semiconductor design, and end-user AI SaaS seats largely sit outside SiliconFlow's directly served market. | 中 | SM001, SM025 |
| CM004 | The main substitutes are hyperscaler model platforms, other independent inference APIs, direct model-lab APIs, and internal self-hosting. | 中 | SM001, SM008, SM011, SM016 |
| CM005 | The HKEX filing explicitly distinguishes independent ecosystem token supply platforms from closed ecosystems that bind users to proprietary compute and models. | 高 | SM001, SM024 |
| CM006 | SiliconFlow's own segmentation runs from developers and startups seeking cost-efficient on-demand access to large enterprises needing dedicated performance or supply assurance. | 高 | SM001, SM008 |
| CM007 | BYOK and preselected integrations reduce adoption friction by letting customers keep existing tools and workflows while switching token providers. | 高 | SM001, SM022 |
| CM008 | Multi-model inference platforms increasingly sell one-API access across text, image, video, code, and audio rather than single-modality access. | 高 | SM016, SM018, SM020, SM021 |
| CM009 | IDC says China enterprise MaaS token consumption rose from 114 trillion tokens in 2024 to 1,944 trillion tokens in 2025, roughly a 16x increase. | 中 | SM002 |
| CM010 | IDC projects China token consumption around 40,000 trillion in 2026, about 20x above 2025. | 高 | SM002, SM003 |
| CM011 | IDC sizes China public-cloud MaaS revenue at RMB3.07 billion in 2025. | 中 | SM002 |
| CM012 | IDC projects China public-cloud MaaS revenue reaching RMB18.6 billion in 2026. | 中 | SM002 |
| CM013 | Frost & Sullivan, as cited in the HKEX filing, says China's token supply market grew 1,602.6% from 2024 to 2025. | 中 | SM001 |
| CM014 | The same filing projects China's token supply market to reach approximately 53.2 quintillion tokens by 2030, implying a 638.3% CAGR from 2025 to 2030. | 中 | SM001 |
| CM015 | SiliconFlow held about 1.5% of China token-supply throughput in 2025, ranking fourth overall and first among independent ecosystem platforms. | 高 | SM001, SM024 |
| CM016 | MarketsandMarkets values the global AI inference market at USD106.15 billion in 2025 and USD254.98 billion in 2030, a 19.2% CAGR. | 中 | SM005 |
| CM017 | Grand View Research estimates the global AI inference market at USD97.24 billion in 2024 and USD253.75 billion in 2030, a 17.5% CAGR. | 中 | SM006 |
| CM018 | Fortune Business Insights sizes the global AI inference market at USD103.73 billion in 2025 and USD312.64 billion by 2034. | 中 | SM007 |
| CM019 | Independent analyst pages cluster the broad global AI inference market around roughly USD100 billion in the 2024-2025 base period. | 高 | SM005, SM006, SM007 |
| CM020 | Third-party global estimates agree that Asia Pacific is among the fastest-growing regions, even when they disagree on precise base-year values. | 中 | SM005, SM006, SM007 |
| CM021 | IDC's China MaaS lens and Frost & Sullivan's China token-supply lens agree on hypergrowth but do not use identical market definitions or baselines. | 中 | SM001, SM002 |
| CM022 | Buyer segments span individual developers, startups, enterprise platform teams, and regulated large organizations. | 高 | SM001, SM008, SM015 |
| CM023 | Dedicated performance, stable supply, latency, and private-environment deployment become important as buyers move from experiments to production-grade enterprise use. | 高 | SM001, SM011, SM015 |
| CM024 | AWS positions Bedrock as serving more than 100,000 organizations worldwide, from startups to global enterprises. | 中 | SM008 |
| CM025 | Microsoft Foundry packages models, agents, governance, RBAC, networking, and policy into one management plane, pointing to enterprise platform owners as the core buyer. | 高 | SM011, SM012 |
| CM026 | Alibaba's Token Plan shows inference demand is also being budgeted as team-productivity seats and shared credit pools rather than only as raw API consumption. | 中 | SM015 |
| CM027 | Together and Fireworks both emphasize OpenAI-compatible APIs and low-friction serverless onboarding, showing that portability is now a core expectation in this market. | 高 | SM016, SM018, SM020 |
| CM028 | SiliconFlow distributes through direct API integration, BYOK, and preselected tools such as LangChain, TRAE, Dify, and Cherry Studio. | 中 | SM001 |
| CM029 | Stanford's 2026 AI Index says frontier-model capability kept accelerating in 2025 and organizational AI adoption reached 88%. | 中 | SM004 |
| CM030 | Stanford also reports the U.S.-China frontier-model performance gap effectively closed by early 2026. | 中 | SM004 |
| CM031 | IDC says competition in China MaaS is shifting from pure price competition toward combined price, performance, and toolchain support. | 高 | SM002, SM003 |
| CM032 | IDC ranks performance, security and compliance, answer quality, platform availability, and cost effectiveness among the top enterprise selection factors, with cost only fifth for now. | 中 | SM002 |
| CM033 | Major platform vendors increasingly market routing, evaluation, observability, privacy, governance, and dedicated throughput as core product features, not extras. | 高 | SM008, SM011, SM012, SM020 |
| CM034 | AWS Bedrock and Microsoft Foundry both frame enterprise security, governance, and cost optimization as central to adoption. | 高 | SM008, SM011 |
| CM035 | Alibaba, Together, and Fireworks all offer packaging beyond simple pay-as-you-go, including seat subscriptions, dedicated endpoints, cached-token discounts, batch discounts, or hourly GPU deployment pricing. | 高 | SM015, SM017, SM019 |
| CM036 | The HKEX filing says SiliconFlow intentionally prioritized market share, user acquisition, and ecosystem building over immediate profitability in its public-cloud business. | 高 | SM001, SM024, SM025 |
| CM037 | Hello China Tech describes a severe 2023-2026 price war in model APIs, with mainstream prices down more than 90% since 2023 and further 2026 cuts from major vendors. | 中 | SM025 |
| CM038 | Compute rental, heterogeneous chip support, and supplier access remain structural constraints that matter for both service reliability and margins. | 中 | SM001, SM003, SM025 |
| CM039 | Global AI inference demand is broadening beyond text chat into edge, real-time, multimodal, and automation-heavy workloads. | 中 | SM005, SM007, SM018 |
| CM040 | Fortune Business Insights explicitly lists high hardware costs and integration challenges as adoption restraints for the AI inference market. | 中 | SM007 |
| CM041 | No retained public source isolates SiliconFlow's serviceable share by geography, customer segment, or category-specific retention. | 低 | SM001, SM024, SM025 |
| CM042 | The most defensible market thesis is a stacked one: use global inference infrastructure for context, China MaaS growth for monetizable demand, and SiliconFlow throughput share only as a competitive signal. | 中 | SM001, SM002, SM005, SM006, SM007 |
| CM043 | The layered market-sizing pyramid is directional only because it mixes revenue, token-throughput, and market-share units rather than one additive denominator. | 中 | SM001, SM002, SM005 |
| CM044 | A common adoption path in this market is direct API experimentation first, then broader tool integration, then dedicated or private deployment once scale or governance requirements rise. | 中 | SM001, SM017, SM020 |
| CP001 | SiliconFlow competes across independent open inference platforms, incumbent cloud platforms, direct model-lab APIs, and internal build substitutes. | 高 | SP001, SP004, SP006, SP009, SP012, SP016, SP021 |
| CP002 | SiliconFlow ranks fourth in China token-supply throughput and first among independent ecosystem platforms in 2025. | 高 | SP001, SP023 |
| CP003 | Together AI positions itself around serverless and dedicated access to open models rather than a proprietary closed model stack. | 中 | SP012, SP013, SP014, SP015 |
| CP004 | Fireworks positions itself as a specialized training-and-inference platform for open models and says it processes 40T+ tokens per day. | 中 | SP016, SP017, SP018, SP020 |
| CP005 | OpenRouter positions itself as a unified API and broker that routes requests across hundreds of models and providers. | 中 | SP021, SP022 |
| CP006 | AWS Bedrock, Microsoft Foundry, and Alibaba Model Studio are incumbent substitutes with broader governance and control-plane depth than most independent peers. | 高 | SP004, SP006, SP007, SP008, SP009 |
| CP007 | Alibaba Model Studio combines official Qwen ownership with third-party model access and OpenAI-compatible APIs. | 高 | SP009, SP010, SP011 |
| CP008 | Retained sources here disclose far more about product scope and pricing than about current funding or revenue for Together, Fireworks, or OpenRouter. | 低 | SP012, SP016, SP021 |
| CP009 | Internal build and direct model-lab APIs remain practical substitutes because many rivals expose portable, OpenAI-like integration paths. | 中 | SP013, SP020, SP021, SP025 |
| CP010 | OpenAI-compatible or drop-in API paths are common across SiliconFlow, OpenRouter, Fireworks, Together, and Alibaba. | 高 | SP009, SP015, SP020, SP021, SP025 |
| CP011 | Broad model-catalog competition is intense: SiliconFlow advertises 200+ models, Bedrock 100+, OpenRouter hundreds, and Foundry 1,900+ to 11,000+ access surfaces. | 中 | SP004, SP007, SP021, SP025 |
| CP012 | Together and Fireworks both offer a serverless-to-dedicated progression, while hyperscalers offer reserved or managed-compute paths for heavier production use. | 中 | SP005, SP008, SP013, SP014, SP018 |
| CP013 | Enterprise incumbents differentiate with guardrails, RBAC, policies, network isolation, observability, and managed-agent features. | 高 | SP004, SP006, SP007, SP008 |
| CP014 | OpenRouter differentiates on provider routing, fallbacks, price/throughput/latency sorting, and data-retention-aware controls. | 中 | SP021, SP022 |
| CP015 | Together differentiates on open-model access, a shared serverless API, and dedicated GPU deployments that reuse the same inference API. | 中 | SP012, SP013, SP014, SP015 |
| CP016 | Fireworks differentiates on serverless paths, prompt caching, on-demand infrastructure, and a combined training/inference posture. | 中 | SP016, SP018, SP019, SP020 |
| CP017 | SiliconFlow differentiates on heterogeneous compute, domestic-chip adaptation, and neutral multi-model positioning inside China. | 中 | SP001, SP003, SP025 |
| CP018 | Together lists Qwen 3.7 Max at $1.25 input and $3.75 output per 1M tokens, and DeepSeek V4 Pro at $1.74 input and $3.48 output. | 中 | SP015 |
| CP019 | Fireworks lists DeepSeek V4 Pro at $1.74 input / $0.145 cached input / $3.48 output and GPT OSS 20B at $0.07 input / $0.035 cached / $0.30 output. | 中 | SP019 |
| CP020 | AWS Bedrock pricing spans premium models such as Claude Opus 4.8 at $6 input / $30 output and lower-cost DeepSeek variants around $0.62 / $1.85 in listed regions. | 中 | SP005 |
| CP021 | Alibaba posts flagship Qwen list pricing with temporary regional discounts and also sells seat-based credit plans starting at $30 per seat per month. | 高 | SP010, SP011 |
| CP022 | OpenRouter makes pricing logic part of the product by letting customers sort providers by price, latency, or throughput and set max_price or performance thresholds. | 中 | SP022 |
| CP023 | Together and Fireworks both discount cached tokens and batch workloads, showing that heavy users are expected to demand lower effective pricing at scale. | 中 | SP013, SP019 |
| CP024 | API-surface switching costs are low because many rivals support OpenAI-compatible integration paths or drop-in migration. | 高 | SP006, SP009, SP020, SP021, SP025 |
| CP025 | OpenRouter is optimized for multi-homing rather than single-provider lock-in. | 中 | SP021, SP022 |
| CP026 | SiliconFlow's BYOK and preselected-tool distribution help adoption but also make customer multi-homing plausible. | 中 | SP001, SP025 |
| CP027 | Hyperscalers counter low API switching costs with identity, policy, networking, procurement, and broader platform integration. | 高 | SP004, SP006, SP007, SP009 |
| CP028 | AWS, Microsoft, and Alibaba can cross-sell inference from much broader cloud estates than independent startups can. | 中 | SP004, SP006, SP009 |
| CP029 | SiliconFlow's strongest retained moat candidate is China-local neutrality plus heterogeneous chip and deployment coverage. | 中 | SP001, SP003, SP025 |
| CP030 | Together and Fireworks compete more on speed, deployment control, and open-model operations than on exclusive model ownership. | 中 | SP013, SP014, SP016, SP018 |
| CP031 | OpenRouter's moat is routing intelligence and provider liquidity, but the model is inherently less lock-in-oriented than compute-owning platforms. | 中 | SP021, SP022 |
| CP032 | The top three China token suppliers above SiliconFlow are hyperscaler divisions, leaving SiliconFlow meaningfully smaller than the largest incumbents. | 高 | SP001, SP024 |
| CP033 | Hello China Tech argues mainstream model-API prices have fallen more than 90% since 2023, highlighting commoditization pressure. | 中 | SP024 |
| CP034 | SiliconFlow's filing shows that public-cloud market-share growth can coexist with negative gross margins and high compute-rental pressure. | 高 | SP001, SP023, SP024 |
| CP035 | Competitive risk is highest where buyers treat model access as commodity and can switch on price, latency, or uptime. | 中 | SP015, SP019, SP022 |
| CP036 | Competitive risk is lower where buyers need China-specific compute options, private deployment, or neutral multi-model support outside a single cloud. | 中 | SP001, SP009, SP025 |
| CP037 | No retained source proves customer lock-in for SiliconFlow comparable to hyperscaler IAM or network lock-in. | 低 | SP004, SP006, SP009, SP025 |
| CP038 | No retained source here proves peer funding superiority or current margin superiority for SiliconFlow versus Together, Fireworks, or OpenRouter. | 低 | |
| CP039 | The competitive field is crowded partly because the same buyer job can be solved by a cloud platform, a neutral platform, a router, or internal build. | 中 | SP001, SP004, SP006, SP021 |
| CP040 | The best positioning map for this market places independent platforms high on openness and hyperscalers high on enterprise control and procurement strength. | 中 | SP004, SP006, SP009, SP012, SP016, SP021, SP025 |
| CP041 | Capability overlap is highest on basic API access and model breadth, while divergence is greatest on routing, governance, and deployment-control depth. | 中 | SP006, SP009, SP014, SP018, SP022, SP025 |
| CP042 | Moat readiness in this market depends more on supply access, governance, and operating efficiency than on simple model-count marketing. | 中 | SP001, SP002, SP004, SP006, SP024 |
| CP043 | Fireworks' vendor-authored 40T+ tokens/day claim indicates that some independent peers already operate at very large throughput scale. | 中 | SP016 |
| CP044 | A common adoption path is prototype on serverless or a broker, then shift toward dedicated capacity or deeper cloud control once scale and governance requirements rise. | 中 | SP005, SP013, SP014, SP018 |
| CP045 | SiliconFlow's competitive verdict is positive on relevance but still unproven on long-term durability versus hyperscaler bundling and category-wide price compression. | 中 | SP001, SP024, SP025 |
| CP046 | Because basic API access and model breadth overlap heavily, competitive selection often shifts to surrounding control surfaces, routing logic, and deployment guarantees. | 中 | SP006, SP009, SP015, SP020, SP025 |
| CI001 | SiliconFlow has two primary revenue lines: public cloud-based services and on-premise deployment solutions. | 中 | SI001, SI004 |
| CI002 | Public cloud-based services include serverless token services and dedicated instances. | 中 | SI001 |
| CI003 | On-premise deployment solutions install inference software in customer environments and currently carry much higher gross margins than public cloud. | 高 | SI001, SI004 |
| CI004 | In 2025 public cloud generated RMB29.261 million, or 52.9% of total revenue, while on-premise generated RMB26.069 million, or 47.1%. | 高 | SI001, SI004 |
| CI005 | SiliconFlow's public pricing surface is usage-based and model-level rather than seat-based. | 中 | SI002 |
| CI006 | List pricing in the inference market is shaped by cached-token discounts, batch discounts, and migrations toward dedicated capacity at scale. | 高 | SI006, SI007, SI010, SI012, SI014 |
| CI007 | Together explicitly frames serverless as cheaper for low or bursty traffic and dedicated replicas as cheaper when utilization stays high. | 中 | SI006 |
| CI008 | AWS, Fireworks, and Together all offer roughly 50% batch discounts in at least part of their pricing stack. | 高 | SI007, SI012, SI014, SI015 |
| CI009 | Alibaba monetizes the category through both pay-as-you-go model calls and seat-based subscription credits. | 高 | SI019, SI020, SI021 |
| CI010 | Official list pricing is useful for market context but is insufficient to infer SiliconFlow's realized pricing or margin. | 中 | SI002, SI019, SI024 |
| CI011 | SiliconFlow reported RMB55.33 million of revenue in 2025, up from RMB7.346 million in 2024. | 高 | SI001, SI004, SI005 |
| CI012 | SiliconFlow's 2025 blended gross margin was -24.0%. | 高 | SI001, SI004 |
| CI013 | The 2025 public-cloud gross loss margin was -119.0%, after -271.6% in 2024. | 高 | SI001, SI004 |
| CI014 | The 2025 on-premise deployment gross margin was 82.5%. | 中 | SI001 |
| CI015 | Compute rental fees accounted for 86.9% of 2025 cost of sales. | 高 | SI001, SI005 |
| CI016 | SiliconFlow explicitly prioritized market share, user acquisition, and ecosystem building over immediate profitability in public cloud. | 高 | SI001, SI004 |
| CI017 | Serverless paying accounts increased from 2,455 to 716,000 in 2025. | 中 | SI004, SI005 |
| CI018 | Using year-end paying-account count as a crude divisor, RMB14.3 million of serverless revenue implies very low annual spend density per paying account. | 中 | SI004 |
| CI019 | More than 64% of 2025 sales and marketing expense went to promotional compute credits. | 中 | SI004 |
| CI020 | Independent analyses argue that SiliconFlow is operating as a compute-renting middle layer inside a price-war environment. | 中 | SI004, SI005 |
| CI021 | IDC says cost effectiveness matters to buyers, but it trails performance, compliance, answer quality, and platform usability as a current selection factor. | 中 | SI022 |
| CI022 | Usage metrics such as registered users and paying-account growth prove demand but do not by themselves prove revenue quality. | 中 | SI001, SI004, SI005 |
| CI023 | At 2025 year-end SiliconFlow held RMB171 million of cash and cash equivalents plus RMB100 million of time deposits. | 高 | SI001, SI005 |
| CI024 | Net cash used in operating activities was RMB172 million in 2025. | 高 | SI001, SI005 |
| CI025 | On a simple backward-looking lens, SiliconFlow was not self-funding before the 2026 financings. | 中 | SI001, SI023, SI005 |
| CI026 | The filing discloses about RMB1.48 billion of post-2025 cash consideration across the A+, B, and B+ rounds. | 中 | SI001 |
| CI027 | SiliconFlow does not directly purchase chips; it mainly leases computing resources through partners. | 高 | SI005, SI001 |
| CI028 | Major purchases are concentrated among the top five suppliers, tying capital adequacy to supplier terms as well as to cash balances. | 高 | SI001, SI005 |
| CI029 | The retained public record does not support a precise current runway calculation as of 2026-07-22. | 低 | SI001, SI005 |
| CI030 | The 2026 financings improved capital adequacy, but they did not by themselves solve revenue-quality or margin-path questions. | 中 | SI001, SI004, SI005 |
| CI031 | The public record does not reveal net revenue retention, gross retention, or customer concentration. | 低 | SI001, SI002 |
| CI032 | The public record does not reveal realized effective pricing by model family or customer cohort. | 低 | SI002, SI019, SI024 |
| CI033 | The public record does not reveal a current monthly cash-burn figure or a management-backed runway target as of run date. | 低 | SI001, SI005 |
| CI034 | On-premise revenue appears higher quality on margin, but public sources are insufficient to show whether it can scale enough to change the whole-company profile. | 中 | SI001, SI004 |
| CI035 | Public-cloud token revenue remains the core growth engine but currently appears subsidy-heavy and compute-rental-heavy. | 高 | SI001, SI004, SI005 |
| CI036 | List pricing across the category increasingly pushes heavy users toward dedicated or reserved capacity once workloads stabilize. | 中 | SI006, SI013, SI014, SI018, SI024 |
| CI037 | Prompt caching is economically important because it can lower effective token costs or reduce compute wasted on repeated context. | 高 | SI008, SI010, SI016, SI025 |
| CI038 | Fine-tuning and dedicated hosting create additional monetization paths in the category, but retained sources do not show SiliconFlow currently monetizing those paths separately. | 中 | SI009, SI013, SI018 |
| CI039 | The clearest financial bridge is from usage into three revenue-quality tiers—serverless, dedicated, and on-premise—rather than into one uniform SaaS line. | 中 | SI001, SI002, SI004 |
| CI040 | The divergence between demand growth and financial quality is the central financial thesis of SiliconFlow as of 2026-07-22. | 中 | SI001, SI004, SI005 |
| CE001 | SiliconFlow is an API-first inference platform rather than a single-model application. | 高 | SE002, SE008 |
| CE002 | Its public product menu spans ready-to-use model APIs, reserved instances, inference acceleration services, and private deployment. | 高 | SE001, SE021, SE008 |
| CE003 | Serverless APIs are the core developer-facing surface. | 高 | SE001, SE002, SE004 |
| CE004 | Private deployment and BYOC are presented as enterprise options for privacy-sensitive workloads. | 中 | SE001, SE006 |
| CE005 | Users follow a simple path of model selection, API key creation, and endpoint integration. | 高 | SE002, SE004 |
| CE006 | The catalog spans text, speech, image, video, vector, reranking, and multimodal model classes. | 高 | SE002, SE013 |
| CE007 | Reserved instances are positioned for enterprise core inference scenarios needing dedicated capacity and cost optimization. | 高 | SE001, SE014 |
| CE008 | Inference acceleration is marketed both for open-source models and for self-developed models. | 高 | SE001, SE002 |
| CE009 | The product bundle reduces the need for customers to stitch together separate API, deployment, and optimization vendors. | 中 | SE001, SE002, SE019 |
| CE010 | SiliconFlow exposes model choice, streaming, context-window controls, tool calling, and request tracing through its chat API surface. | 中 | SE004 |
| CE011 | The platform uses a largely OpenAI-style integration pattern, lowering adoption friction for developers already familiar with that schema. | 中 | SE004, SE018 |
| CE012 | The docs show SiliconFlow regularly updates model availability and service capabilities over time. | 高 | SE003, SE004 |
| CE013 | Developers can stay on serverless or graduate toward reserved instances and private deployments as workloads mature. | 中 | SE001, SE014, SE016 |
| CE014 | SiliconFlow claims self-developed efficient operators, optimization frameworks, and a leading inference acceleration engine. | 高 | SE001, SE002 |
| CE015 | OneDiff is an acceleration library for diffusion models with optimized GPU kernels and compiler tooling. | 中 | SE010 |
| CE016 | OneDiff release history shows ongoing engineering maintenance rather than a one-off repository dump. | 中 | SE011 |
| CE017 | The API docs expose x-siliconcloud-trace-id response headers for request tracing and troubleshooting. | 中 | SE004 |
| CE018 | The strongest external engineering proof currently available is around inference-optimization tooling, not independently audited production reliability. | 中 | SE010, SE011, SE004 |
| CE019 | Public sources do not provide audited SLOs, public error-rate dashboards, or third-party performance benchmarks for the whole platform. | 低 | SE001, SE002, SE004 |
| CE020 | SiliconFlow differentiates through a combination of catalog breadth, integration simplicity, deployment flexibility, and optimization claims. | 高 | SE001, SE002, SE005, SE013 |
| CE021 | The company continues to add current model launches, as shown by public launch messaging around recent models such as GLM-5.2 and Kimi K3. | 中 | SE001, SE012 |
| CE022 | The product moat is operational and integration-driven rather than based on ownership of a proprietary frontier model. | 中 | SE008, SE023, SE024 |
| CE023 | The platform depends on upstream model-provider availability and on continued access to external compute capacity. | 中 | SE003, SE008, SE023 |
| CE024 | The public product roadmap appears to be driven heavily by fast model onboarding and packaging cadence. | 中 | SE003, SE012, SE013 |
| CE025 | Private deployment increases appeal to regulated buyers but also increases implementation complexity and service-delivery burden. | 中 | SE001, SE006, SE014 |
| CE026 | The Kimi K3 launch post confirms a live developer surface with a standard API base URL. | 中 | SE012 |
| CE027 | Public maturity is strongest for the serverless API surface and weaker for enterprise assurance artifacts. | 中 | SE001, SE002, SE006, SE007 |
| CE028 | OneDiff release cadence is a useful roadmap proxy for SiliconFlow's optimization work, but it is not a substitute for a full platform changelog. | 中 | SE011, SE010 |
| CE029 | SiliconFlow publicly claims BYOC deployment and compute/network/storage isolation. | 中 | SE001 |
| CE030 | SiliconFlow maintains distinct China and international service surfaces with separate contractual language. | 高 | SE007, SE021 |
| CE031 | The privacy policy confirms the international surface is operated through siliconflow.com. | 中 | SE006 |
| CE032 | Trace identifiers in API responses are a real supportability control, though not a substitute for published reliability reporting. | 中 | SE004 |
| CE033 | The retained source set does not show a public SOC 2 report, ISO certificate list, or detailed public incident history. | 低 | SE001, SE006, SE007 |
| CE034 | Large enterprise underwriting still requires third-party assurance artifacts that are not visible in the retained public set. | 低 | SE006, SE007, SE001 |
| CE035 | The company's security and compliance messaging is plausible but currently stronger in claims than in public evidence depth. | 中 | SE001, SE006, SE007 |
| CE036 | SiliconFlow is built for supportability and enterprise packaging, but public evidence is insufficient to conclude best-in-class enterprise readiness. | 中 | SE001, SE004, SE006, SE007 |
| CU001 | SiliconFlow serves multiple customer types rather than a single homogeneous user base. | 高 | SU001, SU018, SU019 |
| CU002 | Public 2026 coverage cites more than 10 million users and more than 10,000 enterprise customers. | 中 | SU002, SU023 |
| CU003 | The company explicitly serves both developers and enterprises. | 高 | SU001, SU019 |
| CU004 | Enterprise demand includes reserved instances, private deployment, and infrastructure-style deployments rather than only lightweight API experimentation. | 高 | SU005, SU006, SU008 |
| CU005 | Community and developer-tool integrations prove API onboarding relevance across coding, search/RAG, translation, and agent workflows. | 高 | SU004, SU009, SU010, SU011 |
| CU006 | Telecom / compute partners are a distinct strategic customer surface for SiliconFlow. | 高 | SU007, SU003 |
| CU007 | The filing and public coverage imply that self-serve paying accounts and larger enterprise deployments coexist inside the customer base. | 中 | SU001, SU015 |
| CU008 | Broad user reach does not by itself prove high-quality revenue because developer and community adoption can be monetically light. | 中 | SU002, SU004, SU017 |
| CU009 | SiliconFlow appears especially strong in infra-heavy buyer scenarios where inference performance, private deployment, or国产化 adaptation matters. | 中 | SU005, SU006, SU007, SU018 |
| CU010 | Guizhou Mobile is the clearest named institutional counterparty in the retained customer corpus. | 高 | SU007, SU003 |
| CU011 | The Guizhou Mobile partnership covers inference framework deployment, compute coordination, token services, and joint operational systems. | 中 | SU007 |
| CU012 | The Guizhou Mobile relationship is described as a deepening of earlier 2025 cooperation rather than a first-touch pilot. | 中 | SU007 |
| CU013 | The reserved-instance case study shows at least one customer whose coding-agent workload reached 100 billion daily tokens in 2026. | 中 | SU008 |
| CU014 | MindSearch is a named public project with a documented SiliconFlow integration path. | 高 | SU009, SU004 |
| CU015 | Continue is a named public integration that positions SiliconFlow for coding-assistant workflows inside IDEs. | 高 | SU010, SU004 |
| CU016 | Cline is a named public integration that uses SiliconFlow through an OpenAI-compatible API pattern. | 高 | SU011, SU004 |
| CU017 | The named developer-tool proofs are useful evidence of adoption breadth but weak evidence of contract value or exclusivity. | 中 | SU009, SU010, SU011 |
| CU018 | Anonymous enterprise case studies add operational depth but limit concentration and logo-quality analysis because the customer names are withheld. | 高 | SU005, SU006, SU008 |
| CU019 | Named customer proof in this chapter is therefore stronger on breadth than on directly observable commercial weight. | 中 | SU007, SU009, SU010, SU011 |
| CU020 | SiliconFlow shows a land-and-expand pattern from pay-as-you-go usage toward reserved or deeper infrastructure deployment for some accounts. | 中 | SU008, SU025, SU001 |
| CU021 | The retained public set does not disclose NRR, GRR, or logo-retention rates. | 中 | SU001, SU015 |
| CU022 | The retained public set does not disclose average contract length, renewal rates, or top-customer concentration. | 中 | SU001, SU015, SU017 |
| CU023 | Public evidence for repeat usage is architectural and behavioral rather than cohort-based: heavy workloads, reserved-instance upgrades, and deeper partner cooperation. | 中 | SU007, SU008, SU025 |
| CU024 | Private deployment and国产化 case studies suggest deeper embedment than simple API testing, but the public record still lacks renewal proof. | 中 | SU005, SU006, SU021 |
| CU025 | Revenue concentration could still be high even with 10,000+ enterprise customers if a small number of token-heavy accounts dominate spend. | 中 | SU002, SU008, SU017 |
| CU026 | Strategic telecom or infrastructure partners may become concentrated routes to growth, creating bargaining-power and dependency risk. | 中 | SU007, SU021, SU017 |
| CU027 | Broad community integration can inflate awareness and usage without proving durable paid conversion. | 中 | SU004, SU009, SU010, SU011 |
| CU028 | The customer corpus supports diversification by use case, but not yet diversification by revenue contribution. | 中 | SU004, SU005, SU006, SU007 |
| CU029 | SiliconFlow has enough public evidence to show real adoption, not merely claimed logos. | 高 | SU002, SU007, SU009, SU010, SU011 |
| CU030 | Customer durability is materially less proven than customer breadth. | 中 | SU021, SU022, SU015 |
| CU031 | The strongest public customer signal is that some workloads are mission-critical enough to justify dedicated or private infrastructure decisions. | 高 | SU005, SU006, SU008 |
| CU032 | The cleanest public customer journey is discover the API, prove value in workflow, then deepen into reserved or private deployment as usage scales. | 中 | SU019, SU008, SU025 |
| CU033 | For diligence, the biggest remaining customer question is not acquisition but quality of retained spend. | 中 | SU002, SU015, SU017 |
| CU034 | A retention cohort figure can be shown only as a visibility proxy, not as a real renewal chart. | 中 | SU001, SU015 |
| CU035 | Developer adoption and enterprise adoption reinforce each other strategically, but they should not be valued as equivalent commercial proof. | 中 | SU004, SU007, SU010, SU011 |
| CU036 | The right customer underwriting request is a segmented cohort pack covering active customers, spend concentration, renewal, and upgrade from API to reserved/private deployment. | 中 | SU001, SU015, SU017 |
| CR001 | SiliconFlow's domestic terms expressly limit the service to developer-oriented content-generation scenarios and exclude automatic control, medical information, psychological counseling, and critical information infrastructure uses. | 中 | SR002 |
| CR002 | The domestic privacy policy says mainland-law requirements can trigger real-name verification and identity collection for service activation. | 中 | SR003 |
| CR003 | SiliconFlow's domestic terms explicitly reference the AI-generated-content labeling measures and prohibit malicious deletion or tampering with labels. | 中 | SR002 |
| CR004 | The domestic site publicly displays ICP and telecom-license identifiers, indicating regulated internet-service operations rather than a purely offshore surface. | 中 | SR004 |
| CR005 | SiliconFlow maintains separate China and international legal surfaces, reducing but not eliminating cross-surface governance complexity. | 高 | SR005, SR006 |
| CR006 | China's 2025 AI labeling measures create mandatory explicit and implicit labeling obligations for generated synthetic content. | 中 | SR007, SR032 |
| CR007 | Those labeling rules create operational risk because providers and platforms must implement metadata, logging, and downstream labeling workflows rather than merely publish a policy. | 中 | SR002, SR007, SR031 |
| CR008 | The public terms show compliance awareness, but they do not by themselves prove SiliconFlow's implementation quality for labeling, logging, or auditability. | 中 | SR002, SR007 |
| CR009 | SiliconFlow's domestic privacy policy says user information generated in operating the domestic site is stored within the PRC. | 中 | SR003 |
| CR010 | Mainland onboarding and identity checks can increase friction for some customers and enlarge the company's privacy-handling obligations. | 高 | SR002, SR003 |
| CR011 | The legal risk is a continuous compliance burden across content, privacy, telecom, and sector-specific use restrictions rather than a single one-time license hurdle. | 高 | SR002, SR003, SR004, SR007 |
| CR012 | SiliconFlow serves infrastructure-like workloads where outages, latency swings, or deployment failures can matter more than consumer feature gaps. | 高 | SR015, SR016, SR017 |
| CR013 | Public sources do not furnish a robust public uptime history, incident archive, or detailed SLO record. | 低 | SR015, SR018 |
| CR014 | Public sources do not furnish SOC 2, ISO certification details, or equivalent third-party assurance artifacts in the retained corpus. | 低 | SR004, SR005, SR006 |
| CR015 | The retained corpus suggests enterprise-grade operational aspirations, but not enough third-party evidence to conclude best-in-class operational maturity. | 中 | SR015, SR016, SR018 |
| CR016 | Privacy, labeling, and identity controls interact as a compound execution risk because each adds workflow and data-handling obligations to the platform. | 高 | SR002, SR003, SR007 |
| CR017 | The platform's own API reference warns that model availability and capability can change over time. | 中 | SR020 |
| CR018 | Rapid model churn creates support and migration risk for enterprise workloads even when the platform benefits from a broad model catalog. | 中 | SR019, SR020 |
| CR019 | Monitoring, fault tolerance, and traceability are publicly claimed mitigations, but retained sources do not independently validate their effectiveness at scale. | 中 | SR018, SR015 |
| CR020 | Buyer expectations in critical or regulated environments are rising toward explicit trustworthy-AI risk-management practices, even where frameworks such as NIST AI RMF are voluntary. | 高 | SR010, SR021 |
| CR021 | Operational-quality risk is therefore most acute where high-volume or regulated workloads demand both performance and documented assurance. | 中 | SR015, SR016, SR020 |
| CR022 | SiliconFlow depends on upstream model providers for catalog breadth and competitive relevance. | 高 | SR019, SR020, SR029 |
| CR023 | SiliconFlow depends heavily on leased compute and supplier relationships rather than on fully owned semiconductor supply. | 高 | SR001, SR011, SR012 |
| CR024 | U.S. export controls continue to target PRC access to advanced computing semiconductors, creating a real external dependency risk for Chinese AI-infrastructure players. | 高 | SR008, SR009 |
| CR025 | Domestic-chip adaptation and private-deployment narratives mitigate but do not eliminate supply-side risk. | 中 | SR016, SR017, SR024 |
| CR026 | Strategic telecom or compute partnerships can broaden distribution but also create concentration and execution dependencies. | 中 | SR014, SR023 |
| CR027 | Supplier / capacity planning is an execution-critical function because service quality and margin both depend on it. | 高 | SR001, SR011, SR015 |
| CR028 | Private deployment, reserved instances, and domestic-chip optimization increase implementation complexity and require strong enterprise delivery operations. | 高 | SR015, SR016, SR017 |
| CR029 | Independent analyses already frame SiliconFlow's economics as vulnerable to price war and compute-rental pressure. | 中 | SR011, SR012, SR013 |
| CR030 | Competitive pressure from hyperscalers and other inference platforms can intensify the margin impact of compliance and compute-supply costs. | 高 | SR025, SR026, SR027, SR028, SR029, SR030 |
| CR031 | Hidden concentration of token-heavy customers remains a real model risk because public breadth metrics do not reveal spend distribution. | 中 | SR001, SR013, SR011 |
| CR032 | SiliconFlow already has visible mitigations: separate legal surfaces, BYOC/private deployment, domestic adaptation, and strategic partner operations. | 中 | SR005, SR014, SR016, SR017 |
| CR033 | The public record does not yet prove whether SiliconFlow has sufficient second-line compliance, delivery, and supplier-management depth to scale safely. | 低 | SR011, SR012, SR014 |
| CR034 | A formal regulatory event around labeling, privacy, or telecom compliance would be an immediate thesis-break trigger. | 中 | SR002, SR003, SR007 |
| CR035 | A material compute-capacity shock or supplier loss would directly threaten both reliability and gross margin. | 中 | SR001, SR008, SR023 |
| CR036 | Proof that enterprise retention or revenue concentration is worse than implied by public customer counts would materially weaken the investment case. | 中 | SR001, SR011, SR013 |
| CR037 | Failure of public-cloud margins to improve despite scale would indicate that growth is not translating into a durable business model. | 中 | SR001, SR011, SR029 |
| CR038 | The strongest risk transmission path is from regulation or supply shock into operating complexity, then into revenue quality, margin, financing need, and valuation. | 中 | SR007, SR023, SR029 |
| CR039 | The most likely early diligence flashpoint is weak evidence around enterprise assurance, retention, and supplier concentration rather than a near-term demand collapse. | 中 | SR011, SR012, SR015 |
| CR040 | SiliconFlow's risk profile is investable only if the company can keep converting scale into better margins faster than regulatory and supply complexity increase. | 中 | SR001, SR011, SR024 |
| CV001 | On public evidence alone, the right current recommendation is Track rather than Buy. | 中 | SV001, SV004, SV005, SV006, SV009 |
| CV002 | The risk rating for SiliconFlow at the current price is high. | 中 | SV001, SV025, SV026 |
| CV003 | Confidence in the recommendation is only medium-low because the price is public but the core underwriting metrics remain incomplete. | 中 | SV001, SV004, SV005 |
| CV004 | SiliconFlow is strategically relevant enough to merit continued diligence and tracking. | 高 | SV014, SV015, SV021, SV022 |
| CV005 | The company has real customer and usage proof rather than a purely narrative AI story. | 高 | SV002, SV023, SV024 |
| CV006 | Recent financing and strategic investors improve the company's optionality and time to execute. | 中 | SV002, SV003, SV022 |
| CV007 | Current public evidence is insufficient to justify a strong Buy at the disclosed June 2026 price. | 中 | SV001, SV004, SV005, SV025 |
| CV008 | The valuation stance is therefore price-sensitive rather than categorically negative. | 中 | SV002, SV006, SV009 |
| CV009 | A lower entry price or stronger proof package could move the recommendation upward. | 中 | SV001, SV004, SV006, SV009 |
| CV010 | The positive thesis begins with category positioning: SiliconFlow sits in a growing inference market with real product breadth and deployment flexibility. | 高 | SV014, SV015, SV021, SV028 |
| CV011 | The positive thesis also includes ecosystem and strategic-channel leverage, not only direct API usage. | 中 | SV002, SV003, SV022 |
| CV012 | The anti-thesis begins with disclosed weak economics, especially negative public-cloud gross margins. | 高 | SV001, SV004, SV005 |
| CV013 | Compute-rental dependence and supplier concentration weaken the claim that current scale already represents a mature infrastructure moat. | 高 | SV001, SV005, SV026 |
| CV014 | Compared with Western inference peers, SiliconFlow's disclosed revenue scale is much smaller. | 中 | SV001, SV006, SV009, SV012 |
| CV015 | SiliconFlow may still be strategically important even if it deserves a valuation discount to faster-scaling Western peers. | 中 | SV006, SV009, SV012, SV013 |
| CV016 | The current US$1.2B valuation is not obviously absurd in absolute terms, but it is not clearly supported by public underwriting data either. | 中 | SV002, SV004, SV005, SV006 |
| CV017 | The market-size thesis is real, but current evidence does not yet show that SiliconFlow has converted strategic relevance into high-quality economics. | 中 | SV001, SV014, SV028 |
| CV018 | Regulatory and supply-side risks deserve valuation weight because they can affect both growth and margin simultaneously. | 高 | SV025, SV026, SV029 |
| CV019 | The disclosed price already asks investors to underwrite future improvement rather than current financial quality. | 中 | SV001, SV002, SV004 |
| CV020 | The bull case requires better enterprise retention, richer dedicated/private-deployment mix, and meaningful margin improvement. | 中 | SV001, SV021, SV023 |
| CV021 | The base case assumes SiliconFlow remains strategically relevant and grows, but revenue quality improves only gradually. | 中 | SV001, SV004, SV028 |
| CV022 | The bear case assumes price competition, regulation, or supply constraints prevent margin inflection and weaken financing terms. | 中 | SV004, SV005, SV025, SV026 |
| CV023 | Together AI is a relevant direct comparable because it combines open-model access with inference infrastructure. | 高 | SV006, SV007 |
| CV024 | Together AI's 2026 $8.3B valuation comes with public evidence of >$1.15B annual bookings, a scale far above SiliconFlow's disclosed 2025 revenue base. | 中 | SV006 |
| CV025 | Fireworks AI is a relevant direct comparable because it is an inference-cloud platform rather than a generic hyperscaler. | 高 | SV009, SV010, SV030 |
| CV026 | Fireworks AI's 2026 $17.5B valuation is paired with public evidence of >$1B annualized revenue, making SiliconFlow look less obviously cheap despite its much lower headline price. | 高 | SV009, SV010 |
| CV027 | Baseten is a relevant inference-platform comparable because it combines developer entry with enterprise deployment options. | 中 | SV011, SV012 |
| CV028 | Baseten's 2026 ~$13B valuation and ~US$600M annualized revenue indicate that specialized inference platforms can justify high values—but with much stronger monetization evidence than SiliconFlow currently discloses. | 中 | SV012 |
| CV029 | CoreWeave is an infrastructure-adjacent valuation reference that shows both the upside and the concentration / capital-intensity risks of AI compute businesses. | 中 | SV013, SV026 |
| CV030 | Public comparables therefore support interest in the category, but not an automatic premium view on SiliconFlow at current evidence quality. | 中 | SV006, SV009, SV012, SV013 |
| CV031 | Retention, concentration, and supplier-risk visibility are the most important missing inputs preventing a stronger call. | 中 | SV001, SV004, SV005, SV026 |
| CV032 | A formal regulatory enforcement event would be a thesis-break trigger. | 中 | SV025, SV029 |
| CV033 | A material compute-supply shock or worsening supplier concentration would be a thesis-break trigger. | 中 | SV001, SV026 |
| CV034 | Failure of public-cloud economics to improve would be a thesis-break trigger. | 中 | SV001, SV004, SV005 |
| CV035 | Proof of weak enterprise retention or heavy whale concentration would be a thesis-break trigger. | 中 | SV001, SV004, SV023 |
| CV036 | The most important valuation diligence ask is current 2026 revenue run-rate and mix by product line. | 中 | SV001, SV004 |
| CV037 | The next most important asks are enterprise retention, concentration, and API-to-dedicated expansion metrics. | 中 | SV001, SV023, SV024 |
| CV038 | Security / assurance and supplier contingency are also valuation-critical because they affect enterprise quality and downside risk. | 中 | SV025, SV026, SV027 |
| CV039 | Price alone cannot rescue the investment case if the missing diligence items point to structural weakness rather than temporary opacity. | 中 | SV004, SV005, SV026 |
| CV040 | As of 2026-07-22, the best public-evidence call is Track: strategic relevance is clear, but upside from the current price is not. | 中 | SV001, SV002, SV006, SV009 |
| 编号 | 出版方 | 标题 | 引文 |
|---|---|---|---|
| SO001 | SiliconFlow | 硅基流动 SiliconFlow - 致力于成为全球领先的 AI 能力提供商 | 地址:北京市海淀区中关村东路 1 号院 8 号楼 D 座 23 层 2301。 |
| SO002 | SiliconFlow | SiliconFlow – AI Infrastructure for LLMs & Multimodal Models | One Platform All Your AI Inference Needs. |
| SO003 | SiliconFlow | About Us - SiliconFlow | Global AI Infrastructure Provider | We strive to become the world's most influential provider of AI infrastructure. |
| SO004 | SiliconFlow | SiliconFlow – AI Infrastructure for LLMs & Multimodal Models / Models | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SO005 | SiliconFlow Docs | Quickstart - SiliconFlow | Currently, the platform supports login via SMS, email, as well as OAuth login through GitHub and Google. |
| SO006 | SiliconFlow Docs | Chat - SiliconFlow API Reference | We periodically update our models to enhance service quality. Changes may include model on/offlining or capability adjustments. |
| SO007 | SiliconFlow Docs | Product introduction - SiliconFlow | SiliconFlow is committed to providing developers with faster, more comprehensive, and seamlessly integrated model APIs. |
| SO008 | SILICONFLOW TECHNOLOGY PTE. LTD. | Terms of Use - SiliconFlow | The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn. |
| SO009 | SILICONFLOW TECHNOLOGY PTE. LTD. | Privacy Policy - SiliconFlow | If you want to contact us... please contact us at contact@siliconflow.com. |
| SO010 | GitHub | SiliconFlow · GitHub | SiliconFlow builds scalable, standardized, and high-performance AI infrastructure. |
| SO011 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | As of April 30, 2026, our platform had over 10 million registered users. |
| SO012 | HKEX | Announcement on overall coordinator appointment for Beijing SiliconFlow Technology Co., Ltd. | The Company has further appointed China Renaissance Securities (Hong Kong) Limited as its overall coordinator on July 14, 2026. |
| SO013 | Caixin Global | SiliconFlow Raises $294 Million as China's AI Inference Demand Surges | Chinese AI model inference startup SiliconFlow has raised more than 2 billion yuan ($294 million) in a Series B funding round. |
| SO014 | Sina Finance | 硅基流动完成新一轮超 20 亿元融资 | 过去一年,公司在企业级市场实现爆发式增长。 |
| SO015 | Sina Finance | 硅基流动完成新一轮超 20 亿元融资(转载) | IDC预测,2026年中国市场的Token消费量将达到40,000万亿。 |
| SO016 | 投资界 / InvestorsCN | 硅基流动完成新一轮超20亿元融资,华兴资本担任独家财务顾问 | 近日,硅基流动已完成超20亿元B轮融资。 |
| SO017 | 36Kr Europe | Silicon Flow Completes New Round of Over 2-Billion-Yuan Financing | By providing efficient MaaS through the Token factory model, the daily average Token call volume has reached trillions. |
| SO018 | KrASIA | Surging users, widening losses, and leased compute: Behind SiliconFlow's IPO filing | But the prospectus also shows the cost of that growth. SiliconFlow is selling tokens at a loss. |
| SO019 | Hello China Tech | SiliconFlow IPO: What China's Token Boom Really Costs | Its public cloud service ... generated 52.9% of 2025 revenue ... Its gross margin was negative 119%. |
| SO020 | Tracxn | SiliconFlow - 2026 Company Profile & Team - Tracxn | SiliconFlow is a funded company based in Beijing (China), founded in 2023 by Pan YANG and Jinhui Yuan. |
| SO021 | CB Insights | SiliconFlow Stock Price, Funding, Valuation, Revenue & Financial Statements | SiliconFlow has raised $316.54M over 8 rounds. |
| SO022 | SiliconFlow | Kimi K3 is now live on SiliconFlow | SiliconFlow provides OpenAI- and Anthropic-compatible APIs for Kimi K3. |
| SO023 | SiliconFlow | Kimi K3 on SiliconFlow API | Current support includes image input, tool calling, JSON Mode, streaming, and reasoning output. |
| SO024 | SiliconFlow Cloud | SiliconFlow 统一登录 / Welcome to SiliconFlow | Blazing-fast, cost-effective Generative AI cloud services. |
| SO025 | SiliconFlow | 模型与产品页 - 硅基流动 | 支持 BYOC 部署,全面保护数据隐私与业务安全。 |
| SM001 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | In 2025, we were the fourth largest token supply platform in China in terms of annual token throughput, with a market share of 1.5% and ranked first among all independent ecosystem token supply platforms in China. |
| SM002 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 2025年中国公有云MaaS市场的规模达到30.7亿元人民币。IDC预计2026年全年Token消耗量约为40,000万亿次。 |
| SM003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | IDC predicts that the Token consumption in the Chinese market will reach 40,000 trillion in 2026, an increase of about 20 times compared to 2025. |
| SM004 | Stanford HAI | 2026 AI Index Report | Industry produced over 90% of notable frontier models in 2025... Organizational adoption reached 88%. |
| SM005 | MarketsandMarkets | AI Inference Market Size, Share & Trends | The AI inference market size was valued at USD 106.15 billion in 2025 and is projected to reach USD 254.98 billion by 2030, growing at a CAGR of 19.2%. |
| SM006 | Grand View Research | AI Inference Market Summary | The global AI inference market size was estimated at USD 97.24 billion in 2024 and is projected to reach USD 253.75 billion by 2030, growing at a CAGR of 17.5%. |
| SM007 | Fortune Business Insights | AI Inference Market | The global AI inference market size was valued at USD 103.73 billion in 2025 and is projected to grow from USD 117.80 billion in 2026 to USD 312.64 billion by 2034. |
| SM008 | Amazon Web Services | Amazon Bedrock | Amazon Bedrock powers generative AI for more than 100,000 organizations worldwide—from startups to global enterprises across every industry. |
| SM009 | Amazon Web Services | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models... for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SM010 | AWS Docs | What is Amazon Bedrock? | Bedrock supports 100+ foundation models from industry-leading providers, including Amazon, Anthropic, DeepSeek, Moonshot AI, MiniMax, and OpenAI. |
| SM011 | Microsoft Learn | What is Microsoft Foundry? | Foundry unifies agents, models, and tools under a single management grouping with built-in enterprise-readiness capabilities including tracing, monitoring, evaluations, and customizable enterprise setup configurations. |
| SM012 | Microsoft Azure | Microsoft Foundry Pricing | Foundry Models: Access more than 11,000 foundation, open, reasoning, multimodal and industry-specific models. |
| SM013 | Microsoft Azure | Foundry Models Pricing | Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure... Choose from A100, H100, H200, and MI300 GPU families. |
| SM014 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SM015 | Alibaba Cloud | Token Plan (Team Edition) overview | Token Plan (Team Edition) is a monthly AI model subscription billed in Credits... and includes team management, data privacy, and dedicated throughput. |
| SM016 | Together AI | Serverless Inference | Access all the top open-source models in one place. |
| SM017 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SM018 | Together AI Docs | Available models - Together AI docs | Serverless models are the fastest way to run inference on Together... Pay only for the tokens you process. |
| SM019 | Fireworks AI | Pricing | Additionally, please note that cached input tokens are by default priced at 50% for all text and vision language models... batch inference is priced at 50% of our serverless pricing. |
| SM020 | Fireworks Docs | Models & Inference | Fireworks provides fast, cost-effective access to leading open-source text models through OpenAI-compatible APIs. |
| SM021 | SiliconFlow | Models - SiliconFlow | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SM022 | SiliconFlow Docs | Quickstart - SiliconFlow | Currently, the platform supports login via SMS, email, as well as OAuth login through GitHub and Google. |
| SM023 | SiliconFlow Docs | Chat Completions API Reference | We periodically update our models to enhance service quality. Changes may include model on/offlining or capability adjustments. |
| SM024 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | According to third-party data cited in its prospectus, SiliconFlow was China's largest independent ecosystem token supplier by annual token throughput in 2025 and ranked among the top five token suppliers overall. |
| SM025 | Hello China Tech | SiliconFlow IPO: token economics | As a middle-layer platform, SiliconFlow's main cost pressure comes from leasing computing power... In 2025, computing resource costs totaled RMB 59.627 million, accounting for 86.9% of cost of sales. |
| SP001 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | In 2025, we were the fourth largest token supply platform in China in terms of annual token throughput, with a market share of 1.5% and ranked first among all independent ecosystem token supply platforms in China. |
| SP002 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | IDC认为,MaaS厂商的竞争焦点正在从过去单纯的价格比拼,转向“价格、性能与工具链支持”的综合能力竞争。 |
| SP003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The latest "China AI Software Market Semi-Annual Tracker, 2025H2" released by IDC shows that SiliconFlow is the only startup company among the top four in the market share of China's public cloud MaaS. |
| SP004 | Amazon Web Services | Amazon Bedrock | Amazon Bedrock powers generative AI for more than 100,000 organizations worldwide—from startups to global enterprises across every industry. |
| SP005 | Amazon Web Services | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models ... for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SP006 | Microsoft Learn | What is Microsoft Foundry? | Foundry unifies agents, models, and tools under a single management grouping with built-in enterprise-readiness capabilities including tracing, monitoring, evaluations, and customizable enterprise setup configurations. |
| SP007 | Microsoft Azure | Microsoft Foundry Pricing | Foundry Models: Access more than 11,000 foundation, open, reasoning, multimodal and industry-specific models. |
| SP008 | Microsoft Azure | Foundry Models Pricing | Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure to deploy and run open-source and custom AI models on enterprise-grade GPUs. |
| SP009 | Alibaba Cloud | What is Model Studio? | Alibaba Cloud Model Studio is a one-stop model service platform. It provides the full Qwen series and mainstream third-party LLMs through official Qwen APIs and OpenAI-compatible APIs. |
| SP010 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SP011 | Alibaba Cloud | Token Plan (Team Edition) overview | Token Plan (Team Edition) is a monthly AI model subscription billed in Credits ... and includes team management, data privacy, and dedicated throughput. |
| SP012 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SP013 | Together AI Docs | Serverless overview | Run any supported model through a shared, per-token API with no provisioning and no minimums. |
| SP014 | Together AI Docs | Dedicated model inference overview | Dedicated model inference bills per minute by hardware while a deployment runs, regardless of model or request volume. |
| SP015 | Together AI Docs | Available models - Together AI docs | Serverless models are the fastest way to run inference on Together. You call any supported model through a shared per-token API, with no provisioning, no replicas to size, and no minimum cost. |
| SP016 | Fireworks AI | Fireworks home | Fireworks processes 40T+ tokens per day. |
| SP017 | Fireworks AI | Pricing | Pricing to seamlessly scale from idea to enterprise. |
| SP018 | Fireworks Docs | Serverless overview | Serverless is multi-tenant inference for popular open models running on Fireworks-managed infrastructure. |
| SP019 | Fireworks Docs | Serverless pricing | Per-token serverless pricing for text, vision, and embedding models, including Priority and Fast serving paths. |
| SP020 | Fireworks Docs | Models & Inference | Fireworks provides fast, cost-effective access to leading open-source text models through OpenAI-compatible APIs. |
| SP021 | OpenRouter Docs | Quickstart | OpenRouter provides a unified API that gives you access to hundreds of AI models through a single endpoint, while automatically handling fallbacks and selecting the most cost-effective options. |
| SP022 | OpenRouter Docs | Provider routing | OpenRouter routes requests to the best available providers for your model. By default, requests are load balanced across the top providers to maximize uptime. |
| SP023 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | SiliconFlow was China's largest independent ecosystem token supplier by annual token throughput in 2025 and ranked among the top five token suppliers overall. |
| SP024 | Hello China Tech | SiliconFlow IPO: token economics | The top three token suppliers by throughput, Volcengine, Alibaba Cloud, and Baidu AI Cloud, are all hyperscaler divisions. |
| SP025 | SiliconFlow | Models - SiliconFlow | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SI001 | SiliconFlow / HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | For our public cloud-based services, the gross loss margin was 271.6% in 2024 and 119.0% in 2025. |
| SI002 | SiliconFlow | Models - SiliconFlow | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SI003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | By providing efficient MaaS through the Token factory model, the daily average Token call volume has reached trillions. |
| SI004 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Its public cloud service... generated 52.9% of 2025 revenue. Its gross margin was negative 119%. |
| SI005 | Hello China Tech | SiliconFlow IPO: token economics | In 2025, computing resource costs totaled RMB 59.627 million, accounting for 86.9% of cost of sales. |
| SI006 | Together AI Docs | Dedicated model inference pricing | Dedicated model inference bills based on the hardware your deployments run on, regardless of model or request volume. |
| SI007 | Together AI Docs | Batch processing overview | Run asynchronous batch workloads at up to 50% lower cost. |
| SI008 | Together AI Docs | Send requests on dedicated endpoints | Prompt caching is enabled by default for dedicated model inference. No configuration is required. |
| SI009 | Together AI Docs | Fine-tuning pricing | Fine-tuning is billed per token processed, scaled by model size, training method, and training type. |
| SI010 | Fireworks Docs | Prompt caching | For serverless models, cached prompt tokens are discounted compared to regular prompt tokens. The default discount is 50%. |
| SI011 | Fireworks Docs | Serverless serving paths | Priority tier is for workloads that require higher reliability during peak traffic periods, at a higher price point. |
| SI012 | Fireworks Docs | Serverless pricing | Batch inference is billed at 50% of serverless pricing on both input and output. |
| SI013 | Fireworks AI | Pricing | H100 80 GB GPU $7.00 per hour. |
| SI014 | Amazon Web Services | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SI015 | AWS Docs | Batch inference in Amazon Bedrock | With batch inference, you can submit multiple prompts and generate responses asynchronously. |
| SI016 | AWS Docs | Prompt caching in Amazon Bedrock | Prompt caching can help when you have workloads with long and repeated contexts ... you're charged at a reduced rate for tokens read from cache. |
| SI017 | Microsoft Azure | Microsoft Foundry Pricing | The Microsoft Agent pre-purchase plan allows you to save on Microsoft Foundry and Copilot Credit costs by purchasing Agent Commit Units up front. |
| SI018 | Microsoft Azure | Foundry Models Pricing | Managed Compute in Microsoft Foundry Models gives you dedicated GPU infrastructure... Choose from A100, H100, H200, and MI300 GPU families. |
| SI019 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SI020 | Alibaba Cloud | Token Plan overview | Standard USD 30/seat/month ... Pro USD 100/seat/month ... Max USD 200/seat/month. |
| SI021 | Alibaba Cloud | What is Model Studio? | Activating Model Studio is free. Costs apply only when you invoke models. |
| SI022 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素依次是:模型性能、安全合规要求、回答质量、在AI平台可用性以及成本效益。 |
| SI023 | Fireworks Docs | Models & Inference | Every response includes token usage information and performance metrics for debugging and observability. |
| SI024 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SI025 | AWS | Amazon Bedrock | Features like Model Distillation, Prompt caching, and Intelligent Prompt Routing can reduce expenses while maintaining performance. |
| SE001 | SiliconFlow | SiliconFlow China homepage | 覆盖语言、语音、图片、视频等场景,一站式提供大模型 API 服务,按量计费,助力应用快速上线。 |
| SE002 | SiliconFlow Docs | Product introduction | As a one-stop cloud service platform integrating top-tier large language models, SiliconFlow is committed to providing developers with faster, more comprehensive, and seamlessly integrated model APIs. |
| SE003 | SiliconFlow Docs | API reference home | Corresponding Model Name. To better enhance service quality, we will make periodic changes to the models provided by this service. |
| SE004 | SiliconFlow Docs | Chat API reference | The response header contains the x-siliconcloud-trace-id field, which serves as a unique identifier for tracing requests. |
| SE005 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SE006 | SiliconFlow Docs | Privacy Policy | When you register, log in, and use the services provided through the https://siliconflow.com platform, we will collect and store the relevant information in accordance with this policy. |
| SE007 | SiliconFlow Docs | Terms of Use | The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn. |
| SE008 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | We offer API services for developers and enterprises ... and on-premise deployment solutions. |
| SE009 | GitHub | SiliconFlow organization | High-performance inference for every SOTA model. |
| SE010 | GitHub | OneDiff repository | onediff is an out-of-the-box acceleration library for diffusion models |
| SE011 | GitHub | OneDiff releases | 1.2.0 ... support diffusers sd3 speedup ... add diffusers nexfort example |
| SE012 | SiliconFlow Blog | Kimi K3 API launch post | Base URL: https://api.siliconflow.com/v1 |
| SE013 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The platform brings together over 400 models spanning text, image, audio, and video generation. |
| SE014 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Its public cloud service includes both a pay-as-you-go serverless model and dedicated computing instances. |
| SE015 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素 ... 在AI平台可用性以及成本效益。 |
| SE016 | Together AI Docs | Serverless overview | Use serverless inference to run models without managing infrastructure. |
| SE017 | Fireworks AI | Home | The Generative AI Platform for production-ready workloads. |
| SE018 | OpenRouter Docs | Quickstart | OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers. |
| SE019 | Alibaba Cloud | What is Model Studio? | Model Studio provides model APIs and application development capabilities. |
| SE020 | AWS | Amazon Bedrock | Features like Model Distillation, Prompt caching, and Intelligent Prompt Routing can reduce expenses while maintaining performance. |
| SE021 | SiliconFlow | SiliconFlow international site | Get your Model API fast |
| SE022 | SiliconFlow Cloud | Cloud model catalog | Model catalog and pricing surface for cloud users. |
| SE023 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SE024 | China Money Network | SiliconFlow's role in China's AI ecosystem | The platform supports over 400 models and serves thousands of enterprise customers. |
| SE025 | GitHub | OneFlow releases asset page reference in OneDiff README | please install OneFlow by the links below. |
| SU001 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | Our public cloud-based services primarily target developers and small- and medium-sized enterprises. |
| SU002 | QQ News / Taimei | 问AI · Token工厂模式为何吸引全产业链巨头联手投资? | 过去一年,硅基流动日均Token调用量达数万亿,服务超1000万用户和1万家企业客户。 |
| SU003 | SiliconFlow | News listing | 贵州移动与硅基流动深化战略合作,加速“Token 工厂”建设 |
| SU004 | SiliconFlow Docs | Scenarios and application cases | Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases. |
| SU005 | SiliconFlow | Airline group AI infrastructure case | 该企业引入硅基流动(SiliconFlow)推理加速框架,实现三层技术突破。 |
| SU006 | SiliconFlow | Energy SOE AI infrastructure case | 某能源央企正处于转型关键阶段。 |
| SU007 | SiliconFlow | Guizhou Mobile strategic cooperation | 贵州移动与硅基流动 ... 构建“推理框架 + 算力供给 + 模型服务 + 场景应用”全栈协同体系。 |
| SU008 | SiliconFlow | Reserved instances support 100B-token daily workload | 2026 年以来,其 Coding Agent 单日 Token 消耗冲上千亿量级。 |
| SU009 | SiliconFlow Docs | Use SiliconCloud in MindSearch | After adding this configuration, you can execute the relevant commands to start MindSearch. |
| SU010 | SiliconFlow Docs | Use Continue with SiliconFlow APIs | By integrating SiliconFlow APIs into Continue, you can get access to 200+ open-source models. |
| SU011 | SiliconFlow Docs | Use Cline with SiliconFlow APIs | we’ll show you how to integrate SiliconFlow’s APIs into Cline |
| SU012 | SiliconFlow | About us | From open source to enterprise deployment, we accelerate what matters. |
| SU013 | InforCapital | SiliconFlow company profile | SiliconFlow - AI Infrastructure |
| SU014 | CB Insights | SiliconFlow company profile | SiliconFlow - Products, Competitors, Financials, Employees |
| SU015 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing |
| SU016 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The platform brings together over 400 models spanning text, image, audio, and video generation. |
| SU017 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SU018 | SiliconFlow | China homepage | 面向不同行业及需求场景,提供灵活的解决方案 |
| SU019 | SiliconFlow Docs | Product introduction | Our platform empowers developers and enterprises to focus on product innovation |
| SU020 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models |
| SU021 | China Money Network | SiliconFlow's role in China's AI ecosystem | SiliconFlow's role in China's AI ecosystem |
| SU022 | SiliconFlow | International site | Get your Model API fast |
| SU023 | Tencent News | Funding and ecosystem article | 客户名单涵盖能源、金融、交通等核心行业的头部央企 |
| SU024 | GitHub | SiliconFlow organization | High-performance inference for every SOTA model. |
| SU025 | SiliconFlow | Reserved instances page | 预留实例 |
| SR001 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | Our major purchases are concentrated with the top five suppliers. |
| SR002 | SiliconFlow Docs CN | 服务协议 | 服务适用于面向开发者的内容生成服务场合,不适用于自动控制、医疗信息服务、心理咨询和关键信息基础设施的场合。 |
| SR003 | SiliconFlow Docs CN | 隐私政策 | 根据中华人民共和国大陆地区相关法律法规的规定,我们需要对您进行实名认证。 |
| SR004 | SiliconFlow CN | China homepage | 京ICP备2024051511号-1 增值电信业务经营许可证:京B2-20242084 |
| SR005 | SiliconFlow Docs EN | Terms of Use | The Services are not available for users located in the territory of mainland China and users in China may access our services through https://siliconflow.cn. |
| SR006 | SiliconFlow Docs EN | Privacy Policy | When you register, log in, and use the services provided through the https://siliconflow.com platform, we will collect and store the relevant information. |
| SR007 | Regulations.ai | Measures for the Identification of AI-Generated (Synthetic) Content | The Measures set a mandatory national baseline requiring that AI-generated or AI-synthesized content ... be clearly identified to users via explicit and implicit identification. |
| SR008 | U.S. BIS | Commerce strengthens restrictions on advanced computing semiconductors | These rules reinforce and build on the October 7, 2022, October 17, 2023, and December 2, 2024, controls to restrict the PRC's ability to obtain certain high-end chips. |
| SR009 | U.S. Department of Commerce | Guidance on Advanced Computing Items (May 2026) | a license is required to export advanced computing items to entities headquartered in Country Group D:5 or Macau |
| SR010 | NIST | AI Risk Management Framework | On April 7, 2026, NIST released a concept note for an AI RMF Profile on Trustworthy AI in Critical Infrastructure. |
| SR011 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing |
| SR012 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SR013 | QQ News / Taimei | 问AI · Token工厂模式为何吸引全产业链巨头联手投资? | 当前市场在快速膨胀,但竞争也在加剧。 |
| SR014 | SiliconFlow | Guizhou Mobile strategic cooperation | 双方在 2025 年战略合作基础上再升级 |
| SR015 | SiliconFlow | Reserved instances support 100B-token daily workload | 配合企业级交付与运行保障、明确的 SLA 与完善的售后服务 |
| SR016 | SiliconFlow | Energy SOE AI infrastructure case | 实现完整的性能追踪与故障预警,保障复杂生产环境下的高可用与安全性。 |
| SR017 | SiliconFlow | Airline group AI infrastructure case | 模型更新周期从“数周”压缩至“3 天内” |
| SR018 | SiliconFlow Docs | Product introduction | Provides comprehensive monitoring and fault tolerance mechanisms to guarantee service capabilities. |
| SR019 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SR020 | SiliconFlow Docs | API reference home | we will make periodic changes to the models provided by this service |
| SR021 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素依次是:模型性能、安全合规要求、回答质量、在AI平台可用性以及成本效益。 |
| SR022 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The daily average token call volume has reached trillions. |
| SR023 | China Money Network | SiliconFlow's role in China's AI ecosystem | SiliconFlow's role in China's AI ecosystem |
| SR024 | SiliconFlow Docs | Scenarios and application cases | Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases. |
| SR025 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SR026 | Fireworks AI | Pricing | H100 80 GB GPU $7.00 per hour. |
| SR027 | AWS | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SR028 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SR029 | OpenRouter Docs | Quickstart | OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers. |
| SR030 | Together AI Docs | Dedicated model inference pricing | Dedicated model inference bills based on the hardware your deployments run on, regardless of model or request volume. |
| SR031 | Regulations.ai | Provisions on the Administration of Deep Synthesis of Internet Information Services | Providers must implement real-identity verification before allowing publishing privileges, and all synthetic content must carry clear technical marks indicating its origin. |
| SR032 | China Law Translate | Measures for Labeling of AI-Generated Synthetic Content | Service providers shall clearly explain the methods, styles, and other such specifications for labeling generated synthetic content in user service agreements, and notify users to carefully read and understand the corresponding labeling management requirements. |
| SV001 | HKEX | Application Proof of Beijing SiliconFlow Technology Co., Ltd. | For our public cloud-based services, the gross loss margin was 119.0% in 2025. |
| SV002 | QQ News / Taimei | 问AI · Token工厂模式为何吸引全产业链巨头联手投资? | 6月16日 ... 硅基流动宣布完成超20亿元B轮融资。 |
| SV003 | 36Kr Europe | SiliconFlow completes funding round exceeding RMB 2 billion | The platform brings together over 400 models spanning text, image, audio, and video generation. |
| SV004 | KrASIA | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing | Surging users, widening losses, and leased compute behind SiliconFlow's IPO filing |
| SV005 | Hello China Tech | SiliconFlow IPO: token economics | It is, in essence, renting the hardware, bundling the software, and wrapping the models. |
| SV006 | Reuters / U.S. News | Together AI raises $800 million at $8.3 billion valuation | Together AI ... raised $800 million ... at $8.3 billion valuation ... annual bookings crossed $1.15 billion last quarter. |
| SV007 | BusinessWire | Together AI raises $800 million at $8.3 billion valuation | Together AI Raises $800 Million at $8.3 Billion Valuation |
| SV008 | BusinessWire | Fireworks AI raises $250M Series C | Fireworks AI Raises $250M Series C to Lead the AI Inference Market |
| SV009 | CNBC | Fireworks hits $17.5 billion valuation and $1B in annualized revenue | The company said ... it has exceeded $1 billion in annualized revenue, and it has now raised a $1.5 billion round at a $17.5 billion valuation. |
| SV010 | Sacra | Fireworks AI revenue, valuation & funding | At more than $1B in annualized revenue in July 2026, the latest valuation implies an approximate 17.5× revenue multiple. |
| SV011 | BusinessWire | Baseten raises $300M at a $5B valuation | Baseten Raises $300M at a $5B Valuation to Power a Multi-Model Future |
| SV012 | Sacra | Baseten revenue, valuation & funding | Baseten hit $600M in annualized revenue in March 2026 ... valued at $13B following its $1.5B Series F in June 2026. |
| SV013 | Sacra | CoreWeave revenue, valuation & funding | CoreWeave guided full-year 2026 revenue of $12B–$13B. |
| SV014 | SiliconFlow | Models catalog | One API to run inference on 200+ cutting-edge AI models, and deploy in seconds. |
| SV015 | SiliconFlow Docs | Product introduction | Our platform empowers developers and enterprises to focus on product innovation while eliminating concerns about exorbitant computational costs. |
| SV016 | Together AI | Pricing | Most teams start with serverless inference and move to dedicated endpoints at scale. |
| SV017 | Fireworks AI | Pricing | H100 80 GB GPU $7.00 per hour. |
| SV018 | AWS | Amazon Bedrock Pricing | Amazon Bedrock offers select foundation models for batch inference at a 50% lower price compared to on-demand inference pricing. |
| SV019 | Alibaba Cloud | Model Studio model pricing | Model API calls are billed on a pay-as-you-go basis by default. |
| SV020 | OpenRouter Docs | Quickstart | OpenRouter provides an OpenAI-compatible completion API to more than 400 models & providers. |
| SV021 | SiliconFlow | China homepage | 大模型云服务 ... 预留实例 ... 私有化大模型服务平台 |
| SV022 | SiliconFlow | Guizhou Mobile strategic cooperation | 双方在 2025 年战略合作基础上再升级 |
| SV023 | SiliconFlow | Reserved instances support 100B-token daily workload | 其 Coding Agent 单日 Token 消耗冲上千亿量级。 |
| SV024 | SiliconFlow Docs | Scenarios and application cases | Easily integrate SiliconFlow platform large model capabilities into various scenarios and application cases. |
| SV025 | Regulations.ai | Measures for the Identification of AI-Generated (Synthetic) Content | The Measures set a mandatory national baseline requiring that AI-generated or AI-synthesized content ... be clearly identified. |
| SV026 | U.S. BIS | Commerce strengthens restrictions on advanced computing semiconductors | These rules ... restrict the PRC's ability to obtain certain high-end chips critical for military advantage. |
| SV027 | NIST | AI Risk Management Framework | On April 7, 2026, NIST released a concept note for an AI RMF Profile on Trustworthy AI in Critical Infrastructure. |
| SV028 | IDC China | 中国MaaS市场进入高速增长期,Token经济从概念走向规模 | 影响大模型落地的Top5因素 ... 在AI平台可用性以及成本效益。 |
| SV029 | SiliconFlow Docs CN | 服务协议 | 服务适用于面向开发者的内容生成服务场合,不适用于 ... 关键信息基础设施的场合。 |
| SV030 | CNBC | Fireworks valuation and industry context | By managing computing infrastructure for models, Fireworks does business in the inference cloud market, alongside startups such as Baseten and Together AI. |