Dataiku
规模化的受治理企业 AI,具备 IPO 选择权,但经济性仍不透明
Dataiku 像是真正进入后期阶段的企业 AI 龙头,当前估值也有合理支撑;但利润率、留存和现金披露不足,结论更适合观察,而不是买入。
封面要素
公司概况
Dataiku 是一家 2013 年创立于法国、总部位于 New York 的企业 AI 平台公司。公司把自己定位为一层受治理的编排平台,让分析师、工程师、数据科学家和业务用户能在云端、本地与混合环境中协作构建分析、机器学习模型和 AI 代理。公开证据显示,到 2025 年 10 月,公司 ARR 已超过 $350M,客户超过 750 家,员工超过 1,250 人,同时在准备可能的美国 IPO。
- 成立时间
- 2013-01-01
- 创始人
- Florian Douetteau, Clément Stenac, Thomas Cabrol, Marc Batty
- 创立地点
- Paris, France
- 总部
- New York, New York, USA
- 产品
- 一个统一平台,覆盖分析、机器学习、AI 治理和代理式 AI 编排,把无代码、低代码和全代码工作流与云和模型选择权结合起来。
- 客户
- 医疗健康、制造、金融服务、物流,以及其他受监管或运营复杂的大型企业。
- 商业模式
- 企业软件订阅模式,围绕平台广度、治理和部署灵活性销售,并借助合作伙伴主导的转型项目和大客户内部扩张来支撑增长。
- 阶段
- Series F (private)
- 融资情况
- 2022 年 12 月以 $3.7B 估值完成 $200M Series F;公开渠道尚未确认更新的一级融资,但 Reuters 2025 年 10 月报道称,Dataiku 已聘请 Morgan Stanley 和 Citigroup 准备潜在美国 IPO。
执行摘要
主要优势
- $350M+ ARR、750+ 客户和 1,250+ 员工,支撑它具备真实后期企业级规模。
- 治理优先的定位,切中企业对可信、可审计 AI 和智能体工作流的需求。
- 作为私有基础设施厂商,Dataiku 的客户验证异常扎实,Michelin、Novartis、Roche、Prologis、Standard Chartered 和 SLB 都有具名结果。
- 虽然 $3.7B 估值锚点已经不新,但放到上市 AI / 数据软件可比公司倍数里看,大致处于中段,而不是泡沫高点。
主要风险
- 毛利率、NRR、流失、现金消耗和客户集中度仍未披露,压住了估值信心。
- Databricks、超大规模云厂商和相邻上市数据 / AI 平台的竞争,可能压缩扩张假设和估值倍数。
- EU AI Act 和其他企业合规要求抬高门槛,安全、隐私和 AI 治理义务还在变宽。
- IPO 时点可能放大任何实施、监管或客户留存失误的影响。
未决问题
- 除 2025 年 10 月 $350M+ ARR 里程碑外,没有当前经审计财务披露,也没有 2026 ARR 衔接数据。
- 没有公开披露毛利率、NRR、GRR、流失或客户集中度。
- 合作伙伴带来的销售管线、服务收入占比,以及真实实施强度都缺少公开清晰度。
- 当前现金、烧钱速度,或未来融资可能带来的清算优先权压力,都没有公开视图。
目录
01公司概览
1.1 身份、总部,以及 Dataiku 到底卖什么
Dataiku 是一家后期私营企业软件公司,2013 年创立,如今总部位于 New York;不过它的根基和创始团队明显来自法国。公司仍用推动早期增长的同一个核心想法描述自己:把数据和 AI 变成日常运营能力,而不是专家项目。到 2026 年,公司话术已从较早的「Everyday AI」转向「The Platform for AI Success」,但运营主张没有变——用一个控制平面,让业务用户、分析师、数据科学家、工程师和风险负责人共同构建分析、机器学习和 AI 代理工作流。这个定位很重要,因为 Dataiku 卖的不是单一模型工具,也不是单点 AutoML 产品;它卖的是客户数据栈之上的编排、治理和多人协作。归档的方案页面和当前合作伙伴材料显示,Dataiku 有三条稳定部署路径:托管 SaaS、客户自管云,以及本地 / 私有云安装。这种部署灵活性,加上无代码、低代码和全代码工作流,是公司价值主张的核心,也解释了为什么买方往往是基础设施异构的大型企业,而不是想要一键式 AI 的 SMB。[CO001, CO002, CO003, CO004, CO005, CO008]
| 指标 | 数值 / 状态 | 日期 | 置信度 | 备注 / 缺口 |
|---|---|---|---|---|
| 成立 | 2013 | 2013 | 高 | 官方叙事与 Gartner 档案年份一致 |
| 总部 | New York, NY | 2026 | 高 | 法国创立,但当前总部在 New York |
| 私营 / 上市状态 | 私营;有 IPO 准备报道,未见申报文件 | 2026 | 中 | Reuters 报道的是选定投行,不是已完成 IPO |
| 最近披露估值 | $3.7B | Dec 2022 | 高 | Series F 估值;未公开确认更新的一级融资轮 |
| 融资总额 | ~$846.8M | Nov 2025 研究 | 中 | Sacra 按已披露轮次估算 |
| ARR | Jan 2025 为 $300M+;Oct 2025 为 $350M+ | 2025 | 中 | 公司披露的经营指标 |
| 客户 | Jan 2025 为 700+;Oct 2025 为 750+ | 2025 | 中 | 公司披露的经营指标 |
| 员工 / 覆盖 | Jan 2025 为 1,100+;Oct 2025 达 1,250+,并有 13 个办公室 | 2025 | 中 | 官方新闻稿,非经审计员工数据 |
经营指标来自公司新闻稿;估值和总融资额来自后期轮次报道与二级研究,而非经审计文件。
[CO001, CO004, CO005, CO014, CO016, CO021]Dataiku 把企业数据资产、跨职能建设者和治理要求接到同一个 AI 控制平面。
[CO008, CO009, CO010, CO044, CO045, CO046]私营公司能公开到这一层 KPI,信号很强,但仍依赖公司披露和选择性二级分析。
[CO014, CO016, CO026, CO027, CO028, CO029]1.2 创始人、领导梯队与治理信号
公开领导层图景中,创始核心和更新后的商业化梯队最清晰。Florian Douetteau 仍是明确关键人物,担任联合创始人兼 CEO;Clément Stenac 仍是技术锚点,担任联合创始人兼 CTO。公开创始人资料也持续把 Thomas Cabrol 和 Marc Batty 列在其中,但他们当前运营角色的可见度远低于 Douetteau 和 Stenac。这对成熟私营公司并不少见,却仍让治理细节达不到公开市场投资人想要的厚度。2026 年最相关的管理层事件,是 6 月聘任 Maxwell Long 出任 President 兼 CRO。这一动作很像典型后期扩张:Long 曾帮助 Smartsheet 跨过 $1B ARR 并走向 $8.4B 退出,Dataiku 聘他专门管理全球销售、客户成功和合作伙伴团队。叠加 2025 年引入一位高知名度 CMO,以及公司持续强调合作伙伴,领导层模式符合一家企业在潜在流动性事件前把商业引擎专业化的路径。仍披露不足的是董事会构成、独立性,以及后期投资人与经营团队之间的精确权力分工。[CO002, CO006, CO007, CO031, CO032, CO044]
| 人物 | 当前 / 已知角色 | 重要性 | 公开可见缺口 |
|---|---|---|---|
| Florian Douetteau | 联合创始人兼 CEO | 创始愿景提出者,也是 Dataiku 持续面向公众的代表;战略和 IPO 叙事的关键人物 | 未公开持股或董事会控制细节 |
| Clément Stenac(联合创始人) | 联合创始人兼 CTO | 平台架构和长期产品完整性的技术守门人 | CTO 以下接班梯队披露有限 |
| Thomas Cabrol | 联合创始人 | 官方和投资者档案始终将其列为创始人 | 当前运营职责披露不清 |
| Marc Batty | 联合创始人 | 官方和投资者档案始终将其列为创始人 | 当前运营职责披露不清 |
| Maxwell Long | 总裁兼 CRO(2026 年 6 月加入) | 负责销售、客户成功和合作伙伴;后期扩张操盘手 | 上任时间很短,尚无 Dataiku 任内执行记录 |
| Mark Abramowitz | 首席营销官(2025 年宣布) | 显示公司在下一阶段前强化企业品牌和需求生成 | 公开披露尚未量化需求影响 |
创始人身份有充分证据,但当前董事会构成、独立性和股权控制仍明显披露不足。
[CO002, CO006, CO007, CO031, CO032, CO050]| 利益相关方 | 角色 | 公开意义 | 当前尽调问题 |
|---|---|---|---|
| Wellington Management | 2022 年 Series F 领投方 | 领投设定当前披露 $3.7B 估值的那轮融资 | 确认 Wellington 是否仍锚定 IPO 准备预期 |
| CapitalG | 成长轮投资者 | CapitalG 入局标志着 2019 年独角兽地位和 Alphabet 关联 | 厘清当前持股和董事会影响力 |
| Tiger Global | 后期投资者 | 出现在 2020/2021 年融资历史中,是高可见度成长轮支持者 | 了解内部估值标记变化或流动性压力 |
| FirstMark | 早期投资者 | 仍公开强调创始故事和品类论点 | 确认后续轮次带来的持股稀释 |
| Snowflake | 战略生态合作伙伴 | 300+ 联合客户,并在市场渠道 / GTM 上重要 | 衡量销售管线中合作伙伴来源与直销来源的占比 |
| KPMG | 咨询联盟伙伴 | 代表企业转型交易中的系统集成商渠道可信度 | 验证该联盟能否产出重复性企业实施项目 |
表内同时纳入财务利益相关方和 GTM 利益相关方,因为 Dataiku 的公开记录在生态影响上更丰富,在正式董事会权利或二级持仓上更稀薄。
[CO015, CO018, CO019, CO034, CO035, CO036]1.3 融资历史、估值,以及已有的公开规模指标
Dataiku 的融资史显示,公司在 2018-2022 年软件牛市中激进扩张,此后选择不再宣布新的一级融资。第三方轮次历史显示,公司估值逐级抬升:2018 年完成 $101M Series C,2019 年在 CapitalG 支持下进入独角兽行列,2020 年完成 $100M Series D,2021 年以 $4.6B 估值完成 $400M Series E,并在 2022 年 12 月以 $3.7B 估值完成 $200M Series F。Sacra 2025 年研究把累计融资额估在约 $846.8M;虽然私营公司的股权结构数据仍不完整,这个数与已披露时间线方向一致。作为一家私营 AI 平台供应商,Dataiku 的公开运营规模异常可见。公司官方称,2025 年 1 月 ARR 达 $300M,2025 年 10 月超过 $350M;同期客户数从 700+ 增至 750+,员工数从 1,100+ 增至 1,250+。公司还称拥有 13 个办公室,并渗透了 Forbes Global 2000 中的四分之一。这些都是强劲的后期软件指标,但仍是新闻稿指标,不是审计文件。因此,本章把它们视为有二级研究相互印证的公司说法,而非达到上市公司质量的披露。[CO013, CO014, CO015, CO016, CO017, CO018]
| 日期 | 事件 | 类型 | 金额 / 状态 | 参与方 | 影响 |
|---|---|---|---|---|---|
| 2013 | 公司成立 | 创立 | 成立年份已验证 | Douetteau、Stenac、Cabrol、Batty | 后续历史的标准起点 |
| 2015 | 建立美国布局 | 规模化 | 美国扩张见报道 | Dataiku | 释放最终转向以 New York 为中心的 GTM 信号 |
| 2018-12 | 宣布 Series C | 融资 | $101M | 二级资料显示由 ICONIQ 领投 | 标志着融资迈入后期成长阶段 |
| 2019-12 | CapitalG 投资 / 独角兽地位 | 融资 | 估值 $1.4B | CapitalG 与 Dataiku | 确认品类突破和 Alphabet 关联 |
| 2020-08 | 宣布 Series D | 融资 | $100M | Stripes、Tiger Global | 在企业 AI 快速建设期补充资本 |
| 2021-08 | 宣布 Series E | 融资 | $400M 融资,估值 $4.6B | Tiger Global 及其他投资者 | 披露估值达到牛市峰值 |
| 2022-12 | 宣布 Series F | 融资 | $200M 融资,估值 $3.7B | Wellington 领投轮 | 估值下调,但融资通道仍打开 |
| 2025-01 | 披露 ARR 里程碑 | 规模化 | $300M+ ARR;700+ 客户;1,100+ 员工 | Dataiku | 在私募不透明下仍显示规模延续 |
| 2025-10 | 披露 ARR 里程碑和 IPO 准备 | 规模化 / 治理 | $350M+ ARR;Reuters IPO 准备报道 | Dataiku;Morgan Stanley;Citigroup | 标志着进入潜在上市观察名单阶段 |
| 2026-06 | Maxwell Long 出任总裁兼 CRO | 治理 | 商业领导层任命 | Dataiku | 为下一增长阶段补强高管梯队 |
2022 年前的融资事件依赖公开二级资料;2024 年后的运营里程碑来自公司新闻稿和 Reuters 来源的 IPO 报道。
[CO001, CO013, CO014, CO017, CO018, CO019]从 2013 年创立到 2026 年领导层强化,公开记录显示 Dataiku 走出了一条典型的后期企业软件扩张弧线。
[CO001, CO013, CO014, CO018, CO019, CO020]1.4 IPO 路径、生态强度与主要反向解读
Dataiku 进入 IPO 观察期的最强信号,是 Reuters 2025 年 10 月报道称 Morgan Stanley 和 Citigroup 已受聘,为最早可能在 2026 年上半年进行的美国上市做准备。截至运行日期,公司仍没有公开申报文件,所以公平表述应是「正在准备,但尚未上市」。这个细节重要,因为公司的商业化和生态故事显然在增强:Snowflake 材料指向 300+ 共同客户和重要联合销售动能,KPMG 在 2024 年公开与 Dataiku 结盟,Dataiku 自己的合作伙伴目录也显示,它在超大规模云厂商、数据平台和集成商中覆盖很深。但并非所有信号都干净利好。Gartner 评论页包含对私有云集成的明确批评,更广泛的市场分析也警告,企业代理式 AI 采用仍受信任、集成和治理缺口限制。换句话说,Dataiku 很适合承接下一波企业 AI 预算,但这波预算释放有多快、规模是否足以支撑其 IPO 溢价高于 2022 年估值标记,仍取决于一个比供应商叙事更难落地的采用环境。[CO034, CO035, CO036, CO037, CO038, CO039]
1.5 图表
02市场分析
2.1 市场边界——Dataiku 实际参与的位置
定义 Dataiku 市场,最清楚的方式不是把所有 AI 软件都算进去,甚至也不是把所有机器学习都算进去。Dataiku 位于受治理的企业 AI 编排层:这类软件让组织准备数据、开发分析和模型、把 AI 工作流投入运营,并越来越多地跨团队和基础设施管理基于代理的系统。因此,相关市场包括传统数据科学与机器学习平台、MLOps 工具、AI 治理软件,以及新兴企业代理平台栈的一部分。它不包括原始云基础设施、对独立基础模型供应商的商品化 API 调用,也不包括不需要工作流编排、多人治理或部署管理的轻量单点 copilot 工具。这个区分很重要,因为最宽口径的市场研究会给出巨大的 TAM,但其中包含 Dataiku 不直接变现的区域。现实替代集合也很混杂:超大规模云厂商捆绑原生工具,Databricks 销售湖仓加 AI 控制平面的替代方案,DataRobot、H2O.ai、Alteryx 等供应商覆盖相邻的自动化或低代码分析场景。当客户需要一个受治理的运营层,横跨这些碎片化组件,而不是再多一个孤立工具时,Dataiku 才会赢。[CM001, CM015, CM016, CM017, CM018, CM019]
| 细分 / 品类 | 纳入支出 | 排除支出 | 主要买方 / 付费方 | Dataiku 适配度 |
|---|---|---|---|---|
| 广义 ML 软件 | 模型开发、数据准备、部署、分析工具 | 原始云算力和通用 API 消耗 | CIO / CDO / 分析预算负责人 | 可作为上限背景,但单独看过宽 |
| DSML 平台 | 协作分析、notebook、可视化管道、ML 生命周期 | 没有工作流层的单点模型托管工具 | 数据科学负责人 / 平台负责人 | Dataiku 的核心传统品类 |
| MLOps | 模型部署、监控、血缘、注册表、可复现性 | 没有生产控制的纯实验工具 | ML 工程 / 平台团队 | 与 Dataiku 有重要重叠,但比完整产品更窄 |
| AI 治理 | 风险、合规、血缘、审批、监控、政策控制 | 与 AI 工作流无关的通用网络安全或 GRC 支出 | 风险、合规、AI 治理办公室 | 与 Dataiku 差异化对齐的快速增长细分 |
| 企业智能体平台 | 智能体构建、编排、工具使用、受治理执行 | 消费级副驾和通用聊天订阅 | 创新办公室、平台工程、AI CoE | Dataiku 最新的扩张区域 |
| 邻近低代码分析 | 工作流分析、准备、仪表盘自动化 | 重代码 ML 工程栈 | BI / 运营负责人 | Alteryx 等工具在更简单用例上竞争的邻近区 |
关键纪律是把 Dataiku 视为受治理的编排和协作层,而不是把它当作所有 AI 基础设施支出或企业内每个 LLM token 的代理。
[CM001, CM015, CM016, CM017, CM018, CM019]分层视角显示,Dataiku 的估值应对标更窄的编排和治理类别,而不只是宽泛的 ML TAM。
[CM001, CM035, CM038]2.2 规模测算口径——市场很大,但有多个合理分母
面向 Dataiku 机会的公开市场规模测算跨度异常大,因为分析师测的是相互重叠但并不相同的类别。最宽口径下,Fortune Business Insights 预计 2026 年全球机器学习软件支出为 $65.28B,Precedence Research 则给出同年 $126.91B。严格说,这两个数字都不算错;它们只是口径很宽。更窄的类别镜头对承销 Dataiku 更有用。MarketsandMarkets 预计 MLOps 到 2027 年为 $5.9B、AI 治理到 2029 年为 $5.78B、AI Studio 到 2029 年为 $32.7B。这些窄口径更贴近 Dataiku 变现的受治理开发、部署和监控界面。正确的分析结论不是挑一个 TAM 并为它辩护,而是保留区间,并明确说明 Dataiku 为什么可能捕获三层中的部分份额。估值工作中,本章把宽口径 ML 市场研究当作天花板背景,把 MLOps 和 AI 治理研究当作更接近公司当前 SAM 的代理,把 AI Studio 框架视为桥接类别,用来解释代理和治理融合后市场为什么会变宽。[CM002, CM003, CM004, CM005, CM006, CM007]
| 视角 | 发布方 | 基准年 / 预测年 | 数值 | 增长 | 意义 | 局限 |
|---|---|---|---|---|---|---|
| 广义 ML 市场 | Fortune Business Insights | 2026 | USD 65.28B | 26.7% CAGR 至 2034 年 | 企业 AI 软件需求的上限背景 | 过宽;包含许多 Dataiku 并不直接变现的工作负载 |
| 广义 ML 市场 | Precedence Research | 2026 | USD 126.91B | 33.66% CAGR 至 2035 年 | 显示宽口径定义如何让 TAM 假设翻倍 | 方法论与 Fortune 差异很大,不能直接比较 |
| MLOps 市场 | MarketsandMarkets | 2027 | USD 5.9B | 41.0% CAGR | 更接近 Dataiku 参与竞争的部署 / 监控层代理 | 预测年份与其他视角不同 |
| AI 治理市场 | MarketsandMarkets | 2029 | USD 5.78B | 45.3% CAGR | 贴合 Dataiku 以治理为核心的企业销售叙事 | 仍窄于完整协作 / 编排平台 |
| AI Studio 市场 | MarketsandMarkets | 2029 | USD 32.7B | 38.4% CAGR | 连接 DSML、MLOps 和新兴智能体工具的桥接品类 | 厂商格局宽泛且异质 |
| 企业 AI 编排层 | 作者综合 | 2026 | 未直接拆分 | 未直接拆分 | 概念上最接近 Dataiku 的真实 SAM | 仅靠公开数据无法可信量化 |
不同发布方测算的是不同品类切面;应保留区间,不要强行归并为一个 TAM。
[CM003, CM004, CM005, CM006, CM007, CM038]预测增长和规模估算会因所选类别定义不同而大幅变化。
[CM003, CM004, CM005, CM006, CM007]2.3 买方、用户与企业采用路径
Dataiku 的买方地图在结构上明显偏大型企业。广泛研究显示,大型企业主导 ML 支出;Dataiku 自身客户表面也印证了这一点:制药、银行、物流、制造、保险和交易所运营商在公开资料中反复出现。典型经济买方是 CIO、CDO、分析或数据平台负责人,或掌管受治理数据预算的转型负责人。用户群比买方群更宽。Snowflake、AWS 和 Google Cloud 的合作材料都强调业务与领域专家,而不只是数据科学家;这意味着该平台类别由中心化预算购买,却通过广泛内部使用变现。Dataiku 的归档打包方式透露了标准的先落地再扩张路径:免费或小团队入口可行,但当客户需要自动化、部署、审批工作流和更广治理时,真正产品价值才会释放。系统集成商很重要,因为买方往往需要改变运营模式,而不只是安装软件。KPMG 联盟与 Snowflake 的 300+ 共同客户说法都表明,生态能把产品带入大客户转型项目;这些项目采购周期长,且跨职能。[CM007, CM008, CM009, CM010, CM011, CM012]
| 细分 / 垂直 | 买方 | 主要用户 | 付费方 / 预算负责人 | 采用触发 | Dataiku 适配原因 |
|---|---|---|---|---|---|
| 生命科学 / 制药 | 数字化 / 分析负责人 | 科学家、分析师、工程师 | 转型预算 | 需要受治理的 GenAI、分析和可重复工作流 | Novartis 和 Roche 式用例提供公开证明 |
| 金融服务 / 银行 | CIO / 运营发起人 | 风险与运营团队 | 平台或 COO 预算 | 需要血缘、合规和生产级分析 | 高度契合信任 / 治理叙事和 Standard Chartered 案例 |
| 制造 / 工业 | 运营或数字化负责人 | 工程师和分析师 | 运营预算 | 需要跨站点分析和模型运营化 | 匹配 Mitsubishi Electric 和 Michelin 式部署 |
| 物流 / 供应链 | 运营分析负责人 | 共享服务分析师、支持团队 | 运营或共享数据预算 | 需要工作流自动化和多源数据准备 | 在 Geodis 和更广泛物流案例中可见 |
| 零售 / CPG | 商业洞察负责人 | 分析师和业务用户 | 商业分析预算 | 需要带监督的自助服务 | 契合低代码 / 无代码协作卖点 |
| 企业级 AI CoE | 首席数据官或平台负责人 | 业务与技术混合团队 | 中央 AI 平台预算 | 需要跨云、模型和团队的单一控制平面 | 这是 Dataiku 的典型高价值部署 |
Dataiku 由中心化团队采购,却靠广泛内部使用变现,因此买方画像和用户画像差异明显。
[CM007, CM010, CM011, CM012, CM013, CM014]按买方细分,比较治理需求、技术复杂度、业务用户覆盖、伙伴依赖和预算核心度的相对强度。
[CM011, CM012, CM031, CM032, CM036]企业 AI 采用从试验走向受治理的多智能体生产环境时会急剧收窄。
[CM024, CM027]2.4 增长驱动与采用约束
最主要的市场驱动力,是 AI 从实验转向受治理的生产使用。Dataiku 自己的发布明确这么表述,合作伙伴页面也说明原因:企业想要 GenAI 和代理能力,但不想失去对成本、数据血缘或部署标准的控制。因此,信任和治理不只是风险控制,也是需求创造器。但类别的主要约束同样可见。Deloitte 2025 年调查发现,在财务与会计中已使用代理式 AI 的比例仍只是少数,信任是首要障碍。SiliconANGLE 和 2026 年 arXiv 行业研究从不同角度强化了同一信息:数据质量、集成、验证和人工监督仍是瓶颈。Observer 进一步给出经济解读:架构复杂时,回报可能要按年而不是按季度兑现。对 Dataiku 来说,这意味着市场有吸引力,恰恰因为问题很难;但也意味着销售周期、验证要求和实施摩擦会继续很重。公司受益于治理和编排需求,但同一需求也会放慢类别渗透;如果没有内部赢单率和预算数据,完全受约束的 SAM/SOM 分析仍不完整。[CM022, CM023, CM024, CM025, CM026, CM027]
| 驱动 / 约束 | 方向 | 时点 | 证据 | Dataiku 启示 | 尽调追问 |
|---|---|---|---|---|---|
| 企业从试验转向运营化 | 正向 | 当前 | Dataiku 2025 年 10 月发布 | 利好重治理的编排平台 | 核实成交项目是否从试点扩展为平台标准 |
| 信任、血缘与可解释性需求 | 正向 | 当前 | Deloitte 信任门槛 + Dataiku 定位 | 治理会创造需求,不只是合规成本 | 衡量治理驱动胜率与功能驱动胜率 |
| 支出集中在大型企业 | 正向 | 持久 | Fortune 大企业份额 | 支撑 Dataiku 聚焦企业客户的 GTM | 核查 Fortune 2000 客户还剩多少空白 |
| 多云 / 厂商中立需求 | 正向 | 持久 | AWS、Google Cloud、Databricks 伙伴页面 | 印证 Dataiku 把编排做在技术栈之上的叙事 | 询问赢单中多云或混合部署占比 |
| 集成复杂度与数据就绪缺口 | 负向 | 当前 | SiliconANGLE、arXiv、Gartner 评述 | 销售周期和落地摩擦仍高 | 要求提供投产中位时间和专业服务依赖度 |
| 早期智能体 AI 渗透较慢 | 负向 | 当前 | Deloitte 13.5% 使用数据 | 智能体上行空间真实存在,但短期品类变现可能追不上热度 | 询问传统分析 / ML 与智能体的商机管线占比 |
| 复杂部署 ROI 回收期长 | 负向 | 中期 | Observer 部署分析 | 即使有战略兴趣,也可能拖慢预算审批 | 要求提供已测算回收周期的标杆客户 |
| SI 与云伙伴渠道杠杆 | 正向 | 当前 | KPMG 联盟;Snowflake 300+ 客户 | 伙伴可降低销售摩擦、扩大触达 | 按伙伴衡量来源商机和附加率 |
多项约束也是 Dataiku 机会的另一面:治理和集成痛点会创造需求,也会拖慢采用、拉长周期。
[CM022, CM023, CM024, CM025, CM026, CM027]2.5 图表
03竞争对手
3.1 竞争格局与最接近的对手
Dataiku 的竞争集合比一份 DSML 候选清单更宽。现实买方替代方案包括 Databricks 这类统一数据与 AI 平台,Amazon SageMaker、Azure Machine Learning、Vertex AI 这类超大规模云厂商原生 ML 栈,以及 DataRobot、H2O.ai、Alteryx 这类更窄的专家型供应商;后者用不同部署和定价假设解决相邻任务。最重要的区别在于,Dataiku 想成为既有企业数据资产之上的中立控制层,而不是存储、计算和模型服务发生的唯一地点。这让 Databricks 成为最接近的广义平台同行,因为它现在把数据、治理、MLOps 和代理工具放在一个产品家族里销售;超大规模云厂商则通过把 AI 做成既有云采购的另一个功能来竞争。专家型供应商同样重要,因为它们显示,在更窄的见效速度、AutoML 或分析自动化路径上,预算仍可能从完整编排平台流走。[CP001, CP002, CP003, CP004, CP005, CP006]
| 公司 | 类别 | 规模 / 融资信号 | 目标客群 | 核心差异化 | 核心限制 |
|---|---|---|---|---|---|
| Dataiku | 中立的企业 AI 编排平台 | 2025 年 10 月 ARR 超过 $350M;750+ 家客户 | 拥有多角色、受治理 AI 工作流的大型企业 | 基础设施中立的协作、治理与部署灵活性 | 公开价格透明度有限,规模小于 Databricks |
| Databricks | 统一数据 + AI 平台 | 2026 年年化收入 $6.9B;2025 年 12 月估值 $134B | 整合数据工程、分析与 AI 的企业 | 同时掌握数据和 AI 工作流界面,智能体路线图强 | 中立性较弱,因为它也想成为核心数据平台 |
| Amazon SageMaker | 超大云厂商原生 ML 技术栈 | AWS 级采购通道和细颗粒度用量计费 | 以 AWS 为中心的工程和 ML 团队 | 原生 AWS 集成,按工作负载计量定价 | 更像组件化工具,不像跨异构技术栈的中立层 |
| Azure Machine Learning | 超大云厂商原生 ML 技术栈 | Microsoft 企业协议杠杆和按量付费选项 | Azure 优先企业,尤其是既有 Microsoft 体系 | Azure 体系内的企业 MLOps 与负责任 AI 框架 | 经济性和路线图仍绑定 Azure 消费选择 |
| Vertex AI / Agent Platform 平台 | 超大云厂商原生 ML 与智能体技术栈 | Google 模型、训练与推理价格公开披露 | 以 GCP 为中心的数据和 GenAI 团队 | Gemini 原生工具,加上细颗粒度模型运营定价 | 仍锚定 Google Cloud,而非跨云中立 |
| DataRobot | 专精型企业 AI 套件 | 2024 年收入约 $285M;此前估值峰值 $6.3B | 希望不用重建整个平台、优先获得引导式企业 AI 交付的团队 | 可在 on-prem、VPC 和 SaaS 间选择部署,并主打集成套件 | 反向证据显示,相比打包平台,品类韧性较弱 |
| H2O.ai | 专精型混合 / AutoML 平台 | 2021 年 $100M Series E 轮,估值 $1.7B;声称覆盖 20,000 家组织 | 需要开源血统、混合部署和 AutoML 的用户 | 开源根基和混合云灵活性 | 已披露资本基础远小于 Databricks 或 Dataiku |
| Alteryx | 相邻的分析自动化替代品 | 2023 年以 $4.4B 被收购;8,000+ 家客户 | 商业分析和低代码自动化团队 | 大众化分析和工作流自动化品牌强 | 对端到端 ML / 智能体生命周期深度聚焦较少 |
这些画像按买方预算流向分组:哪些平台最可能吸走 Dataiku 瞄准的同一笔预算或工作流主导权。
[CP001, CP003, CP005, CP007, CP008, CP009]当坐标轴是基础设施中立性和受治理工作流广度时,Dataiku 得分最高;但 Databricks 掌握更多相邻数据平台预算,正在缩小差距。
坐标来自对已审阅源材料包的序数综合。X 轴代表基础设施中立性 / 部署灵活性;Y 轴代表受治理 AI 工作流广度。
[CP002, CP003, CP005, CP007, CP008, CP009]3.2 能力与打包对比
公开打包表面显示出清晰分野。超大规模云厂商原生选项暴露细粒度用量定价,而 Dataiku 和多数专家型同行仍销售企业合同和架构决策,而不是简单的标价 SKU。Dataiku 的归档方案页面清楚展示了它过去如何描述产品:当团队从免费或小团队使用走向企业级采用时,更深的连接器覆盖、自动化和受治理部署能力才会解锁。Databricks 透明度稍高,因为它公布 SKU 组价目表;但即便如此,买方仍必须把用量映射到具体云服务和折扣。AWS、Azure 和 Vertex AI 把计量经济性明示出来,这有助于在功能层面对比,但也把成本风险推向架构和运行时选择。DataRobot、H2O.ai 和 Alteryx 更接近演示驱动或联系销售的路径,这符合企业软件常态,却让外部仅凭公开信息很难做同口径 TCO 对比。[CP012, CP013, CP014, CP015, CP016, CP017]
| 采购标准 | Dataiku | Databricks | SageMaker | Azure ML | Vertex AI | DataRobot | H2O.ai | Alteryx |
|---|---|---|---|---|---|---|---|---|
| 跨云 / 本地部署灵活性 | 强 | 中等 | 仅限 AWS | 仅限 Azure | 仅限 GCP | 强 | 强 | 中等 |
| 业务用户可用性 | 强 | 中等 | 弱至中等 | 中等 | 中等 | 中等 | 中等 | 强 |
| 受治理的端到端工作流广度 | 强 | 强 | 中等至强 | 中等至强 | 中等至强 | 中等 | 中等 | 中等 |
| 原生数据平台控制权 | 弱 | 强 | AWS 内中等 | Azure 内中等 | GCP 内中等 | 弱 | 弱 | 弱 |
| 伙伴 / 渠道杠杆 | 强 | 强 | 强 | 强 | 强 | 中等 | 中等 | 中等 |
| 公开定价透明度 | 低 | 中等 | 高 | 中等 | 高 | 低 | 低 | 低 |
评分是对已审阅产品、伙伴和定价页面的序数综合;未获支持的实际成本判断有意不作推断。
[CP002, CP012, CP013, CP014, CP015, CP016]| 公司 | 公开合同模式 | 可见包含内容 | 未知项 | 含义 |
|---|---|---|---|---|
| Dataiku | 企业分层;存档资料显示 free / discover / business / enterprise 框架 | 协作、连接器、自动化、部署选项、治理深度 | 当前实际成交价格、折扣和云托管加价不公开 | 采购由架构牵引,而非由自助价格牵引 |
| Databricks | 基于用量的 SKU 定价,并有分云价格表 | 平台服务按标价 SKU 和组合销售 | 实际折扣和按工作负载计算的完整 TCO 仍不公开 | 公开透明度高于多数同业,但财务团队仍不容易理解 |
| Amazon SageMaker | 按实例、时长、存储和推理模式,对功能逐项计量定价 | Notebook、训练、推理、特征库、处理、MLflow 等服务 | 完整账单取决于架构、实例选择和工作负载强度 | 可以小规模起步,但一旦计算密集型场景跑通,支出可能不可预测地上升 |
| Azure Machine Learning | 报价驱动的 Azure 服务,支持按量付费、预留和节省计划 | 叠加在 Azure 计算选择之上的端到端 ML 生命周期服务 | 企业实际价格取决于 Microsoft 协议条款和所选基础设施 | 对既有 Azure 体系很强;外部比较时透明度较低 |
| Vertex AI / Agent Platform 平台 | 训练、部署、AutoML 和预测按量定价 | 模型运营按小时计费,无最低使用时长,预测按次数分层 | 混合成本仍取决于模型选择、端点设计和 GCP 使用模式 | 对突发式试验有吸引力,但原生云锁定仍在 |
| DataRobot、H2O.ai 与 Alteryx | 多数是演示驱动或联系销售的企业销售动作 | 套件定位、部署选项和包装线索公开 | 已抓取页面没有详细企业标价 | 与 Dataiku 的公开 TCO 对比在结构上不完整 |
本表比较已抓取来源包实际披露的内容,而不是私下谈判合同最终可能长什么样。
[CP012, CP013, CP014, CP015, CP016, CP017]Dataiku 在中立性和混合用户画像工作流广度上领先;超大规模云厂商赢在原生采购,Databricks 赢在相邻平台所有权。
这些值是 1-5 的序数评分,来自抓取的产品、伙伴和定价页面,而不是单一第三方基准。
[CP012, CP013, CP014, CP015, CP016, CP017]公开规模指标解释了为什么 Databricks 是最严峻的竞争威胁;DataRobot 和 Alteryx 则显示,相邻类别可能走出很不同的经济路径。
KPI 条带混合了公司披露的经营指标和独立规模标记。它用于概括竞争就绪度,不代表市场份额。
[CP019, CP020, CP021, CP022, CP033, CP034]3.3 转换成本与分发力量
Dataiku 真正的护城河不是单一模型或专有数据资产,而是在异构工具、云和用户类型之间提供受治理协作的运营便利。这一点在已经并行运行 Snowflake、Databricks、AWS、Google Cloud 和内部代码工具的企业中最关键。在这种环境下,Dataiku 可以作为横跨技术栈的编排和治理层取胜。但同一架构也带来主要战略脆弱性:超大规模云厂商可以从买方既有合同、身份系统和云端数据引力出发;Databricks 则可以从数据与计算工作流本身的所有权出发。公开合作伙伴页面显示,Dataiku 更倾向于共存,而不是推倒重来式替换;这扩大了分发,也降低了孤立风险。专家型供应商在买方想要较短见效周期或低代码自动化时仍有空间,但它们通常缺少最大平台竞争者的生态杠杆广度或已披露资本规模。这也意味着销售竞争往往由集成可信度和变革管理舒适度决定,而不只是原始模型功能。[CP018, CP023, CP024, CP025, CP026, CP027]
3.4 护城河耐久性与反向解读
反向证据并不是说 Dataiku 弱;它说明当独立 AI 平台面对捆绑基础设施或更广系统记录时失去差异化,这个类别会非常残酷。Databricks 是最高严重度威胁,因为它扩张更快、资本实力强得多,并持续从数据基础设施扩展到治理、AI 代理和应用界面。超大规模云厂商是第二个结构性威胁,因为它们能让原生工具在云承诺内部显得「足够免费」。DataRobot 提供了最清楚的警示案例:曾经炙手可热的独立 AI 公司,在原生云工具和范式变化改变买方关注点时,战略相关性会迅速流失。H2O.ai 和 Alteryx 说明相邻领域仍有价值,但也说明并非每个相邻类别都能拿到 AI 平台倍数。Dataiku 最好的防守,是继续在那些即便原生 AI 功能改善后仍会保持多云、多工具的账户中,成为中立、受治理的工作流层。[CP031, CP032, CP033, CP034, CP035, CP036]
| 护城河主张 | 威胁 | 严重性 | 可信原因 | 缓释 / 尽调追问 |
|---|---|---|---|---|
| 基础设施中立治理层 | 超大云厂商把原生工具做到在既有云合同内「够用」 | 高 | AWS、Azure 和 Google 都提供第一方生命周期和智能体工具,并公开用量定价 | 要求按云和部署模式提供赢单 / 输单数据 |
| 覆盖业务和技术用户的宽工作流 | Databricks 持续从数据平台扩展到 AI 智能体和受治理交付 | 高 | Databricks 现在在其数据平台之上营销智能体构建、治理和模型生命周期 | 测试当 Databricks 已是数据标准时,Dataiku 是否还能赢 |
| 伙伴主导分销 | 伙伴可能把预算导向自身原生服务 | 中 | Dataiku 与 AWS、Google Cloud、NVIDIA、Databricks 和 Snowflake 联合销售,而这些伙伴都有自己的盘算 | 量化来源商机、影响商机和伙伴依赖度 |
| 特定用例中的专精工具价值兑现速度优势 | 企业平台标准化前,范围更窄的工具可能先拿下部门预算 | 中 | DataRobot、H2O.ai 和 Alteryx 仍在营销易用性、混合部署或低代码自动化 | 核查试点流失是否输在简单性而非能力 |
| 企业 AI 品类热情 | 当价值看起来已被打包或被过度炒作,独立 AI 平台叙事可能被压缩 | 高 | DataRobot 的不利轨迹和 Alteryx 截然不同的公开市场结果表明,该品类可被快速重估 | 要求提供历史定价压力、续约和治理模块附加率 |
风险清单聚焦差异化耐久性的威胁,而非市场分析章节已覆盖的通用市场风险。
[CP024, CP025, CP027, CP032, CP033, CP034]3.5 图表
04财务
4.1 收入模式与变现
公开记录足以识别 Dataiku 的商业形态,尽管还不足以精确建模。Dataiku 显然是一家经常性企业软件公司,不是广告支持产品、市场平台,也不是纯粹按用量计费的 API 供应商。公司 2025 年发布以 ARR 而不是订单额或服务收入来描述增长;其归档方案页面也显示,从免费或小团队入口,到更深的自动化、部署、安全和治理能力,产品存在结构化升级路径。当前产品页面进一步强化了平台围绕广度销售:编排、治理、代理和企业数据控制。也就是说,买方购买的是一个跨职能控制平面,其价值随部署范围扩大而增加,而不是单模块点解决方案。关键细节在于,公开材料没有披露实际标价或具体收入确认机制。与 Databricks 和超大规模云厂商相比,Dataiku 看起来更偏合同导向,用量定价透明度较低;这可能提升预算可预测性,但也让外部仅凭公开信息无法对实际经济性做基准比较。[CI001, CI002, CI003, CI004, CI005, CI006]
| 收入流 | 机制 | 单位 | 当前价值 / 状态 | 质量 | 尽调追问 |
|---|---|---|---|---|---|
| 核心平台订阅 / ARR | 周期性企业平台合同 | ARR / 年合同价值 | 2025 年 1 月 ARR 超过 $300M;2025 年 10 月 ARR 超过 $350M | 存在性置信度高,当前收入规模置信度中等 | 要求提供季度 ARR 拆解、队列扩张和合同期限组合 |
| 部署 / 托管变现 | 平台可由 Dataiku 或客户环境托管 | 合同收入加可能的托管加价 | 公开部署选项可见;实际托管收入未披露 | 中 | 要求提供云托管收入占比与客户自管部署对比 |
| 治理 / 智能体广度变现 | 平台广度扩大企业账户内变现面 | 向上销售 / 版本扩展 | 当前产品页面强调智能体、治理、编排和共享控制平面 | 中 | 要求按能力族提供模块附加率和扩张 |
| 专业服务 / 培训 | 落地、咨询和赋能可能存在,但公开未分项列示 | 服务费 | 公开证据显示服务存在,但未披露收入拆分 | 低 | 要求提供服务收入占比、毛利率,以及伙伴交付与直接交付拆分 |
| 伙伴交付服务 | 系统集成商和云伙伴可围绕核心平台交付实施 | 间接服务影响 | 伙伴生态广泛,但经济性未披露 | 中 | 要求提供来源商机、服务附加和伙伴补偿模型 |
| 支持 / 维护 | 企业支持很可能打包进核心合同 | 已包含支持服务 | 公开材料暗示企业支持打包提供,而非单独定价维护 | 低 | 要求提供支持负担、续约条款和单个大客户支持成本 |
本表区分已清晰披露的周期性软件牵引力,以及透明度较低的实施或模块级收入构成。
[CI001, CI003, CI004, CI005, CI006, CI007]| 公司 / 模式 | 价格 / 单位 / 合同 | 标价 vs. 实际成交价 | 折扣 / 未知项 | 含义 |
|---|---|---|---|---|
| Dataiku | 企业合同模式;存档分层从 free 到 enterprise | 存档标价结构可见,当前实际成交价不公开 | 当前合同条款、折扣和托管加价未知 | 预算可能比纯用量模式更可预测,但公开基准较弱 |
| Databricks | 按用量定价的 SKU 和分云价格表 | 标价公开,实际商业条款靠谈判确定 | 折扣阶梯和完整工作负载 TCO 仍不公开 | 按企业软件标准算透明,但仍取决于架构 |
| Amazon SageMaker | 按实例、时长、存储和推理配置计量 | 公开到功能层级的定价 | 最终账单取决于运行时选择和工作负载强度 | 自助可见度强,但事前预算确定性较弱 |
| Azure Machine Learning | Azure 服务定价,包含按量付费和预留选项 | 公开费率卡结构,加上企业报价语境 | 实际价格取决于 Azure 协议和所选基础设施 | 在 Microsoft 客户体系内,企业采购议价能力强 |
| Vertex AI | 模型运维、训练、部署和预测定价 | 公开计量定价 | 模型、端点和用量选择决定实际成本 | 适合突发式实验,但成本可见度取决于架构 |
| C3.ai 对照项 | 订阅费叠加基于用量的运行时和托管费用,后者嵌入订阅 | 申报文件没有简单公开费率卡 | 客户定制结构和专业服务存在差异 | 显示 AI 平台可把承诺订阅与用量挂钩元素结合起来 |
官方定价揭示同行如何变现;但不披露 Dataiku 的实际合同经济性,后者仍属私有信息。
[CI006, CI010, CI011, CI013, CI016, CI017]Dataiku 通过企业 AI 广度变现:客户先采用平台,再扩展到更广的治理、部署和智能体用例,支撑经常性 ARR。
这张图根据公开 ARR 披露、产品页面和归档包装抽象商业动作。它不意味着已披露的转化率或附加率。
[CI001, CI003, CI004, CI005, CI006, CI007]4.2 销售动作与单位经济代理指标
Dataiku 的产品和客户模式指向经典的企业级先落地再扩张模式,但硬性的单位经济字段仍未披露。大型客户、合作伙伴生态和云无关部署都指向长销售周期、高 ACV 与多利益相关方采购;同时 2025 年 ARR 里程碑显示,该模式可以越过试点阶段继续放大。最好的公开代理指标来自相邻上市公司。C3.ai 的申报文件和 FY2026 业绩显示,企业 AI 软件可以高度由订阅驱动,但仍带有服务和优先工程工作,以支撑部署和路线图加速。Snowflake 年报展示了变现光谱的另一端:消费驱动模式、很高的产品毛利、可观的剩余履约义务和强劲经营现金生成,但客户可以优化用量,因此收入可见性不那么线性。合在一起,这些可比公司说明,对 Dataiku 的真正承销问题不是公司有没有牵引力,而是服务强度、云成本和合作伙伴经济性是否能在 ARR 扩大后让业务收敛到强软件利润率。[CI010, CI011, CI012, CI013, CI014, CI015]
| 指标 | 数值 / null | 置信度 | 重要性 | 尽调要求 |
|---|---|---|---|---|
| ARR | 2025 年 10 月披露 $350M+ | 中 | 验证公司规模和企业市场相关性 | 要求提供截至 runDate 的月度 ARR 桥接表 |
| 收入增长 | 过去三年 ARR 增长超过一倍;当前确切增长率未公开 | 中 | 用于检验经营杠杆和估值支撑 | 要求提供年度 ARR、收入和订单额历史 |
| 毛利率 | 未公开披露 | 低 | 决定 Dataiku 能否走向有吸引力的软件经济性 | 要求按软件、托管和服务拆分毛利率 |
| NRR / 扩张率 | 未公开披露 | 低 | 是评估先落地再扩张模型的关键 | 要求按细分市场提供同期群留存和美元口径扩张率 |
| CAC 回收期 / 销售效率 | 未公开披露 | 低 | 用于判断企业销售效率 | 要求按地区和渠道提供销售产能、CAC 和回收期 |
| 服务收入占比 | 未公开披露 | 低 | 服务强度高会压住利润率和现金转化 | 要求拆分直接交付与伙伴交付服务占比 |
| 合同积压 / RPO | 未公开披露 | 低 | RPO 可显示前瞻可见度和续约质量 | 要求提供递延收入和 RPO 时间表 |
| 现金转化 / 经营现金流 | 未公开披露 | 低 | 决定资本需求和自筹资金能力 | 要求提供历史经营现金流和自由现金流 |
每个 null 字段都是真正的承销障碍,不是格式遗漏。
[CI001, CI002, CI003, CI020, CI030, CI032]可观察的经济路径从企业获客走向经常性 ARR,但几个关键单位经济检查点仍未披露。
节点为定性,因为 Dataiku 未公开披露 CAC、回本周期、NRR 或毛利率。
[CI020, CI030, CI031, CI032, CI035, CI036]公开证据给出 Dataiku 牵引力底线的有界视角,也给出毛利率和平台规模的宽比较样本。
[CI001, CI003, CI014, CI015, CI022, CI033]4.3 成本结构与资本充足性
外部对 Dataiku 实际成本结构的可见度很差,因此本章必须把可观察事实和只能推断的内容分开。可观察的是:Dataiku 2025 年 1 月 ARR 越过 $300M,2025 年 10 月 ARR 越过 $350M;自 2022 年 Series F 后没有公开宣布新融资;并据报道在 2025 年末准备可能的美国 IPO。Sacra 估计其累计融资约 $846.8M,2025 年 9 月 ARR 约 $342.5M,整体与公司官方披露吻合。可推断的是:一家以这种规模运营的公司,很可能在企业销售、客户成功、产品开发、云基础设施和合作伙伴赋能上投入很重。但推断不足以支撑承销。公开资料没有现金余额、现金消耗历史、有意义的债务披露、留存指标,也没有直接毛利披露。因此,公平的资本充足性解读应保持谨慎:公开记录中没有任何内容强烈指向困境,但也没有任何内容能让投资人验证现金可支撑时间或融资依赖。IPO 选择权看起来像战略选择,而不是已被证明的必要性。[CI021, CI022, CI023, CI024, CI025, CI026]
| 字段 | 公开状态 | 重要性 | 当前解读 | 尽调要求 |
|---|---|---|---|---|
| 最近披露的一级融资 | 2022 年 12 月完成 $200M Series F,估值 $3.7B | 锚定历史资本结构,但不能证明当前流动性 | 之后未公开宣布新的一级融资 | 要求提供 Series F 以来按季度列示的当前现金 |
| 融资总额 | Sacra 估算累计融资约 $846.8M | 提供稀释和融资历史语境 | 从方向上看,作为私有软件公司资本储备较充足 | 要求完整股权结构表和债务明细 |
| 在手现金 | 未公开披露 | 跑道判断的核心输入 | Unknown | 要求提供非受限和受限现金余额 |
| 烧钱额 / 跑道 | 未公开披露 | 用于判断融资依赖度 | 未知;公开信息未显示困境信号 | 要求提供月度烧钱额、预算和情景计划 |
| 资本市场可选项 | Reuters 2025 年 10 月报道称其正筹备 IPO 投行工作;2026 年 IPO 市场背景改善但仍阶段性波动 | 影响流动性和未来融资灵活性 | 该选项看起来偏战略性,尚不能证明迫切 | 要求提供董事会批准的融资计划和 IPO 准备预算 |
| 债务 / 项目融资义务 | 未发现重大公开债务或项目融资义务 | 债务会迅速改变风险画像 | 无公开证据显示沉重债务负担 | 要求提供债务额度、契约条款和表外承诺 |
本表关注前瞻资本充足性,不重复公司概览已覆盖的逐轮融资历史。
[CI021, CI022, CI023, CI024, CI025, CI026]公开证据指向一个有企业销售和云交付需求的软件业务,而不是资本密集型硬件或项目融资模式。
这张图识别可见现金用途和融资选项,不是披露预算。
[CI021, CI022, CI024, CI026, CI027, CI028]4.4 财务结论与尽调阻塞项
正面结论很直接:Dataiku 有真实规模、经常性收入和足够的类别可信度,可以被放在 IPO 候选公司中讨论,而不是私营实验项目。负面结论同样直接:公开记录仍不足以让人有信心承销利润率路径、现金效率或稀释风险。ARR 里程碑、客户广度和平台定位都支持它是一家高质量企业软件公司的判断。但缺失字段正是区分「有意思的私营公司」和「可投资公司」的字段:净收入留存、毛利、服务收入占比、CAC 回收期、销售效率、当前现金、现金消耗和合同积压。上市公司比较只能当护栏。Snowflake 展示强软件经济性可以是什么样;C3.ai 展示执行问题和服务强度如何压缩利润率。Dataiku 可能位于两者之间,但公开证据无法判断精确位置。因此,正确的承销态度不是看空牵引力,而是对缺失经济性保持纪律。[CI020, CI024, CI025, CI030, CI032, CI034]
| 缺失的私有指标 | 影响 | 公开证据为何不足 | 精确尽调路径 |
|---|---|---|---|
| 按收入流拆分的毛利率 | 没有它,毛利率路径只能靠猜 | ARR 公告未披露软件与服务毛利率 | 要求提供按收入流和部署模式拆分的经审计毛利率桥接表 |
| NRR、流失率和同期群扩张 | 仅凭 ARR 里程碑无法判断收入质量 | 公开材料没有同期群数据或 NRR | 要求按年份、细分市场和地域提供同期群表 |
| 当前现金和烧钱额 | 无法验证跑道 | 私有公司没有公开资产负债表 | 要求提供月度现金瀑布表和 12 个月经营计划 |
| 服务结构和合作伙伴经济性 | 实施强度可能实质影响毛利率 | 合作伙伴覆盖可见,经济条款不可见 | 要求提供直接服务结构、伙伴附加率和工作说明书经济性 |
| 合同期限、递延收入和 RPO | 积压订单和收入可见度仍未知 | 无公开申报披露合同积压订单 | 要求提供递延收入明细和 RPO 披露 |
| 销售效率和 CAC 回收期 | 难以判断增长是高效,还是只是昂贵 | 管道转化或回收期没有公开披露 | 要求提供销售产能模型、CAC、回收期和配额达成率 |
这些是认真给出估值或融资建议前最低限度必须补齐的指标。
[CI024, CI025, CI030, CI032, CI035, CI036]4.5 图表
05产品与技术
5.1 产品实际交付什么
与其把 Dataiku 理解成单一建模功能,不如把它看成企业 AI 工作的共享运营层。当前产品页围绕三项功能描述平台:人来构建,编排来连接,治理来保护。这个框架与底层文档一致。DSS 结合了可视化数据准备、notebook 式代码工作、自动化、部署和 API 访问。Govern 增加了一个独立监督节点,用于跟踪 AI 项目、审批、登记和审计工作流。围绕代理和受治理 AI 的当前营销显示,Dataiku 现在想成为企业设计、监控和路由代理式工作负载的地方,而不只是训练经典 ML 模型的地方。因此,产品横跨多类角色:分析师、数据科学家、ML 工程师、平台负责人、治理团队和业务利益相关方。重要的技术含义是,Dataiku 的价值不在某一个算法,而在一个受控环境里协调异构的人、数据、模型和审批工作流。[CE001, CE002, CE003, CE004, CE005, CE006]
| 模块 / 资产 | 主要用户 | 状态 / 成熟度 | 差异化 | 尽调缺口 |
|---|---|---|---|---|
| 核心 DSS 平台 | 分析师、数据科学家、工程师 | 成熟核心平台 | 在一个环境中结合可视化和代码工作流 | 无公开基准性能数据 |
| 自动化和部署 | ML 工程师、平台负责人 | 成熟且持续维护 | 串起构建、部署、监控和运营工作流 | 抓取材料中看不到公开运行时间 / SLA 细节 |
| Dataiku Govern | 治理、风险、AI 监督团队 | 成熟附加节点,功能进阶 | 中央登记簿、签核规则、工作流跟踪、审计时间线 | Govern 当前客户采用率未公开量化 |
| LLM Mesh / 生成式 AI 层 | AI 平台团队和应用构建者 | 在 v14 发布线中持续扩展 | 模型抽象、路由、安全控制和 GenAI 工作流支持 | 公开文档未量化延迟、路由成本或准确率提升 |
| AI 智能体 / 智能体管理 | AI 应用构建者和治理负责人 | 较新,但明显是活跃优先领域 | 企业级受治理智能体构建、管理和监控 | 生产结果的公开证据仍有限 |
| API 和开发者工具 | 开发者、集成商、管理员 | 已成体系且外部可见 | Python API、开发者指南、客户端工具、自动化接口 | 代码仓库活跃度证据有限;更广泛的外部开发者足迹不清楚 |
模块拆分反映产品页、文档和发布说明中可见的内容,而非内部 SKU 粒度。
[CE001, CE002, CE003, CE005, CE006, CE007]| 用户任务 | 当前工作流问题 | Dataiku 方案 | 可衡量收益 | 限制 |
|---|---|---|---|---|
| 构建跨职能 AI 项目 | 业务、数据和工程团队之间工作割裂 | 共享平台,兼具可视化和代码界面 | 可能加快协作并管控交接 | 抓取材料中没有公开价值兑现时间基准 |
| 将受治理 GenAI 落到生产 | 企业需要模型路由、控制和可见性 | LLM Mesh、智能体工具和治理功能 | 集中控制智能体工作流 | 公开证据未量化可靠性或成本节省 |
| 跟踪 AI 倡议和审批 | Shadow AI 和审计轨迹难以手工管理 | Dataiku Govern 工作流、登记簿、签核规则、告警 | 提升审计就绪度和政策执行 | 企业实际流程采用率未披露 |
| 将平台接入云技术栈 | 团队希望在既有云和数据资产上运行 AI | 由合作伙伴牵引,配合 AWS、Google Cloud、Databricks、NVIDIA 等部署 | 降低推倒重来替换既有基础设施的需求 | 在部分环境中,集成复杂度看来仍是真实风险 |
| 用代码自动化平台动作 | 团队需要可重复的程序化操作 | Python API、开发者指南、API 客户端、场景、代码配方 | 支持扩展性和自动化 | 从公开信号看,外部社区深度并不清晰 |
收益表述偏保守,因为公开页面更强调能力,而非量化 ROI。
[CE002, CE003, CE004, CE007, CE013, CE016]Dataiku 在共享构建界面之下叠加代码 API、治理、部署运营,以及外部云或模型生态。
[CE001, CE003, CE005, CE007, CE011, CE015]典型 Dataiku 工作流从数据和项目设置出发,经过模型或智能体创建、治理评审、部署,再进入受监控的复用。
[CE002, CE004, CE005, CE016, CE017, CE020]5.2 架构、部署与 API
官方页面、归档打包、合作伙伴页面和技术文档中最一致的产品主题,是灵活性。文档显示,Dataiku 支持 SaaS、客户自管云,以及本地或私有云运营模式。它为非代码用户提供可视化界面,也通过 API、代码配方、notebook、场景和客户端工具提供完整开发者界面。Python API 参考明确表示,Dataiku 工具可以在 DSS 内任何运行代码的地方使用;开源的 GitHub API client 和 README 也表明,针对平台的外部自动化并非只限内部。与 AWS、Google Cloud、Databricks 和 NVIDIA 的合作页面进一步说明,Dataiku 的架构是坐在客户基础设施和外部 AI 栈之上,而不是替代它们。这一点有战略意义:平台的技术身份首先是编排,其次是集成。它也意味着,对云、模型和合作伙伴生态的依赖是设计特征,不是偶然副产品。[CE007, CE008, CE011, CE012, CE013, CE018]
| 层 / 组件 | 作用 | 依赖 | 风险 |
|---|---|---|---|
| 可视化工作流和 UI 层 | 让数据准备、建模、仪表盘和智能体设计更易上手 | 依赖 DSS 核心和发布节奏 | 如果运营证据不足,容易滑向营销式广度 |
| 代码和 API 层 | 支持 notebook、配方、自动化和外部集成 | 依赖 Python API、客户端库和开发者工具 | 平台铺得越宽,版本管理和集成复杂度可能越高 |
| 治理节点 | 跟踪资产、审批、模板和登记簿 | 依赖 Govern 实例设置和政策设计 | 只有客户把治理流程跑起来,强项才成立 |
| LLM / 智能体编排层 | 路由模型、工具和智能体工作流 | 依赖模型提供商、API、云基础设施和安全护栏 | 依赖格局变化很快,可能带来变更管理负担 |
| 部署和运行时层 | 将项目推入生产和监控流程 | 依赖客户选择云端、本地或托管部署 | 不同部署模式下运营负担差异很大 |
| 合作伙伴和生态层 | 将 Dataiku 接入云、数据和加速器生态 | 依赖第三方合作伙伴优先级和兼容性 | 伙伴依赖可改善分销,但也会增加技术耦合 |
这是基于官方文档、合作伙伴页面和发布说明形成的运营架构综合判断,不是供应商发布的块图。
[CE007, CE011, CE012, CE013, CE015, CE021]Dataiku 拥有控制平面,但产品的重要部分依赖云平台、模型提供商、伙伴生态、安全态势和客户基础设施。
[CE013, CE014, CE021, CE022, CE030, CE033]5.3 信任、安全与治理控制
在 Dataiku 的公开叙事里,治理不是薄薄一层勾选框,而是主要产品支柱之一。Govern 产品页和 Govern 文档都描述了集中式项目跟踪、资产登记、工作流审批、签核规则和审计时间线。公开材料还把这些控制直接连接到监管压力,明确提到 EU AI Act 准备度和减少 shadow AI。技术文档进一步说明 Standard 和 Advanced govern 许可证;Advanced 功能覆盖 GenAI 登记库、自定义治理模板、自定义动作和脚本。安全证据更混杂,但仍有可信度。Dataiku 安全页面和安全文档显示公司有活跃安全运营;同时,一页 2026 年警告文档为影响 Dataiku 环境的 Linux 本地提权漏洞给出响应指引。公司的 SOC 2 公告时间较早,不应被视作完整当前认证清单,但它仍提供证据:正式合规工作多年来一直是产品和运营故事的一部分。[CE005, CE006, CE014, CE015, CE016, CE017]
| 控制项 / 认证 / 质量信号 | 状态 | 范围 | 缺口 |
|---|---|---|---|
| Govern 签核规则 | 已有文档 | 审批工作流可在满足要求前阻止部署 | 没有公开证据显示客户在生产中使用得多广 |
| 审计时间线和登记簿 | 已有文档 | Bundle、模型和 LLM 登记簿,加上审计就绪时间线 | 没有关于审计效率或误报减少的公开指标 |
| EU AI Act 准备度表述 | 官方 Govern 页面已有文档 | 合规加速明确写进产品卖点 | 公开法律映射细节仍停留在高层 |
| 安全文档和补丁指引 | 已有文档 | 文档和 2026 年安全警告显示其提供主动响应指引 | 详细安全架构和测试证据未充分公开 |
| SOC 2 合规公告 | 历史证明点 | 显示正式合规工作和企业信任信号 | 公告时间较早,且不是完整当前认证清单 |
抓取材料证明存在治理和安全流程界面,但不能证明其量化效果。
[CE014, CE015, CE016, CE017, CE027, CE033]核心工作流和部署能力看起来成熟,治理和智能体界面更新一些,但明显活跃并在扩张。
评分是 1-5 的序数综合,来自官方文档、发布说明和公开开发者界面,而非第三方基准。
[CE007, CE008, CE009, CE017, CE019, CE025]5.4 路线图、依赖与产品风险
发布说明清楚显示,Dataiku 按持续的企业节奏出货,而不是躺在较老的 DSS 核心上。Version 14 发布流显示,2026 年 6 月和 7 月,公司在代理式 AI 与 RAG、LLM Mesh、治理、MLOps、AI 助手、数据质量、Git、安全和代码工具上反复更新。这支撑了成熟度判断:Dataiku 不是试点时代的平台。与此同时,路线图证据也凸显依赖和风险。发布说明提到 Python 版本变化、Java 最低版本上调、Llama 模型移除、容器镜像和 OS 变化,以及云栈注意事项。这些都是真实企业平台的正常信号,但也确认 Dataiku 的产品依赖底层云、OS、模型和开源生态保持稳定。外部产品证明比官方文档薄。Gartner 评论证据仍至少包含一条关于私有云集成的批评,公开表面也没有提供经过基准测试的性能、uptime SLA 或量化代理准确率。尽调中,主要产品问题不是广度,而是运营深度和实施摩擦。[CE009, CE010, CE022, CE023, CE024, CE026]
| 日期 / 阶段 | 功能 / 里程碑 | 状态 | 含义 | 来源 |
|---|---|---|---|---|
| 版本 14.7.2 – 2026 年 7 月 10 日 | Agentic AI 与 RAG、LLM Mesh、治理、AI 助手、编码和 API、插件 | 已发布 | 显示路线图在 GenAI 和平台运营两端都保持活跃广度 | v14 发布说明 |
| 版本 14.7.1 – 2026 年 7 月 1 日 | AI 服务、Cobuild、Spark、Git、安全 | 已发布 | 表明平台持续硬化,开发者界面也在推进 | v14 发布说明 |
| 版本 14.7.0 – 2026 年 6 月 18 日 | 新功能:Cobuild、图表、数据质量、性能 | 已发布 | 显示投入不只围绕 GenAI 功能,也覆盖其他平台能力 | v14 发布说明 |
| 版本 14.6.2 – 2026 年 6 月 11 日 | MLOps、数据集和连接、场景和自动化、Code Studio | 已发布 | 强化其在部署和运营上的成熟度,而不只是实验 | v14 发布说明 |
| 2026 年安全警告 | 面向 Dataiku 环境的 LPE 漏洞响应指引 | 已发布指引 | 显示其安全运营活跃,也要承担依赖管理责任 | 官方安全页面 |
路线图证据来自发布说明,因此比营销承诺更能反映已交付工作。
[CE009, CE010, CE014, CE022, CE028, CE034]5.5 图表
06客户
6.1 客户分层与买方地图
Dataiku 的客户足迹明显由大型企业主导,而不是 SMB 主导;公开证据覆盖多个受监管且运营复杂的垂直行业。公司称,2025 年 1 月服务超过 700 家组织,2025 年 10 月超过 750 家;2024 年 7 月与 KPMG 的联盟发布则称,Dataiku 已拥有超过 600 家客户和 200 家 Forbes Global 2000 客户。Dataiku 客户页面可见的具名客户集中在医疗健康与生命科学(Johnson & Johnson、Novartis、Roche)、制造与工业(Michelin、Mitsubishi Electric、SLB)、金融服务与资本市场(Standard Chartered、Euronext)、物流与交通(Geodis、Prologis)、食品 / 农业(Perdue Farms)。用户很少只是单个数据科学家。更常见的模式是一个跨职能联盟,把数据科学家、分析师、工程师、运营专家和业务领域用户放在同一个受治理平台上。这让经济买方更可能来自中心化的数据、分析、数字化转型或业务平台预算,而不是按席位购买的部门预算。因此,分层证据对谁在使用 Dataiku、什么类型企业会采用它很强,但对按垂直、地理或账户规模分段的精确收入结构仍很弱。[CU001, CU002, CU003, CU004, CU005, CU006]
| 细分市场 | 示例客户 | 买方 / 用户 / 付款方 | 主要用例 | 收入 / 战略价值 | 关键不确定性 |
|---|---|---|---|---|---|
| 医疗健康与生命科学 | Johnson & Johnson、Novartis、Roche | 数据 / AI 负责人 + 业务团队 + 受监管利益相关方 | GenAI 助手、市场研究、专利分析、分析标准化 | 高价值受监管客户,验证治理需求 | 未公开按细分垂直领域拆分的收入结构或续约数据 |
| 制造与工业 | Michelin、Mitsubishi Electric、SLB | 工程、运营、卓越制造、数字化项目 | 工厂分析、根因分析、能源优化、现场工程 | 运营场景支撑席位和工作流广泛扩张 | 部署究竟全球标准化,还是只在部分区域落地,尚不清楚 |
| 金融服务与资本市场 | Standard Chartered、Euronext | 分析卓越中心、产品团队、银行技术团队 | 市占率分析、FP&A、投资组合分析、治理 | 与有治理、可解释的企业分析高度契合 | 未披露单客户 ARR 或银行业集中度 |
| 物流、房地产与交通运输 | Geodis、Prologis | 运营、IT 支持、数据与分析团队 | 工单分流、预测、地理空间分析、企业聊天 | 证明平台不只服务传统 ML 团队 | 结果数据扎实,但合同规模未披露 |
| 食品与农业 | Perdue Farms | 食品安全运营与业务用户 | 报告自动化、数据准备、自助分析 | 证明 Dataiku 能在技术密集型买家之外,将运营数据工作流商业化 | 单一公开案例不足以证明品类深度 |
| 全企业转型项目 | 跨客户模式 | 中央数据 / AI 平台负责人 + 分散的业务用户 | 业务自助数据科学、治理、智能体 AI、共享工作流 | 最能解释客户数增长的同时,账户内用例宽度为何也在扩大 | 公开证据偏定性,而非基于队列 |
这些细分来自具名公开案例和官方客户数披露,而不是公司发布的收入分部表。
[CU001, CU002, CU003, CU004, CU005, CU006]Dataiku 通常从高价值工作流切入,再扩展到更广的受治理分析、GenAI 和业务用户赋能。
旅程综合了 Michelin、Prologis、Roche、Standard Chartered 和 J&J 案例研究中反复出现的模式;并非每个客户都完全按这些阶段推进。
[CU008, CU009, CU026, CU039]6.2 采用轨迹与规模信号
最强的商业解读是,Dataiku 看起来正从一个大型企业平台,走向更广安装基础和更深内部使用。Sacra 估计 2023 年底约有 500 家客户;这个基准与官方披露方向一致,即 2025 年 1 月客户超过 700 家,2025 年 10 月超过 750 家。多个案例研究展示的不只是拿下客户 logo,还有内部采用扩大。Johnson & Johnson 的 Vision 组织称,超过 80 名分析和数据科学专业人员采用 Dataiku,页面标题还写明 650+ 名员工使用 Dataiku。Michelin 从 2021 年的 35 名用户,扩展到 2025 年中 50+ 家工厂的 1,500+ 名用户,其中 80% 用户被描述为业务专家。Standard Chartered 称,自 2020 年以来已有 518 名员工完成公民数据科学项目,700 多人完成在线 Dataiku 培训课程;Prologis 称,超过 2,000 名用户通过其用 Dataiku 构建的企业 ChatGPT 部署使用 AI。这些不是完美的留存指标,但它们是有意义的部署规模信号,表明 Dataiku 往往先作为平台落地,然后在客户组织内部扩展。[CU010, CU011, CU012, CU013, CU014, CU015]
| 指标 | 数值 | 截至 | 来源依据 | 含义 |
|---|---|---|---|---|
| 客户 | ~500 | 2023 年末 | Sacra 估算 | 2025 年前客户基础已相当可观 |
| 客户 | 700+ | 2025 年 1 月 | 官方 ARR 新闻稿 + Reuters | 显示 2025 年新 Logo 增长清晰 |
| 客户 | 750+ | 2025 年 10 月 | 官方 ARR 新闻稿 | 即便未披露新的一级融资轮,规模仍进一步扩大 |
| Global 2000 渗透 | 200 家客户 | 2024 年 7 月 | KPMG 联盟新闻稿 | 2025 年激增前,大企业重心已经确立 |
| Johnson & Johnson 用户 | 650+ 名员工;80+ 名分析 / 数据科学采用者 | 案例研究 | 客户佐证 | 证明内部采用范围较广,而不是小型试点 |
| Michelin 用户 | 50+ 家工厂的 1,500+ 名用户 | 2025 年中 | 客户佐证 | 工业运营部署深度高 |
| Standard Chartered 赋能 | 培训 518 人;在线课程完成 700+ 次 | 2020–2022 年起 | 客户佐证 | 培训和业务自助赋能是扩张模型的一部分 |
| Prologis 使用情况 | 2,000+ 名用户;60+ 个项目已进生产 | 案例研究 | 客户佐证 | Dataiku 可以成为内部 AI 运营层 |
| 分享案例的客户 | 2024 年登台 100+ 家 | 2025 年 1 月 | 官方 ARR 新闻稿 | 可引用客户基础足够宽,能支撑营销和同业验证 |
轨迹表结合官方客户数和账户内部更深的采用信号,因为 Dataiku 不披露留存或队列数据。
[CU010, CU011, CU012, CU013, CU014, CU015]公开证据从宽口径客户数披露,收窄到数量更少但记录深入的企业部署案例。
相对值是基于公开证据密度的示意权重,不是披露的转化漏斗或留存曲线。
[CU002, CU010, CU029, CU036]6.3 具名客户证明与结果
作为一家私营基础设施软件公司,Dataiku 的公开证明集异常丰富,因为许多客户故事包含具名高管、具体工作流和量化结果。最好的证据是运营性的,而不是炫耀 logo。Novartis 称,一个 GenAI 用例的洞察时间缩短 90%,基于电子表格的数据摄取提速 600%。Michelin 称,一个过去最长需要六个月的根因分析流程,现在约一小时就能运行;同时,十多家工厂的 600 多名工程师和技术人员依赖一个基于 Dataiku 的参数分析解决方案。Prologis 称,它从 5 个已生产化数据科学产品,推进到 30 个项目和 30 个 API 正在活跃使用,并把 60 多个 AI/ML 项目投入生产。Roche 描述了新 GenAI 项目构建时间从数月压缩到数天、每年节省六位数律师工时,并计划从 80 名欧洲专利专业人员扩展到全球最多 250 名用户。Geodis、Euronext、Perdue Farms、Standard Chartered、SLB 和 Mitsubishi Electric 都发布了类似具体的生产率或决策支持改善。关键限制在于,这些是供应商策划的客户故事;它们证明具名账户确有真实部署和结果,但不证明完整客户群的典型体验。[CU016, CU017, CU018, CU019, CU020, CU021]
| 客户 | 部署 / 用例 | 生产环境与试点 | 量化结果 | 背书质量 | 局限 |
|---|---|---|---|---|---|
| Johnson & Johnson Vision | GenAI/LLM 培训、黑客松、通用分析平台 | 生产平台 + 原型活动 | 650+ 名员工使用 Dataiku;<2 天搭出可运行原型 | 高(具名高管、指标、直接案例) | 案例研究更强调赋能,而不是合同经济性 |
| Novartis | 医疗健康市场研究聊天机器人和预测自动化 | 生产环境 / 内部规模化使用 | 洞察产出速度提升 90%;数据摄取速度提升 600% | 高 | 供应商筛选过的成功案例;未披露席位数或支出 |
| Michelin | 工厂 AI、质量管理、时间序列副驾驶 | 规模化生产环境 | 50+ 家工厂的 1,500+ 名用户;分析时间从最长 6 个月缩到约 1 小时 | 高 | 无收入或续约细节 |
| Euronext | 市场分析助手智能体 | 生产工作流部署 | 经常性查询时间最多缩短 20% | 高 | 收益围绕节省时间展开,而非财务 ROI |
| Geodis | ServiceNow 中的 AI IT 支持智能体 | 早期生产环境 / 运营推广 | 分派速度提升 60%;每张工单节省约 30 分钟 | 高 | 案例研究暗示推广在推进,但不能证明全企业饱和 |
| Roche | 面向律师的专利研究和智能体 AI | 已进生产并在扩张 | 每年节省律师时间价值 $100K-$250K;避免咨询支出 $375K-$475K | 高 | 法务团队用例未必能外推到 Roche 更广泛采用 |
| Prologis | 企业 AI/ML 与 GenAI 平台标准化 | 规模化生产环境 | 60+ 个 AI/ML 项目已进生产;30 个项目 + 30 个 API 活跃;2,000 名用户 | 高 | 未披露收入或合同期限 |
| SLB | 井筒建造、油藏分析、HR 留存 | 多职能生产环境 | > $10B 投标项目完成评估;投标分析速度提升 25x;压力分析速度提升 76% | 高 | 用例多个,但没有统一采用分母 |
这些行优先选择有具体工作流和结果的具名部署;并不枚举 Dataiku 的完整客户基础。
[CU016, CU017, CU018, CU019, CU020, CU021]具名公开参考案例的证据质量不一,但多个旗舰账号显示了生产状态、量化成效和可识别的高管背书。
该矩阵只评估公开标杆的质量;不代表这些客户就是 Dataiku 最大或最赚钱的客户。
[CU017, CU018, CU020, CU030, CU037]6.4 扩张动作与生态杠杆
公开材料中可见的商业动作,是通过平台标准化、业务用户赋能和合作伙伴辅助的企业转型来先落地再扩张。案例研究反复从单个运营痛点开始,但最终走向更广的受治理分析和 AI 采用。Michelin 从数字化制造起步,随后扩展到工厂、R&D 和公司职能。Prologis 从描述性分析扩展到地理空间分析、预测建模和企业 GenAI。Roche 从专利搜索试点起步,再把多个 GenAI 项目整合进一个代理式界面。KPMG 联盟说明 Dataiku 也通过现代化和治理项目销售,而不只是直接软件采购;2025 年 Frontrunner Awards 则显示,客户因多个行业中的代理式 AI、治理和生产率用例而获认可。这套扩张逻辑重要,因为它说明 Dataiku 最好的账户之所以粘性强,不只因为模型进入生产,更因为平台变成了组织工作流、培训和治理层。仍未知的是,有多少销售管线来自合作伙伴,多少收入受市场平台或服务影响,以及扩张是在长尾中广泛发生,还是集中在少数旗舰账户。[CU026, CU027, CU028, CU029, CU030, CU035]
| 驱动因素 / 风险 | 方向 | 公开证据 | 含义 | 尽调路径 |
|---|---|---|---|---|
| 平台标准化 | 正向 | 客户常把工作流、治理和 AI 项目整合到一个平台 | 支撑账户内多团队扩张 | 衡量扩张账户贡献的 ARR 占比 |
| 业务自助赋能与培训 | 正向 | Standard Chartered 和 J&J 显示由培训带动扩散 | 业务用户上手后,Dataiku 更难被替换 | 索取各账户活跃用户随时间变化 |
| 伙伴辅助渠道 | 正向 / 风险 | KPMG 联盟和以伙伴为中心的转型表述显示渠道杠杆 | 可加速采购,但可能掩盖服务依赖 | 披露伙伴来源管线和服务收入结构 |
| 背书账户集中度 | 风险 | 大多数公开证据来自少数旗舰企业 | 可能意味着收入集中在头部客户 | 索取前 10 和前 20 大客户 ARR 占比 |
| 实施复杂度 | 风险 | 负面评论标题提到云集成问题 | 大型部署可能价值兑现更慢,或推广失败 | 索取实施周期和扩张转化数据 |
| 留存不透明 | 风险 | 未公开 NRR、GRR、流失或合同期限 | 仅凭案例研究无法充分验证耐久性 | 索取续约队列和降级明细 |
扩张在定性层面可见,但集中度和伙伴依赖仍未解决,因为 Dataiku 不披露客户经济性细节。
[CU026, CU027, CU028, CU034, CU035, CU040]6.5 耐久性、留存与反向解读
客户质量的公开证据,最强在采用广度和结果轶事,最弱在续约机制。Dataiku 不公开披露净收入留存、总留存、流失、合同期限、头部客户集中度,或前十大账户贡献的 ARR 份额。Gartner 2026 年评论页方向上偏正面,本次抓取页面显示五星占 75%、四星占 23%;Dataiku 2025 年 1 月发布也引用了 96% 的 Gartner 愿意推荐得分。但同一 Gartner 页面也包含一条批评性评论标题,描述云集成问题;这很重要,因为困难的企业集成正是 AI 平台落地容易卡住的地方。FeaturedCustomers 和 TrustRadius 确认了可见的评论与案例研究语料,但不能解决续约或集中度这些核心承销问题。公平解读是,Dataiku 在成功企业部署中很可能享有有意义的转换成本,因为工作流、治理和非技术用户习惯会随时间积累;不过,在私有留存分组和集中度表披露前,这种耐久性仍只能给中等置信度。[CU031, CU032, CU033, CU034, CU036, CU037]
| 指标 / 信号 | 数值 / 状态 | 置信度 | 重要性 | 尽调追问 |
|---|---|---|---|---|
| 净收入留存 | 未公开披露 | 低 | 检验先落地再扩张耐久性的核心指标 | 索取按细分和地区拆分的 NRR |
| 毛留存 / 流失 | 未公开披露 | 低 | 需要区分扩张和 Logo 流失 | 索取队列流失桥表 |
| 合同期限 / 续约节奏 | 未公开披露 | 低 | 企业 AI 软件看似粘性强,但仍可能在压力下按年续约 | 索取标准合同条款和续约率 |
| 满意度信号 | 方向上正面 | 中 | Gartner 页面和公司引用的推荐意愿暗示客户满意 | 用原始调查样本和队列拆分验证 |
| 重复使用 / 内部广度 | 在具名客户可见 | 中 | J&J、Michelin、Prologis、Standard Chartered 和 Roche 显示多用户或多项目深度 | 梳理前 20 大客户的活跃用户增长 |
| 实施摩擦 | 真实存在但未量化 | 中 | 一条批评性 Gartner 评论标题提到私有云集成问题 | 索取实施失败、延期和回滚率 |
| 可引用客户基础 | 范围广但经过筛选 | 中 | FeaturedCustomers 和 TrustRadius 显示评论量可见,但不能说明续约真相 | 要求投资人自行挑选独立客户背书 |
公开留存证据大多是间接的;正面信号来自部署深度和评论平台,但实际续约数据仍缺失。
[CU031, CU032, CU033, CU034, CU036, CU037]公开证据在部署案例上最强,在续约和集中度上最弱。
该矩阵评估证据可见度,不评估表现质量。「独立佐证」指来自非供应商来源的佐证;在留存和集中度上,这类证据有限。
[CU031, CU032, CU033, CU034, CU040]07风险
7.1 监管、隐私与法律风险
Dataiku 的法律风险与其说来自某个可见诉讼,不如说来自它身处不断扩大的合规边界。按私营软件标准,公司法律和隐私表面较成熟:当前隐私政策、云条款栈、2026 年 DPA,以及明确讨论隐私保护设计、负责任 AI 和集成管理体系的信任页面。这些是真实缓释因素,尤其有利于企业采购。它们也是真实义务。隐私政策称,Dataiku 对其在该政策下收集的数据担任控制者;云 DPA 则把 Dataiku 定义为 SaaS 场景中客户个人数据的处理者,并把运营连接到 GDPR、CCPA、SCCs 和事件处理。监管背景只会更难,不会更容易。EU AI Act 现在禁止某些做法,对高风险系统施加严格义务,并将在 2026 年 8 月让生成式 AI 透明度规则生效。NIST 的 AI RMF 和 GenAI profile 也强化了市场预期:企业 AI 供应商必须把可追溯性、监督和风险管理运营起来,而不是只拿来营销。因此,主要法律风险不是 Dataiku 已知遭遇公开执法,而是广泛企业部署、第三方模型供应商链条,或客户在敏感工作流中误用,可能让公司面对更慢销售周期、更高赔偿谈判成本,或未来监管审查。[CR001, CR002, CR003, CR004, CR005, CR006]
| 风险 | 法域 / 规则 | 状态 | 可能性 | 严重性 | 缓释措施 | 剩余敞口 | 尽调路径 |
|---|---|---|---|---|---|---|---|
| AI 治理合规缺口 | EU AI Act / 客户行业规则 | 规则分阶段生效;透明度义务延续至 2026 年 8 月 | 中 | 高 | 治理优先的产品定位;信任计划;文档 | 高,因为客户可将 Dataiku 部署到敏感工作流 | 测试高风险用途中的智能体治理、日志、人类监督和客户指引 |
| 隐私 / 数据处理错误 | GDPR、CCPA、SCC、DPA 义务 | 持续 | 中 | 高 | 隐私政策、DPA、处理方条款、事件表述 | 中-高,原因是跨境处理和 AI 提供商链条 | 审查 DPA 红线、子处理方名单、数据驻留控制和事件流程 |
| 第三方 AI 提供商责任外溢 | 合同与隐私义务 | 当前 | 中 | 中-高 | 客户指令、合同责任分配、未经同意不训练声明 | 中,因为启用的 AI 服务可能把内容路由给第三方 | 审查 AI 服务条款、特定提供商隐私通知和退出控制 |
| 营销 / AI 声称审查 | FTC / 消费者保护与不公平行为背景 | 持续 | 低-中 | 中 | 采购驱动的企业销售降低零售营销风险 | 中,因为企业 AI 声称正受到更多审查 | 审查治理、可解释性和安全声称的依据 |
| 未披露或潜在执法敞口 | 公开执法追踪器和案例库 | 已抓取来源中未浮现明显匹配 | 低 | 中 | 未浮现公开罚款;但法律暴露面已经成熟 | 未知,因为未出现在追踪器中不等于不存在 | 由律师主导检索诉讼、执法和投诉数据库 |
各行按对收入和估值的实际风险排序,而非法律新颖性。最大威胁是监管负担在客户部署内部扩张,而不是今天已有的头条案件。
[CR001, CR002, CR003, CR004, CR007, CR008]Dataiku 最高的残余风险,不是单一生死线,而是监管范围、安全补丁执行力、客户与财务不透明叠在一起。
该热力图是基于公开证据综合出的投资视角,不是管理层发布、带校准概率的风险登记册。
[CR007, CR014, CR024, CR026, CR038, CR040]7.2 安全与运营可靠性风险
Dataiku 不是狭窄功能产品,因此运营风险画像偏高。它横跨数据访问、notebook、API、自动化、模型部署、治理,以及现在的代理式 AI 编排,攻击面天然很宽。OpenCVE 显示 Dataiku DSS 曾有重要已披露漏洞,包括 2025 年 6 月披露的一个 9.8 分关键认证绕过问题,以及一系列较早的访问控制和信息泄露问题。这本身不能证明安全文化薄弱——每个成熟企业平台都必须修补问题——但确实证明补丁纪律至关重要。产品的运营形态还带来分裂的风险模型。在自管部署中,Dataiku 称默认不处理或存储客户数据,这降低了供应商数据保管暴露,但也把更多配置、可用性和安全负担转移到客户环境。在 Dataiku Cloud 中,公司成为事件处理和子处理方管理的更直接控制点,同时也继承云提供商基础设施风险。Gartner 评论内容补充了另一个实务风险:抓取页面上至少一条批评性评论标题指向私有云集成摩擦。这一点重要,因为常常不是软件缺陷,而是实施复杂度,把一个技术上可靠的平台变成商业上痛苦的落地。[CR014, CR015, CR016, CR017, CR018, CR019]
| 失效模式 | 可能性 | 严重性 | 缓释成熟度 | 剩余敞口 | 未解决缺口 |
|---|---|---|---|---|---|
| 关键软件漏洞或认证绕过 | 中 | 高 | 中 | 重大 | 需要补丁 SLA、客户升级节奏和事件复盘证据 |
| 私有云或混合集成摩擦 | 中 | 中-高 | 中 | 重大 | 需要实施失败率和扩张转化指标 |
| 云提供商宕机或控制失效,影响 Dataiku Cloud | 低-中 | 高 | 中 | 重大 | 需要架构、冗余和事件沟通细节 |
| 自托管客户配置错误导致责任转移 | 中 | 中 | 低-中 | 重大 | 需要支持边界、参考架构和升级工具 |
| 生产业务流程中的智能体 / 模型工作流故障 | 中 | 高 | 中 | 重大 | 需要护栏覆盖、回退模式和客户回滚数据 |
运营风险既是传统软件安全风险,也是企业落地风险;后者对续约同样关键。
[CR014, CR015, CR016, CR017, CR018, CR019]7.3 依赖、客户与执行风险
Dataiku 的 GTM 依赖本身也是风险传导通道。KPMG 联盟说明,公司能嵌入更大的现代化和治理项目里拿单;这有利于放大规模, 但也让合作伙伴质量和服务经济性更关键。客户案例反复强调,它能与 Snowflake、Azure、ServiceNow、PowerBI 等系统互通。 互操作性是优势,但连接器故障、云政策变化或模型供应商争端,都可能传导成客户不满。客户基础质量高,却很难尽调定价。 公开证明覆盖医疗、银行、物流、制造和资本市场;受监管买方重视治理,战略上很有吸引力。但运营上也危险:控制、文档或事件响应一旦短板, 这类买方不会宽容。公开来源仍未披露头部客户集中度、合作伙伴来源销售管线或 NRR。丰富证据大多来自供应商筛选过的标杆案例, 而不是独立的分群披露。人员和执行层面,公司已扩张到 1,250 多名员工,并在筹备潜在 IPO 的同时补入资深商业领导层。 纪律性可能因此改善,但产品、客户成功或合作伙伴管理的执行滑坡,也更可能在最不合适的时间点暴露出来。[CR021, CR022, CR023, CR024, CR031, CR033]
| 依赖 | 对手方 / 类别 | 作用 | 集中度 | 失效情境 | 严重性 | 缓释措施 | 剩余敞口 |
|---|---|---|---|---|---|---|---|
| 云基础设施 | 底层云服务商 | 承载 Dataiku Cloud,并影响可用性 / 安全边界 | Unknown | 宕机、政策变化或定价调整损害服务经济性 | 高 | 多环境部署选项 | 中-高 |
| 系统集成商渠道 | KPMG 等类似合作伙伴 | 企业现代化和实施杠杆 | Unknown | 合作伙伴交付不达标或截留经济收益 | 中 | 直销叠加合作伙伴生态 | 中 |
| 第三方 AI 服务商 | 模型供应商 | 处理 AI 服务内容,并支撑 GenAI 功能 | Unknown | 服务商宕机、政策变化或数据使用疑虑打断工作流 | 高 | 模型无关定位和客户侧控制 | 中-高 |
| 外部数据平台和应用 | Snowflake、Azure、ServiceNow、PowerBI 等 | 连接器和工作流目的地 | 客户层面高 | 集成中断会拖慢扩张或导致流失 | 高 | 广泛连接器覆盖和共享工作流 | 中-高 |
| 标杆企业客户 | 大型参考客户 | 收入、验证和 IPO 叙事 | 未披露 | 少数大客户停滞或降级 | 高 | 客户 logo 数量多,但权重未知 | 未知-高 |
最大的依赖风险不是单一供应商,而是生态系统;一旦失灵,问题会在客户价值兑现中暴露。
[CR021, CR022, CR023, CR031, CR033, CR038]| 角色 / 职能 | 依赖或缺口 | 可能性 | 严重性 | 缓释措施 | 尽调路径 |
|---|---|---|---|---|---|
| 客户成功 / 实施 | 需要把试点和迁移转成规模化续约 | 中 | 高 | 合作伙伴杠杆和可复用模板 | 要求提供实施周期、价值实现时间和扩张转化数据 |
| 产品 / 安全工程 | 必须修补漏洞,并在不引入回归的情况下发布新的智能体 / 治理功能 | 中 | 高 | 正式安全计划和认证 | 要求提供漏洞管理指标和人员深度 |
| 销售与合作伙伴领导层 | 新领导层和广泛生态必须拿出有纪律的后期增长 | 中 | 中-高 | 规模和品牌动能 | 要求提供配额达成、合作伙伴贡献管线和分群流失数据 |
| 监管 / 信任运营 | 必须跟上 AI 治理、隐私和审计预期 | 中 | 中-高 | IMS、信任页面和隐私内建姿态 | 要求提供审计日程、政策例外和整改积压 |
| 处在 IPO 审视下的高管团队 | IPO 准备度越高,任何失误都会更显眼 | 中 | 高 | ARR 规模和投行参与 | 审阅董事会材料、准备度里程碑和内控计划 |
Dataiku 已经过了只靠产品契合度就能抵消流程不一致的创业阶段,因此执行风险抬高。
[CR024, CR025, CR026, CR037, CR039]Dataiku 的投资逻辑靠云、模型供应商、集成商和旗舰企业客户共同撑着;任一层失效,投资论点都会变弱。
[CR021, CR022, CR023, CR034, CR038]7.4 财务不透明与 IPO 传导风险
最后一层风险在于传导:技术、监管或客户问题会怎样影响融资和估值。Reuters 报道,Dataiku 于 2025 年 10 月聘请 Morgan Stanley 和 Citigroup,为潜在美国 IPO 做准备;这意味着公司面对的隐含审视门槛,高于普通后期私营软件厂商。与此同时,公开证据仍未披露毛利率、 NRR、现金消耗或头部客户占比等核心承保字段。缺失字段很关键,因为它们决定 Dataiku 能不能扛住冲击。毛利强、现金深、扩张分散的公司, 能熬过交易窗口延后或一次实施失误;缺少这些缓冲的公司,会很快失去谈判筹码。因此,正确的风险框架不只是「IPO 窗口风险」。 安全缺陷、合规失误、合作伙伴摩擦或客户扩张停滞,都可能传导成续约假设下调、公开市场可比公司倍数下降,或可信上市前更长的私有持有期。 Dataiku 的管控姿态和产品贴合度是实质缓释因素,但剩余暴露仍高于普通 SaaS,因为平台直接嵌在企业 AI 治理和运营工作流里; 这类场景一旦失败,会很快传到信任、采购和估值。[CR024, CR025, CR026, CR030, CR039, CR040]
| 风险 | 可监控触发项 | 阈值 / 事件 | 行动含义 |
|---|---|---|---|
| 安全补丁纪律 | 未解决的关键漏洞 | 关键漏洞披露后 >30 天仍未修补,或客户大范围未升级 | 暂停投资判断,直到补丁节奏和客户升级覆盖得到证明 |
| 监管姿态 | 重大执法事件 | 涉及 Dataiku 或大型客户部署、且归因于 Dataiku 平台的 GDPR / FTC / 行业监管行动或同意令 | 立即重切风险评级和估值区间 |
| 客户韧性 | 续约 / 扩张恶化 | 私有 NRR 低于企业软件预期,或头部客户出现明显降级 | 建议转向等待 / 重定价 |
| 合作伙伴依赖 | 渠道集中 | 新 ARR 或服务交付中 >40% 依赖单一渠道伙伴或单一路云路径 | 对利润率路径和执行信心打折 |
| 实施复杂度 | 价值实现慢 | 平均企业实施周期显著长于管理层指引,或扩张转化下滑 | 在证伪前,把客户案例视为不具代表性 |
| IPO 准备度 | 融资延误 | IPO 时间表延后,同时增长或治理指标走弱 | 假设私有持有期更长,下轮估值下调风险更高 |
这些终止标准刻意设计成可监控项;它们把宽泛担忧转成足以改变投资决策的阈值。
[CR024, CR025, CR026, CR030, CR038, CR040]安全、监管和落地一旦失手,可能先打到信任与扩张,再拖累估值和融资选择权。
[CR024, CR025, CR026, CR030, CR040]08估值
8.1 投资论点与反论点
Dataiku 的多头逻辑很直接。公司已有真实企业级规模、已披露 ARR 里程碑、蓝筹客户名单、契合当前企业 AI 时点的重治理产品叙事, 以及一个尚未像部分 AI 公司那样明显跑在收入前面的私募估值锚。2025 年 1 月,Dataiku 披露 $300M+ ARR 和 700+ 客户; 2025 年 10 月,公司披露 $350M+ ARR、750+ 客户和 1,250+ 员工。这是真实的运营体量。反论点不是公司没有业务牵引力, 而是投资者仍不够了解牵引力背后的经济性。毛利率、NRR、现金消耗、服务交付强度和头部客户集中度仍未披露,同时 Databricks 和超大云厂商的竞争还在加剧。如果业务比公开叙事更重服务、更依赖渠道,或粘性更弱,公允估值会很快变成高估。[CV001, CV002, CV003, CV004, CV005, CV006]
| 维度 | 多头论点 | 空头反论点 |
|---|---|---|
| 规模 | 350M+ ARR 和 750+ 客户支撑品类领导地位 | 只看规模、不披露利润率或 NRR,可能误判价值质量 |
| 产品 | 治理优先的 AI 平台契合企业需求 | Databricks 和超大云厂商正在复制治理和智能体界面 |
| 客户 | 蓝筹企业客户名单显示需求契合度持久 | 客户案例经过筛选,集中度未披露 |
| 估值 | 过时的 $3.7B 估值标记目前只对应约 10-12x ARR | 私有资产流动性差、信息不透明,应较高溢价上市可比公司打折 |
| 退出可选性 | Reuters 报道的 IPO 准备保留了公开市场路径 | 披露或市场窗口走弱时,IPO 时间可能延后 |
| 证据质量 | 多个公开证据点支撑叙事 | 关键经济性仍需私有尽调 |
投资论点主要取决于 Dataiku 更接近高溢价、软件质量的基础设施,还是更慢、更偏服务的企业平台。
[CV001, CV003, CV007, CV020, CV021, CV022]建议一边来自真实经营规模和客户证据,另一边受经济性缺口与执行风险牵制。
[CV004, CV006, CV020, CV027, CV028]8.2 融资背景与公开可比公司桥接
Dataiku 最近一次披露的私募估值,仍是 2022 年 12 月 Series F 的 $3.7B。此后,公开运营里程碑继续改善,但估值锚没有公开重设; 这不常见,却有分析价值。按上次披露估值计算,Dataiku 在 2025 年 1 月 $300M ARR 里程碑上约为 12.3x ARR,在 2025 年 10 月 $350M ARR 里程碑上约为 10.6x ARR;Sacra 对 2025 年 9 月约 $342.5M 的估算,对应约 10.8x。这些倍数不便宜, 但也没到公开市场顶级 AI 或数据平台资产的峰值区间。公开市值 / 收入代理指标区间很宽:Palantir 约 58.2x、Datadog 25.0x、 Snowflake 19.4x、MongoDB 11.2x、Confluent 9.6x、ServiceNow 8.0x、C3.ai 4.6x。因此,Dataiku 滞后的私募倍数 更接近公开区间中部,而不是两端。估值问题不是公司是否配得上溢价,而是一家私营、不透明、面向企业治理的 AI 平台, 应该落在这个宽区间的哪里。[CV001, CV002, CV003, CV004, CV005, CV006]
| 可比对象 | 指标口径 | 估值 / 市值 | 收入口径 | 市值 / 收入代理倍数 | 适配度 / 局限 |
|---|---|---|---|---|---|
| Dataiku | 私有估值标记对比官方 ARR | $3.7B | 350M+ ARR(Oct 2025) | ~10.6x | 最接近的锚点,但过时且私有 |
| Dataiku | 私有估值标记对比官方 ARR | $3.7B | 300M+ ARR(Jan 2025) | ~12.3x | 显示收入扩张后倍数压缩 |
| Dataiku | 私有估值标记对比 Sacra 估算 | $3.7B | ~342.5M ARR(Sep 2025 估) | ~10.8x | 估算值,未审计 |
| Palantir | 公开市值 / TTM 收入 | $303.95B | $5.22B | ~58.2x | 高溢价 AI / 政府软件异常值;公开度和流动性高得多 |
| Datadog | 公开市值 / TTM 收入 | $91.67B | $3.67B | ~25.0x | 高质量基础设施软件的上沿参照 |
| Snowflake | 公开市值 / TTM 收入 | $90.61B | $4.68B | ~19.4x | 消费驱动的高溢价数据平台 |
| MongoDB | 公开市值 / TTM 收入 | $27.51B | $2.46B | ~11.2x | 披露更成熟的成长型软件公司 |
| ServiceNow | 公开市值 / TTM 收入 | $111.08B | $13.96B | ~8.0x | 具备规模和盈利能力的工作流软件基准 |
| Confluent | 公开市值 / TTM 收入 | $11.13B | $1.16B | ~9.6x | 相比顶级 AI 公司溢价更低的基础设施软件 |
| C3.ai | 公开市值 / TTM 收入 | $1.39B | $0.30B | ~4.6x | AI 软件经济性较弱时的反向下限可比 |
这些是来自公开市值和收入页面的市值 / 收入代理倍数,不是干净的企业价值倍数。覆盖面仍足以把 Dataiku 放进宽泛的上市区间。
[CV004, CV005, CV006, CV010, CV011, CV012]上市可比公司的市值 / 收入代理倍数显示,Dataiku 的估值判断会随选取的参考组快速变化。
这些是由公开市值和收入页面推算的市值 / 收入代理倍数,不是按净现金或债务调整后的企业价值倍数。
[CV010, CV011, CV012, CV013, CV014, CV015]8.3 情景估值与建议
我们的基准情景假设,2025 年 10 月披露的 $350M+ ARR 在方向上足够接近当前,可作为估值锚;但在毛利、留存和现金效率证据改善前, 投资者仍应给出私营公司折扣。基于此,公允的基准估值区间约为 $3.5B-$4.5B,整体接近或略高于上次披露的 $3.7B。 如果 Dataiku 能证明可信 AI 治理、750+ 企业客户和 IPO 准备度能转化为持久扩张和软件式毛利,多头情景可支持约 $5.0B-$6.5B。 如果增长放缓、留存不及预期,或公开市场投资者因不透明和实施风险给出更重折扣,空头情景约为 $2.5B-$3.2B。 因此,当前公开证据支持「跟踪」建议、中等置信度、中高风险和公允估值立场。公司真实到不能忽视,也不透明到不该以溢价激进承保。[CV018, CV019, CV023, CV024, CV025, CV026]
| 维度 | 评估 | 依据 |
|---|---|---|
| 建议 | 跟踪 | 规模和客户验证强,但经济性披露仍不足 |
| 置信度 | 中 | 估值依赖过时私有估值标记和上市可比公司代理指标 |
| 风险评级 | 中-高 | 安全、监管和财务不透明风险仍然重大 |
| 估值立场 | 合理 | 当前 $3.7B 估值标记落在基准情景区间内 |
| 综合评分 | 7.4 / 10 | 业务质量好,但投资判断数据不完整 |
| 入场纪律 | 要求更新经济性 | 溢价上行需要利润率和留存证明 |
建议本身对价格和证据都敏感,不是泛泛的业务质量评分。
[CV026, CV027, CV028, CV040]| 情景 | 概率信号 | 核心假设 | 估值区间 | 相对 $3.7B 估值标记回报 |
|---|---|---|---|---|
| 乐观 | ~25% | 增长保持强劲,可信 AI 护城河守住,公开披露支持溢价倍数 | $5.0B-$6.5B | +35% to +76% |
| 基准 | ~50% | 当前 ARR 规模真实,但经济性只算中等软件化 | $3.5B-$4.5B | -5% to +22% |
| 悲观 | ~25% | 增长放缓,留存或利润率不及预期,IPO 路径延后 | $2.5B-$3.2B | -32% to -14% |
区间是基于上次披露估值、当前 ARR 披露和上市可比公司代理指标锚定的情景估算;不是管理层指引。
[CV023, CV024, CV025, CV026, CV031, CV034]情景区间说明,当前估值看起来公允,但谈不上明显便宜。
区间为作者估算,锚定最近披露估值、当前 ARR 披露和上市可比公司代理倍数。
[CV023, CV024, CV025, CV026, CV031, CV034]IC 式汇报的核心指标。
KPI 汇总投资建议和估值分析;ARR 倍数使用最近披露的 $3.7B 估值与 2025 年 10 月 ARR 里程碑。
[CV026, CV027, CV028, CV040]8.4 最终尽调与论点失效触发项
要给出明确买入判断,剩下的不是理念问题,而是材料问题。投资者需要截至 runDate 的当前 ARR 桥接、毛利率拆分、NRR 和流失分群、 头部客户集中度、合作伙伴来源销售管线、服务组合、现金和消耗,以及未来融资或 IPO 时的优先权结构。这些问题重要, 因为倍数争论已经不再是学术题。Dataiku 已经足够大,主导风险的不再是需求缺失,而是经济性缺失。若留存明显弱于客户故事暗示, 若安全或监管问题打断 IPO 时间表,若 ARR 增速低于 $3.7B+ 估值锚所隐含的预期,或若合作伙伴依赖被证明是在弥补直接产品杠杆不足, 论点就会失效。在这些问题得到回答前,正确姿态是有纪律地贴近跟踪,而不是给出最高确信度。二级市场或跨界投资者也不应把没有坏消息, 等同于优质经济性的证据。[CV028, CV032, CV036, CV037, CV038, CV039]
| 触发器 | 阈值 / 事件 | 对论点的传导 | 行动含义 |
|---|---|---|---|
| 留存不足 | 私有 NRR 或续约分群显著低于企业软件常模 | 削弱落地扩张和溢价倍数逻辑 | 下调为等待 / 重定价 |
| 毛利率偏弱 | 软件毛利率显著低于强平台预期 | 暗示服务或托管拖累大于市场叙事 | 转向悲观情景 |
| 安全 / 监管事件 | IPO 进程前后出现重大泄露、执法行动或重大控制失效 | 同时打击信任、采购和退出时间 | 下调公允价值区间 |
| 增长减速 | 更新后的 ARR 桥显示增长慢于 2025 年披露所隐含水平 | 削弱溢价倍数支撑 | 用较低区间可比公司重新做投资判断 |
| 客户集中 | 头部客户占比或合作伙伴依赖过高 | 收入质量不如 logo 数量暗示的那样持久 | 施加流动性和集中度折价 |
| IPO 延误 | 披露仍弱时 IPO 路径延后 | 拉长私有持有期,并抬高估值标记风险 | 继续跟踪,不支付溢价 |
这些触发器设计成能在尽调、季度更新或未来公开文件中观察到,而不是停留在模糊的定性担忧。
[CV033, CV034, CV037, CV038, CV039, CV040]| 主题 | 缺失证据 | 重要性 | 负责人 / 路径 |
|---|---|---|---|
| ARR 桥 | Oct 2025 至 runDate 的月度或季度 ARR | 确认增长是否延续或显著放缓 | 公司财务团队 |
| 毛利率 | 软件、托管和服务毛利率拆分 | 判断 Dataiku 落在上市可比公司区间的哪个位置 | 公司财务团队 |
| NRR / 流失 | 分群留存和降级历史 | 区分真实平台韧性和筛选后的案例 | 公司 RevOps |
| 现金与消耗 | 现金余额、消耗、跑道和融资计划 | 检验 IPO 时间延后时的下行韧性 | 公司财务团队 / 董事会 |
| 客户集中度 | 前 10 大客户收入占比和分部组合 | 验证 750+ 客户暗示的覆盖广度 | 公司销售 / 财务团队 |
| 合作伙伴经济性 | 合作伙伴来源预订额和服务附加率 | 判断合作伙伴杠杆是在提升还是损害增长质量 | 公司合作伙伴团队 |
| 优先权结构 | 当前清算优先权和摊薄悬顶 | 影响后期入场经济性 | 公司法务 / 财务团队 |
这些是从跟踪立场推进到明确价格判断所需的最低缺失数据。
[CV028, CV032, CV036, CV040]8.5 附录
免责声明
本报告仅供参考,反映截至 2026-07-13 可获得的公开来源,不构成投资建议。任何投资决策前,都应独立核验私营公司估值、ARR 数据和可比倍数桥接。
证据索引
| 编号 | 陈述 | 可信度 | 来源 |
|---|---|---|---|
| CO001 | Dataiku was founded in 2013. | 高 | SO001, SO010 |
| CO002 | Dataiku was founded by Florian Douetteau, Clément Stenac, Thomas Cabrol, and Marc Batty. | 高 | SO001, SO015 |
| CO003 | Independent profiles describe Dataiku as beginning in Paris before later scaling in the United States. | 中 | SO017, SO016 |
| CO004 | Dataiku is headquartered in New York, New York. | 高 | SO010, SO009 |
| CO005 | Gartner's 2026 company profile still categorizes Dataiku as a private company. | 中 | SO010 |
| CO006 | Florian Douetteau is Dataiku's co-founder and CEO. | 高 | SO001, SO005 |
| CO007 | Clément Stenac is Dataiku's co-founder and CTO. | 高 | SO001, SO015 |
| CO008 | Dataiku is positioned as an enterprise orchestration layer for analytics, machine learning, and AI agents. | 中 | SO002, SO005, SO015 |
| CO009 | The platform is designed for no-code, low-code, and full-code collaboration between business and technical users. | 中 | SO002, SO009 |
| CO010 | Dataiku emphasizes cloud and model optionality rather than locking customers into one infrastructure stack. | 中 | SO002, SO023 |
| CO011 | An archived official plans page shows Dataiku offered hosted SaaS, customer-cloud, and on-premises deployment options. | 中 | SO019 |
| CO012 | The archived enterprise plan highlighted unlimited instances, full deployment capabilities, and resource governance. | 中 | SO019, SO020 |
| CO013 | Dataiku raised $200 million in a Series F round in December 2022. | 高 | SO001, SO006 |
| CO014 | The Series F round valued Dataiku at $3.7 billion. | 高 | SO001, SO006, SO009 |
| CO015 | Wellington Management led Dataiku's 2022 Series F financing. | 中 | SO003, SO006 |
| CO016 | Sacra estimates Dataiku's total funding at about $846.8 million. | 中 | SO008, SO009 |
| CO017 | Zippia reports a $101 million Series C in December 2018. | 中 | SO016 |
| CO018 | Zippia reports CapitalG joined in December 2019 when Dataiku reached unicorn status at a $1.4 billion valuation. | 中 | SO016 |
| CO019 | Zippia reports Dataiku raised an additional $100 million Series D in August 2020 led by Stripes and Tiger Global. | 中 | SO016 |
| CO020 | Zippia reports Dataiku raised $400 million in August 2021 at a $4.6 billion valuation. | 中 | SO016 |
| CO021 | Dataiku's January 2025 press release said ARR had surpassed $300 million. | 高 | SO003, SO009 |
| CO022 | The January 2025 release said Dataiku had grown ARR by more than 2x over the prior three years. | 中 | SO003 |
| CO023 | The January 2025 release said Dataiku served more than 700 organizations worldwide. | 中 | SO003, SO012 |
| CO024 | The January 2025 release said Dataiku employed more than 1,100 people worldwide. | 中 | SO003 |
| CO025 | The January 2025 release said more than 20% of customers were already using Dataiku for GenAI workflows. | 中 | SO003 |
| CO026 | The January 2025 release said Dataiku ranked No. 34 on the 2024 Forbes Cloud 100. | 中 | SO003 |
| CO027 | Dataiku's October 2025 release said ARR had surpassed $350 million. | 中 | SO004, SO008 |
| CO028 | The October 2025 release said Dataiku served more than 750 organizations worldwide. | 中 | SO004, SO005 |
| CO029 | The October 2025 release said Dataiku employed more than 1,250 people across 13 offices and remote locations. | 中 | SO004 |
| CO030 | The October 2025 release said Dataiku was trusted by one in four of the world's top companies in the 2024 Forbes Global 2000. | 中 | SO004 |
| CO031 | June 2026 leadership changes added Maxwell Long as President and CRO overseeing sales, customer success, and partnerships. | 中 | SO005 |
| CO032 | Maxwell Long arrived from Smartsheet after helping scale it past $1 billion ARR and through its $8.4 billion acquisition. | 中 | SO005 |
| CO033 | By June 2026 Dataiku highlighted Agent Management, Cobuild, and Reasoning Systems as new enterprise AI products. | 中 | SO005, SO002 |
| CO034 | The Snowflake partnership page says Dataiku and Snowflake have helped 300+ customers operationalize AI together. | 中 | SO022 |
| CO035 | Dataiku's Snowflake blog says it was Snowflake's #1 AI/ML partner by platform consumption and marketplace revenue in 2025. | 中 | SO013 |
| CO036 | KPMG and Dataiku announced a 2024 alliance to modernize analytics and accelerate enterprise AI adoption. | 中 | SO014 |
| CO037 | Reuters reported in October 2025 that Dataiku hired Morgan Stanley and Citigroup to prepare for a U.S. IPO. | 高 | SO006, SO007 |
| CO038 | Reuters reported the IPO could come as soon as the first half of 2026, but timing remained subject to change. | 高 | SO006, SO007 |
| CO039 | The latest accessible public evidence shows IPO preparation rather than a completed listing. | 中 | SO006, SO010 |
| CO040 | FeaturedCustomers lists 182 reviews/testimonials, 150 case studies, and 63 customer videos for Dataiku. | 中 | SO011 |
| CO041 | Dataiku's customer directory highlights reference customers across finance, pharma, manufacturing, logistics, retail, and energy. | 中 | SO012 |
| CO042 | Gartner's 2026 product page includes a critical review citing private-cloud integration issues despite praising low-code AI strengths. | 中 | SO010 |
| CO043 | SiliconANGLE argues enterprises remain years away from broad agentic AI deployment because data, integration, and governance foundations are still missing. | 中 | SO025 |
| CO044 | Dataiku's partner directory highlights an ecosystem spanning Snowflake, Databricks, AWS, Google Cloud, NVIDIA, Accenture, and many regional integrators. | 中 | SO021 |
| CO045 | The Google Cloud partner page says Dataiku integrates with BigQuery, Vertex AI, Gemini-family models, and governance controls for GenAI. | 中 | SO023 |
| CO046 | The NVIDIA partner page says Dataiku can self-host open-source LLMs on NVIDIA GPUs through its LLM Mesh and NIM-related integrations. | 中 | SO024 |
| CO047 | FirstMark describes Dataiku as a collaboration layer connecting data repositories, algorithms, models, and people. | 中 | SO015 |
| CO048 | Zippia says Dataiku established itself in the United States in 2015. | 中 | SO016 |
| CO049 | PM Insights exposes only teaser-level valuation charts and cap-table references publicly, underscoring how limited open secondary-market visibility remains. | 低 | SO018 |
| CO050 | Public sources still leave board composition, cash balance, and current secondary pricing materially under-disclosed. | 中 | SO018, SO010, SO007 |
| CM001 | Dataiku participates at the intersection of enterprise data-science platforms, MLOps, AI governance, and agent-orchestration software rather than in one pure-play category. | 中 | SM001, SM011, SM013 |
| CM002 | The broadest comparable market lens is general machine-learning software, which Fortune Business Insights sizes at $65.28B in 2026. | 中 | SM018 |
| CM003 | A more expansive machine-learning market estimate from Precedence Research places 2026 spend at $126.91B, highlighting how category boundaries can materially widen TAM claims. | 中 | SM019 |
| CM004 | MarketsandMarkets sizes the narrower MLOps market at $5.9B by 2027 with a 41.0% CAGR. | 中 | SM017 |
| CM005 | MarketsandMarkets sizes the AI governance market at $5.78B by 2029 with a 45.3% CAGR. | 中 | SM017 |
| CM006 | MarketsandMarkets sizes the adjacent AI Studio market at $32.7B by 2029 with a 38.4% CAGR. | 中 | SM017 |
| CM007 | Fortune says large enterprises account for 55.61% of the machine-learning market in 2026. | 中 | SM018 |
| CM008 | Fortune says cloud deployment accounts for 53.14% of the machine-learning market in 2026. | 中 | SM018 |
| CM009 | Fortune says North America held a 32.5% share of the global machine-learning market in 2025. | 中 | SM018 |
| CM010 | Dataiku's public customer set spans life sciences, logistics, retail, manufacturing, energy, financial services, software, and technology. | 中 | SM001, SM004 |
| CM011 | Snowflake partnership materials explicitly position business and domain experts — not just technical specialists — as builders in this category. | 中 | SM006 |
| CM012 | The Google Cloud partner page positions Dataiku as a front-end for Vertex AI, BigQuery, and Gemini-backed governed application building. | 中 | SM007 |
| CM013 | The AWS partner page positions Dataiku as a collaborative visual layer on top of AWS ML, AI, and elastic cloud infrastructure. | 中 | SM008 |
| CM014 | The Databricks partner page positions Dataiku as a business-user layer on top of governed Databricks data. | 中 | SM009 |
| CM015 | Databricks markets a unified lakehouse-plus-governance stack and is a direct alternative for enterprises consolidating data engineering, analytics, and AI on one platform. | 中 | SM010 |
| CM016 | SageMaker markets a unified studio for generative AI, model training, AI ops, governance, lineage, and lakehouse analytics inside AWS. | 中 | SM011 |
| CM017 | Azure Machine Learning markets a centralized studio with a 99.9% SLA and pay-for-compute economics layered on the wider Azure stack. | 中 | SM012 |
| CM018 | Google Cloud's Agent Platform markets 200+ models, notebooks, pipelines, model registry, vector search, and usage-based pricing under one platform. | 中 | SM013 |
| CM019 | DataRobot emphasizes an all-in-one enterprise AI suite with on-prem, VPC, and SaaS deployment options. | 中 | SM014 |
| CM020 | H2O AI Cloud emphasizes managed cloud, hybrid cloud, AutoML, and no-code accessibility for enterprise ML teams. | 中 | SM015 |
| CM021 | Alteryx One represents a simpler analytics-and-automation adjacent substitute rather than a perfect like-for-like governed AI platform. | 中 | SM016 |
| CM022 | Dataiku's October 2025 press release frames the market shift as moving from AI experimentation to operationalization and trusted execution. | 中 | SM003 |
| CM023 | Dataiku's January 2025 press release said more than 20% of customers were already integrating GenAI into workflows, showing real but still partial production adoption. | 中 | SM002 |
| CM024 | Deloitte's July 2025 poll found only 13.5% of respondents were already using agentic AI in finance and accounting. | 中 | SM020 |
| CM025 | Deloitte found trust in the underlying data and programming was the top barrier to agentic AI adoption at 21.3% of responses. | 中 | SM020 |
| CM026 | SiliconANGLE argues 2025 will not be the year of broad enterprise agentic AI because data quality, integration, and governance gaps remain unresolved. | 中 | SM021 |
| CM027 | The 2026 arXiv interview study found seven of twelve companies were still only at the AI-assistant maturity level and only one had reached multi-agent orchestration. | 中 | SM022 |
| CM028 | The same arXiv study identified a capability-deployment verification gap where higher-level experimental AI cannot be trusted in production without human verification. | 中 | SM022 |
| CM029 | Observer argues early agentic AI deployments are primarily architectural projects whose attributable returns often take two to four years in complex environments. | 中 | SM025 |
| CM030 | Gartner review evidence shows even satisfied users can still encounter private-cloud integration friction, underscoring how hard deployment environments remain. | 中 | SM024 |
| CM031 | Dataiku's archived plans suggest the product can land with small teams but monetizes highest when governance, automation, and unlimited-scale deployment matter. | 中 | SM005 |
| CM032 | KPMG's 2024 alliance shows systems integrators remain important demand multipliers in the enterprise AI platform market. | 中 | SM023 |
| CM033 | Dataiku's public customer roster repeatedly surfaces pharma, financial services, industrials, logistics, insurance, and exchange operators. | 中 | SM004 |
| CM034 | The Snowflake partner page says Dataiku and Snowflake have helped 300+ customers operationalize AI, supporting a co-sell-led enterprise adoption model. | 中 | SM006 |
| CM035 | The real SAM for Dataiku should exclude raw cloud infrastructure spend, general-purpose LLM consumption, and lightweight point copilots that do not require governed multi-user orchestration. | 中 | SM010, SM011, SM013 |
| CM036 | Dataiku's fit is strongest where buyers want one governed control plane across multiple vendors rather than a single-cloud native toolchain. | 中 | SM001, SM007, SM008 |
| CM037 | Hyperscaler stacks pressure Dataiku on distribution and bundling, but also validate the demand for integrated AI development and governance. | 中 | SM011, SM012, SM013 |
| CM038 | The wide spread between broad ML market estimates and narrower MLOps/AI-governance estimates means later valuation work should use multiple market lenses rather than one headline TAM. | 中 | SM017, SM018, SM019 |
| CP001 | Dataiku competes across three practical classes of alternatives: unified data-and-AI platforms, hyperscaler-native ML stacks, and specialist AI or analytics-automation vendors. | 中 | SP001, SP008, SP012, SP014, SP016, SP018, SP021, SP023 |
| CP002 | Dataiku's product and partner surfaces position it as a neutral control layer that can sit across multiple clouds and partner ecosystems rather than as a single-vendor full stack. | 中 | SP001, SP003, SP005, SP006, SP007 |
| CP003 | Databricks is Dataiku's closest broad-platform rival because it combines data, governance, model lifecycle, and agent tooling in one integrated platform family. | 高 | SP008, SP009, SP011 |
| CP004 | Dataiku and Databricks are simultaneously competitors and collaborators, implying that some accounts adopt Dataiku as a workflow layer on top of a Databricks-centered data architecture. | 中 | SP003, SP004 |
| CP005 | Amazon SageMaker is strongest where the buyer wants first-party AWS procurement and lifecycle tooling rather than a neutral orchestration layer. | 中 | SP005, SP012, SP013 |
| CP006 | SageMaker pricing is granular and usage-metered by instance type, duration, storage, and inference configuration, reinforcing the economics of a native cloud service rather than a seat-based platform. | 高 | SP012, SP013 |
| CP007 | Azure Machine Learning competes through enterprise MLOps and responsible-AI positioning inside the broader Azure estate. | 高 | SP014, SP015 |
| CP008 | Vertex AI couples model-development breadth with explicit training, deployment, and prediction pricing, making it a strong option for GCP-centric AI programs. | 高 | SP016, SP017 |
| CP009 | DataRobot still markets a full enterprise AI suite with on-premise, VPC, and SaaS deployment choices rather than a pure single-cloud service. | 中 | SP018 |
| CP010 | H2O.ai differentiates through hybrid-cloud deployment, open-source lineage, and AutoML-led accessibility, but it remains much smaller in disclosed capital scale than Dataiku or Databricks. | 中 | SP021, SP022, SP025 |
| CP011 | Alteryx is a meaningful substitute for analytics automation and low-code data work, but it is not positioned as broadly around end-to-end enterprise ML and agent governance as Dataiku or Databricks. | 中 | SP023, SP024 |
| CP012 | Databricks is more publicly transparent on list pricing than most enterprise AI peers because it publishes a price list for SKU groups, even though actual realized economics still depend on cloud and discount structure. | 中 | SP010, SP011 |
| CP013 | Amazon SageMaker exposes detailed public price components across notebooks, training, inference, feature store, processing, and MLflow surfaces. | 高 | SP012, SP013 |
| CP014 | Azure Machine Learning offers pay-as-you-go, reservation, and savings-plan choices rather than a single public platform fee. | 高 | SP014, SP015 |
| CP015 | Vertex AI pricing includes hourly model-operation charges and no minimum usage duration for training and prediction, favoring bursty experimentation on GCP. | 高 | SP016, SP017 |
| CP016 | Dataiku's archived plans page shows packaging that expands from free or small-team usage into deeper automation, deployment, security, and governance capabilities for larger teams. | 中 | SP001, SP002 |
| CP017 | DataRobot, H2O.ai, and Alteryx emphasize demos, downloads, or contact-sales enterprise motions more than detailed self-serve enterprise list pricing. | 中 | SP018, SP021, SP023 |
| CP018 | Dataiku's ecosystem breadth across Databricks, AWS, Google Cloud, and NVIDIA reduces channel isolation and lets it sell into accounts already standardized on adjacent platforms. | 中 | SP003, SP004, SP005, SP006, SP007 |
| CP019 | Databricks has a scale advantage Dataiku cannot match publicly today, with Sacra estimating $6.9B in annualized revenue for 2026. | 中 | SP011 |
| CP020 | Public company-stat reporting tracked by Latka places DataRobot at roughly $285M of revenue in 2024, far smaller than Databricks and only modestly below Dataiku's last disclosed 2025 ARR markers. | 中 | SP019 |
| CP021 | H2O.ai's funding announcement says it serves 20,000 organizations and was valued at $1.7B after its 2021 Series E round. | 高 | SP022, SP025 |
| CP022 | Alteryx disclosed more than 8,000 customers globally and agreed to a $4.4B take-private transaction in December 2023. | 高 | SP023, SP024 |
| CP023 | Dataiku's strongest buying-criteria advantage is governed workflow breadth across technical and business personas without forcing one cloud or data platform choice. | 中 | SP001, SP002, SP003, SP006 |
| CP024 | Hyperscalers can undercut standalone platform value because ML tooling rides existing cloud identity, data gravity, and procurement paths. | 中 | SP012, SP014, SP016, SP013, SP015, SP017 |
| CP025 | Databricks benefits from owning both data and AI workflow surfaces, which raises switching costs once customers consolidate multiple workloads on the platform. | 高 | SP008, SP009, SP011 |
| CP026 | Dataiku's route to market appears coexistence-first rather than rip-and-replace, because its partner set includes companies whose native stacks also compete with it. | 中 | SP003, SP004, SP005, SP006, SP007 |
| CP027 | Specialists can still win budgets where buyers want faster time-to-value, narrower AutoML workflows, or low-code automation without replatforming the full data estate. | 中 | SP018, SP021, SP023 |
| CP028 | Alteryx remains most credible when the buyer's job is business analytics automation rather than governed multi-stage ML and agent deployment. | 中 | SP023, SP024 |
| CP029 | Metered hyperscaler pricing can reduce entry friction but also makes realized spend highly sensitive to model architecture and inference intensity. | 中 | SP013, SP015, SP017 |
| CP030 | Dataiku's lack of current public list pricing likely matters less in large-enterprise evaluations than architecture fit and governance requirements, but it still limits public TCO benchmarking. | 低 | SP001, SP002, SP003 |
| CP031 | Dataiku's most durable moat claim is infrastructure neutrality combined with governed workflow breadth across mixed personas. | 中 | SP001, SP002, SP003, SP007 |
| CP032 | That moat is vulnerable if Databricks and the hyperscalers keep closing the governance and agent-functionality gap inside native environments. | 中 | SP008, SP009, SP012, SP014, SP016 |
| CP033 | Databricks is the highest-severity competitive threat because it combines adjacent budget ownership, a rapid AI roadmap, and far larger disclosed financial scale than Dataiku. | 高 | SP008, SP009, SP011 |
| CP034 | DataRobot is a cautionary adverse case for the category because analyses of its decline explicitly cite hyperscaler bundling, valuation compression, layoffs, and shifting buyer priorities around AI. | 中 | SP019, SP020 |
| CP035 | H2O.ai shows that hybrid and open-source-led challengers still have room, but its smaller disclosed funding base likely constrains global distribution compared with Dataiku or Databricks. | 中 | SP021, SP022, SP025 |
| CP036 | The Alteryx take-private underscores that analytics-automation value exists, but adjacent categories can be priced and financed very differently from AI-platform growth narratives. | 中 | SP023, SP024 |
| CP037 | Dataiku's partner-heavy distribution strategy partly mitigates displacement risk because it can ride ecosystems that also compete with it. | 中 | SP003, SP005, SP006, SP007 |
| CP038 | Public materials do not support a clean apples-to-apples realized-TCO comparison across Dataiku and peers because negotiated discounts, services mix, and cloud commitments remain private. | 中 | SP010, SP013, SP015, SP017 |
| CP039 | A material share of Dataiku's competition is effectively internal build plus cloud-native services, because large enterprises can assemble AI workflows from first-party tools without buying a neutral umbrella platform. | 中 | SP008, SP012, SP014, SP016 |
| CP040 | The coexistence of Dataiku with Databricks and other partners suggests it often competes more for workflow governance and collaboration ownership than for raw storage or compute budget. | 中 | SP003, SP004 |
| CI001 | Dataiku reported surpassing $300M of ARR in January 2025. | 高 | SI002, SI006 |
| CI002 | The January 2025 release said ARR had more than doubled over the prior three years and that more than 20% of customers were already using Dataiku for GenAI workflows. | 中 | SI002 |
| CI003 | Dataiku reported surpassing $350M of ARR in October 2025 while enterprises accelerated trusted-AI deployments. | 中 | SI003 |
| CI004 | Public Dataiku materials position the product as an enterprise AI platform whose value comes from orchestration, governance, agents, and deployment breadth rather than from a single standalone feature. | 中 | SI001, SI004 |
| CI005 | Archived Dataiku packaging showed a progression from free or small-team use into business and enterprise tiers with deeper automation, deployment, and governance capability. | 中 | SI005 |
| CI006 | Dataiku’s public record supports an enterprise contract model, but does not disclose current realized list pricing, discounting, or contract term mix. | 中 | SI001, SI005 |
| CI007 | Because Dataiku emphasizes cloud- and model-agnostic deployment rather than a single native cloud runtime, its monetization likely depends more on platform contract value than on raw compute resell. | 中 | SI002, SI004, SI024, SI025 |
| CI008 | Dataiku’s broad partner ecosystem suggests some implementation and enablement work can be carried by partners rather than fully in-house delivery teams. | 中 | SI008, SI024, SI025 |
| CI009 | Public sources do not disclose Dataiku’s revenue mix across software, support, hosting, and services. | 中 | SI002, SI003, SI005 |
| CI010 | C3.ai’s FY2026 disclosures show an enterprise AI platform can be overwhelmingly subscription-led: 91% of total FY2026 revenue was subscription revenue. | 高 | SI016, SI017 |
| CI011 | C3.ai’s 10-K says revenue consists of subscriptions and professional services, with software licenses, SaaS, stand-ready support, usage-based runtime, and hosting charges embedded inside subscription revenue. | 高 | SI015, SI016 |
| CI012 | C3.ai explicitly says it relies on partners for larger or continuing professional-services presence in order to maintain margin flexibility. | 中 | SI016 |
| CI013 | Snowflake’s FY2026 annual report describes a customer-centric, consumption-based pricing model in which revenue is recognized on customer consumption rather than ratably over a subscription term. | 中 | SI018 |
| CI014 | Snowflake reported 67% total gross margin and 72% product gross margin in FY2026, while professional services and other gross margin remained negative 31%. | 中 | SI018 |
| CI015 | C3.ai’s FY2026 results showed 31% GAAP gross margin and 46% non-GAAP gross margin, highlighting how enterprise AI-platform margins can vary materially with services intensity and execution quality. | 高 | SI016, SI017 |
| CI016 | Official pricing pages for SageMaker, Azure Machine Learning, and Vertex AI show that major adjacent rivals monetize primarily through compute, duration, and deployed-model activity rather than opaque seat pricing. | 中 | SI012, SI013, SI014 |
| CI017 | Databricks publishes usage-based price lists and Sacra describes its pay-as-you-go model as aligned to cloud consumption, reinforcing that much of the adjacent category prices infrastructure-linked usage. | 中 | SI010, SI011 |
| CI018 | Compared with usage-centric rivals, Dataiku’s packaging likely offers more budget predictability if contracts are primarily subscription based, but public evidence does not reveal realized terms or cost-to-serve. | 中 | SI005, SI010, SI012, SI013, SI014 |
| CI019 | Dataiku’s CFO said in January 2025 that the company’s growth and financial efficiency differentiated it from OpEx-heavy business models elsewhere in the AI ecosystem. | 中 | SI002 |
| CI020 | Public evidence is strong on top-line traction but weak on revenue quality because there is no disclosed retention, gross-margin, or cash-conversion data. | 中 | SI002, SI003, SI007 |
| CI021 | Reuters-reported IPO preparation indicates Dataiku had not publicly raised a new primary round after the December 2022 Series F as of late 2025. | 中 | SI006 |
| CI022 | Sacra estimated roughly $342.5M of ARR in September 2025 and roughly $846.8M of lifetime funding for Dataiku. | 中 | SI007 |
| CI023 | Sacra also estimated that Dataiku’s 2022 $3.7B valuation represented roughly 18.5x ARR on a $200M ARR base. | 中 | SI007 |
| CI024 | As of the run date, no public source in the reviewed pack discloses Dataiku’s current cash balance, burn rate, runway, or debt facilities. | 中 | SI002, SI003, SI006, SI007 |
| CI025 | The absence of a public cash and burn disclosure does not imply strength or weakness by itself; it simply leaves capital adequacy unverified. | 中 | SI006, SI007 |
| CI026 | EY says IPO markets gained momentum in 1H 2026, but execution windows remain episodic and can be shaped by mega-IPOs and geopolitics. | 中 | SI020 |
| CI027 | Forge describes 2025 as a modest but meaningful IPO reopening that created a cautiously stronger setup for 2026 private-company listings. | 中 | SI021 |
| CI028 | For Dataiku, an IPO path in 2026 would likely be about liquidity, recruiting currency, and financing optionality as much as about immediate survival capital. | 中 | SI006, SI020, SI021 |
| CI029 | Scaled ARR growth plus no announced follow-on round suggests no obvious public distress signal, but it still does not prove that Dataiku is self-funding or flush with cash. | 中 | SI003, SI006, SI007 |
| CI030 | Public-company comparators show that professional services or implementation-heavy work can materially dilute a software-margin narrative. | 中 | SI016, SI017, SI018 |
| CI031 | Snowflake and C3.ai together show a wide economics range for adjacent AI/data platforms, from low-30s GAAP margins to low-70s product margins. | 中 | SI016, SI017, SI018 |
| CI032 | Because Dataiku is private, serious financial underwriting currently relies on proxies rather than direct disclosure for margin, retention, and cash efficiency. | 中 | SI007, SI016, SI018 |
| CI033 | C3.ai reported $575.4M of cash, cash equivalents, and marketable securities at FY2026 year-end and $673M shortly after results, illustrating the balance-sheet buffer some enterprise AI vendors need while execution remains uneven. | 高 | SI016, SI017 |
| CI034 | C3.ai’s FY2027 guidance of $210M-$240M revenue after FY2026 $250.3M revenue is an adverse reminder that scaled enterprise AI vendors can still face growth and profitability pressure. | 中 | SI017 |
| CI035 | Snowflake disclosed about $9.8B of remaining performance obligations with about 46% expected to convert within 12 months, underscoring the kind of backlog visibility that Dataiku does not provide publicly. | 中 | SI018 |
| CI036 | Dataiku’s partner footprint and enterprise positioning imply longer sales cycles and larger ACVs than self-serve AI tooling, which makes revenue quality more dependent on retention and expansion than on pure sign-up volume. | 中 | SI008, SI009, SI024, SI025 |
| CI037 | Snowflake’s consumption model shows how optimization by customers can reduce near-term visibility even in a large-scale data platform, a risk Dataiku may partly avoid if contract commitments are firmer. | 中 | SI018 |
| CI038 | C3.ai’s 10-K explicitly notes that customers can reduce usage, renew on less favorable terms, or allow RPO to decline, illustrating why renewal and contracted backlog are critical diligence items for AI platforms. | 中 | SI016 |
| CI039 | Current Dataiku product pages emphasize governed AI agents, orchestration, and enterprise data controls, implying that monetization is tied to platform breadth rather than single-seat utility. | 中 | SI004 |
| CI040 | Without a public filing, Dataiku’s exact revenue-recognition policy, deferred revenue balance, contractual term mix, and remaining performance obligations remain unknown. | 中 | SI005, SI015, SI018 |
| CE001 | Dataiku positions the platform as a combination of people, orchestration, and governance rather than as a single-purpose ML tool. | 中 | SE001, SE002 |
| CE002 | Public product and documentation surfaces show that Dataiku supports both visual workflows and code-based work inside the same platform. | 中 | SE002, SE006, SE008 |
| CE003 | Current product messaging places enterprise AI agents and governed reasoning workflows near the center of Dataiku’s product story. | 中 | SE002, SE025 |
| CE004 | The product page says Dataiku can centralize agent creation, collaboration, lifecycle management, and orchestration across agents, models, and tools. | 中 | SE002 |
| CE005 | The govern product surface and documentation both present Govern as a dedicated layer for tracking AI initiatives, approvals, and audit readiness. | 高 | SE003, SE010 |
| CE006 | Govern documentation describes Standard and Advanced licenses, with advanced features such as GenAI registries, custom actions, blueprint design, custom pages, and scripting. | 中 | SE010 |
| CE007 | The Python API documentation says DSS APIs can be used anywhere code can run inside Dataiku, including recipes, notebooks, scenarios, and webapps. | 中 | SE009 |
| CE008 | Dataiku maintains a public Python API client repository and README, providing at least a modest external developer signal beyond marketing pages. | 高 | SE012, SE013 |
| CE009 | The DSS 14 release stream in June and July 2026 spans agentic AI and RAG, LLM Mesh, governance, MLOps, AI assistants, security, Git, and code tooling. | 中 | SE007 |
| CE010 | The frequency of 14.6.x and 14.7.x releases in mid-2026 suggests active enterprise product maintenance rather than a static legacy platform. | 中 | SE007 |
| CE011 | The public architecture reads as a control plane layered over build surfaces, automation, governance, and deployment operations rather than a monolithic closed system. | 中 | SE001, SE002, SE006, SE009 |
| CE012 | Archived packaging and partner pages indicate that Dataiku supports SaaS, customer-managed cloud, and on-prem or private-cloud styles of operation. | 中 | SE019, SE020, SE024 |
| CE013 | Partner pages with AWS, Google Cloud, Databricks, and NVIDIA show that Dataiku is intentionally built to integrate with large external infrastructure and AI ecosystems. | 中 | SE019, SE020, SE021, SE022 |
| CE014 | Dataiku’s public security surfaces document response guidance for 2026 Linux local-privilege-escalation vulnerabilities affecting Dataiku environments. | 中 | SE004, SE011 |
| CE015 | Govern is documented as an additional node integrated into the broader Dataiku platform rather than as a simple label inside the core UI. | 中 | SE010 |
| CE016 | The govern product page explicitly links the product to shadow-AI reduction, EU AI Act readiness, audit readiness, qualification, and signoff workflows. | 中 | SE003 |
| CE017 | Govern’s public materials identify bundle, model, and LLM registries as specialized registries inside the governance system. | 中 | SE003 |
| CE018 | Reference docs, developer guide, API docs, community links, and academy references together show a substantial enablement surface for platform users and builders. | 中 | SE006, SE008, SE009 |
| CE019 | The API documentation and public client indicate that Dataiku exposes programmable interfaces beyond the visual UI, which is important for enterprise automation and integration. | 中 | SE009, SE012, SE013 |
| CE020 | Dataiku’s core technical differentiation is the combination of business-user accessibility, code extensibility, orchestration, and governance in one shared platform. | 中 | SE001, SE002, SE006, SE008, SE010 |
| CE021 | Critical dependencies in the product design include customer cloud or on-prem infrastructure, external data systems, model ecosystems, and implementation partners. | 中 | SE019, SE020, SE021, SE022 |
| CE022 | Release-note items such as OS upgrades, Python version changes, Java minimum-version bumps, and model removals show that Dataiku actively manages a changing dependency stack. | 中 | SE007 |
| CE023 | Some public URLs collapse back to high-level marketing pages, which limits how much deep architecture detail can be independently validated from the public web alone. | 中 | SE002, SE025 |
| CE024 | Independent review evidence includes at least one critical readthrough around private-cloud integration, suggesting implementation friction can still matter even when users like the overall platform concept. | 中 | SE023 |
| CE025 | Govern’s advanced features and separate documentation imply that governance is a meaningful product area, not a superficial afterthought. | 高 | SE003, SE010 |
| CE026 | Public materials strongly suggest that Dataiku supports model abstraction, agent tools, and governed GenAI workflows, but they do not publish benchmarked performance or agent reliability metrics. | 中 | SE002, SE003, SE025 |
| CE027 | The SOC 2 announcement is useful historical trust evidence, but it does not by itself prove the full current certification inventory or present-day scope. | 中 | SE005, SE004 |
| CE028 | The v14 release stream supports the view that Dataiku is a mature, continuously shipped enterprise platform rather than an experimental toolset. | 中 | SE007 |
| CE029 | Dataiku’s product surface is explicitly built for both coders and non-coders, combining visual logic blocks and code-first tooling in one environment. | 中 | SE002, SE008 |
| CE030 | The mix of partner pages and competitor platforms implies that Dataiku is integration-oriented and orchestration-oriented rather than vertically locked to one proprietary infrastructure stack. | 中 | SE014, SE015, SE016, SE017, SE019, SE020, SE021 |
| CE031 | The public API-client repo is a modest but real developer-signal that third-party or customer-side automation against Dataiku is supported. | 高 | SE012, SE013 |
| CE032 | The fetched public pack does not expose benchmarked performance, uptime SLA terms, or quantified agent accuracy claims for Dataiku. | 中 | SE002, SE006, SE011 |
| CE033 | Public governance controls include qualification workflows, signoff rules, alerts, registries, and audit timelines. | 高 | SE003, SE010 |
| CE034 | Module maturity appears uneven but healthy: the DSS core looks mature, while governance and agent surfaces are newer yet clearly active and shipping. | 中 | SE002, SE007, SE010, SE025 |
| CE035 | Current public materials do not fully enumerate certifications, detailed security architecture, or quantified implementation outcomes, leaving trust and quality diligence incomplete. | 中 | SE004, SE005, SE011, SE023 |
| CE036 | Overall, Dataiku looks like a broad, mature enterprise AI platform with active roadmap velocity, governance depth, and programmable interfaces, but external proof on operational depth remains thinner than the breadth of the official story. | 中 | SE002, SE007, SE010, SE012, SE023 |
| CU001 | Dataiku said in January 2025 that it had grown its customer base to more than 700 organizations worldwide. | 高 | SU003, SU005 |
| CU002 | Dataiku said in October 2025 that the platform powered initiatives at more than 750 organizations worldwide. | 中 | SU004 |
| CU003 | A July 2024 KPMG alliance release said Dataiku already had more than 600 customers, including 200 Forbes Global 2000 companies. | 中 | SU009 |
| CU004 | Named public customer references clearly include healthcare and life-sciences organizations such as Johnson & Johnson, Novartis, and Roche. | 高 | SU015, SU016, SU022 |
| CU005 | Named public customer references also include manufacturing and industrial organizations such as Michelin, Mitsubishi Electric, and SLB. | 高 | SU017, SU024, SU025 |
| CU006 | Financial-services and capital-markets proof is visible through Standard Chartered and Euronext customer stories. | 中 | SU018, SU023 |
| CU007 | Logistics, transportation, real-estate, and food/agriculture adoption is visible through Geodis, Prologis, and Perdue Farms. | 中 | SU019, SU020, SU021 |
| CU008 | The public stories consistently show Dataiku being used by cross-functional groups that mix technical practitioners with business or operations users. | 中 | SU015, SU017, SU021, SU023, SU025 |
| CU009 | The likely economic buyer is a centralized data, analytics, digital-transformation, or platform budget rather than an isolated single-team software seat purchase. | 中 | SU009, SU017, SU021, SU023 |
| CU010 | The public record supports a customer-growth path from roughly 500 customers at the end of 2023 to 700-plus in January 2025 and 750-plus in October 2025. | 高 | SU006, SU003, SU004 |
| CU011 | Johnson & Johnson Vision's case study says 650-plus employees use Dataiku and more than 80 analytics and data-science professionals in the organization had adopted it. | 中 | SU015 |
| CU012 | Michelin expanded from 35 users in 2021 to more than 1,500 users across 50-plus factories by mid-2025, with 80% of users described as business experts. | 中 | SU017 |
| CU013 | Standard Chartered said 518 staff completed its citizen-data-science training program and more than 700 completed Dataiku online training courses. | 中 | SU023 |
| CU014 | Prologis said it had put over 60 AI/ML projects into production, maintained 30 projects and 30 APIs in active use, and had more than 2,000 users leveraging AI through its enterprise ChatGPT platform. | 中 | SU021 |
| CU015 | Mitsubishi Electric's story shows Dataiku being treated as a foundational component of the Serendie platform rather than a narrow point solution. | 中 | SU025 |
| CU016 | Johnson & Johnson used a two-day Dataiku event to build working generative-AI and LLM prototypes and then emphasized common-platform standardization afterward. | 中 | SU015 |
| CU017 | Novartis reported a 90% reduction in time to insights for a GenAI use case and a 600% acceleration in spreadsheet data-ingestion time with Dataiku. | 中 | SU016 |
| CU018 | Michelin said a defect-analysis workflow that previously took up to six months can now be completed in about an hour with a Dataiku-based solution. | 中 | SU017 |
| CU019 | Euronext reported up to a 20% reduction in time spent on recurring market-share queries after deploying a Dataiku-based analytics agent. | 中 | SU018 |
| CU020 | Geodis reported a 60% reduction in ticket-assignment time and about 30 minutes saved per ticket through a Dataiku IT support agent. | 中 | SU019 |
| CU021 | Perdue Farms reported that tasks which once took days are now completed in hours and highlighted more than six hours of monthly labor savings from automated reporting workflows. | 中 | SU020 |
| CU022 | Roche said Dataiku cut the time to build new GenAI projects from months to days and generated an estimated $100K-$250K of annual attorney time savings plus $375K-$475K of avoided consulting cost. | 中 | SU022 |
| CU023 | SLB said Dataiku-supported workflows assessed more than $10 billion of well-construction tenders, cut tender analysis from eight hours to twenty minutes, and made reservoir-pressure analysis 76% faster. | 中 | SU024 |
| CU024 | Standard Chartered said Dataiku-enabled FP&A workflows made analysts roughly 30 times more productive and supported a space-planning initiative expected to reduce annual property cost by $34 million. | 中 | SU023 |
| CU025 | Mitsubishi Electric reported a 60% workload reduction from data preparation to reporting, an 80% reduction in visualization time versus Python, and faster collaboration through shared flows. | 中 | SU025 |
| CU026 | Across multiple case studies, Dataiku's adoption pattern looks like a land-and-expand motion in which a first workflow broadens into governance, training, and wider business use. | 中 | SU017, SU021, SU022, SU023 |
| CU027 | The KPMG alliance indicates Dataiku can be sold as part of larger modernization, cloud, and AI-governance programs rather than only as standalone software. | 中 | SU009 |
| CU028 | Because partner-assisted transformation is part of the public story, some portion of Dataiku's commercial success likely depends on ecosystem leverage and implementation partners. | 中 | SU009, SU014 |
| CU029 | The 2025 Frontrunner Awards show active, public customer engagement around agentic AI, governance, productivity, and ROI use cases in multiple industries. | 中 | SU014 |
| CU030 | The strongest public Dataiku customer evidence is operational proof with quantified outcomes, not merely logo placement on a customers page. | 中 | SU015, SU016, SU017, SU018, SU019, SU021, SU022, SU024 |
| CU031 | Independent customer-satisfaction evidence is directionally positive: the Gartner page fetched for this run showed a review distribution of 75% five-star and 23% four-star ratings, while Dataiku's January 2025 release cited a 96% willingness-to-recommend score in Gartner Peer Insights. | 高 | SU007, SU003 |
| CU032 | The same Gartner review surface also included a critical review headline describing private-cloud integration issues, showing that implementation friction is not zero even for a well-regarded platform. | 中 | SU007 |
| CU033 | FeaturedCustomers and TrustRadius confirm that Dataiku has a visible public corpus of reviews and case-study references, but those directories do not establish renewal quality or average deployment success. | 中 | SU008, SU013 |
| CU034 | Dataiku does not publicly disclose NRR, GRR, churn, contract duration, or top-customer concentration in the sources reviewed for this run. | 中 | SU003, SU004, SU005, SU006 |
| CU035 | Comparing Sacra's 2023 estimate with 2025 official disclosures suggests customer growth is accompanied by a broader and likely more heterogeneous account base, not simply a small set of giant accounts growing in place. | 中 | SU006, SU003, SU004 |
| CU036 | Public evidence on Dataiku customers is strongest on named deployment outcomes and weakest on renewal, churn, and concentration. | 中 | SU015, SU017, SU021, SU007, SU003, SU006 |
| CU037 | Most of the strongest customer proof is vendor-hosted, which means the existence of deployments is credible but the representativeness of outcomes across the full base remains uncertain. | 中 | SU015, SU016, SU017, SU018, SU019, SU020, SU021, SU022, SU023, SU024, SU025 |
| CU038 | Customer-count disclosures moved from 600-plus in mid-2024 to 700-plus in early 2025 and 750-plus by October 2025, supporting a continued new-logo or expanded-disclosure momentum story. | 高 | SU009, SU003, SU004 |
| CU039 | The named customer base spans several regulated or complex sectors including healthcare, banking, capital markets, insurance-adjacent operations, logistics, and industrial manufacturing, which supports Dataiku's governance-centric enterprise positioning. | 高 | SU015, SU016, SU018, SU022, SU023, SU024 |
| CU040 | Customer durability should therefore be rated medium-confidence: switching costs are plausibly meaningful where Dataiku becomes the shared workflow and governance layer, but public retention and concentration data are insufficient to fully underwrite that durability. | 中 | SU017, SU021, SU023, SU007, SU006 |
| CR001 | Dataiku's privacy policy revised August 18, 2025 says the policy applies to visitors and users and that Dataiku acts as a controller for personal data collected under that policy. | 中 | SR002 |
| CR002 | The privacy policy says that when AI Services are enabled, content may be processed by Dataiku and the applicable third-party AI provider, and Dataiku says it will not use customer data to train its models without consent. | 中 | SR002 |
| CR003 | The February 2026 Dataiku Cloud DPA defines Dataiku as processor for customer personal data and explicitly references GDPR, CCPA, SCCs, subprocessors, and security incidents. | 高 | SR005, SR006 |
| CR004 | Dataiku's legal hub centralizes privacy, acceptable-use, installed-software, cloud-legal, and modern-slavery documents, indicating a relatively mature contractual surface for enterprise procurement. | 中 | SR001, SR004 |
| CR005 | Dataiku's trust page says self-managed deployments do not cause Dataiku to process or store client data by default unless the customer explicitly grants access. | 中 | SR007 |
| CR006 | The same trust page says Dataiku Cloud is a managed SaaS offering, available in multi-tenant or single-tenant form, and that it leverages cloud-provider infrastructure controls. | 中 | SR007 |
| CR007 | Dataiku publicly cites ISO 27001, ISO 27701, ISO 9001, SOC 1/SOC 2 assessments, a HIPAA compliance report, and GxP readiness as trust mitigants. | 高 | SR007, SR010 |
| CR008 | The EU AI Act subjects high-risk AI systems to obligations such as risk mitigation, logging, documentation, human oversight, and robustness, and its generative-AI transparency rules come into effect in August 2026. | 中 | SR011 |
| CR009 | NIST's AI RMF and its GenAI profile establish a strong market expectation that enterprise AI systems incorporate trustworthiness and structured risk management throughout design, deployment, and evaluation. | 中 | SR012 |
| CR010 | Dataiku's large legal and trust surface mitigates enterprise risk but also expands the contractual and operational obligations the company must consistently honor. | 中 | SR001, SR002, SR005, SR007 |
| CR011 | The FTC legal library shows that U.S. regulators continue to pursue privacy, false-advertising, and consumer-harm cases aggressively, underscoring that AI and data-platform claims face real enforcement risk. | 中 | SR013 |
| CR012 | The GDPR Enforcement Tracker page fetched for this run reported 3,202 tracked enforcement actions and €6.31B of total fines, confirming that privacy failures are financially material in Europe. | 中 | SR014 |
| CR013 | No Dataiku-specific FTC case or obvious GDPR fine surfaced in the public trackers fetched for this run, but absence from these surfaces is not proof of no latent exposure. | 中 | SR013, SR014 |
| CR014 | OpenCVE lists CVE-2023-51717, updated June 16, 2025, as a critical 9.8 incorrect-access-control flaw that could lead to full authentication bypass in Dataiku DSS before versions 11.4.5 and 12.4.1. | 中 | SR015 |
| CR015 | OpenCVE also lists historical medium and high-severity issues involving file access, Jupyter notebook permissions, metadata manipulation, and REST API information exposure in older DSS versions. | 中 | SR015 |
| CR016 | The Gartner review surface fetched for this run includes a critical review headline that specifically cites private-cloud integration issues. | 中 | SR017 |
| CR017 | UpGuard's public Dataiku vendor-risk page shows continuous external monitoring across 330+ checks, which is useful context but not a substitute for direct technical diligence. | 中 | SR016 |
| CR018 | Because Dataiku spans notebooks, APIs, connectors, automation, governance, and agent workflows, its operational attack surface is materially broader than that of a narrow single-purpose analytics tool. | 中 | SR008, SR015, SR022 |
| CR019 | Self-managed deployments reduce vendor data-custody exposure but shift more configuration, availability, and security responsibility into the customer environment. | 中 | SR007, SR008 |
| CR020 | Cloud delivery increases Dataiku's direct responsibility for incident handling, access controls, and subprocessor management while also inheriting risk from underlying cloud providers. | 中 | SR005, SR007 |
| CR021 | The KPMG alliance demonstrates that partner-led modernization and AI-governance programs are part of Dataiku's route to market, making services quality and channel economics relevant risks. | 中 | SR021 |
| CR022 | Customer stories repeatedly reference external systems such as Snowflake, Azure, ServiceNow, and PowerBI, which means interoperability is both a moat and a dependency chain. | 高 | SR025, SR026, SR028, SR030 |
| CR023 | Michelin, Prologis, Standard Chartered, and Geodis each describe Dataiku as part of a broader operational stack rather than an isolated tool, so connector reliability directly affects realized customer value. | 高 | SR025, SR026, SR028, SR030 |
| CR024 | Reuters reported in October 2025 that Dataiku hired Morgan Stanley and Citigroup to prepare for a possible U.S. IPO, increasing the valuation consequences of any compliance, security, or customer-retention stumble. | 高 | SR018, SR019 |
| CR025 | Even at $350M+ ARR scale, public sources still do not disclose core underwriting metrics such as NRR, GRR, gross margin, burn, or top-customer concentration. | 中 | SR020, SR022, SR023 |
| CR026 | That financial opacity is itself a material risk because investors cannot verify runway, renewal quality, or margin resilience if an IPO or financing window closes. | 中 | SR018, SR020, SR022 |
| CR027 | The EU AI Act's high-risk categories include employment and certain essential-service uses, so a general enterprise AI platform like Dataiku can become entangled in sensitive customer workflows even if Dataiku is not the end-use decision-maker. | 中 | SR011, SR024, SR028 |
| CR028 | Because the privacy policy contemplates AI Services using applicable third-party AI providers, model-vendor policy changes or data-handling concerns can transmit into Dataiku customer risk. | 中 | SR002, SR004 |
| CR029 | The DPA and cloud terms are real mitigants, but they also imply procurement friction because sophisticated enterprise buyers will review subprocessors, transfers, and incident obligations carefully. | 中 | SR005, SR006, SR021 |
| CR030 | No major public breach or enforcement event surfaced in the sources reviewed, but the disclosed vulnerability history means security patch cadence remains a core diligence item. | 中 | SR013, SR014, SR015, SR016 |
| CR031 | The public customer-proof set is rich, but most of the strongest risk-reducing evidence comes from vendor-hosted case studies rather than independent retention or concentration disclosures. | 中 | SR024, SR025, SR026, SR027, SR028, SR029, SR030 |
| CR032 | Customer stories imply that Dataiku often requires workflow redesign, training, and operational standardization, which increases implementation effort and therefore risk of slow time-to-value. | 中 | SR025, SR026, SR027, SR028 |
| CR033 | Regulated-industry strength is both a moat and a risk amplifier because healthcare, banking, and capital-markets customers require stricter validation, auditability, and incident response than ordinary SaaS buyers. | 中 | SR024, SR027, SR028 |
| CR034 | Prologis, Michelin, Standard Chartered, and Roche show that Dataiku can become embedded in business processes, which raises switching costs but also raises the business-interruption cost of failure. | 中 | SR025, SR026, SR027, SR028 |
| CR035 | The August 2026 EU AI Act transparency milestone increases pressure on Dataiku's governance narrative precisely as enterprises expand agentic-AI deployments. | 中 | SR011, SR022 |
| CR036 | Certifications and trust documentation improve the mitigation story, but they do not eliminate the need for rapid patching, careful access management, or customer-specific deployment diligence. | 中 | SR007, SR010, SR015 |
| CR037 | Execution risk remains meaningful because a 1,250+ employee, pre-IPO enterprise platform company must coordinate product delivery, customer success, security, and partner operations at a much higher bar than an earlier-stage startup. | 中 | SR018, SR022 |
| CR038 | Customer concentration and partner-sourced pipeline remain unresolved because public sources show reference accounts and alliances but not revenue weighting or channel mix. | 中 | SR020, SR021, SR024 |
| CR039 | Dataiku's 2025 trusted-AI and agent-management positioning helps the mitigation case, but it also raises expectations that the company can operationalize governance safely at scale for customers. | 中 | SR022, SR011, SR012 |
| CR040 | Overall residual risk should be considered medium-high: Dataiku has real controls and market relevance, but security vulnerabilities, regulatory expansion, ecosystem dependence, and financial opacity still create several plausible thesis-break paths. | 中 | SR015, SR018, SR020, SR022 |
| CV001 | Dataikus last disclosed primary valuation anchor is the $3.7B December 2022 Series F led by Wellington Management. | 中 | SV001 |
| CV002 | Dataiku disclosed surpassing $300M ARR in January 2025. | 高 | SV002, SV004 |
| CV003 | Dataiku disclosed surpassing $350M ARR in October 2025. | 中 | SV003 |
| CV004 | The stale $3.7B mark implies about 12.3x ARR against the January 2025 $300M milestone. | 高 | SV001, SV002 |
| CV005 | The stale $3.7B mark implies about 10.6x ARR against the October 2025 $350M milestone. | 高 | SV001, SV003 |
| CV006 | Using Sacras roughly $342.5M September 2025 ARR estimate, the stale $3.7B mark implies about {mult["dikusacra"]}x ARR. | 中 | SV001, SV005 |
| CV007 | Reuters reported in October 2025 that Dataiku hired Morgan Stanley and Citigroup to prepare for a possible U.S. IPO. | 高 | SV004, SV007, SV008 |
| CV008 | Forbes and official customer materials support that Dataiku is a large private AI/data platform with a blue-chip enterprise customer base. | 中 | SV006, SV029 |
| CV009 | Public evidence still does not disclose Dataikus gross margin, NRR, cash burn, or top-customer concentration. | 中 | SV003, SV005 |
| CV010 | Palantirs public market-cap-to-revenue proxy is about {mult["pal"]}x based on the fetched July 2026 market-cap and revenue pages. | 中 | SV009, SV010 |
| CV011 | Snowflakes public market-cap-to-revenue proxy is about {mult["snow"]}x based on the fetched July 2026 market-cap and revenue pages. | 中 | SV013, SV014 |
| CV012 | Datadogs public market-cap-to-revenue proxy is about {mult["ddog"]}x based on the fetched July 2026 market-cap and revenue pages. | 中 | SV015, SV016 |
| CV013 | MongoDBs public market-cap-to-revenue proxy is about {mult["mdb"]}x based on the fetched July 2026 market-cap and revenue pages. | 中 | SV017, SV018 |
| CV014 | ServiceNows public market-cap-to-revenue proxy is about {mult["now"]}x based on the fetched July 2026 market-cap and revenue pages. | 中 | SV019, SV020 |
| CV015 | Confluents public market-cap-to-revenue proxy is about {mult["cflt"]}x based on the fetched July 2026 market-cap and revenue pages. | 中 | SV021, SV022 |
| CV016 | C3.ais public market-cap-to-revenue proxy is about {mult["c3"]}x based on the fetched July 2026 market-cap and revenue pages. | 中 | SV011, SV012 |
| CV017 | Compared with public comps, Dataikus stale roughly 10-12x ARR multiple sits above weaker AI software names like C3.ai and around MongoDB/Confluent-ServiceNow territory, while still below Snowflake, Datadog, and far below Palantir. | 中 | SV001, SV003, SV009, SV010, SV011, SV012, SV013, SV014, SV015, SV016, SV017, SV018, SV019, SV020, SV021, SV022 |
| CV018 | Snowflakes FY2026 filing shows what premium platform economics can look like, with 67% total gross margin and 72% product gross margin. | 中 | SV024 |
| CV019 | C3.ais FY2026 disclosure and results remind investors that enterprise AI software can trade at much lower multiples when growth and margins disappoint. | 中 | SV026, SV027 |
| CV020 | Because Dataiku is private, illiquid, and economically opaque, it deserves a discount to the very highest public AI software multiples even if its category position is strong. | 中 | SV004, SV005, SV017 |
| CV021 | Conversely, Dataikus enterprise scale, governance positioning, and reference customer quality argue against valuing it at the lowest public AI software multiple tier. | 中 | SV003, SV006, SV029, SV030 |
| CV022 | The valuation debate is therefore not whether Dataiku deserves a premium at all, but where inside a very wide public comp band it belongs. | 中 | SV005, SV017 |
| CV023 | A reasonable bear-case valuation range is about $2.5B-$3.2B if growth slows, IPO timing slips, or hidden economics disappoint. | 中 | SV004, SV005, SV016, SV022, SV027 |
| CV024 | A reasonable base-case valuation range is about $3.5B-$4.5B if ARR scale is real and economics are decent but not elite. | 中 | SV003, SV005, SV017, SV018 |
| CV025 | A reasonable bull-case valuation range is about $5.0B-$6.5B if Dataiku can prove premium retention, strong margins, and credible IPO readiness. | 中 | SV003, SV004, SV006, SV017 |
| CV026 | The current disclosed $3.7B mark sits inside the base-case range, so the fairest present stance is fair rather than cheap or wildly stretched. | 中 | SV001, SV003 |
| CV027 | The evidence supports a track recommendation with medium confidence and a medium-high risk rating. | 中 | SV004, SV005, SV026 |
| CV028 | The missing economics—especially gross margin, NRR, cash burn, services mix, and concentration—prevent a stronger buy call despite the companys quality. | 中 | SV003, SV005, SV026 |
| CV029 | Dataikus strongest pro-valuation signals are scale, customer quality, and governance-first enterprise positioning. | 中 | SV003, SV006, SV029, SV030 |
| CV030 | Dataikus strongest anti-valuation signals are competitive intensity, implementation complexity, regulatory expansion, and economic opacity. | 中 | SV004, SV005, SV027, SV028 |
| CV031 | The scenario weighting used here is roughly 25% bull, 50% base, and 25% bear. | 中 | SV005, SV017 |
| CV032 | Price-sensitive upside from todays mark therefore depends on evidence that would move the company from the base case toward the bull case, not merely on continued category excitement. | 中 | SV004, SV024, SV025 |
| CV033 | If Dataiku can combine strong 2025 ARR momentum with credible IPO disclosures, the public-market path could unlock multiple expansion above the stale private mark. | 中 | SV003, SV004, SV006 |
| CV034 | If IPO timing slips or public-market investors demand cleaner economics, the valuation could compress below the last disclosed mark even without an outright business failure. | 中 | SV004, SV007, SV008, SV027 |
| CV035 | The fetched public comp set spans roughly 4.6x at the low end to 58.2x at the high end, proving that the right answer for Dataiku must be scenario-driven rather than a single-point multiple. | 中 | SV009, SV010, SV011, SV012, SV013, SV014, SV015, SV016, SV017, SV018, SV019, SV020, SV021, SV022 |
| CV036 | The final diligence asks should prioritize ARR recency, gross margin, NRR/churn, cash and burn, concentration, partner economics, and preference stack. | 中 | SV005, SV026 |
| CV037 | A thesis-break trigger would be materially weak retention or unexpectedly high customer concentration once private data is disclosed. | 中 | SV005, SV029 |
| CV038 | A second thesis-break trigger would be a major security or regulatory event that interrupts IPO timing or weakens trust with large enterprise customers. | 中 | SV004, SV027, SV029 |
| CV039 | A third thesis-break trigger would be an updated ARR bridge showing that growth has slowed meaningfully below what investors infer from the 2025 milestones. | 中 | SV002, SV003, SV005 |
| CV040 | Overall evidence quality is medium because the valuation analysis still relies on a stale private mark and public-comp proxies rather than current audited company disclosures. | 中 | SV001, SV005, SV024, SV026 |