Dataiku
Governed enterprise AI at scale, with IPO optionality but still-opaque economics
Dataiku looks like a real late-stage enterprise AI leader with fair current valuation support, but the lack of public margin, retention, and cash disclosure keeps it in track rather than buy territory.
Cover facts
Company profile
Dataiku is a French-founded, New York-headquartered enterprise AI platform company founded in 2013. It positions itself as a governed orchestration layer where analysts, engineers, data scientists, and business users can collaboratively build analytics, machine-learning models, and AI agents across cloud, on-prem, and hybrid environments. Public evidence shows the company scaled to $350M+ ARR, 750+ customers, and 1,250+ employees by October 2025 while preparing for a possible U.S. IPO.
- Website
- www.dataiku.com
- Founded
- 2013-01-01
- Founders
- Florian Douetteau, Clément Stenac, Thomas Cabrol, Marc Batty
- Founding location
- Paris, France
- Headquarters
- New York, New York, USA
- Product
- A unified platform for analytics, machine learning, AI governance, and agentic-AI orchestration, combining no-code, low-code, and full-code workflows with cloud and model optionality.
- Customers
- Large enterprises across healthcare, manufacturing, financial services, logistics, and other regulated or operationally complex sectors.
- Business model
- Enterprise software subscription model sold around platform breadth, governance, and deployment flexibility, supported by partner-led transformation and customer expansion inside large accounts.
- Stage
- Series F (private)
- Funding status
- $200M Series F at a $3.7B valuation in December 2022; no newer primary round is publicly confirmed, but Reuters reported in October 2025 that Dataiku hired Morgan Stanley and Citigroup to prepare for a potential U.S. IPO.
Executive summary
Top strengths
- $350M+ ARR, 750+ customers, and 1,250+ employees support genuine late-stage enterprise scale.
- Governance-first positioning fits enterprise demand for trusted, auditable AI and agent workflows.
- Customer proof is unusually strong for a private infrastructure vendor, with named results across Michelin, Novartis, Roche, Prologis, Standard Chartered, and SLB.
- A stale $3.7B valuation anchor now screens as roughly mid-band against public AI/data software comp multiples rather than peak-bubble territory.
Top risks
- Gross margin, NRR, churn, cash burn, and concentration remain undisclosed, capping valuation confidence.
- Competition from Databricks, hyperscalers, and adjacent public data/AI platforms can compress expansion assumptions and multiples.
- Security, privacy, and AI-governance obligations are widening under the EU AI Act and other enterprise compliance expectations.
- IPO timing can amplify the impact of any implementation, regulatory, or customer-retention stumble.
Open gaps
- No current audited financial disclosure or 2026 ARR bridge beyond the October 2025 $350M+ milestone.
- No public gross-margin, NRR, GRR, churn, or customer-concentration disclosure.
- No public clarity on partner-sourced pipeline, services mix, or the true degree of implementation intensity.
- No public view into current cash, burn, or liquidation-preference overhang if a future financing occurs.
Contents
01Company Overview
1.1 Identity, headquarters, and what Dataiku actually sells
Dataiku is a late-stage private enterprise software company founded in 2013 and now headquartered in New York, even though its roots and founding team are unmistakably French. The company still describes itself through the same core idea that powered its early growth: turning data and AI into an everyday operating capability rather than a specialist project. In 2026 its own language has shifted from the older "Everyday AI" framing toward "The Platform for AI Success," but the operating proposition is consistent — one control plane where business users, analysts, data scientists, engineers, and risk owners can build analytics, machine learning, and AI-agent workflows together. That positioning matters because Dataiku is not selling a single-model tool or a point AutoML product; it is selling orchestration, governance, and multi-user collaboration on top of customer data stacks. Archived plans pages and current partner materials show three durable deployment paths: hosted SaaS, customer-managed cloud, and on-prem/private-cloud installation. That deployment flexibility, plus no-, low-, and full-code workflows, is central to the company's value proposition and also explains why buyers tend to be large enterprises with heterogeneous infrastructure rather than SMBs looking for one-click AI.[CO001, CO002, CO003, CO004, CO005, CO008]
| Metric | Value / status | Date | Confidence | Note / gap |
|---|---|---|---|---|
| Founded | 2013 | 2013 | high | Official story and Gartner profile agree on the year |
| Headquarters | New York, NY | 2026 | high | French-founded, but current HQ is New York |
| Private / public status | Private; IPO-prep reported, no filing observed | 2026 | medium | Reuters reported banker selection, not a completed IPO |
| Last disclosed valuation | $3.7B | Dec 2022 | high | Series F valuation; no newer primary round publicly confirmed |
| Total raised | ~$846.8M | Nov 2025 research | medium | Sacra estimate based on disclosed rounds |
| ARR | $300M+ in Jan 2025; $350M+ in Oct 2025 | 2025 | medium | Company-claimed operating metrics |
| Customers | 700+ in Jan 2025; 750+ in Oct 2025 | 2025 | medium | Company-claimed operating metrics |
| Employees / footprint | 1,100+ in Jan 2025; 1,250+ and 13 offices by Oct 2025 | 2025 | medium | Official releases, not audited workforce data |
Operating metrics come from company press releases; valuation and total funding come from late-stage round reporting and secondary research rather than audited filings.
[CO001, CO004, CO005, CO014, CO016, CO021]Dataiku connects enterprise data estates, cross-functional builders, and governance requirements into one AI control plane.
[CO008, CO009, CO010, CO044, CO045, CO046]The public KPI surface is strong for a private company, but still relies on company disclosures and selected secondary analysis.
[CO014, CO016, CO026, CO027, CO028, CO029]1.2 Founders, leadership bench, and governance signals
The public leadership picture is strongest around the founding core and the newer go-to-market bench. Florian Douetteau remains the clear key person as co-founder and CEO, while Clément Stenac remains the technical anchor as co-founder and CTO. Public founder references consistently list Thomas Cabrol and Marc Batty alongside them, but their current operating roles are much less visible than Douetteau and Stenac, which is typical for a mature private company but still leaves governance detail thinner than a public-market investor would want. The most relevant 2026 management event is the June hiring of Maxwell Long as President and CRO. That move reads as classic late-stage scaling behavior: Long was hired specifically to run global sales, customer success, and partnership teams after helping Smartsheet scale past $1 billion ARR and into an $8.4 billion exit. Combined with the 2025 addition of a high-profile CMO and the continued emphasis on partners, the leadership pattern is consistent with a company professionalizing its commercial engine ahead of a potential liquidity event. What remains under-disclosed is board composition, independence, and the precise division of authority among late-stage investors and operators.[CO002, CO006, CO007, CO031, CO032, CO044]
| Person | Current / known role | Why they matter | Publicly visible gap |
|---|---|---|---|
| Florian Douetteau | Co-founder & CEO | Founding visionary and continuing public face of Dataiku; key person for strategy and IPO narrative | No public share ownership or board-control detail |
| Clément Stenac | Co-founder & CTO | Technical steward for the platform architecture and long-term product integrity | Limited public disclosure on succession depth below CTO |
| Thomas Cabrol | Co-founder | Consistently listed as a founder in official and investor profiles | Current operating remit is not clearly disclosed |
| Marc Batty | Co-founder | Consistently listed as a founder in official and investor profiles | Current operating remit is not clearly disclosed |
| Maxwell Long | President & CRO (joined June 2026) | Runs sales, customer success, and partnerships; late-stage scale-up operator | Very new in role, so no execution track record at Dataiku yet |
| Mark Abramowitz | Chief Marketing Officer (announced 2025) | Signals sharpening of enterprise brand and demand generation ahead of next phase | Public disclosures do not quantify demand impact yet |
Founder identities are well supported, but current board composition, independence, and equity control remain materially under-disclosed.
[CO002, CO006, CO007, CO031, CO032, CO050]| Stakeholder | Role | Public significance | Current diligence ask |
|---|---|---|---|
| Wellington Management | Lead investor in 2022 Series F | Led the round that set the current disclosed $3.7B valuation | Confirm whether Wellington still anchors any IPO preparation expectations |
| CapitalG | Growth investor | CapitalG involvement marked 2019 unicorn status and Alphabet adjacency | Clarify present ownership and board influence |
| Tiger Global | Late-stage investor | Appears in 2020/2021 round history and is a high-visibility growth backer | Understand any internal mark changes or liquidity pressure |
| FirstMark | Early investor | Still publicly highlights the founding story and category thesis | Confirm ownership dilution across later rounds |
| Snowflake | Strategic ecosystem partner | 300+ joint customers and major marketplace/go-to-market relevance | Measure how much pipeline is partnership-sourced vs direct |
| KPMG | Consulting alliance partner | Represents systems-integrator channel credibility for enterprise transformation deals | Test whether the alliance produces recurring enterprise implementations |
This table mixes financial stakeholders and go-to-market stakeholders because Dataiku's public record is richer on ecosystem influence than on formal board rights or secondary holdings.
[CO015, CO018, CO019, CO034, CO035, CO036]1.3 Funding history, valuation, and the public scale markers that do exist
Dataiku's funding history shows a company that scaled aggressively through the 2018-2022 software bull market, then chose not to announce a fresh primary round afterward. Third-party round histories show step-ups from a $101 million Series C in 2018 to CapitalG-backed unicorn status in 2019, a $100 million Series D in 2020, a $400 million Series E in 2021 at a $4.6 billion valuation, and a $200 million Series F in December 2022 at a $3.7 billion valuation. Sacra's 2025 research pegs lifetime funding at roughly $846.8 million, which is directionally consistent with the disclosed chronology even if private-company cap-table data remains incomplete. Public operating scale is unusually visible for a private AI platform vendor. Dataiku officially reported $300 million ARR in January 2025 and more than $350 million ARR by October 2025, while customer count moved from 700+ to 750+ and employee count from 1,100+ to 1,250+ over the same period. The company also claimed 13 offices and one-in-four penetration of the Forbes Global 2000. Those are strong late-stage software metrics, but they are still press-release metrics rather than audited filings, so the chapter treats them as company claims backed by corroborating secondary analysis instead of as public-company-quality disclosures.[CO013, CO014, CO015, CO016, CO017, CO018]
| Date | Event | Type | Amount / status | Participants | Implication |
|---|---|---|---|---|---|
| 2013 | Company founded | founding | Founding year verified | Douetteau, Stenac, Cabrol, Batty | Canonical starting point for all later history |
| 2015 | Established U.S. presence | scale | U.S. expansion reported | Dataiku | Signals eventual shift toward New York-centered GTM |
| 2018-12 | Series C announced | financing | $101M | ICONIQ-led round per secondary history | Marked step-up into late-stage growth financing |
| 2019-12 | CapitalG investment / unicorn status | financing | $1.4B valuation | CapitalG and Dataiku | Confirmed category breakout and Alphabet adjacency |
| 2020-08 | Series D announced | financing | $100M | Stripes, Tiger Global | Added capital during rapid enterprise-AI buildout |
| 2021-08 | Series E announced | financing | $400M at $4.6B valuation | Tiger Global and investors | Peak disclosed bull-market valuation |
| 2022-12 | Series F announced | financing | $200M at $3.7B valuation | Wellington-led round | Reset valuation lower but kept capital access open |
| 2025-01 | ARR milestone disclosed | scale | $300M+ ARR; 700+ customers; 1,100+ employees | Dataiku | Showed continued scale despite private opacity |
| 2025-10 | ARR milestone and IPO prep disclosed | scale / governance | $350M+ ARR; Reuters IPO-prep report | Dataiku; Morgan Stanley; Citigroup | Marked transition into public-offering watchlist territory |
| 2026-06 | Maxwell Long joins as President & CRO | governance | Commercial leadership hire | Dataiku | Strengthened senior bench for the next growth phase |
Funding events before 2022 rely on public secondary histories; post-2024 operating milestones come from company releases and Reuters-sourced IPO reporting.
[CO001, CO013, CO014, CO017, CO018, CO019]From 2013 founding through 2026 leadership reinforcement, the public record shows a classic late-stage enterprise software scaling arc.
[CO001, CO013, CO014, CO018, CO019, CO020]1.4 IPO trajectory, ecosystem strength, and the main adverse readthroughs
The strongest sign that Dataiku has entered an IPO-watched phase is Reuters' October 2025 report that Morgan Stanley and Citigroup were hired to prepare for a U.S. listing that could come as soon as the first half of 2026. As of the run date there is still no public filing, so the fair reading is "preparing, not yet public." That nuance matters because the company's go-to-market and ecosystem story is clearly strengthening: Snowflake materials point to 300+ shared customers and major co-sell momentum, KPMG publicly aligned with Dataiku in 2024, and Dataiku's own partner directory shows deep coverage across hyperscalers, data platforms, and integrators. But not every signal is cleanly bullish. Gartner's review surface includes explicit criticism around private-cloud integration, and broader market analysis warns that enterprise agentic AI adoption remains constrained by trust, integration, and governance gaps. In other words, Dataiku is well positioned for the next enterprise-AI budget wave, but the speed of that wave — and whether it is large enough to justify an IPO premium above the 2022 mark — still depends on an adoption environment that remains more operationally difficult than vendor narratives suggest.[CO034, CO035, CO036, CO037, CO038, CO039]
1.5 Exhibits
02Market Analysis
2.1 Market boundary — where Dataiku actually plays
The cleanest way to define Dataiku's market is not all AI software and not even all machine learning. Dataiku sits in the governed enterprise AI orchestration layer: the software that lets organizations prepare data, develop analytics and models, operationalize AI workflows, and increasingly manage agent-based systems across multiple teams and infrastructures. That means the relevant market includes classic data-science-and-machine-learning platforms, MLOps tooling, AI-governance software, and parts of the emerging enterprise agent platform stack. It excludes raw cloud infrastructure, commodity API calls to standalone foundation-model providers, and lightweight point copilots that do not require workflow orchestration, multi-user governance, or deployment management. This distinction matters because the broadest market studies produce enormous TAMs, but those numbers include areas where Dataiku does not directly monetize. The practical substitute set is also mixed: hyperscalers bundle native tools, Databricks sells a lakehouse-plus-AI-control-plane alternative, and vendors like DataRobot, H2O.ai, and Alteryx cover adjacent automation or low-code analytics use cases. Dataiku wins when customers need one governed operating layer across these fragmented components rather than one more isolated tool.[CM001, CM015, CM016, CM017, CM018, CM019]
| Segment / category | Included spend | Excluded spend | Primary buyer / payer | Relevance to Dataiku |
|---|---|---|---|---|
| Broad ML software | Model development, data prep, deployment, analytics tooling | Raw cloud compute and generic API consumption | CIO / CDO / analytics budget owner | Useful ceiling context but too broad on its own |
| DSML platforms | Collaborative analytics, notebooks, visual pipelines, ML lifecycle | Point model-hosting tools without workflow layer | Data science leaders / platform owners | Core legacy category for Dataiku |
| MLOps | Model deployment, monitoring, lineage, registries, reproducibility | Pure experimentation tools with no production controls | ML engineering / platform teams | Important overlap for Dataiku but narrower than the whole product |
| AI governance | Risk, compliance, lineage, approval, monitoring, policy controls | General cyber or GRC spend unrelated to AI workflows | Risk, compliance, AI governance office | Fast-growing segment aligned with Dataiku differentiation |
| Enterprise agent platforms | Agent building, orchestration, tool use, governed execution | Consumer copilots and generic chat subscriptions | Innovation office, platform engineering, AI CoE | Newest expansion zone for Dataiku |
| Adjacent low-code analytics | Workflow analytics, prep, dashboard automation | Heavy-code ML engineering stacks | BI / operations leaders | Adjacency where Alteryx and similar tools compete for simpler use cases |
The key discipline is to treat Dataiku as a governed orchestration and collaboration layer, not as a proxy for every dollar of AI infrastructure or every LLM token spent in the enterprise.
[CM001, CM015, CM016, CM017, CM018, CM019]Layered lens showing why Dataiku should be valued against narrower orchestration and governance categories, not only broad ML TAM.
[CM001, CM035, CM038]2.2 Sizing lenses — big market, but multiple legitimate denominators
Public market sizing for Dataiku's opportunity spans an unusually wide range because analysts are measuring overlapping but non-identical categories. At the broadest level, Fortune Business Insights puts global machine-learning software spend at $65.28 billion in 2026, while Precedence Research puts it at $126.91 billion in the same year. Neither number is wrong in a strict sense; they are simply broad. Narrower category lenses are more useful for underwriting Dataiku. MarketsandMarkets sizes MLOps at $5.9 billion by 2027, AI governance at $5.78 billion by 2029, and AI Studio at $32.7 billion by 2029. Those narrower lenses more closely map to the governed development, deployment, and monitoring surfaces where Dataiku monetizes. The right analytical conclusion is not to choose one TAM and defend it, but to preserve the range and explicitly show why Dataiku may capture pieces of all three layers. For valuation work, this chapter treats broad ML market studies as ceiling context, MLOps and AI-governance studies as closer proxies for the company's current SAM, and the AI-studio framing as a bridge category that explains why the market can widen as agents and governance converge.[CM002, CM003, CM004, CM005, CM006, CM007]
| Lens | Publisher | Base year / forecast year | Value | Growth | Why it matters | Limitation |
|---|---|---|---|---|---|---|
| Broad ML market | Fortune Business Insights | 2026 | USD 65.28B | 26.7% CAGR to 2034 | Upper-bound context for enterprise AI software demand | Too broad; includes many workloads Dataiku does not directly monetize |
| Broad ML market | Precedence Research | 2026 | USD 126.91B | 33.66% CAGR to 2035 | Shows how expansive definitions can double TAM assumptions | Methodology differs sharply from Fortune and is not directly comparable |
| MLOps market | MarketsandMarkets | 2027 | USD 5.9B | 41.0% CAGR | Closer proxy for deployment/monitoring layer where Dataiku competes | Forecast year differs from other lenses |
| AI governance market | MarketsandMarkets | 2029 | USD 5.78B | 45.3% CAGR | Relevant to Dataiku's governance-led enterprise pitch | Still narrower than the full collaboration/orchestration platform |
| AI Studio market | MarketsandMarkets | 2029 | USD 32.7B | 38.4% CAGR | Bridge category between DSML, MLOps, and emerging agent tooling | Vendor landscape is broad and heterogeneous |
| Enterprise AI orchestration layer | Author synthesis | 2026 | Not directly isolated | Not directly isolated | Best describes Dataiku's real SAM conceptually | Cannot be credibly quantified from public data alone |
Different publishers are sizing different category cuts; preserve the spread instead of forcing one reconciled TAM.
[CM003, CM004, CM005, CM006, CM007, CM038]Forecast growth and size estimates vary materially depending on the category definition chosen.
[CM003, CM004, CM005, CM006, CM007]2.3 Buyers, users, and the enterprise adoption path
The buyer map for Dataiku is structurally enterprise-heavy. Broad research says large enterprises dominate ML spending, and Dataiku's own customer surface confirms it: pharma, banking, logistics, manufacturing, insurance, and exchange operators recur across public references. The usual economic buyer is a CIO, CDO, analytics or data-platform lead, or a transformation owner sitting over a governed data budget. The user base is wider than that buyer base. Snowflake, AWS, and Google Cloud partnership materials all emphasize business and domain experts, not only data scientists, which means the platform category is bought centrally but monetized through broad internal usage. Dataiku's archived packaging hints at the standard land-and-expand motion: free or small-team entry is possible, but the real product value emerges when customers need automation, deployment, approval workflows, and broader governance. Systems integrators matter because buyers often need operating-model change, not just software installation. KPMG's alliance and Snowflake's 300+ joint-customer claim both show that ecosystems help carry the product into large-account transformation programs where procurement cycles are long and cross-functional.[CM007, CM008, CM009, CM010, CM011, CM012]
| Segment / vertical | Buyer | Primary users | Payer / budget owner | Adoption trigger | Why Dataiku fits |
|---|---|---|---|---|---|
| Life sciences / pharma | Digital / analytics leader | Scientists, analysts, engineers | Transformation budget | Need governed GenAI, analytics, and repeatable workflows | Public proof from Novartis and Roche-style use cases |
| Financial services / banking | CIO / operations sponsor | Risk and ops teams | Platform or COO budget | Need lineage, compliance, and production analytics | Strong fit for trust/governance narrative and Standard Chartered proof |
| Manufacturing / industrial | Operations or digital lead | Engineers and analysts | Operations budget | Need cross-site analytics and model operationalization | Matches Mitsubishi Electric and Michelin-style deployments |
| Logistics / supply chain | Operations analytics lead | Shared-services analysts, support teams | Operations or shared-data budget | Need workflow automation and multi-source data prep | Visible in Geodis and broader logistics references |
| Retail / CPG | Commercial insights lead | Analysts and business users | Commercial analytics budget | Need self-service with oversight | Fits low-/no-code collaboration pitch |
| Enterprise-wide AI CoE | Chief Data Officer or platform owner | Mixed business and technical teams | Central AI platform budget | Need one control plane across clouds, models, and teams | This is the archetypal high-value Dataiku deployment |
Dataiku is bought centrally but monetized through broad internal usage, which is why buyer and user personas differ materially.
[CM007, CM010, CM011, CM012, CM013, CM014]Relative intensity of governance, technical complexity, business-user breadth, partner reliance, and budget centrality by buyer segment.
[CM011, CM012, CM031, CM032, CM036]Enterprise AI adoption narrows sharply from experimentation to governed multi-agent production.
[CM024, CM027]2.4 Growth drivers and adoption constraints
The primary market driver is the shift from AI experimentation to governed production use. Dataiku's own releases frame that move explicitly, and the partner pages show why: enterprises want GenAI and agent capabilities without losing control over cost, data lineage, or deployment standards. Trust and governance are therefore not just risk controls; they are demand creators. But the category's main constraints are equally visible. Deloitte's 2025 poll found only a small minority already using agentic AI in finance and accounting, with trust as the top barrier. SiliconANGLE and the 2026 arXiv industry study reinforce the same message from different angles: data quality, integration, verification, and human oversight remain bottlenecks. Observer adds the economic readthrough that returns may take years, not quarters, when architectures are complex. For Dataiku, this means the market is attractive precisely because the problem is hard — but also that sales cycles, proof requirements, and implementation friction will remain significant. The company benefits from the need for governance and orchestration, yet that same need slows category penetration and keeps a fully constrained SAM/SOM analysis incomplete without internal win-rate and budget data.[CM022, CM023, CM024, CM025, CM026, CM027]
| Driver / constraint | Direction | Timing | Evidence | Implication for Dataiku | Diligence ask |
|---|---|---|---|---|---|
| Enterprise shift from experimentation to operationalization | Positive | Now | Dataiku Oct 2025 release | Favors governance-heavy orchestration platforms | Verify whether deals are expanding beyond pilots to platform standards |
| Need for trust, lineage, and explainability | Positive | Now | Deloitte trust barrier + Dataiku positioning | Governance is a demand creator, not only a compliance tax | Measure governance-led win rates vs feature-led win rates |
| Large-enterprise concentration of spend | Positive | Durable | Fortune large-enterprise share | Supports Dataiku's enterprise-focused GTM | Check how much whitespace remains in Fortune-2000 accounts |
| Multi-cloud / vendor-agnostic demand | Positive | Durable | AWS, Google Cloud, Databricks partner pages | Validates Dataiku's orchestration-above-the-stack pitch | Ask what percentage of wins are multi-cloud or hybrid |
| Integration complexity and data readiness gaps | Negative | Now | SiliconANGLE, arXiv, Gartner review | Sales cycles and implementation friction remain high | Request median time-to-production and professional-services dependency |
| Slow early agentic AI penetration | Negative | Now | Deloitte 13.5% usage figure | Agent upside is real but near-term category monetization may lag hype | Ask for pipeline split between classic analytics/ML and agents |
| Long ROI payback in complex deployments | Negative | Medium term | Observer deployment analysis | Could slow budget approvals despite strategic interest | Request reference accounts with measured payback timelines |
| SI and cloud-partner channel leverage | Positive | Now | KPMG alliance; Snowflake 300+ customers | Partners can reduce selling friction and expand reach | Measure sourced pipeline and attach rates by partner |
Several constraints are the flip side of Dataiku's opportunity: governance and integration pain create demand, but they also slow adoption and elongate cycles.
[CM022, CM023, CM024, CM025, CM026, CM027]2.5 Exhibits
03Competitors
3.1 Landscape and the closest rivals
Dataiku's competitive set is wider than a single DSML shortlist. The practical buyer alternative set includes unified data-and-AI platforms such as Databricks, hyperscaler-native ML stacks such as Amazon SageMaker, Azure Machine Learning, and Vertex AI, plus narrower specialists such as DataRobot, H2O.ai, and Alteryx that solve adjacent jobs with different deployment and pricing assumptions. The most important distinction is that Dataiku is trying to be a neutral control layer across existing enterprise data estates, not the only place where storage, compute, and model serving happen. That makes Databricks the closest broad-platform peer because it now sells data, governance, MLOps, and agent tooling in one product family, while the hyperscalers compete by making AI another feature of existing cloud procurement. The specialists matter because they show where narrower time-to-value, AutoML, or analytics-automation motions can still divert budgets away from a full orchestration platform.[CP001, CP002, CP003, CP004, CP005, CP006]
| Company | Category | Scale / funding signal | Target segment | Key differentiation | Key limitation |
|---|---|---|---|---|---|
| Dataiku | Neutral enterprise AI orchestration platform | $350M+ ARR in Oct. 2025; 750+ customers | Large enterprises with multi-person governed AI workflows | Infrastructure-neutral collaboration, governance, and deployment flexibility | Limited public pricing transparency and smaller scale than Databricks |
| Databricks | Unified data + AI platform | $6.9B annualized revenue in 2026; $134B valuation in Dec. 2025 | Enterprises consolidating data engineering, analytics, and AI | Owns both data and AI workflow surfaces with strong agent roadmap | Less neutral because it also seeks to be the core data platform |
| Amazon SageMaker | Hyperscaler-native ML stack | AWS-scale procurement and granular usage billing | AWS-centric engineering and ML teams | Native AWS integration and metered pricing by workload | Can feel componentized rather than neutral across heterogeneous stacks |
| Azure Machine Learning | Hyperscaler-native ML stack | Microsoft enterprise agreement leverage and pay-as-you-go options | Azure-first enterprises, especially existing Microsoft estates | Enterprise MLOps and responsible-AI framing inside Azure estate | Economics and roadmap stay tied to Azure consumption choices |
| Vertex AI / Agent Platform | Hyperscaler-native ML and agent stack | Google model, training, and inference pricing disclosed publicly | GCP-centric data and GenAI teams | Gemini-native tooling plus granular model operations pricing | Still anchored to Google Cloud rather than cross-cloud neutrality |
| DataRobot | Specialist enterprise AI suite | ~$285M revenue in 2024; prior $6.3B peak valuation | Teams prioritizing guided enterprise AI delivery without full platform rebuild | Deployment choice across on-prem, VPC, and SaaS with integrated suite pitch | Adverse evidence shows weaker category durability versus bundled platforms |
| H2O.ai | Specialist hybrid / AutoML platform | $100M Series E at $1.7B valuation in 2021; 20,000 organizations claimed | Users wanting open-source lineage, hybrid deployment, and AutoML | Open-source roots and hybrid-cloud flexibility | Much smaller disclosed capital base than Databricks or Dataiku |
| Alteryx | Adjacent analytics automation substitute | Acquired for $4.4B in 2023; 8,000+ customers | Business analytics and low-code automation teams | Strong democratized analytics and workflow automation brand | Less focused on end-to-end ML / agent lifecycle depth |
Profiles group buyers into the platforms most likely to absorb the same budget line item or workflow ownership that Dataiku targets.
[CP001, CP003, CP005, CP007, CP008, CP009]Dataiku scores highest when the axes are infrastructure neutrality and governed workflow breadth, but Databricks closes the gap by owning more adjacent data-platform budget.
Coordinates are ordinal synthesis from the reviewed source pack. X-axis represents infrastructure neutrality / deployment flexibility; Y-axis represents governed AI workflow breadth.
[CP002, CP003, CP005, CP007, CP008, CP009]3.2 Capability and packaging comparison
The public packaging surfaces show a clear divide. Hyperscaler-native options expose granular usage pricing, while Dataiku and most specialist peers still sell an enterprise contract and architecture decision rather than a simple list-price SKU. Dataiku's archived plans page shows exactly how it historically framed the product: broader connector depth, automation, and governed deployment capabilities unlocked as teams move from free or small-team use into enterprise-scale adoption. Databricks is somewhat more transparent because it publishes a price list for SKU groups, but even there the buyer still has to map usage onto cloud-specific services and discounts. AWS, Azure, and Vertex AI make metered economics explicit, which helps comparison at the feature level but also shifts cost risk toward architecture and runtime choices. DataRobot, H2O.ai, and Alteryx stay closer to demo-led or contact-sales motions, which is typical for enterprise software but makes apples-to-apples TCO comparisons difficult from public information alone.[CP012, CP013, CP014, CP015, CP016, CP017]
| Buying criterion | Dataiku | Databricks | SageMaker | Azure ML | Vertex AI | DataRobot | H2O.ai | Alteryx |
|---|---|---|---|---|---|---|---|---|
| Cross-cloud / on-prem flexibility | Strong | Moderate | Limited to AWS | Limited to Azure | Limited to GCP | Strong | Strong | Moderate |
| Business-user accessibility | Strong | Moderate | Weak-to-moderate | Moderate | Moderate | Moderate | Moderate | Strong |
| Governed end-to-end workflow breadth | Strong | Strong | Moderate-to-strong | Moderate-to-strong | Moderate-to-strong | Moderate | Moderate | Moderate |
| Native data-platform ownership | Weak | Strong | Moderate within AWS | Moderate within Azure | Moderate within GCP | Weak | Weak | Weak |
| Partner / channel leverage | Strong | Strong | Strong | Strong | Strong | Moderate | Moderate | Moderate |
| Public pricing transparency | Low | Moderate | High | Moderate | High | Low | Low | Low |
Scores are ordinal synthesis from reviewed product, partner, and pricing surfaces; unsupported realized-cost claims are intentionally not inferred.
[CP002, CP012, CP013, CP014, CP015, CP016]| Company | Public contract model | What is visibly included | What remains unknown | Implication |
|---|---|---|---|---|
| Dataiku | Enterprise tiering; archived free/discover/business/enterprise framing | Collaboration, connectors, automation, deployment options, governance depth | Current realized pricing, discounts, and cloud-hosting uplift are not public | Procurement is architecture-led rather than self-serve price-led |
| Databricks | Usage-based SKU pricing with cloud-specific price lists | Platform services sold as list-price SKUs and groups | Realized discounts and full workload-specific TCO remain private | Better public transparency than most peers, but not simple for finance teams |
| Amazon SageMaker | Feature-level metered pricing by instance, duration, storage, and inference mode | Notebook, training, inference, feature store, processing, MLflow and other services | Full bill depends on architecture, instance choices, and workload intensity | Can start small, but compute-heavy success can raise spend unpredictably |
| Azure Machine Learning | Quote-driven Azure service with pay-as-you-go, reservations, and savings plans | End-to-end ML lifecycle service layered onto Azure compute choices | Effective enterprise price depends on Microsoft agreement terms and chosen infrastructure | Strong for existing Azure estates; less transparent for outsider comparison |
| Vertex AI / Agent Platform | Metered training, deployment, AutoML, and prediction pricing | Hourly model operations, no minimum usage duration, per-count forecasting tiers | Blended cost still depends on model choice, endpoint design, and GCP usage pattern | Attractive for bursty experimentation but native-cloud lock-in remains |
| DataRobot / H2O.ai / Alteryx | Mostly demo-led or contact-sales enterprise motion | Suite positioning, deployment options, and packaging cues are public | No detailed enterprise list pricing in fetched surfaces | Public TCO comparison with Dataiku is structurally incomplete |
This table compares what the fetched source pack actually exposes, not what a private negotiated contract might ultimately look like.
[CP012, CP013, CP014, CP015, CP016, CP017]Dataiku leads on neutrality and mixed-persona workflow breadth, whereas hyperscalers win on native procurement and Databricks wins on adjacent platform ownership.
Values are ordinal 1-5 scores derived from fetched product, partner, and pricing pages rather than from one third-party benchmark.
[CP012, CP013, CP014, CP015, CP016, CP017]Public scale markers show why Databricks is the most severe competitive threat, while DataRobot and Alteryx illustrate how adjacent categories can follow very different economic paths.
KPI strip mixes company-disclosed operating metrics and independent scale markers. It is intended to summarize competitive readiness, not market share.
[CP019, CP020, CP021, CP022, CP033, CP034]3.3 Switching costs and distribution power
Dataiku's real moat is not a single model or proprietary data asset; it is the operating convenience of governed collaboration across heterogeneous tools, clouds, and user types. That matters most in enterprises that already run Snowflake, Databricks, AWS, Google Cloud, and internal code-based tooling in parallel. In that context, Dataiku can win as the orchestration and governance layer that spans the stack. But the same architecture creates its main strategic vulnerability: hyperscalers can start from the buyer's existing contract, identity system, and cloud data gravity, while Databricks can start from ownership of the data-and-compute workflow itself. Public partner pages show Dataiku leaning into coexistence rather than rip-and-replace, which broadens distribution and reduces isolation. Specialists still have room where the buyer wants narrow time-to-value or low-code automation, but they generally lack the same breadth of ecosystem leverage or disclosed capital scale as the largest platform rivals. It also means the sales contest is often decided by integration credibility and change-management comfort, not just by raw model features.[CP018, CP023, CP024, CP025, CP026, CP027]
3.4 Moat durability and adverse readthroughs
The adverse evidence does not say Dataiku is weak; it says the category is unforgiving when a standalone AI platform loses differentiation against bundled infrastructure or a broader system of record. Databricks is the highest-severity threat because it is scaling faster, is far better capitalized, and keeps widening from data infrastructure into governance, AI agents, and application surfaces. The hyperscalers are the second structural threat because they can make native tooling feel ‘free enough’ inside a cloud commitment. DataRobot offers the clearest warning case: a once-hot standalone AI company can lose strategic relevance quickly when native cloud tools and paradigm shifts change what buyers care about. H2O.ai and Alteryx show that adjacencies remain valuable, but also that not every adjacent category earns AI-platform multiples. Dataiku's best defense is to keep being the neutral, governed workflow layer in accounts that will remain multi-cloud and multi-tool even as native AI features improve.[CP031, CP032, CP033, CP034, CP035, CP036]
| Moat claim | Threat | Severity | Why it is credible | Mitigation / diligence ask |
|---|---|---|---|---|
| Infrastructure-neutral governance layer | Hyperscalers make native tooling good-enough inside existing cloud contracts | High | AWS, Azure, and Google all expose first-party lifecycle and agent tooling with public usage pricing | Request win/loss data by cloud and by deployment model |
| Broad workflow coverage across business and technical users | Databricks keeps expanding from data platform into AI agents and governed delivery | High | Databricks now markets agent building, governance, and model lifecycle on top of its data platform | Test whether Dataiku wins when Databricks is already the data standard |
| Partner-led distribution | Partners can steer budgets toward their own native services | Medium | Dataiku co-sells with AWS, Google Cloud, NVIDIA, Databricks, and Snowflake, all of which have their own agendas | Quantify sourced pipeline, influenced pipeline, and partner dependency |
| Specialist time-to-value advantage in selected use cases | Narrower tools can win department budgets before enterprise platform standardization | Medium | DataRobot, H2O.ai, and Alteryx still market ease of use, hybrid deployment, or low-code automation | Check whether pilot losses occur on simplicity rather than capability |
| Category enthusiasm around enterprise AI | Standalone AI-platform narratives can compress when value looks bundled or overhyped | High | DataRobot's adverse trajectory and Alteryx's very different public-market outcome show the category can re-rate quickly | Request historical pricing pressure, renewals, and attach rates for governance modules |
The risk register focuses on threats to differentiation durability, not on generic market risk already covered in the market-analysis chapter.
[CP024, CP025, CP027, CP032, CP033, CP034]3.5 Exhibits
04Financials
4.1 Revenue model and monetization
The public record is strong enough to identify Dataiku’s commercial shape even if it is not strong enough to model it precisely. Dataiku is clearly a recurring enterprise software business, not an ad-supported product, a marketplace, or a purely usage-metered API vendor. Its own 2025 releases frame growth in ARR, not bookings or services revenue, and its archived plans page shows a structured progression from free or small-team entry into deeper automation, deployment, security, and governance capability. Current product pages reinforce that the platform is sold around breadth: orchestration, governance, agents, and enterprise data controls. That suggests the buyer is purchasing a cross-functional control plane whose value increases with deployment scope, not a single-module point solution. The key nuance is that public materials do not disclose realized list pricing or exact revenue recognition mechanics. Compared with Databricks and the hyperscalers, Dataiku appears more contract-oriented and less transparently usage-priced, which may improve budget predictability but also leaves outsiders unable to benchmark realized economics from public information alone.[CI001, CI002, CI003, CI004, CI005, CI006]
| Stream | Mechanism | Unit | Current value / status | Quality | Diligence ask |
|---|---|---|---|---|---|
| Core platform subscription / ARR | Recurring enterprise platform contracts | ARR / annual contract value | $300M+ ARR in Jan. 2025; $350M+ ARR in Oct. 2025 | High confidence on existence, medium confidence on current run rate | Request quarterly ARR bridge, cohort expansion, and term mix |
| Deployment / hosting monetization | Platform can be hosted by Dataiku or customer environments | Contract plus possible hosting uplift | Public deployment choices are visible; realized hosting revenue not disclosed | Medium | Request cloud-hosted revenue mix versus customer-managed deployments |
| Governance / agent breadth monetization | Platform breadth expands monetization surface inside enterprise account | Upsell / edition expansion | Current product pages emphasize agents, governance, orchestration, and shared control plane | Medium | Request module attach rates and expansion by capability family |
| Professional services / training | Implementation, consulting, and enablement likely exist but are not itemized publicly | Services fees | Public evidence indicates services exist, but no revenue split is disclosed | Low | Request services share of revenue, gross margin, and partner-versus-direct delivery split |
| Partner-delivered services | Systems integrators and cloud partners can deliver implementation around the core platform | Indirect services influence | Partner ecosystem is extensive, but economics are undisclosed | Medium | Request sourced pipeline, services attach, and partner compensation model |
| Support / maintenance | Enterprise support is likely bundled into core contracts | Included support services | Public materials imply bundled enterprise support rather than separately priced maintenance | Low | Request support burden, renewal terms, and support cost per large account |
The table distinguishes between clearly disclosed recurring software traction and less transparent implementation or module-level revenue components.
[CI001, CI003, CI004, CI005, CI006, CI007]| Company / model | Price / unit / contract | List vs. realized pricing | Discounts / unknowns | Implication |
|---|---|---|---|---|
| Dataiku | Enterprise contract model; archived tiering from free to enterprise | Archived list structure visible, current realized pricing not public | Current contract terms, discounts, and hosting uplift unknown | Budgeting may be more predictable than consumption-only models, but public benchmarking is weak |
| Databricks | Usage-priced SKUs and cloud-specific price lists | List prices are public, realized economics are negotiated | Discount ladders and full workload TCO remain private | Transparent by enterprise-software standards, but still architecture-dependent |
| Amazon SageMaker | Metered by instance, duration, storage, and inference configuration | Public feature-level pricing | Final bill depends on runtime choices and workload intensity | Strong self-serve visibility, weaker ex ante budget certainty |
| Azure Machine Learning | Azure service pricing with pay-as-you-go and reserved options | Public rate-card structure plus enterprise quote context | Actual price depends on Azure agreement and chosen infrastructure | Enterprise procurement leverage is strong inside Microsoft estates |
| Vertex AI | Model-operation, training, deployment, and prediction pricing | Public metered pricing | Model, endpoint, and usage choices drive realized cost | Good for bursty experimentation, but cost visibility depends on architecture |
| C3.ai comparator | Subscription plus usage-based runtime and hosting charges embedded in subscriptions | No simple public rate card in filing | Customer-specific structures and professional services vary | Shows how AI platforms can combine committed subscriptions with usage-linked elements |
Official pricing reveals how peers monetize; it does not reveal Dataiku’s realized contract economics, which remain private.
[CI006, CI010, CI011, CI013, CI016, CI017]Dataiku monetizes enterprise AI breadth: customers start with platform adoption, then expand into broader governance, deployment, and agent use cases that support recurring ARR.
This figure abstracts the commercial motion from public ARR disclosures, product pages, and archived packaging. It does not imply a disclosed conversion rate or attach rate.
[CI001, CI003, CI004, CI005, CI006, CI007]4.2 Sales motion and unit-economics proxies
Dataiku’s product and customer pattern imply a classic enterprise land-and-expand motion, but the hard unit-economics fields remain undisclosed. Large customers, partner ecosystems, and cloud-agnostic deployment all point to long sales cycles, high ACVs, and multistakeholder procurement, while the 2025 ARR milestones show that the model can scale beyond pilot stage. The best public proxies come from adjacent public companies. C3.ai’s filings and FY2026 results show how enterprise AI software can remain overwhelmingly subscription-led while still carrying services and prioritized engineering work that supports deployment and roadmap acceleration. Snowflake’s annual report shows the opposite end of the monetization spectrum: a consumption-led model with very high product margins, substantial remaining performance obligations, and strong operating cash generation, but also less linear revenue visibility because customers can optimize usage. Taken together, those comps suggest the real underwriting question for Dataiku is not whether the company has traction; it is whether services intensity, cloud costs, and partner economics allow the business to converge toward strong software margins as ARR scales.[CI010, CI011, CI012, CI013, CI014, CI015]
| Metric | Value / null | Confidence | Why it matters | Diligence ask |
|---|---|---|---|---|
| ARR | $350M+ disclosed in Oct. 2025 | Medium | Validates scale and enterprise relevance | Request monthly ARR bridge through runDate |
| Revenue growth | ARR more than doubled over prior three years; exact current growth rate not public | Medium | Needed to test operating leverage and valuation support | Request yearly ARR, revenue, and bookings history |
| Gross margin | Not publicly disclosed | Low | Determines whether Dataiku is converging toward attractive software economics | Request gross margin split by software, hosting, and services |
| NRR / expansion rate | Not publicly disclosed | Low | Critical for land-and-expand underwriting | Request cohort retention and dollar-based expansion by segment |
| CAC payback / sales efficiency | Not publicly disclosed | Low | Needed to judge enterprise-sales efficiency | Request sales productivity, CAC, and payback by region and channel |
| Services share of revenue | Not publicly disclosed | Low | High services intensity can cap margin and cash conversion | Request direct versus partner-delivered services share |
| Contracted backlog / RPO | Not publicly disclosed | Low | RPO would show forward visibility and renewal quality | Request deferred revenue and RPO schedule |
| Cash conversion / operating cash flow | Not publicly disclosed | Low | Determines capital needs and self-funding capacity | Request historical operating cash flow and free cash flow |
Every null field is a real underwriting blocker rather than a formatting omission.
[CI001, CI002, CI003, CI020, CI030, CI032]The observable economic path runs from enterprise acquisition to recurring ARR, but several critical unit-economics checkpoints remain undisclosed.
Nodes are qualitative because Dataiku does not disclose CAC, payback, NRR, or gross margin publicly.
[CI020, CI030, CI031, CI032, CI035, CI036]Public evidence gives a bounded view of Dataiku’s traction floor and a wide comparator envelope for margins and platform scale.
[CI001, CI003, CI014, CI015, CI022, CI033]4.3 Cost structure and capital adequacy
Public visibility on Dataiku’s actual cost structure is poor, so the chapter has to separate what is observable from what is only inferable. Observable: Dataiku crossed $300M ARR in January 2025 and $350M ARR in October 2025, has not publicly announced a new round since the 2022 Series F, and was reportedly preparing for a possible U.S. IPO in late 2025. Sacra estimates roughly $846.8M of lifetime funding and about $342.5M ARR in September 2025, which broadly fits the company’s official disclosures. Inferable: a company operating at this scale likely spends heavily on enterprise sales, customer success, product development, cloud infrastructure, and partner enablement. But inferable is not good enough for underwriting. There is no public cash balance, no burn history, no debt disclosure of note, no retention metric, and no direct gross margin disclosure. The fair capital-adequacy readthrough is therefore cautious: nothing in the public record screams distress, but nothing lets an investor verify runway or financing dependency either. IPO optionality looks like a strategic choice, not a proven necessity.[CI021, CI022, CI023, CI024, CI025, CI026]
| Field | Public status | Why it matters | Current readthrough | Diligence ask |
|---|---|---|---|---|
| Last disclosed primary capital | $200M Series F in Dec. 2022 at $3.7B valuation | Anchors historical capitalization without proving current liquidity | No later primary round publicly announced | Request current cash by quarter since Series F |
| Total funding | Sacra estimates roughly $846.8M lifetime funding | Sets dilution and financing-history context | Directionally well capitalized for a private software company | Request full cap table and debt schedule |
| Cash on hand | Not publicly disclosed | Core runway input | Unknown | Request unrestricted and restricted cash balances |
| Burn / runway | Not publicly disclosed | Needed to judge financing dependency | Unknown; no public distress signal | Request monthly burn, budget, and scenario plan |
| Capital-market optionality | Reuters reported IPO-bank preparation in Oct. 2025; 2026 IPO market backdrop improved but episodic | Matters for liquidity and future financing flexibility | Optionality appears strategic, not provably urgent | Request board-approved financing plan and IPO readiness budget |
| Debt / project-finance obligations | No material public debt or project-finance obligation surfaced | Debt can change risk profile quickly | No public evidence of heavy debt burden | Request debt facilities, covenants, and off-balance-sheet commitments |
This table focuses on forward capital adequacy, not on repeating the round-by-round history already covered in Company Overview.
[CI021, CI022, CI023, CI024, CI025, CI026]Public evidence points to a software business with enterprise-sales and cloud-delivery needs, but not to a capital-hungry hardware or project-finance model.
The map identifies visible cash uses and financing options, not a disclosed budget.
[CI021, CI022, CI024, CI026, CI027, CI028]4.4 Financial verdict and diligence blockers
The positive verdict is straightforward: Dataiku has real scale, recurring revenue, and enough category credibility to be discussed alongside IPO candidates rather than private experiments. The negative verdict is equally straightforward: the public record still does not reveal enough to underwrite margin path, cash efficiency, or dilution risk with confidence. ARR milestones, customer breadth, and platform positioning all support the idea of a high-quality enterprise software business. Yet the missing fields remain the ones that separate an interesting private company from an investable one: net revenue retention, gross margin, services mix, CAC payback, sales productivity, current cash, burn, and contractual backlog. Public-company comparisons are useful only as guardrails. Snowflake shows what strong software economics can look like; C3.ai shows how execution problems and services intensity can compress them. Dataiku plausibly sits somewhere between those poles, but the exact position is unknowable from public evidence. The right underwriting stance is therefore not bearish on traction, but disciplined on missing economics.[CI020, CI024, CI025, CI030, CI032, CI034]
| Missing private metric | Impact | Why public evidence is insufficient | Exact diligence path |
|---|---|---|---|
| Gross margin by stream | Without it, margin-path underwriting is speculative | ARR releases do not disclose software versus services margin | Request audited gross-margin bridge by stream and by deployment model |
| NRR, churn, and cohort expansion | Revenue quality cannot be judged from ARR milestones alone | No cohort data or NRR appears in public materials | Request cohort tables by vintage, segment, and geography |
| Current cash and burn | Runway cannot be verified | No public balance sheet for the private company | Request monthly cash waterfall and 12-month operating plan |
| Services mix and partner economics | Implementation intensity may materially affect gross margin | Partner footprint is visible, economics are not | Request direct-service mix, partner attach, and statement-of-work economics |
| Contract duration, deferred revenue, and RPO | Backlog and revenue visibility remain unknown | No public filing exposes contractual backlog | Request deferred-revenue schedule and RPO disclosure |
| Sales efficiency and CAC payback | Hard to assess whether growth is efficient or simply expensive | No public disclosure on pipeline conversion or payback | Request sales-capacity model, CAC, payback, and quota attainment |
These are the minimum missing metrics required before a serious valuation or financing recommendation should be finalized.
[CI024, CI025, CI030, CI032, CI035, CI036]4.5 Exhibits
05Product & Technology
5.1 What the product actually delivers
Dataiku is best understood as a shared operating layer for enterprise AI work rather than as a single modeling feature. The current product page frames the platform around three functions: people build, orchestration connects, and governance protects. That framing matches the underlying documentation. DSS combines visual data preparation, notebook-style code work, automation, deployment, and API access. Govern adds a separate oversight node for tracking AI initiatives, approvals, registries, and audit workflows. Current marketing around agents and governed AI highlights that Dataiku now wants to be where enterprises design, monitor, and route agentic workloads rather than just train classical ML models. The product therefore spans multiple personas: analysts, data scientists, ML engineers, platform owners, governance teams, and business stakeholders. The important technical implication is that Dataiku’s value is less about any one algorithm and more about coordinating heterogeneous human, data, model, and approval workflows in one controlled environment.[CE001, CE002, CE003, CE004, CE005, CE006]
| Module / asset | Primary user | Status / maturity | Differentiation | Diligence gap |
|---|---|---|---|---|
| Core DSS platform | Analysts, data scientists, engineers | Mature core platform | Combines visual and code workflows in one environment | No public benchmarked performance data |
| Automation and deployment | ML engineers, platform owners | Mature and actively maintained | Bridges build, deploy, monitor, and operate workflows | Public uptime / SLA detail not visible in fetched pack |
| Dataiku Govern | Governance, risk, AI oversight teams | Mature add-on node with advanced features | Central registries, signoff rules, workflow tracking, audit timeline | Current customer adoption of Govern not publicly quantified |
| LLM Mesh / generative AI layer | AI platform teams and app builders | Actively expanding in v14 release stream | Model abstraction, routing, safety controls, and GenAI workflow support | Public docs do not quantify latency, routing cost, or accuracy uplift |
| AI agents / agent management | AI app builders and governance owners | Newer but clearly active priority area | Governed agent building, management, and monitoring at enterprise scale | Public proof of production outcomes remains limited |
| APIs and developer tooling | Coders, integrators, admins | Established and externally visible | Python APIs, developer guide, client tooling, automation surface | Repo activity proof is modest; broader external developer footprint unclear |
The module split reflects what is visible across product pages, docs, and release notes, not internal SKU granularity.
[CE001, CE002, CE003, CE005, CE006, CE007]| User job | Current workflow problem | Dataiku solution | Measurable benefit | Limitation |
|---|---|---|---|---|
| Build cross-functional AI project | Work is fragmented across business, data, and engineering teams | Shared platform with visual and code interfaces | Potentially faster collaboration and controlled handoffs | No public time-to-value benchmark in fetched pack |
| Operationalize governed GenAI | Enterprises need model routing, controls, and visibility | LLM Mesh, agent tooling, and governance features | Centralized control of agentic workflows | Public evidence does not quantify reliability or cost savings |
| Track AI initiatives and approvals | Shadow AI and audit trails are hard to manage manually | Dataiku Govern workflows, registries, signoff rules, alerts | Improved audit readiness and policy enforcement | Actual enterprise process adoption rates undisclosed |
| Integrate platform into cloud stack | Teams want AI on existing cloud and data estates | Partner-led deployment with AWS, Google Cloud, Databricks, NVIDIA and others | Lower need to rip and replace existing infrastructure | Integration complexity still appears to be a real risk in some environments |
| Automate platform actions via code | Teams need repeatable programmatic operations | Python APIs, developer guide, API client, scenarios, code recipes | Supports extensibility and automation | External community depth is not clearly visible from public signals |
Benefits are stated conservatively because public pages emphasize capability more than quantified ROI.
[CE002, CE003, CE004, CE007, CE013, CE016]Dataiku layers a shared build surface over code APIs, governance, deployment operations, and external cloud or model ecosystems.
[CE001, CE003, CE005, CE007, CE011, CE015]A typical Dataiku workflow moves from data and project setup through model or agent creation, governed review, deployment, and monitored reuse.
[CE002, CE004, CE005, CE016, CE017, CE020]5.2 Architecture, deployment, and APIs
The most consistent product theme across official pages, archived packaging, partner pages, and technical docs is flexibility. Dataiku is documented as supporting SaaS, customer-managed cloud, and on-prem or private-cloud styles of operation. It exposes visual interfaces for non-coders but also a full developer surface through APIs, code recipes, notebooks, scenarios, and client tooling. The Python API reference explicitly says Dataiku tools can be used anywhere code runs inside DSS, while the open GitHub API client and README show that external automation against the platform is not an internal-only capability. Partnership pages with AWS, Google Cloud, Databricks, and NVIDIA further imply that Dataiku is architected to sit on top of customer infrastructure and external AI stacks rather than replace them. That is strategically important: the platform’s technical identity is orchestration-first and integration-first. It also means that dependencies on cloud, model, and partner ecosystems are a feature of the design, not an accidental by-product.[CE007, CE008, CE011, CE012, CE013, CE018]
| Layer / component | Role | Dependency | Risk |
|---|---|---|---|
| Visual workflow and UI layer | Makes data prep, modeling, dashboards, and agent design accessible | Depends on DSS core and release cadence | Can drift into marketing breadth if operational proof is thin |
| Code and API layer | Enables notebooks, recipes, automation, and external integration | Depends on Python APIs, client libraries, and developer tooling | Versioning and integration complexity can rise with platform breadth |
| Governance node | Tracks assets, approvals, templates, and registries | Depends on Govern instance setup and policy design | Strong only if customers operationalize governance processes |
| LLM / agent orchestration layer | Routes models, tools, and agent workflows | Depends on model providers, APIs, cloud infrastructure, and guardrails | Fast-moving dependency landscape can create change-management burden |
| Deployment and runtime layer | Moves projects into production and monitoring flows | Depends on customer cloud, on-prem, or hosted deployment choices | Operational burden varies widely by deployment model |
| Partner and ecosystem layer | Connects Dataiku to cloud, data, and accelerator ecosystems | Depends on third-party partner priorities and compatibility | Partner dependency can improve distribution but add technical coupling |
This is an operating-architecture synthesis from official docs, partner pages, and release notes rather than a vendor-published block diagram.
[CE007, CE011, CE012, CE013, CE015, CE021]Dataiku owns the control plane, but important parts of the product depend on cloud platforms, model providers, partner ecosystems, security posture, and customer infrastructure.
[CE013, CE014, CE021, CE022, CE030, CE033]5.3 Trust, security, and governance controls
Governance is not a thin checkbox layer in the public Dataiku story; it is one of the main product pillars. The govern product page and the Govern documentation both describe centralized project tracking, asset registries, workflow approvals, signoff rules, and audit timelines. Public materials also connect these controls directly to regulatory pressure, explicitly naming EU AI Act readiness and shadow-AI reduction. The technical docs go further by describing Standard and Advanced govern licenses, with advanced features for GenAI registries, custom governance templates, custom actions, and scripting. Security evidence is more mixed but still credible. Dataiku’s security page and security documentation show active security operations, while a 2026 warning page documents response guidance for Linux local-privilege-escalation vulnerabilities affecting Dataiku environments. The company’s SOC 2 announcement is dated and should not be treated as a full current certification inventory, but it still adds evidence that formal compliance work has been part of the product and operating story for years.[CE005, CE006, CE014, CE015, CE016, CE017]
| Control / certification / quality signal | Status | Scope | Gap |
|---|---|---|---|
| Govern signoff rules | Documented | Approval workflows can block deployment until requirements are met | No public evidence on how widely customers use them in production |
| Audit timeline and registries | Documented | Bundle, model, and LLM registries plus audit-ready timelines | No public metrics on audit efficiency or false-positive reduction |
| EU AI Act readiness messaging | Documented on official govern page | Compliance acceleration is explicitly part of product pitch | Public legal mapping detail remains high level |
| Security documentation and patch guidance | Documented | Docs and 2026 security warning show active response guidance | Detailed security architecture and testing evidence not fully exposed publicly |
| SOC 2 compliance announcement | Historical proof point | Shows formal compliance work and enterprise trust signaling | Announcement is dated and not a full current certification inventory |
The fetched pack proves governance and security process surfaces exist, but not their quantitative effectiveness.
[CE014, CE015, CE016, CE017, CE027, CE033]Core workflow and deployment capabilities look mature, while governance and agent surfaces appear newer but clearly active and expanding.
Scores are ordinal 1-5 syntheses from official docs, release notes, and public developer surfaces rather than from a third-party benchmark.
[CE007, CE008, CE009, CE017, CE019, CE025]5.4 Roadmap, dependencies, and product risks
Release notes make clear that Dataiku is shipping on an ongoing enterprise cadence rather than resting on an older DSS core. The version 14 release stream shows repeated updates in June and July 2026 across agentic AI and RAG, LLM Mesh, governance, MLOps, AI assistants, data quality, Git, security, and code tooling. That supports a maturity thesis: Dataiku is not a pilot-era platform. At the same time, the roadmap evidence also highlights dependencies and risk. Release notes mention Python version changes, Java minimum-version bumps, Llama model removals, container-image and OS changes, and cloud-stack cautions. Those are normal signals for a real enterprise platform, but they also confirm that Dataiku’s product depends on underlying cloud, OS, model, and open-source ecosystems remaining stable. External product proof is thinner than official docs. Gartner review evidence still includes at least one criticism around private-cloud integration, and the public surface does not provide benchmarked performance, uptime SLAs, or quantified agent accuracy. For diligence, the main product question is not breadth; it is operational depth and implementation friction.[CE009, CE010, CE022, CE023, CE024, CE026]
| Date / stage | Feature / milestone | Status | Implication | Source |
|---|---|---|---|---|
| Version 14.7.2 – Jul. 10, 2026 | Agentic AI & RAG, LLM Mesh, governance, AI assistants, coding & API, plugins | Released | Shows active roadmap breadth across both GenAI and platform operations | v14 release notes |
| Version 14.7.1 – Jul. 1, 2026 | AI services, Cobuild, Spark, Git, security | Released | Suggests continuing platform hardening and developer-surface work | v14 release notes |
| Version 14.7.0 – Jun. 18, 2026 | New feature: Cobuild, charts, data quality, performance | Released | Signals ongoing investment beyond purely GenAI features | v14 release notes |
| Version 14.6.2 – Jun. 11, 2026 | MLOps, datasets and connections, scenarios and automation, code studio | Released | Reinforces maturity in deployment and operations, not only experimentation | v14 release notes |
| 2026 security warning | LPE vulnerability response guidance for Dataiku environments | Published guidance | Shows active security operations and dependency management responsibilities | official security surface |
Roadmap evidence is release-note-driven, so it reflects shipped work better than marketing promises.
[CE009, CE010, CE014, CE022, CE028, CE034]5.5 Exhibits
06Customers
6.1 Customer segmentation and buyer map
Dataiku's customer footprint is clearly enterprise-led rather than SMB-led, and the public evidence spans several regulated and operationally complex verticals. The company itself said it served more than 700 organizations in January 2025 and more than 750 in October 2025, while a July 2024 alliance release with KPMG said Dataiku already had more than 600 customers and 200 Forbes Global 2000 customers. The named-customer roster visible on Dataiku's customer surfaces clusters into healthcare and life sciences (Johnson & Johnson, Novartis, Roche), manufacturing and industrials (Michelin, Mitsubishi Electric, SLB), financial services and capital markets (Standard Chartered, Euronext), logistics and transportation (Geodis, Prologis), and food/agriculture (Perdue Farms). The user is rarely a single data scientist. Instead, the recurring pattern is a cross-functional coalition that mixes data scientists, analysts, engineers, operations experts, and business-domain users on a common governed platform. That makes the economic buyer more likely a centralized data, analytics, digital-transformation, or business-platform budget rather than a seat-by-seat departmental purchase. The segmentation evidence is therefore strong on who uses Dataiku and what kinds of enterprises adopt it, but still weak on exact revenue mix by vertical, geography, or account size band.[CU001, CU002, CU003, CU004, CU005, CU006]
| Segment | Example customers | Buyer / user / payer | Primary use cases | Revenue / strategic value | Key uncertainty |
|---|---|---|---|---|---|
| Healthcare & life sciences | Johnson & Johnson, Novartis, Roche | Data/AI leaders + business teams + regulated stakeholders | GenAI assistants, market research, patent analysis, analytics standardization | High-value regulated accounts that validate governance needs | No public revenue mix or renewal data by sub-vertical |
| Manufacturing & industrial | Michelin, Mitsubishi Electric, SLB | Engineering, operations, manufacturing excellence, digital programs | Factory analytics, root-cause analysis, energy optimization, field engineering | Operational use cases support broad seat/workflow expansion | Unknown whether deployments are globally standardized or regionally partial |
| Financial services & capital markets | Standard Chartered, Euronext | Analytics centers of excellence, product teams, bank technology groups | Market-share analytics, FP&A, portfolio analytics, governance | Strong fit for governed, explainable enterprise analytics | No disclosed ARR per account or banking concentration |
| Logistics, real estate & transportation | Geodis, Prologis | Operations, IT support, data & analytics teams | Ticket triage, forecasting, geospatial analysis, enterprise chat | Shows platform relevance beyond classic ML teams | Outcome data is strong, but contract scale is undisclosed |
| Food & agriculture | Perdue Farms | Food-safety operations and business users | Reporting automation, data prep, self-service analytics | Proof that Dataiku can monetize operational data workflows outside tech-heavy buyers | Single public case study does not show category depth |
| Enterprise-wide transformation programs | Cross-customer pattern | Central data/AI platform owner with distributed business users | Citizen data science, governance, agentic AI, shared workflows | Best explanation for why customer counts can rise while use-case breadth expands inside accounts | Public evidence is qualitative rather than cohort-based |
Segments are built from named public case studies and official customer-count disclosures rather than from a company-published revenue segmentation table.
[CU001, CU002, CU003, CU004, CU005, CU006]Dataiku typically enters through a high-value workflow, then expands into broader governed analytics, GenAI, and business-user enablement.
The journey synthesizes recurring patterns visible across Michelin, Prologis, Roche, Standard Chartered, and J&J case studies; not every customer follows each stage identically.
[CU008, CU009, CU026, CU039]6.2 Adoption trajectory and scale signals
The strongest commercial readthrough is that Dataiku appears to be moving from a large enterprise platform into a broad installed base with deeper internal usage. Sacra estimated roughly 500 customers at the end of 2023, and that benchmark lines up directionally with official disclosures showing 700-plus customers by January 2025 and 750-plus by October 2025. Several case studies show not just logo acquisition but widening internal adoption. Johnson & Johnson's Vision organization says more than 80 analytics and data-science professionals adopted Dataiku and the page headline states 650-plus employees use Dataiku. Michelin expanded from 35 users in 2021 to 1,500-plus users across 50-plus factories by mid-2025, with 80% of users described as business experts. Standard Chartered reported 518 staff completed its citizen-data-science program since 2020 and more than 700 completed online Dataiku training courses, while Prologis said over 2,000 users leverage AI through its enterprise ChatGPT deployment built with Dataiku. These are not perfect retention metrics, but they are meaningful deployment-scale signals showing that Dataiku often lands as a platform and then broadens within the customer organization.[CU010, CU011, CU012, CU013, CU014, CU015]
| Metric | Value | As of | Source basis | Implication |
|---|---|---|---|---|
| Customers | ~500 | End 2023 | Sacra estimate | Pre-2025 base was already substantial |
| Customers | 700+ | Jan 2025 | Official ARR release + Reuters | Shows clear new-logo growth into 2025 |
| Customers | 750+ | Oct 2025 | Official ARR release | Further scaled despite no new primary round disclosure |
| Global 2000 penetration | 200 customers | Jul 2024 | KPMG alliance release | Large-enterprise focus was established before 2025 surge |
| Johnson & Johnson users | 650+ employees; 80+ analytics/data-science adopters | Case study | Customer-proof | Evidence of broad internal adoption rather than a tiny pilot |
| Michelin users | 1,500+ users across 50+ factories | Mid-2025 | Customer-proof | High deployment depth in industrial operations |
| Standard Chartered enablement | 518 trained; 700+ online-course completions | 2020-2022 onward | Customer-proof | Training and citizen enablement are part of the expansion model |
| Prologis usage | 2,000+ users; 60+ projects in production | Case study | Customer-proof | Dataiku can become an internal AI operating layer |
| Customers sharing stories | 100+ on stage in 2024 | Jan 2025 | Official ARR release | Reference base is wide enough to power marketing and peer validation |
The trajectory table combines official customer counts with deeper inside-account adoption signals because Dataiku does not publish retention or cohort data.
[CU010, CU011, CU012, CU013, CU014, CU015]Public evidence narrows from broad customer-count disclosures to a smaller set of deeply documented enterprise deployments.
Relative values are illustrative weights based on the density of public proof, not a disclosed conversion funnel or retention curve.
[CU002, CU010, CU029, CU036]6.3 Named customer proof and outcomes
Dataiku's public proof set is unusually rich for a private infrastructure-software company because many customer stories include named executives, concrete workflows, and quantified outcomes. The best evidence is operational rather than vanity-logo based. Novartis reported a 90% reduction in time to insight for a GenAI use case and a 600% acceleration in spreadsheet-driven data ingestion. Michelin said a root-cause analysis process that previously took up to six months can now run in about an hour, and that more than 600 engineers and technicians across more than ten factories rely on a Dataiku-based parameters-analyzer solution. Prologis said it moved from five productionalized data-science products to 30 projects and 30 APIs in active use while putting over 60 AI/ML projects into production. Roche described months-to-days build-time compression for new GenAI projects, six-figure annual attorney-hour savings, and a planned expansion from 80 European patent professionals to as many as 250 users globally. Geodis, Euronext, Perdue Farms, Standard Chartered, SLB, and Mitsubishi Electric all publish similarly specific productivity or decision-support improvements. The key caveat is that these are vendor-curated customer stories, so they prove real deployment and outcomes for named accounts but not the typical experience across the full customer base.[CU016, CU017, CU018, CU019, CU020, CU021]
| Customer | Deployment / use case | Production vs. pilot | Quantified outcome | Reference quality | Limitation |
|---|---|---|---|---|---|
| Johnson & Johnson Vision | GenAI/LLM training, hackathon, common analytics platform | Production platform + prototype event | 650+ employees use Dataiku; working prototypes built in <2 days | High (named exec, metrics, direct story) | Case study highlights enablement more than contract economics |
| Novartis | Healthcare market-research chatbot and forecast automation | Production / scaled internal use | 90% faster time to insight; 600% faster data ingestion | High | Vendor-curated success story; no seat count or spend disclosed |
| Michelin | Factory AI, quality management, time-series copilots | Production at scale | 1,500+ users across 50+ factories; analysis cut from up to 6 months to ~1 hour | High | No revenue or renewal detail |
| Euronext | Market analytics assistant agent | Production workflow deployment | Up to 20% reduction in recurring query time | High | Benefit framed around time savings, not financial ROI |
| Geodis | AI IT Support Agent in ServiceNow | Early production / operational rollout | 60% faster assignment; ~30 minutes saved per ticket | High | Case study implies rollout progression but not enterprise-wide saturation |
| Roche | Patent research and agentic AI for attorneys | Production and expanding | $100K-$250K annual attorney time savings; $375K-$475K consulting avoidance | High | Legal-team use case may not generalize to broader Roche adoption |
| Prologis | Enterprise AI/ML and GenAI platform standardization | Production at scale | 60+ AI/ML projects in production; 30 projects + 30 APIs active; 2,000 users | High | Does not disclose revenue or contract length |
| SLB | Well construction, reservoir analysis, HR retention | Production across multiple functions | > $10B tenders assessed; 25x faster tender analysis; 76% faster pressure analysis | High | Multiple use cases, but no consolidated adoption denominator |
These rows prioritize named deployments with concrete workflows and outcomes; they do not enumerate Dataiku's full customer base.
[CU016, CU017, CU018, CU019, CU020, CU021]Named public references differ in evidence quality, but multiple flagship accounts show production status, quantified outcomes, and identifiable executive sponsorship.
The matrix scores only public reference quality; it does not imply these are Dataiku's largest or most profitable customers.
[CU017, CU018, CU020, CU030, CU037]6.4 Expansion motion and ecosystem leverage
The commercial motion visible in public materials is land-and-expand via platform standardization, citizen enablement, and partner-assisted enterprise transformation. Case studies repeatedly start with a single operational pain point, but they end with broader governed analytics and AI adoption. Michelin began with digital manufacturing and then spread across factories, R&D, and corporate functions. Prologis moved from descriptive analytics into geospatial analysis, predictive modeling, and enterprise GenAI. Roche started with patent-search pilots before consolidating multiple GenAI projects into an agentic interface. The KPMG alliance indicates Dataiku is also sold through modernization and governance programs rather than only direct software procurement, while the 2025 Frontrunner Awards show customers being recognized for agentic AI, governance, and productivity use cases across multiple industries. This expansion logic matters because it suggests Dataiku's best accounts are sticky not only because of models in production, but because the platform becomes an organizational workflow, training, and governance layer. What remains unknown is how much pipeline comes through partners, how much revenue is marketplace or services influenced, and whether expansion is broad across the long tail or concentrated in a smaller set of flagship accounts.[CU026, CU027, CU028, CU029, CU030, CU035]
| Driver / risk | Direction | Public evidence | Implication | Diligence path |
|---|---|---|---|---|
| Platform standardization | Positive | Customers often consolidate workflows, governance, and AI projects in one platform | Supports multi-team expansion inside accounts | Measure % of ARR from expanded accounts |
| Citizen enablement and training | Positive | Standard Chartered and J&J show training-led spread | Makes Dataiku harder to displace once business users are onboarded | Request active users per account over time |
| Partner-assisted channel | Positive / risk | KPMG alliance and partner-centric transformation language suggest channel leverage | Can accelerate procurement but may hide services dependence | Disclose partner-sourced pipeline and services mix |
| Reference-account concentration | Risk | Most public proof comes from a subset of flagship enterprises | Could mean revenue concentration at top accounts | Request top-10 and top-20 ARR share |
| Implementation complexity | Risk | Negative review headline cites cloud-integration issues | Large deployments may face slower time-to-value or failed rollouts | Request implementation duration and expansion conversion data |
| Retention opacity | Risk | No public NRR, GRR, churn, or contract-term disclosure | Durability cannot be fully underwritten from case studies | Request renewal cohorts and downgrades schedule |
Expansion is visible qualitatively, but concentration and partner-dependence remain unresolved because Dataiku does not publish customer-economics detail.
[CU026, CU027, CU028, CU034, CU035, CU040]6.5 Durability, retention, and adverse readthroughs
Public evidence on customer quality is strongest on adoption breadth and outcome anecdotes, and weakest on renewal mechanics. Dataiku does not publicly disclose net revenue retention, gross retention, churn, contract lengths, top-customer concentration, or share of ARR from the ten largest accounts. Gartner's 2026 review surface is directionally positive, with a 75% five-star and 23% four-star distribution on the page fetched for this run, and Dataiku's January 2025 release cited a 96% Gartner willingness-to-recommend score. But the same Gartner page also contains a critical review headline describing cloud-integration issues, which matters because difficult enterprise integration is exactly where AI-platform rollouts can stall. FeaturedCustomers and TrustRadius confirm a visible review and case-study corpus, but they do not solve the core underwriting questions around renewals or concentration. The fair readthrough is that Dataiku likely enjoys meaningful switching costs inside successful enterprise deployments because workflows, governance, and nontechnical user habits accumulate over time; however, that durability remains medium-confidence until private retention cohorts and concentration schedules are disclosed.[CU031, CU032, CU033, CU034, CU036, CU037]
| Metric / signal | Value / status | Confidence | Why it matters | Diligence ask |
|---|---|---|---|---|
| Net revenue retention | Not publicly disclosed | Low | Core test of land-and-expand durability | Request NRR by segment and geography |
| Gross retention / churn | Not publicly disclosed | Low | Needed to distinguish expansion from logo churn | Request cohort churn bridge |
| Contract length / renewal cadence | Not publicly disclosed | Low | Enterprise AI software can look sticky but still renew annually under pressure | Request standard contract terms and renewal rates |
| Satisfaction signal | Positive directionally | Medium | Gartner page and company-cited willingness-to-recommend imply customer satisfaction | Validate with raw survey base and cohort split |
| Repeat usage / internal breadth | Visible at named accounts | Medium | J&J, Michelin, Prologis, Standard Chartered and Roche show multi-user or multi-project depth | Map active-user growth at top 20 accounts |
| Implementation friction | Real but unquantified | Medium | Critical Gartner review headline cites private-cloud integration issues | Request implementation failure, delay, and rollback rates |
| Reference base | Broad but curated | Medium | FeaturedCustomers and TrustRadius show visible review volume, but not renewal truth | Request independent customer references chosen by investor |
Public retention evidence is largely indirect; positive signals come from deployment depth and review surfaces, while the actual renewal data remains absent.
[CU031, CU032, CU033, CU034, CU036, CU037]Public visibility is strongest on deployment anecdotes and weakest on renewal and concentration.
This matrix evaluates evidence visibility, not performance quality. 'Independent corroboration' means corroboration from non-vendor sources, which is limited for retention and concentration.
[CU031, CU032, CU033, CU034, CU040]07Risks
7.1 Regulatory, privacy, and legal risk
Dataiku's legal risk is less about one visible lawsuit and more about living inside a widening compliance perimeter. The company's legal and privacy surfaces are mature by private-software standards: a current privacy policy, a cloud terms stack, a 2026 DPA, and a trust page that explicitly discusses privacy-by-design, responsible AI, and integrated management systems. Those are real mitigants, especially for enterprise procurement. They are also real obligations. The privacy policy says Dataiku is a controller for data it collects under that policy, while the cloud DPA defines Dataiku as a processor for customer personal data in the SaaS context and ties operations to GDPR, CCPA, SCCs, and incident handling. The regulatory backdrop is getting harder, not easier. The EU AI Act now bans certain practices, imposes strict obligations on high-risk systems, and brings transparency rules for generative AI into force in August 2026. NIST's AI RMF and GenAI profile reinforce the market expectation that enterprise AI vendors operationalize traceability, oversight, and risk management rather than merely market them. The main legal risk is therefore not a known public enforcement action against Dataiku; it is the possibility that broad enterprise deployments, third-party model-provider chains, or customer misuse in sensitive workflows expose the company to slower sales cycles, higher indemnity negotiation, or future regulatory scrutiny.[CR001, CR002, CR003, CR004, CR005, CR006]
| Risk | Jurisdiction / rule | Status | Likelihood | Severity | Mitigation | Residual exposure | Diligence path |
|---|---|---|---|---|---|---|---|
| AI-governance compliance gap | EU AI Act / customer-sector rules | Rules active in phases; transparency obligations extend into Aug 2026 | Medium | High | Governance-first product positioning; trust program; documentation | High because customers can deploy Dataiku into sensitive workflows | Test agent governance, logs, human oversight, and customer guidance for high-risk uses |
| Privacy / data-processing error | GDPR, CCPA, SCCs, DPA obligations | Ongoing | Medium | High | Privacy policy, DPA, processor terms, incident language | Medium-high due to cross-border processing and AI-provider chains | Review DPA redlines, subprocessor list, data-residency controls, and incident workflow |
| Third-party AI provider liability spillover | Contract and privacy obligations | Current | Medium | Medium-high | Customer instructions, contractual allocation, no-training-without-consent statement | Medium because enabled AI services can route content to third parties | Review AI-services terms, provider-specific privacy notices, and opt-out controls |
| Marketing / AI claim scrutiny | FTC / consumer-protection and unfair-practices backdrop | Ongoing | Low-medium | Medium | Procurement-led enterprise selling reduces retail-marketing risk | Medium because enterprise AI claims are increasingly scrutinized | Review substantiation for governance, explainability, and security claims |
| Undisclosed or latent enforcement exposure | Public enforcement trackers and case libraries | No obvious match surfaced in fetched sources | Low | Medium | No public fine surfaced; mature legal surface exists | Unknown because absence from trackers is not proof of absence | Run counsel-led search across litigation, enforcement, and complaint databases |
Rows are ordered by practical risk to revenue and valuation rather than by legal novelty. The biggest threat is regulatory burden expansion inside customer deployments, not a known headline case today.
[CR001, CR002, CR003, CR004, CR007, CR008]Dataiku's highest residual risks combine regulatory breadth, security patch discipline, and customer/financial opacity rather than a single existential fault line.
The heatmap is a synthesized investment view using public evidence, not a management-issued risk register with calibrated probabilities.
[CR007, CR014, CR024, CR026, CR038, CR040]7.2 Security and operational reliability risk
The operational risk profile is elevated because Dataiku is not a narrow feature product. It spans data access, notebooks, APIs, automation, model deployment, governance, and now agentic-AI orchestration, which means the attack surface is inherently broad. OpenCVE shows that Dataiku DSS has had significant disclosed vulnerabilities, including a critical 9.8 authentication-bypass issue disclosed in June 2025 and a series of older access-control and information-disclosure problems. That does not prove weak security culture by itself—every mature enterprise platform has to patch issues—but it does prove that patch discipline matters materially. The operational shape of the product also creates a split risk model. In self-managed deployments, Dataiku says it does not process or store client data by default, which reduces vendor data-custody exposure but shifts more configuration, availability, and security burden to the customer environment. In Dataiku Cloud, the company becomes a more direct control point for incident handling and subprocessor management while also inheriting risk from cloud-provider infrastructure. Gartner review content adds another practical risk: at least one critical review headline on the fetched page points to private-cloud integration friction. That is important because implementation complexity, not just software defects, is often what converts a technically sound platform into a commercially painful rollout.[CR014, CR015, CR016, CR017, CR018, CR019]
| Failure mode | Likelihood | Severity | Mitigation maturity | Residual exposure | Unresolved gap |
|---|---|---|---|---|---|
| Critical software vulnerability or auth bypass | Medium | High | Medium | Material | Need evidence of patch SLAs, customer upgrade cadence, and incident postmortems |
| Private-cloud or hybrid integration friction | Medium | Medium-high | Medium | Material | Need implementation-failure and expansion-conversion metrics |
| Cloud-provider outage or control failure affecting Dataiku Cloud | Low-medium | High | Medium | Material | Need architecture, redundancy, and incident-communication details |
| Self-managed customer misconfiguration causing blame transfer | Medium | Medium | Low-medium | Material | Need support boundaries, reference architectures, and upgrade tooling |
| Agent / model workflow failure in production business processes | Medium | High | Medium | Material | Need guardrail coverage, fallback patterns, and customer rollback data |
Operational risk is partly classic software security risk and partly enterprise-implementation risk, which matters just as much for renewals.
[CR014, CR015, CR016, CR017, CR018, CR019]7.3 Dependency, customer, and execution risk
Dataiku's go-to-market dependencies are also risk channels. The KPMG alliance shows that the company can win as part of a broader modernization and governance program, which is helpful for scale but can make partner quality and services economics more consequential. Customer stories repeatedly highlight interoperability with Snowflake, Azure, ServiceNow, PowerBI, and other systems. That interoperability is a strength, but it means connector breakage, cloud-policy changes, or model-provider disputes can propagate into customer dissatisfaction. The customer base itself is high quality but hard to underwrite. Public proof spans healthcare, banking, logistics, manufacturing, and capital markets, which is strategically attractive because regulated buyers care about governance. It is also operationally dangerous because these buyers are unforgiving when controls, documentation, or incident response fall short. Public sources still do not reveal top-customer concentration, partner-sourced pipeline, or NRR. Most of the rich evidence comes from vendor-curated flagship stories, not from independent cohort disclosures. On people and execution, the company has grown to more than 1,250 employees and is layering in senior commercial leadership while preparing for a possible IPO. That can improve discipline, but it also raises the odds that execution slippage in product, customer success, or partner management becomes visible at exactly the wrong moment.[CR021, CR022, CR023, CR024, CR031, CR033]
| Dependency | Counterparty / category | Role | Concentration | Failure scenario | Severity | Mitigation | Residual exposure |
|---|---|---|---|---|---|---|---|
| Cloud infrastructure | Underlying cloud providers | Host Dataiku Cloud and influence availability/security envelope | Unknown | Outage, policy change, or pricing shift harms service economics | High | Multi-environment deployment options | Medium-high |
| Systems-integrator channel | KPMG and similar partners | Enterprise modernization and implementation leverage | Unknown | Partner under-delivers or captures economics | Medium | Direct sales plus partner ecosystem | Medium |
| Third-party AI providers | Model vendors | Process AI-service content and power GenAI features | Unknown | Provider outage, policy shift, or data-use concern disrupts workflows | High | Model-agnostic positioning and customer controls | Medium-high |
| External data platforms and apps | Snowflake, Azure, ServiceNow, PowerBI, others | Connectors and workflow destinations | High at customer level | Integration breaks slow expansion or cause churn | High | Broad connector surface and shared workflows | Medium-high |
| Flagship enterprise customers | Large reference accounts | Revenue, proof, and IPO narrative | Undisclosed | A few large customers stall or downgrade | High | Broad logo count but unknown weighting | Unknown-high |
The largest dependency risks are not single suppliers; they are ecosystems whose failures would surface inside customer value realization.
[CR021, CR022, CR023, CR031, CR033, CR038]| Role / function | Dependency or gap | Likelihood | Severity | Mitigation | Diligence path |
|---|---|---|---|---|---|
| Customer success / implementation | Needed to convert pilots and migrations into scaled renewals | Medium | High | Partner leverage and reusable templates | Request implementation cycle-time, time-to-value, and expansion-conversion data |
| Product / security engineering | Must patch vulnerabilities and ship new agent/governance features without regressions | Medium | High | Formal security program and certifications | Request vuln-management metrics and staffing depth |
| Sales and partnerships leadership | New leadership and broad ecosystem must deliver disciplined late-stage growth | Medium | Medium-high | Scale and brand momentum | Request quota attainment, partner-sourced pipeline, and churn by cohort |
| Regulatory / trust operations | Must keep up with AI-governance, privacy, and audit expectations | Medium | Medium-high | IMS, trust page, privacy-by-design posture | Request audit calendar, policy exceptions, and remediation backlog |
| Executive team under IPO scrutiny | Any miss becomes more visible as IPO readiness increases | Medium | High | ARR scale and banker engagement | Review board materials, readiness milestones, and internal controls program |
Execution risk is elevated because Dataiku is past the startup phase where product relevance alone can compensate for process inconsistency.
[CR024, CR025, CR026, CR037, CR039]Dataiku depends on a web of clouds, model providers, integrators, and flagship enterprise accounts; failure in any one layer can weaken the investment case.
[CR021, CR022, CR023, CR034, CR038]7.4 Financial opacity and IPO transmission risk
The final risk layer is transmission: how a technical, regulatory, or customer issue would affect financing and valuation. Reuters reported that Dataiku hired Morgan Stanley and Citigroup in October 2025 to prepare for a possible U.S. IPO, which means the company is operating under a higher implied scrutiny bar than a normal late-stage private software vendor. At the same time, public evidence still does not disclose core underwriting fields such as gross margin, NRR, cash burn, or top-customer share. That matters because these missing fields determine whether Dataiku can absorb a shock. A company with strong margins, deep cash, and diversified expansion can survive a delayed deal window or an implementation stumble; a company without those cushions can quickly lose negotiating leverage. The right risk framing is therefore not simply “IPO window risk.” It is that security flaws, compliance misses, partner friction, or stalled customer expansion could all transmit into weaker renewal assumptions, lower public-market comparables, or a longer private holding period before a credible listing. Dataiku's control posture and product relevance are meaningful mitigants, but the residual exposure remains above ordinary SaaS because the platform sits directly in enterprise AI governance and operational workflows, where failures travel quickly into trust, procurement, and valuation.[CR024, CR025, CR026, CR030, CR039, CR040]
| Risk | Monitorable trigger | Threshold / event | Action implication |
|---|---|---|---|
| Security patch discipline | Critical unresolved vulnerability | Critical vuln remains unpatched or broadly un-upgraded for >30 days after disclosure | Pause underwriting until patch cadence and customer-upgrade coverage are proven |
| Regulatory posture | Material enforcement event | GDPR/FTC/sector regulator action or consent order involving Dataiku or a major customer deployment attributable to the platform | Re-cut risk rating and valuation range immediately |
| Customer durability | Renewal / expansion deterioration | Private NRR below enterprise-software expectations or notable top-customer downgrades | Shift recommendation toward wait / reprice |
| Partner dependence | Channel concentration | >40% of new ARR or services delivery dependent on one channel partner or one cloud route | Discount margin path and execution confidence |
| Implementation complexity | Slow time-to-value | Average enterprise implementation cycle materially longer than management guidance or expansion conversion falls | Treat customer stories as non-representative until disproven |
| IPO readiness | Financing slippage | IPO timetable slips while growth or governance indicators weaken | Assume longer private holding period and higher down-round risk |
These kill criteria are deliberately monitorable; they convert broad concern into thresholds that can change an investment decision.
[CR024, CR025, CR026, CR030, CR038, CR040]Security, regulatory, and implementation failures would likely transmit first into trust and expansion, then into valuation and financing optionality.
[CR024, CR025, CR026, CR030, CR040]08Valuation
8.1 Investment thesis and anti-thesis
The bull case for Dataiku is straightforward. It has genuine enterprise scale, disclosed ARR milestones, a blue-chip customer roster, a governance-heavy product narrative that fits the current enterprise AI moment, and a still-private valuation mark that has not obviously outrun revenue in the way some AI names have. In January 2025 Dataiku disclosed $300M+ ARR and 700+ customers; in October 2025 it disclosed $350M+ ARR, 750+ customers, and 1,250+ employees. That is real operating heft. The anti-thesis is not that the company lacks traction; it is that investors still do not know enough about the economics behind the traction. Gross margin, NRR, cash burn, services intensity, and top-customer concentration remain undisclosed, while competition from Databricks and hyperscalers keeps increasing. A fair valuation can become an overvaluation very quickly if the business is more services-heavy, more channel-dependent, or less sticky than its public narrative suggests.[CV001, CV002, CV003, CV004, CV005, CV006]
| Dimension | Bull thesis | Bear anti-thesis |
|---|---|---|
| Scale | 350M+ ARR and 750+ customers support category leadership | Scale without disclosed margins or NRR can mislead on value quality |
| Product | Governance-first AI platform matches enterprise needs | Databricks and hyperscalers are copying governance and agent surfaces |
| Customers | Blue-chip enterprise roster implies durable relevance | Customer stories are curated and concentration is undisclosed |
| Valuation | Stale $3.7B mark now implies only ~10-12x ARR | Private illiquidity and opacity warrant a discount to premium public comps |
| Exit optionality | Reuters IPO prep keeps public-market path alive | IPO timing can slip if disclosures or market windows weaken |
| Evidence quality | Multiple public proof points support the story | Critical economics still require private diligence |
The thesis turns mainly on whether Dataiku is closer to premium software-quality infrastructure or to a slower, services-heavier enterprise platform.
[CV001, CV003, CV007, CV020, CV021, CV022]The recommendation flows from real operating scale and customer proof on one side, and missing economics plus execution risk on the other.
[CV004, CV006, CV020, CV027, CV028]8.2 Financing context and public comp bridge
Dataiku's most recent disclosed private mark remains the December 2022 Series F at $3.7B. Since then, public operating milestones improved while the valuation anchor did not publicly reset, which is unusual but analytically useful. Using the last disclosed mark, Dataiku traded at about 12.3x ARR on the January 2025 $300M milestone and about 10.6x ARR on the October 2025 $350M milestone; Sacra's roughly $342.5M September 2025 estimate implies about 10.8x. Those are not cheap multiples, but they are also not public-market peak territory for elite AI or data-platform assets. Public market-cap-to-revenue proxies show a wide band: roughly 58.2x for Palantir, 25.0x for Datadog, 19.4x for Snowflake, 11.2x for MongoDB, 9.6x for Confluent, 8.0x for ServiceNow, and 4.6x for C3.ai. Dataiku's stale private multiple therefore sits closer to the middle of the public range than the extremes. The valuation question is not whether the company deserves any premium; it is where inside that wide band a private, opaque, enterprise-governance AI platform should trade.[CV001, CV002, CV003, CV004, CV005, CV006]
| Comparable | Metric basis | Valuation / market cap | Revenue basis | Market-cap-to-revenue proxy | Relevance / limitation |
|---|---|---|---|---|---|
| Dataiku | Private mark vs official ARR | $3.7B | 350M+ ARR (Oct 2025) | ~10.6x | Closest anchor but stale and private |
| Dataiku | Private mark vs official ARR | $3.7B | 300M+ ARR (Jan 2025) | ~12.3x | Shows multiple compression as revenue scaled |
| Dataiku | Private mark vs Sacra estimate | $3.7B | ~342.5M ARR (Sep 2025 est.) | ~10.8x | Estimated, not audited |
| Palantir | Public market cap / TTM revenue | $303.95B | $5.22B | ~58.2x | Premium AI/government software outlier; much more public and liquid |
| Datadog | Public market cap / TTM revenue | $91.67B | $3.67B | ~25.0x | High-quality infra software ceiling reference |
| Snowflake | Public market cap / TTM revenue | $90.61B | $4.68B | ~19.4x | Consumption-led premium data platform |
| MongoDB | Public market cap / TTM revenue | $27.51B | $2.46B | ~11.2x | Growth software with more mature public disclosure |
| ServiceNow | Public market cap / TTM revenue | $111.08B | $13.96B | ~8.0x | Workflow software benchmark with scale and profitability |
| Confluent | Public market cap / TTM revenue | $11.13B | $1.16B | ~9.6x | Infra software with lower premium than elite AI names |
| C3.ai | Public market cap / TTM revenue | $1.39B | $0.30B | ~4.6x | Adverse lower-bound comp for weaker AI software economics |
These are market-cap-to-revenue proxies from public market-cap and revenue pages, not clean enterprise-value multiples. Coverage is nonetheless sufficient to frame Dataiku inside a broad public-band range.
[CV004, CV005, CV006, CV010, CV011, CV012]Public comp market-cap-to-revenue proxies show how quickly the valuation answer changes depending on which reference set best matches Dataiku.
These are market-cap-to-revenue proxies derived from public market-cap and revenue pages, not enterprise-value multiples adjusted for net cash or debt.
[CV010, CV011, CV012, CV013, CV014, CV015]8.3 Scenario valuation and recommendation
Our base case assumes the October 2025 $350M+ ARR disclosure is directionally current enough to anchor value, but that investors should apply a private-company discount until margin, retention, and cash-efficiency evidence improves. On that basis, a fair base-case valuation range is about $3.5B-$4.5B, broadly in line with or modestly above the last disclosed $3.7B mark. The bull case supports roughly $5.0B-$6.5B if Dataiku can show that trusted-AI governance, 750+ enterprise customers, and IPO readiness translate into durable expansion and software-like margins. The bear case supports roughly $2.5B-$3.2B if growth slows, retention disappoints, or public-market investors apply a harsher discount for opacity and implementation risk. The current public evidence therefore supports a track recommendation, medium confidence, medium-high risk, and a fair valuation stance. The company is too real to dismiss and too opaque to underwrite aggressively at a premium.[CV018, CV019, CV023, CV024, CV025, CV026]
| Dimension | Assessment | Basis |
|---|---|---|
| Recommendation | Track | Strong scale and customer proof, but economics remain under-disclosed |
| Confidence | Medium | Valuation rests on stale private mark plus public comp proxies |
| Risk rating | Medium-high | Security, regulatory, and financial-opacity risks remain material |
| Valuation stance | Fair | Current $3.7B mark fits the base-case range |
| Overall score | 7.4 / 10 | Quality business, incomplete underwriting data |
| Entry discipline | Require updated economics | Premium upside needs margin and retention proof |
The recommendation is price-sensitive and evidence-sensitive rather than a generic business-quality score.
[CV026, CV027, CV028, CV040]| Scenario | Probability signal | Core assumptions | Valuation range | Return vs $3.7B mark |
|---|---|---|---|---|
| Bull | ~25% | Growth remains strong, trusted-AI moat holds, public disclosures support premium multiple | $5.0B-$6.5B | +35% to +76% |
| Base | ~50% | Current ARR scale is real but economics remain only moderately software-like | $3.5B-$4.5B | -5% to +22% |
| Bear | ~25% | Growth slows, retention or margin disappoints, IPO path slips | $2.5B-$3.2B | -32% to -14% |
Ranges are scenario estimates anchored to the last disclosed valuation, current ARR disclosures, and public comp proxies; they are not management guidance.
[CV023, CV024, CV025, CV026, CV031, CV034]Scenario ranges show why the current mark looks fair but not obviously cheap.
Ranges are author estimates anchored to the last disclosed valuation, current ARR disclosures, and public comp proxies.
[CV023, CV024, CV025, CV026, CV031, CV034]Headline indicators for an IC-style readout.
KPIs summarize the recommendation and valuation analysis; the ARR multiple uses the last disclosed $3.7B valuation and Oct. 2025 ARR milestone.
[CV026, CV027, CV028, CV040]8.4 Final diligence and thesis-break triggers
The remaining work before a firm buy call is not philosophical; it is documentary. Investors need a current ARR bridge through runDate, a gross-margin decomposition, NRR and churn cohorts, top-customer concentration, partner-sourced pipeline, services mix, cash and burn, and the preference stack if a future round or IPO occurs. These asks matter because the multiple debate is no longer academic. Dataiku is already large enough that missing economics, not missing demand, now dominate the risk. The thesis breaks if retention is materially weaker than the customer stories imply, if security or regulatory issues interrupt IPO timing, if ARR growth slows below the expectations embedded in a $3.7B+ mark, or if partner dependence turns out to be compensating for weak direct product leverage. Until those questions are answered, the correct posture is disciplined proximity rather than maximal conviction. It also means secondary or crossover investors should resist treating the absence of bad news as equivalent to proof of premium economics.[CV028, CV032, CV036, CV037, CV038, CV039]
| Trigger | Threshold / event | Transmission to thesis | Action implication |
|---|---|---|---|
| Retention shortfall | Private NRR or renewal cohorts materially below enterprise-software norms | Undermines land-and-expand and premium-multiple logic | Downgrade to wait / reprice |
| Gross-margin weakness | Software gross margin materially below strong-platform expectations | Suggests services or hosting drag is larger than the market narrative implies | Move toward bear case |
| Security / regulatory event | Material breach, enforcement action, or major control failure near IPO process | Hits trust, procurement, and exit timing simultaneously | Cut fair-value range |
| Growth deceleration | Updated ARR bridge shows slower growth than implied by 2025 disclosures | Reduces premium-multiple support | Re-underwrite using lower-band comps |
| Customer concentration | Top-customer share or partner dependence is too high | Makes revenue quality less durable than logo count implies | Apply liquidity and concentration discount |
| IPO slippage | IPO path slips while disclosures remain weak | Extends private hold and increases valuation-mark risk | Stay on track, do not pay premium |
These triggers are designed to be observable in diligence, quarterly updates, or future public filings rather than vague qualitative worries.
[CV033, CV034, CV037, CV038, CV039, CV040]| Topic | Missing evidence | Why it matters | Owner / path |
|---|---|---|---|
| ARR bridge | Monthly or quarterly ARR from Oct 2025 to runDate | Confirms whether growth continued or slowed materially | Company finance |
| Gross margin | Software, hosting, and services gross margin split | Determines where Dataiku sits inside public comp band | Company finance |
| NRR / churn | Cohort retention and downgrade history | Separates real platform durability from curated case studies | Company RevOps |
| Cash and burn | Cash balance, burn, runway, financing plan | Tests downside resilience if IPO timing slips | Company finance / board |
| Customer concentration | Top-10 revenue share and segment mix | Validates breadth implied by 750+ customers | Company sales / finance |
| Partner economics | Partner-sourced bookings and services attach | Shows whether partner leverage helps or hurts quality of growth | Company partnerships |
| Preference stack | Current liquidation preferences and dilution overhang | Matters for late-stage entry economics | Company legal / finance |
These are the minimum missing data points needed to move from a track stance to a committed price call.
[CV028, CV032, CV036, CV040]8.5 Exhibits
Disclaimer
This report is for informational purposes only, reflects public sources available as of 2026-07-13, and is not investment advice. Private-company valuations, ARR figures, and comparable-multiple bridges should be independently verified before any investment decision.
Evidence index
| ID | Statement | Confidence | Sources |
|---|---|---|---|
| CO001 | Dataiku was founded in 2013. | High | SO001, SO010 |
| CO002 | Dataiku was founded by Florian Douetteau, Clément Stenac, Thomas Cabrol, and Marc Batty. | High | SO001, SO015 |
| CO003 | Independent profiles describe Dataiku as beginning in Paris before later scaling in the United States. | Medium | SO017, SO016 |
| CO004 | Dataiku is headquartered in New York, New York. | High | SO010, SO009 |
| CO005 | Gartner's 2026 company profile still categorizes Dataiku as a private company. | Medium | SO010 |
| CO006 | Florian Douetteau is Dataiku's co-founder and CEO. | High | SO001, SO005 |
| CO007 | Clément Stenac is Dataiku's co-founder and CTO. | High | SO001, SO015 |
| CO008 | Dataiku is positioned as an enterprise orchestration layer for analytics, machine learning, and AI agents. | Medium | SO002, SO005, SO015 |
| CO009 | The platform is designed for no-code, low-code, and full-code collaboration between business and technical users. | Medium | SO002, SO009 |
| CO010 | Dataiku emphasizes cloud and model optionality rather than locking customers into one infrastructure stack. | Medium | SO002, SO023 |
| CO011 | An archived official plans page shows Dataiku offered hosted SaaS, customer-cloud, and on-premises deployment options. | Medium | SO019 |
| CO012 | The archived enterprise plan highlighted unlimited instances, full deployment capabilities, and resource governance. | Medium | SO019, SO020 |
| CO013 | Dataiku raised $200 million in a Series F round in December 2022. | High | SO001, SO006 |
| CO014 | The Series F round valued Dataiku at $3.7 billion. | High | SO001, SO006, SO009 |
| CO015 | Wellington Management led Dataiku's 2022 Series F financing. | Medium | SO003, SO006 |
| CO016 | Sacra estimates Dataiku's total funding at about $846.8 million. | Medium | SO008, SO009 |
| CO017 | Zippia reports a $101 million Series C in December 2018. | Medium | SO016 |
| CO018 | Zippia reports CapitalG joined in December 2019 when Dataiku reached unicorn status at a $1.4 billion valuation. | Medium | SO016 |
| CO019 | Zippia reports Dataiku raised an additional $100 million Series D in August 2020 led by Stripes and Tiger Global. | Medium | SO016 |
| CO020 | Zippia reports Dataiku raised $400 million in August 2021 at a $4.6 billion valuation. | Medium | SO016 |
| CO021 | Dataiku's January 2025 press release said ARR had surpassed $300 million. | High | SO003, SO009 |
| CO022 | The January 2025 release said Dataiku had grown ARR by more than 2x over the prior three years. | Medium | SO003 |
| CO023 | The January 2025 release said Dataiku served more than 700 organizations worldwide. | Medium | SO003, SO012 |
| CO024 | The January 2025 release said Dataiku employed more than 1,100 people worldwide. | Medium | SO003 |
| CO025 | The January 2025 release said more than 20% of customers were already using Dataiku for GenAI workflows. | Medium | SO003 |
| CO026 | The January 2025 release said Dataiku ranked No. 34 on the 2024 Forbes Cloud 100. | Medium | SO003 |
| CO027 | Dataiku's October 2025 release said ARR had surpassed $350 million. | Medium | SO004, SO008 |
| CO028 | The October 2025 release said Dataiku served more than 750 organizations worldwide. | Medium | SO004, SO005 |
| CO029 | The October 2025 release said Dataiku employed more than 1,250 people across 13 offices and remote locations. | Medium | SO004 |
| CO030 | The October 2025 release said Dataiku was trusted by one in four of the world's top companies in the 2024 Forbes Global 2000. | Medium | SO004 |
| CO031 | June 2026 leadership changes added Maxwell Long as President and CRO overseeing sales, customer success, and partnerships. | Medium | SO005 |
| CO032 | Maxwell Long arrived from Smartsheet after helping scale it past $1 billion ARR and through its $8.4 billion acquisition. | Medium | SO005 |
| CO033 | By June 2026 Dataiku highlighted Agent Management, Cobuild, and Reasoning Systems as new enterprise AI products. | Medium | SO005, SO002 |
| CO034 | The Snowflake partnership page says Dataiku and Snowflake have helped 300+ customers operationalize AI together. | Medium | SO022 |
| CO035 | Dataiku's Snowflake blog says it was Snowflake's #1 AI/ML partner by platform consumption and marketplace revenue in 2025. | Medium | SO013 |
| CO036 | KPMG and Dataiku announced a 2024 alliance to modernize analytics and accelerate enterprise AI adoption. | Medium | SO014 |
| CO037 | Reuters reported in October 2025 that Dataiku hired Morgan Stanley and Citigroup to prepare for a U.S. IPO. | High | SO006, SO007 |
| CO038 | Reuters reported the IPO could come as soon as the first half of 2026, but timing remained subject to change. | High | SO006, SO007 |
| CO039 | The latest accessible public evidence shows IPO preparation rather than a completed listing. | Medium | SO006, SO010 |
| CO040 | FeaturedCustomers lists 182 reviews/testimonials, 150 case studies, and 63 customer videos for Dataiku. | Medium | SO011 |
| CO041 | Dataiku's customer directory highlights reference customers across finance, pharma, manufacturing, logistics, retail, and energy. | Medium | SO012 |
| CO042 | Gartner's 2026 product page includes a critical review citing private-cloud integration issues despite praising low-code AI strengths. | Medium | SO010 |
| CO043 | SiliconANGLE argues enterprises remain years away from broad agentic AI deployment because data, integration, and governance foundations are still missing. | Medium | SO025 |
| CO044 | Dataiku's partner directory highlights an ecosystem spanning Snowflake, Databricks, AWS, Google Cloud, NVIDIA, Accenture, and many regional integrators. | Medium | SO021 |
| CO045 | The Google Cloud partner page says Dataiku integrates with BigQuery, Vertex AI, Gemini-family models, and governance controls for GenAI. | Medium | SO023 |
| CO046 | The NVIDIA partner page says Dataiku can self-host open-source LLMs on NVIDIA GPUs through its LLM Mesh and NIM-related integrations. | Medium | SO024 |
| CO047 | FirstMark describes Dataiku as a collaboration layer connecting data repositories, algorithms, models, and people. | Medium | SO015 |
| CO048 | Zippia says Dataiku established itself in the United States in 2015. | Medium | SO016 |
| CO049 | PM Insights exposes only teaser-level valuation charts and cap-table references publicly, underscoring how limited open secondary-market visibility remains. | Low | SO018 |
| CO050 | Public sources still leave board composition, cash balance, and current secondary pricing materially under-disclosed. | Medium | SO018, SO010, SO007 |
| CM001 | Dataiku participates at the intersection of enterprise data-science platforms, MLOps, AI governance, and agent-orchestration software rather than in one pure-play category. | Medium | SM001, SM011, SM013 |
| CM002 | The broadest comparable market lens is general machine-learning software, which Fortune Business Insights sizes at $65.28B in 2026. | Medium | SM018 |
| CM003 | A more expansive machine-learning market estimate from Precedence Research places 2026 spend at $126.91B, highlighting how category boundaries can materially widen TAM claims. | Medium | SM019 |
| CM004 | MarketsandMarkets sizes the narrower MLOps market at $5.9B by 2027 with a 41.0% CAGR. | Medium | SM017 |
| CM005 | MarketsandMarkets sizes the AI governance market at $5.78B by 2029 with a 45.3% CAGR. | Medium | SM017 |
| CM006 | MarketsandMarkets sizes the adjacent AI Studio market at $32.7B by 2029 with a 38.4% CAGR. | Medium | SM017 |
| CM007 | Fortune says large enterprises account for 55.61% of the machine-learning market in 2026. | Medium | SM018 |
| CM008 | Fortune says cloud deployment accounts for 53.14% of the machine-learning market in 2026. | Medium | SM018 |
| CM009 | Fortune says North America held a 32.5% share of the global machine-learning market in 2025. | Medium | SM018 |
| CM010 | Dataiku's public customer set spans life sciences, logistics, retail, manufacturing, energy, financial services, software, and technology. | Medium | SM001, SM004 |
| CM011 | Snowflake partnership materials explicitly position business and domain experts — not just technical specialists — as builders in this category. | Medium | SM006 |
| CM012 | The Google Cloud partner page positions Dataiku as a front-end for Vertex AI, BigQuery, and Gemini-backed governed application building. | Medium | SM007 |
| CM013 | The AWS partner page positions Dataiku as a collaborative visual layer on top of AWS ML, AI, and elastic cloud infrastructure. | Medium | SM008 |
| CM014 | The Databricks partner page positions Dataiku as a business-user layer on top of governed Databricks data. | Medium | SM009 |
| CM015 | Databricks markets a unified lakehouse-plus-governance stack and is a direct alternative for enterprises consolidating data engineering, analytics, and AI on one platform. | Medium | SM010 |
| CM016 | SageMaker markets a unified studio for generative AI, model training, AI ops, governance, lineage, and lakehouse analytics inside AWS. | Medium | SM011 |
| CM017 | Azure Machine Learning markets a centralized studio with a 99.9% SLA and pay-for-compute economics layered on the wider Azure stack. | Medium | SM012 |
| CM018 | Google Cloud's Agent Platform markets 200+ models, notebooks, pipelines, model registry, vector search, and usage-based pricing under one platform. | Medium | SM013 |
| CM019 | DataRobot emphasizes an all-in-one enterprise AI suite with on-prem, VPC, and SaaS deployment options. | Medium | SM014 |
| CM020 | H2O AI Cloud emphasizes managed cloud, hybrid cloud, AutoML, and no-code accessibility for enterprise ML teams. | Medium | SM015 |
| CM021 | Alteryx One represents a simpler analytics-and-automation adjacent substitute rather than a perfect like-for-like governed AI platform. | Medium | SM016 |
| CM022 | Dataiku's October 2025 press release frames the market shift as moving from AI experimentation to operationalization and trusted execution. | Medium | SM003 |
| CM023 | Dataiku's January 2025 press release said more than 20% of customers were already integrating GenAI into workflows, showing real but still partial production adoption. | Medium | SM002 |
| CM024 | Deloitte's July 2025 poll found only 13.5% of respondents were already using agentic AI in finance and accounting. | Medium | SM020 |
| CM025 | Deloitte found trust in the underlying data and programming was the top barrier to agentic AI adoption at 21.3% of responses. | Medium | SM020 |
| CM026 | SiliconANGLE argues 2025 will not be the year of broad enterprise agentic AI because data quality, integration, and governance gaps remain unresolved. | Medium | SM021 |
| CM027 | The 2026 arXiv interview study found seven of twelve companies were still only at the AI-assistant maturity level and only one had reached multi-agent orchestration. | Medium | SM022 |
| CM028 | The same arXiv study identified a capability-deployment verification gap where higher-level experimental AI cannot be trusted in production without human verification. | Medium | SM022 |
| CM029 | Observer argues early agentic AI deployments are primarily architectural projects whose attributable returns often take two to four years in complex environments. | Medium | SM025 |
| CM030 | Gartner review evidence shows even satisfied users can still encounter private-cloud integration friction, underscoring how hard deployment environments remain. | Medium | SM024 |
| CM031 | Dataiku's archived plans suggest the product can land with small teams but monetizes highest when governance, automation, and unlimited-scale deployment matter. | Medium | SM005 |
| CM032 | KPMG's 2024 alliance shows systems integrators remain important demand multipliers in the enterprise AI platform market. | Medium | SM023 |
| CM033 | Dataiku's public customer roster repeatedly surfaces pharma, financial services, industrials, logistics, insurance, and exchange operators. | Medium | SM004 |
| CM034 | The Snowflake partner page says Dataiku and Snowflake have helped 300+ customers operationalize AI, supporting a co-sell-led enterprise adoption model. | Medium | SM006 |
| CM035 | The real SAM for Dataiku should exclude raw cloud infrastructure spend, general-purpose LLM consumption, and lightweight point copilots that do not require governed multi-user orchestration. | Medium | SM010, SM011, SM013 |
| CM036 | Dataiku's fit is strongest where buyers want one governed control plane across multiple vendors rather than a single-cloud native toolchain. | Medium | SM001, SM007, SM008 |
| CM037 | Hyperscaler stacks pressure Dataiku on distribution and bundling, but also validate the demand for integrated AI development and governance. | Medium | SM011, SM012, SM013 |
| CM038 | The wide spread between broad ML market estimates and narrower MLOps/AI-governance estimates means later valuation work should use multiple market lenses rather than one headline TAM. | Medium | SM017, SM018, SM019 |
| CP001 | Dataiku competes across three practical classes of alternatives: unified data-and-AI platforms, hyperscaler-native ML stacks, and specialist AI or analytics-automation vendors. | Medium | SP001, SP008, SP012, SP014, SP016, SP018, SP021, SP023 |
| CP002 | Dataiku's product and partner surfaces position it as a neutral control layer that can sit across multiple clouds and partner ecosystems rather than as a single-vendor full stack. | Medium | SP001, SP003, SP005, SP006, SP007 |
| CP003 | Databricks is Dataiku's closest broad-platform rival because it combines data, governance, model lifecycle, and agent tooling in one integrated platform family. | High | SP008, SP009, SP011 |
| CP004 | Dataiku and Databricks are simultaneously competitors and collaborators, implying that some accounts adopt Dataiku as a workflow layer on top of a Databricks-centered data architecture. | Medium | SP003, SP004 |
| CP005 | Amazon SageMaker is strongest where the buyer wants first-party AWS procurement and lifecycle tooling rather than a neutral orchestration layer. | Medium | SP005, SP012, SP013 |
| CP006 | SageMaker pricing is granular and usage-metered by instance type, duration, storage, and inference configuration, reinforcing the economics of a native cloud service rather than a seat-based platform. | High | SP012, SP013 |
| CP007 | Azure Machine Learning competes through enterprise MLOps and responsible-AI positioning inside the broader Azure estate. | High | SP014, SP015 |
| CP008 | Vertex AI couples model-development breadth with explicit training, deployment, and prediction pricing, making it a strong option for GCP-centric AI programs. | High | SP016, SP017 |
| CP009 | DataRobot still markets a full enterprise AI suite with on-premise, VPC, and SaaS deployment choices rather than a pure single-cloud service. | Medium | SP018 |
| CP010 | H2O.ai differentiates through hybrid-cloud deployment, open-source lineage, and AutoML-led accessibility, but it remains much smaller in disclosed capital scale than Dataiku or Databricks. | Medium | SP021, SP022, SP025 |
| CP011 | Alteryx is a meaningful substitute for analytics automation and low-code data work, but it is not positioned as broadly around end-to-end enterprise ML and agent governance as Dataiku or Databricks. | Medium | SP023, SP024 |
| CP012 | Databricks is more publicly transparent on list pricing than most enterprise AI peers because it publishes a price list for SKU groups, even though actual realized economics still depend on cloud and discount structure. | Medium | SP010, SP011 |
| CP013 | Amazon SageMaker exposes detailed public price components across notebooks, training, inference, feature store, processing, and MLflow surfaces. | High | SP012, SP013 |
| CP014 | Azure Machine Learning offers pay-as-you-go, reservation, and savings-plan choices rather than a single public platform fee. | High | SP014, SP015 |
| CP015 | Vertex AI pricing includes hourly model-operation charges and no minimum usage duration for training and prediction, favoring bursty experimentation on GCP. | High | SP016, SP017 |
| CP016 | Dataiku's archived plans page shows packaging that expands from free or small-team usage into deeper automation, deployment, security, and governance capabilities for larger teams. | Medium | SP001, SP002 |
| CP017 | DataRobot, H2O.ai, and Alteryx emphasize demos, downloads, or contact-sales enterprise motions more than detailed self-serve enterprise list pricing. | Medium | SP018, SP021, SP023 |
| CP018 | Dataiku's ecosystem breadth across Databricks, AWS, Google Cloud, and NVIDIA reduces channel isolation and lets it sell into accounts already standardized on adjacent platforms. | Medium | SP003, SP004, SP005, SP006, SP007 |
| CP019 | Databricks has a scale advantage Dataiku cannot match publicly today, with Sacra estimating $6.9B in annualized revenue for 2026. | Medium | SP011 |
| CP020 | Public company-stat reporting tracked by Latka places DataRobot at roughly $285M of revenue in 2024, far smaller than Databricks and only modestly below Dataiku's last disclosed 2025 ARR markers. | Medium | SP019 |
| CP021 | H2O.ai's funding announcement says it serves 20,000 organizations and was valued at $1.7B after its 2021 Series E round. | High | SP022, SP025 |
| CP022 | Alteryx disclosed more than 8,000 customers globally and agreed to a $4.4B take-private transaction in December 2023. | High | SP023, SP024 |
| CP023 | Dataiku's strongest buying-criteria advantage is governed workflow breadth across technical and business personas without forcing one cloud or data platform choice. | Medium | SP001, SP002, SP003, SP006 |
| CP024 | Hyperscalers can undercut standalone platform value because ML tooling rides existing cloud identity, data gravity, and procurement paths. | Medium | SP012, SP014, SP016, SP013, SP015, SP017 |
| CP025 | Databricks benefits from owning both data and AI workflow surfaces, which raises switching costs once customers consolidate multiple workloads on the platform. | High | SP008, SP009, SP011 |
| CP026 | Dataiku's route to market appears coexistence-first rather than rip-and-replace, because its partner set includes companies whose native stacks also compete with it. | Medium | SP003, SP004, SP005, SP006, SP007 |
| CP027 | Specialists can still win budgets where buyers want faster time-to-value, narrower AutoML workflows, or low-code automation without replatforming the full data estate. | Medium | SP018, SP021, SP023 |
| CP028 | Alteryx remains most credible when the buyer's job is business analytics automation rather than governed multi-stage ML and agent deployment. | Medium | SP023, SP024 |
| CP029 | Metered hyperscaler pricing can reduce entry friction but also makes realized spend highly sensitive to model architecture and inference intensity. | Medium | SP013, SP015, SP017 |
| CP030 | Dataiku's lack of current public list pricing likely matters less in large-enterprise evaluations than architecture fit and governance requirements, but it still limits public TCO benchmarking. | Low | SP001, SP002, SP003 |
| CP031 | Dataiku's most durable moat claim is infrastructure neutrality combined with governed workflow breadth across mixed personas. | Medium | SP001, SP002, SP003, SP007 |
| CP032 | That moat is vulnerable if Databricks and the hyperscalers keep closing the governance and agent-functionality gap inside native environments. | Medium | SP008, SP009, SP012, SP014, SP016 |
| CP033 | Databricks is the highest-severity competitive threat because it combines adjacent budget ownership, a rapid AI roadmap, and far larger disclosed financial scale than Dataiku. | High | SP008, SP009, SP011 |
| CP034 | DataRobot is a cautionary adverse case for the category because analyses of its decline explicitly cite hyperscaler bundling, valuation compression, layoffs, and shifting buyer priorities around AI. | Medium | SP019, SP020 |
| CP035 | H2O.ai shows that hybrid and open-source-led challengers still have room, but its smaller disclosed funding base likely constrains global distribution compared with Dataiku or Databricks. | Medium | SP021, SP022, SP025 |
| CP036 | The Alteryx take-private underscores that analytics-automation value exists, but adjacent categories can be priced and financed very differently from AI-platform growth narratives. | Medium | SP023, SP024 |
| CP037 | Dataiku's partner-heavy distribution strategy partly mitigates displacement risk because it can ride ecosystems that also compete with it. | Medium | SP003, SP005, SP006, SP007 |
| CP038 | Public materials do not support a clean apples-to-apples realized-TCO comparison across Dataiku and peers because negotiated discounts, services mix, and cloud commitments remain private. | Medium | SP010, SP013, SP015, SP017 |
| CP039 | A material share of Dataiku's competition is effectively internal build plus cloud-native services, because large enterprises can assemble AI workflows from first-party tools without buying a neutral umbrella platform. | Medium | SP008, SP012, SP014, SP016 |
| CP040 | The coexistence of Dataiku with Databricks and other partners suggests it often competes more for workflow governance and collaboration ownership than for raw storage or compute budget. | Medium | SP003, SP004 |
| CI001 | Dataiku reported surpassing $300M of ARR in January 2025. | High | SI002, SI006 |
| CI002 | The January 2025 release said ARR had more than doubled over the prior three years and that more than 20% of customers were already using Dataiku for GenAI workflows. | Medium | SI002 |
| CI003 | Dataiku reported surpassing $350M of ARR in October 2025 while enterprises accelerated trusted-AI deployments. | Medium | SI003 |
| CI004 | Public Dataiku materials position the product as an enterprise AI platform whose value comes from orchestration, governance, agents, and deployment breadth rather than from a single standalone feature. | Medium | SI001, SI004 |
| CI005 | Archived Dataiku packaging showed a progression from free or small-team use into business and enterprise tiers with deeper automation, deployment, and governance capability. | Medium | SI005 |
| CI006 | Dataiku’s public record supports an enterprise contract model, but does not disclose current realized list pricing, discounting, or contract term mix. | Medium | SI001, SI005 |
| CI007 | Because Dataiku emphasizes cloud- and model-agnostic deployment rather than a single native cloud runtime, its monetization likely depends more on platform contract value than on raw compute resell. | Medium | SI002, SI004, SI024, SI025 |
| CI008 | Dataiku’s broad partner ecosystem suggests some implementation and enablement work can be carried by partners rather than fully in-house delivery teams. | Medium | SI008, SI024, SI025 |
| CI009 | Public sources do not disclose Dataiku’s revenue mix across software, support, hosting, and services. | Medium | SI002, SI003, SI005 |
| CI010 | C3.ai’s FY2026 disclosures show an enterprise AI platform can be overwhelmingly subscription-led: 91% of total FY2026 revenue was subscription revenue. | High | SI016, SI017 |
| CI011 | C3.ai’s 10-K says revenue consists of subscriptions and professional services, with software licenses, SaaS, stand-ready support, usage-based runtime, and hosting charges embedded inside subscription revenue. | High | SI015, SI016 |
| CI012 | C3.ai explicitly says it relies on partners for larger or continuing professional-services presence in order to maintain margin flexibility. | Medium | SI016 |
| CI013 | Snowflake’s FY2026 annual report describes a customer-centric, consumption-based pricing model in which revenue is recognized on customer consumption rather than ratably over a subscription term. | Medium | SI018 |
| CI014 | Snowflake reported 67% total gross margin and 72% product gross margin in FY2026, while professional services and other gross margin remained negative 31%. | Medium | SI018 |
| CI015 | C3.ai’s FY2026 results showed 31% GAAP gross margin and 46% non-GAAP gross margin, highlighting how enterprise AI-platform margins can vary materially with services intensity and execution quality. | High | SI016, SI017 |
| CI016 | Official pricing pages for SageMaker, Azure Machine Learning, and Vertex AI show that major adjacent rivals monetize primarily through compute, duration, and deployed-model activity rather than opaque seat pricing. | Medium | SI012, SI013, SI014 |
| CI017 | Databricks publishes usage-based price lists and Sacra describes its pay-as-you-go model as aligned to cloud consumption, reinforcing that much of the adjacent category prices infrastructure-linked usage. | Medium | SI010, SI011 |
| CI018 | Compared with usage-centric rivals, Dataiku’s packaging likely offers more budget predictability if contracts are primarily subscription based, but public evidence does not reveal realized terms or cost-to-serve. | Medium | SI005, SI010, SI012, SI013, SI014 |
| CI019 | Dataiku’s CFO said in January 2025 that the company’s growth and financial efficiency differentiated it from OpEx-heavy business models elsewhere in the AI ecosystem. | Medium | SI002 |
| CI020 | Public evidence is strong on top-line traction but weak on revenue quality because there is no disclosed retention, gross-margin, or cash-conversion data. | Medium | SI002, SI003, SI007 |
| CI021 | Reuters-reported IPO preparation indicates Dataiku had not publicly raised a new primary round after the December 2022 Series F as of late 2025. | Medium | SI006 |
| CI022 | Sacra estimated roughly $342.5M of ARR in September 2025 and roughly $846.8M of lifetime funding for Dataiku. | Medium | SI007 |
| CI023 | Sacra also estimated that Dataiku’s 2022 $3.7B valuation represented roughly 18.5x ARR on a $200M ARR base. | Medium | SI007 |
| CI024 | As of the run date, no public source in the reviewed pack discloses Dataiku’s current cash balance, burn rate, runway, or debt facilities. | Medium | SI002, SI003, SI006, SI007 |
| CI025 | The absence of a public cash and burn disclosure does not imply strength or weakness by itself; it simply leaves capital adequacy unverified. | Medium | SI006, SI007 |
| CI026 | EY says IPO markets gained momentum in 1H 2026, but execution windows remain episodic and can be shaped by mega-IPOs and geopolitics. | Medium | SI020 |
| CI027 | Forge describes 2025 as a modest but meaningful IPO reopening that created a cautiously stronger setup for 2026 private-company listings. | Medium | SI021 |
| CI028 | For Dataiku, an IPO path in 2026 would likely be about liquidity, recruiting currency, and financing optionality as much as about immediate survival capital. | Medium | SI006, SI020, SI021 |
| CI029 | Scaled ARR growth plus no announced follow-on round suggests no obvious public distress signal, but it still does not prove that Dataiku is self-funding or flush with cash. | Medium | SI003, SI006, SI007 |
| CI030 | Public-company comparators show that professional services or implementation-heavy work can materially dilute a software-margin narrative. | Medium | SI016, SI017, SI018 |
| CI031 | Snowflake and C3.ai together show a wide economics range for adjacent AI/data platforms, from low-30s GAAP margins to low-70s product margins. | Medium | SI016, SI017, SI018 |
| CI032 | Because Dataiku is private, serious financial underwriting currently relies on proxies rather than direct disclosure for margin, retention, and cash efficiency. | Medium | SI007, SI016, SI018 |
| CI033 | C3.ai reported $575.4M of cash, cash equivalents, and marketable securities at FY2026 year-end and $673M shortly after results, illustrating the balance-sheet buffer some enterprise AI vendors need while execution remains uneven. | High | SI016, SI017 |
| CI034 | C3.ai’s FY2027 guidance of $210M-$240M revenue after FY2026 $250.3M revenue is an adverse reminder that scaled enterprise AI vendors can still face growth and profitability pressure. | Medium | SI017 |
| CI035 | Snowflake disclosed about $9.8B of remaining performance obligations with about 46% expected to convert within 12 months, underscoring the kind of backlog visibility that Dataiku does not provide publicly. | Medium | SI018 |
| CI036 | Dataiku’s partner footprint and enterprise positioning imply longer sales cycles and larger ACVs than self-serve AI tooling, which makes revenue quality more dependent on retention and expansion than on pure sign-up volume. | Medium | SI008, SI009, SI024, SI025 |
| CI037 | Snowflake’s consumption model shows how optimization by customers can reduce near-term visibility even in a large-scale data platform, a risk Dataiku may partly avoid if contract commitments are firmer. | Medium | SI018 |
| CI038 | C3.ai’s 10-K explicitly notes that customers can reduce usage, renew on less favorable terms, or allow RPO to decline, illustrating why renewal and contracted backlog are critical diligence items for AI platforms. | Medium | SI016 |
| CI039 | Current Dataiku product pages emphasize governed AI agents, orchestration, and enterprise data controls, implying that monetization is tied to platform breadth rather than single-seat utility. | Medium | SI004 |
| CI040 | Without a public filing, Dataiku’s exact revenue-recognition policy, deferred revenue balance, contractual term mix, and remaining performance obligations remain unknown. | Medium | SI005, SI015, SI018 |
| CE001 | Dataiku positions the platform as a combination of people, orchestration, and governance rather than as a single-purpose ML tool. | Medium | SE001, SE002 |
| CE002 | Public product and documentation surfaces show that Dataiku supports both visual workflows and code-based work inside the same platform. | Medium | SE002, SE006, SE008 |
| CE003 | Current product messaging places enterprise AI agents and governed reasoning workflows near the center of Dataiku’s product story. | Medium | SE002, SE025 |
| CE004 | The product page says Dataiku can centralize agent creation, collaboration, lifecycle management, and orchestration across agents, models, and tools. | Medium | SE002 |
| CE005 | The govern product surface and documentation both present Govern as a dedicated layer for tracking AI initiatives, approvals, and audit readiness. | High | SE003, SE010 |
| CE006 | Govern documentation describes Standard and Advanced licenses, with advanced features such as GenAI registries, custom actions, blueprint design, custom pages, and scripting. | Medium | SE010 |
| CE007 | The Python API documentation says DSS APIs can be used anywhere code can run inside Dataiku, including recipes, notebooks, scenarios, and webapps. | Medium | SE009 |
| CE008 | Dataiku maintains a public Python API client repository and README, providing at least a modest external developer signal beyond marketing pages. | High | SE012, SE013 |
| CE009 | The DSS 14 release stream in June and July 2026 spans agentic AI and RAG, LLM Mesh, governance, MLOps, AI assistants, security, Git, and code tooling. | Medium | SE007 |
| CE010 | The frequency of 14.6.x and 14.7.x releases in mid-2026 suggests active enterprise product maintenance rather than a static legacy platform. | Medium | SE007 |
| CE011 | The public architecture reads as a control plane layered over build surfaces, automation, governance, and deployment operations rather than a monolithic closed system. | Medium | SE001, SE002, SE006, SE009 |
| CE012 | Archived packaging and partner pages indicate that Dataiku supports SaaS, customer-managed cloud, and on-prem or private-cloud styles of operation. | Medium | SE019, SE020, SE024 |
| CE013 | Partner pages with AWS, Google Cloud, Databricks, and NVIDIA show that Dataiku is intentionally built to integrate with large external infrastructure and AI ecosystems. | Medium | SE019, SE020, SE021, SE022 |
| CE014 | Dataiku’s public security surfaces document response guidance for 2026 Linux local-privilege-escalation vulnerabilities affecting Dataiku environments. | Medium | SE004, SE011 |
| CE015 | Govern is documented as an additional node integrated into the broader Dataiku platform rather than as a simple label inside the core UI. | Medium | SE010 |
| CE016 | The govern product page explicitly links the product to shadow-AI reduction, EU AI Act readiness, audit readiness, qualification, and signoff workflows. | Medium | SE003 |
| CE017 | Govern’s public materials identify bundle, model, and LLM registries as specialized registries inside the governance system. | Medium | SE003 |
| CE018 | Reference docs, developer guide, API docs, community links, and academy references together show a substantial enablement surface for platform users and builders. | Medium | SE006, SE008, SE009 |
| CE019 | The API documentation and public client indicate that Dataiku exposes programmable interfaces beyond the visual UI, which is important for enterprise automation and integration. | Medium | SE009, SE012, SE013 |
| CE020 | Dataiku’s core technical differentiation is the combination of business-user accessibility, code extensibility, orchestration, and governance in one shared platform. | Medium | SE001, SE002, SE006, SE008, SE010 |
| CE021 | Critical dependencies in the product design include customer cloud or on-prem infrastructure, external data systems, model ecosystems, and implementation partners. | Medium | SE019, SE020, SE021, SE022 |
| CE022 | Release-note items such as OS upgrades, Python version changes, Java minimum-version bumps, and model removals show that Dataiku actively manages a changing dependency stack. | Medium | SE007 |
| CE023 | Some public URLs collapse back to high-level marketing pages, which limits how much deep architecture detail can be independently validated from the public web alone. | Medium | SE002, SE025 |
| CE024 | Independent review evidence includes at least one critical readthrough around private-cloud integration, suggesting implementation friction can still matter even when users like the overall platform concept. | Medium | SE023 |
| CE025 | Govern’s advanced features and separate documentation imply that governance is a meaningful product area, not a superficial afterthought. | High | SE003, SE010 |
| CE026 | Public materials strongly suggest that Dataiku supports model abstraction, agent tools, and governed GenAI workflows, but they do not publish benchmarked performance or agent reliability metrics. | Medium | SE002, SE003, SE025 |
| CE027 | The SOC 2 announcement is useful historical trust evidence, but it does not by itself prove the full current certification inventory or present-day scope. | Medium | SE005, SE004 |
| CE028 | The v14 release stream supports the view that Dataiku is a mature, continuously shipped enterprise platform rather than an experimental toolset. | Medium | SE007 |
| CE029 | Dataiku’s product surface is explicitly built for both coders and non-coders, combining visual logic blocks and code-first tooling in one environment. | Medium | SE002, SE008 |
| CE030 | The mix of partner pages and competitor platforms implies that Dataiku is integration-oriented and orchestration-oriented rather than vertically locked to one proprietary infrastructure stack. | Medium | SE014, SE015, SE016, SE017, SE019, SE020, SE021 |
| CE031 | The public API-client repo is a modest but real developer-signal that third-party or customer-side automation against Dataiku is supported. | High | SE012, SE013 |
| CE032 | The fetched public pack does not expose benchmarked performance, uptime SLA terms, or quantified agent accuracy claims for Dataiku. | Medium | SE002, SE006, SE011 |
| CE033 | Public governance controls include qualification workflows, signoff rules, alerts, registries, and audit timelines. | High | SE003, SE010 |
| CE034 | Module maturity appears uneven but healthy: the DSS core looks mature, while governance and agent surfaces are newer yet clearly active and shipping. | Medium | SE002, SE007, SE010, SE025 |
| CE035 | Current public materials do not fully enumerate certifications, detailed security architecture, or quantified implementation outcomes, leaving trust and quality diligence incomplete. | Medium | SE004, SE005, SE011, SE023 |
| CE036 | Overall, Dataiku looks like a broad, mature enterprise AI platform with active roadmap velocity, governance depth, and programmable interfaces, but external proof on operational depth remains thinner than the breadth of the official story. | Medium | SE002, SE007, SE010, SE012, SE023 |
| CU001 | Dataiku said in January 2025 that it had grown its customer base to more than 700 organizations worldwide. | High | SU003, SU005 |
| CU002 | Dataiku said in October 2025 that the platform powered initiatives at more than 750 organizations worldwide. | Medium | SU004 |
| CU003 | A July 2024 KPMG alliance release said Dataiku already had more than 600 customers, including 200 Forbes Global 2000 companies. | Medium | SU009 |
| CU004 | Named public customer references clearly include healthcare and life-sciences organizations such as Johnson & Johnson, Novartis, and Roche. | High | SU015, SU016, SU022 |
| CU005 | Named public customer references also include manufacturing and industrial organizations such as Michelin, Mitsubishi Electric, and SLB. | High | SU017, SU024, SU025 |
| CU006 | Financial-services and capital-markets proof is visible through Standard Chartered and Euronext customer stories. | Medium | SU018, SU023 |
| CU007 | Logistics, transportation, real-estate, and food/agriculture adoption is visible through Geodis, Prologis, and Perdue Farms. | Medium | SU019, SU020, SU021 |
| CU008 | The public stories consistently show Dataiku being used by cross-functional groups that mix technical practitioners with business or operations users. | Medium | SU015, SU017, SU021, SU023, SU025 |
| CU009 | The likely economic buyer is a centralized data, analytics, digital-transformation, or platform budget rather than an isolated single-team software seat purchase. | Medium | SU009, SU017, SU021, SU023 |
| CU010 | The public record supports a customer-growth path from roughly 500 customers at the end of 2023 to 700-plus in January 2025 and 750-plus in October 2025. | High | SU006, SU003, SU004 |
| CU011 | Johnson & Johnson Vision's case study says 650-plus employees use Dataiku and more than 80 analytics and data-science professionals in the organization had adopted it. | Medium | SU015 |
| CU012 | Michelin expanded from 35 users in 2021 to more than 1,500 users across 50-plus factories by mid-2025, with 80% of users described as business experts. | Medium | SU017 |
| CU013 | Standard Chartered said 518 staff completed its citizen-data-science training program and more than 700 completed Dataiku online training courses. | Medium | SU023 |
| CU014 | Prologis said it had put over 60 AI/ML projects into production, maintained 30 projects and 30 APIs in active use, and had more than 2,000 users leveraging AI through its enterprise ChatGPT platform. | Medium | SU021 |
| CU015 | Mitsubishi Electric's story shows Dataiku being treated as a foundational component of the Serendie platform rather than a narrow point solution. | Medium | SU025 |
| CU016 | Johnson & Johnson used a two-day Dataiku event to build working generative-AI and LLM prototypes and then emphasized common-platform standardization afterward. | Medium | SU015 |
| CU017 | Novartis reported a 90% reduction in time to insights for a GenAI use case and a 600% acceleration in spreadsheet data-ingestion time with Dataiku. | Medium | SU016 |
| CU018 | Michelin said a defect-analysis workflow that previously took up to six months can now be completed in about an hour with a Dataiku-based solution. | Medium | SU017 |
| CU019 | Euronext reported up to a 20% reduction in time spent on recurring market-share queries after deploying a Dataiku-based analytics agent. | Medium | SU018 |
| CU020 | Geodis reported a 60% reduction in ticket-assignment time and about 30 minutes saved per ticket through a Dataiku IT support agent. | Medium | SU019 |
| CU021 | Perdue Farms reported that tasks which once took days are now completed in hours and highlighted more than six hours of monthly labor savings from automated reporting workflows. | Medium | SU020 |
| CU022 | Roche said Dataiku cut the time to build new GenAI projects from months to days and generated an estimated $100K-$250K of annual attorney time savings plus $375K-$475K of avoided consulting cost. | Medium | SU022 |
| CU023 | SLB said Dataiku-supported workflows assessed more than $10 billion of well-construction tenders, cut tender analysis from eight hours to twenty minutes, and made reservoir-pressure analysis 76% faster. | Medium | SU024 |
| CU024 | Standard Chartered said Dataiku-enabled FP&A workflows made analysts roughly 30 times more productive and supported a space-planning initiative expected to reduce annual property cost by $34 million. | Medium | SU023 |
| CU025 | Mitsubishi Electric reported a 60% workload reduction from data preparation to reporting, an 80% reduction in visualization time versus Python, and faster collaboration through shared flows. | Medium | SU025 |
| CU026 | Across multiple case studies, Dataiku's adoption pattern looks like a land-and-expand motion in which a first workflow broadens into governance, training, and wider business use. | Medium | SU017, SU021, SU022, SU023 |
| CU027 | The KPMG alliance indicates Dataiku can be sold as part of larger modernization, cloud, and AI-governance programs rather than only as standalone software. | Medium | SU009 |
| CU028 | Because partner-assisted transformation is part of the public story, some portion of Dataiku's commercial success likely depends on ecosystem leverage and implementation partners. | Medium | SU009, SU014 |
| CU029 | The 2025 Frontrunner Awards show active, public customer engagement around agentic AI, governance, productivity, and ROI use cases in multiple industries. | Medium | SU014 |
| CU030 | The strongest public Dataiku customer evidence is operational proof with quantified outcomes, not merely logo placement on a customers page. | Medium | SU015, SU016, SU017, SU018, SU019, SU021, SU022, SU024 |
| CU031 | Independent customer-satisfaction evidence is directionally positive: the Gartner page fetched for this run showed a review distribution of 75% five-star and 23% four-star ratings, while Dataiku's January 2025 release cited a 96% willingness-to-recommend score in Gartner Peer Insights. | High | SU007, SU003 |
| CU032 | The same Gartner review surface also included a critical review headline describing private-cloud integration issues, showing that implementation friction is not zero even for a well-regarded platform. | Medium | SU007 |
| CU033 | FeaturedCustomers and TrustRadius confirm that Dataiku has a visible public corpus of reviews and case-study references, but those directories do not establish renewal quality or average deployment success. | Medium | SU008, SU013 |
| CU034 | Dataiku does not publicly disclose NRR, GRR, churn, contract duration, or top-customer concentration in the sources reviewed for this run. | Medium | SU003, SU004, SU005, SU006 |
| CU035 | Comparing Sacra's 2023 estimate with 2025 official disclosures suggests customer growth is accompanied by a broader and likely more heterogeneous account base, not simply a small set of giant accounts growing in place. | Medium | SU006, SU003, SU004 |
| CU036 | Public evidence on Dataiku customers is strongest on named deployment outcomes and weakest on renewal, churn, and concentration. | Medium | SU015, SU017, SU021, SU007, SU003, SU006 |
| CU037 | Most of the strongest customer proof is vendor-hosted, which means the existence of deployments is credible but the representativeness of outcomes across the full base remains uncertain. | Medium | SU015, SU016, SU017, SU018, SU019, SU020, SU021, SU022, SU023, SU024, SU025 |
| CU038 | Customer-count disclosures moved from 600-plus in mid-2024 to 700-plus in early 2025 and 750-plus by October 2025, supporting a continued new-logo or expanded-disclosure momentum story. | High | SU009, SU003, SU004 |
| CU039 | The named customer base spans several regulated or complex sectors including healthcare, banking, capital markets, insurance-adjacent operations, logistics, and industrial manufacturing, which supports Dataiku's governance-centric enterprise positioning. | High | SU015, SU016, SU018, SU022, SU023, SU024 |
| CU040 | Customer durability should therefore be rated medium-confidence: switching costs are plausibly meaningful where Dataiku becomes the shared workflow and governance layer, but public retention and concentration data are insufficient to fully underwrite that durability. | Medium | SU017, SU021, SU023, SU007, SU006 |
| CR001 | Dataiku's privacy policy revised August 18, 2025 says the policy applies to visitors and users and that Dataiku acts as a controller for personal data collected under that policy. | Medium | SR002 |
| CR002 | The privacy policy says that when AI Services are enabled, content may be processed by Dataiku and the applicable third-party AI provider, and Dataiku says it will not use customer data to train its models without consent. | Medium | SR002 |
| CR003 | The February 2026 Dataiku Cloud DPA defines Dataiku as processor for customer personal data and explicitly references GDPR, CCPA, SCCs, subprocessors, and security incidents. | High | SR005, SR006 |
| CR004 | Dataiku's legal hub centralizes privacy, acceptable-use, installed-software, cloud-legal, and modern-slavery documents, indicating a relatively mature contractual surface for enterprise procurement. | Medium | SR001, SR004 |
| CR005 | Dataiku's trust page says self-managed deployments do not cause Dataiku to process or store client data by default unless the customer explicitly grants access. | Medium | SR007 |
| CR006 | The same trust page says Dataiku Cloud is a managed SaaS offering, available in multi-tenant or single-tenant form, and that it leverages cloud-provider infrastructure controls. | Medium | SR007 |
| CR007 | Dataiku publicly cites ISO 27001, ISO 27701, ISO 9001, SOC 1/SOC 2 assessments, a HIPAA compliance report, and GxP readiness as trust mitigants. | High | SR007, SR010 |
| CR008 | The EU AI Act subjects high-risk AI systems to obligations such as risk mitigation, logging, documentation, human oversight, and robustness, and its generative-AI transparency rules come into effect in August 2026. | Medium | SR011 |
| CR009 | NIST's AI RMF and its GenAI profile establish a strong market expectation that enterprise AI systems incorporate trustworthiness and structured risk management throughout design, deployment, and evaluation. | Medium | SR012 |
| CR010 | Dataiku's large legal and trust surface mitigates enterprise risk but also expands the contractual and operational obligations the company must consistently honor. | Medium | SR001, SR002, SR005, SR007 |
| CR011 | The FTC legal library shows that U.S. regulators continue to pursue privacy, false-advertising, and consumer-harm cases aggressively, underscoring that AI and data-platform claims face real enforcement risk. | Medium | SR013 |
| CR012 | The GDPR Enforcement Tracker page fetched for this run reported 3,202 tracked enforcement actions and €6.31B of total fines, confirming that privacy failures are financially material in Europe. | Medium | SR014 |
| CR013 | No Dataiku-specific FTC case or obvious GDPR fine surfaced in the public trackers fetched for this run, but absence from these surfaces is not proof of no latent exposure. | Medium | SR013, SR014 |
| CR014 | OpenCVE lists CVE-2023-51717, updated June 16, 2025, as a critical 9.8 incorrect-access-control flaw that could lead to full authentication bypass in Dataiku DSS before versions 11.4.5 and 12.4.1. | Medium | SR015 |
| CR015 | OpenCVE also lists historical medium and high-severity issues involving file access, Jupyter notebook permissions, metadata manipulation, and REST API information exposure in older DSS versions. | Medium | SR015 |
| CR016 | The Gartner review surface fetched for this run includes a critical review headline that specifically cites private-cloud integration issues. | Medium | SR017 |
| CR017 | UpGuard's public Dataiku vendor-risk page shows continuous external monitoring across 330+ checks, which is useful context but not a substitute for direct technical diligence. | Medium | SR016 |
| CR018 | Because Dataiku spans notebooks, APIs, connectors, automation, governance, and agent workflows, its operational attack surface is materially broader than that of a narrow single-purpose analytics tool. | Medium | SR008, SR015, SR022 |
| CR019 | Self-managed deployments reduce vendor data-custody exposure but shift more configuration, availability, and security responsibility into the customer environment. | Medium | SR007, SR008 |
| CR020 | Cloud delivery increases Dataiku's direct responsibility for incident handling, access controls, and subprocessor management while also inheriting risk from underlying cloud providers. | Medium | SR005, SR007 |
| CR021 | The KPMG alliance demonstrates that partner-led modernization and AI-governance programs are part of Dataiku's route to market, making services quality and channel economics relevant risks. | Medium | SR021 |
| CR022 | Customer stories repeatedly reference external systems such as Snowflake, Azure, ServiceNow, and PowerBI, which means interoperability is both a moat and a dependency chain. | High | SR025, SR026, SR028, SR030 |
| CR023 | Michelin, Prologis, Standard Chartered, and Geodis each describe Dataiku as part of a broader operational stack rather than an isolated tool, so connector reliability directly affects realized customer value. | High | SR025, SR026, SR028, SR030 |
| CR024 | Reuters reported in October 2025 that Dataiku hired Morgan Stanley and Citigroup to prepare for a possible U.S. IPO, increasing the valuation consequences of any compliance, security, or customer-retention stumble. | High | SR018, SR019 |
| CR025 | Even at $350M+ ARR scale, public sources still do not disclose core underwriting metrics such as NRR, GRR, gross margin, burn, or top-customer concentration. | Medium | SR020, SR022, SR023 |
| CR026 | That financial opacity is itself a material risk because investors cannot verify runway, renewal quality, or margin resilience if an IPO or financing window closes. | Medium | SR018, SR020, SR022 |
| CR027 | The EU AI Act's high-risk categories include employment and certain essential-service uses, so a general enterprise AI platform like Dataiku can become entangled in sensitive customer workflows even if Dataiku is not the end-use decision-maker. | Medium | SR011, SR024, SR028 |
| CR028 | Because the privacy policy contemplates AI Services using applicable third-party AI providers, model-vendor policy changes or data-handling concerns can transmit into Dataiku customer risk. | Medium | SR002, SR004 |
| CR029 | The DPA and cloud terms are real mitigants, but they also imply procurement friction because sophisticated enterprise buyers will review subprocessors, transfers, and incident obligations carefully. | Medium | SR005, SR006, SR021 |
| CR030 | No major public breach or enforcement event surfaced in the sources reviewed, but the disclosed vulnerability history means security patch cadence remains a core diligence item. | Medium | SR013, SR014, SR015, SR016 |
| CR031 | The public customer-proof set is rich, but most of the strongest risk-reducing evidence comes from vendor-hosted case studies rather than independent retention or concentration disclosures. | Medium | SR024, SR025, SR026, SR027, SR028, SR029, SR030 |
| CR032 | Customer stories imply that Dataiku often requires workflow redesign, training, and operational standardization, which increases implementation effort and therefore risk of slow time-to-value. | Medium | SR025, SR026, SR027, SR028 |
| CR033 | Regulated-industry strength is both a moat and a risk amplifier because healthcare, banking, and capital-markets customers require stricter validation, auditability, and incident response than ordinary SaaS buyers. | Medium | SR024, SR027, SR028 |
| CR034 | Prologis, Michelin, Standard Chartered, and Roche show that Dataiku can become embedded in business processes, which raises switching costs but also raises the business-interruption cost of failure. | Medium | SR025, SR026, SR027, SR028 |
| CR035 | The August 2026 EU AI Act transparency milestone increases pressure on Dataiku's governance narrative precisely as enterprises expand agentic-AI deployments. | Medium | SR011, SR022 |
| CR036 | Certifications and trust documentation improve the mitigation story, but they do not eliminate the need for rapid patching, careful access management, or customer-specific deployment diligence. | Medium | SR007, SR010, SR015 |
| CR037 | Execution risk remains meaningful because a 1,250+ employee, pre-IPO enterprise platform company must coordinate product delivery, customer success, security, and partner operations at a much higher bar than an earlier-stage startup. | Medium | SR018, SR022 |
| CR038 | Customer concentration and partner-sourced pipeline remain unresolved because public sources show reference accounts and alliances but not revenue weighting or channel mix. | Medium | SR020, SR021, SR024 |
| CR039 | Dataiku's 2025 trusted-AI and agent-management positioning helps the mitigation case, but it also raises expectations that the company can operationalize governance safely at scale for customers. | Medium | SR022, SR011, SR012 |
| CR040 | Overall residual risk should be considered medium-high: Dataiku has real controls and market relevance, but security vulnerabilities, regulatory expansion, ecosystem dependence, and financial opacity still create several plausible thesis-break paths. | Medium | SR015, SR018, SR020, SR022 |
| CV001 | Dataikus last disclosed primary valuation anchor is the $3.7B December 2022 Series F led by Wellington Management. | Medium | SV001 |
| CV002 | Dataiku disclosed surpassing $300M ARR in January 2025. | High | SV002, SV004 |
| CV003 | Dataiku disclosed surpassing $350M ARR in October 2025. | Medium | SV003 |
| CV004 | The stale $3.7B mark implies about 12.3x ARR against the January 2025 $300M milestone. | High | SV001, SV002 |
| CV005 | The stale $3.7B mark implies about 10.6x ARR against the October 2025 $350M milestone. | High | SV001, SV003 |
| CV006 | Using Sacras roughly $342.5M September 2025 ARR estimate, the stale $3.7B mark implies about {mult["dikusacra"]}x ARR. | Medium | SV001, SV005 |
| CV007 | Reuters reported in October 2025 that Dataiku hired Morgan Stanley and Citigroup to prepare for a possible U.S. IPO. | High | SV004, SV007, SV008 |
| CV008 | Forbes and official customer materials support that Dataiku is a large private AI/data platform with a blue-chip enterprise customer base. | Medium | SV006, SV029 |
| CV009 | Public evidence still does not disclose Dataikus gross margin, NRR, cash burn, or top-customer concentration. | Medium | SV003, SV005 |
| CV010 | Palantirs public market-cap-to-revenue proxy is about {mult["pal"]}x based on the fetched July 2026 market-cap and revenue pages. | Medium | SV009, SV010 |
| CV011 | Snowflakes public market-cap-to-revenue proxy is about {mult["snow"]}x based on the fetched July 2026 market-cap and revenue pages. | Medium | SV013, SV014 |
| CV012 | Datadogs public market-cap-to-revenue proxy is about {mult["ddog"]}x based on the fetched July 2026 market-cap and revenue pages. | Medium | SV015, SV016 |
| CV013 | MongoDBs public market-cap-to-revenue proxy is about {mult["mdb"]}x based on the fetched July 2026 market-cap and revenue pages. | Medium | SV017, SV018 |
| CV014 | ServiceNows public market-cap-to-revenue proxy is about {mult["now"]}x based on the fetched July 2026 market-cap and revenue pages. | Medium | SV019, SV020 |
| CV015 | Confluents public market-cap-to-revenue proxy is about {mult["cflt"]}x based on the fetched July 2026 market-cap and revenue pages. | Medium | SV021, SV022 |
| CV016 | C3.ais public market-cap-to-revenue proxy is about {mult["c3"]}x based on the fetched July 2026 market-cap and revenue pages. | Medium | SV011, SV012 |
| CV017 | Compared with public comps, Dataikus stale roughly 10-12x ARR multiple sits above weaker AI software names like C3.ai and around MongoDB/Confluent-ServiceNow territory, while still below Snowflake, Datadog, and far below Palantir. | Medium | SV001, SV003, SV009, SV010, SV011, SV012, SV013, SV014, SV015, SV016, SV017, SV018, SV019, SV020, SV021, SV022 |
| CV018 | Snowflakes FY2026 filing shows what premium platform economics can look like, with 67% total gross margin and 72% product gross margin. | Medium | SV024 |
| CV019 | C3.ais FY2026 disclosure and results remind investors that enterprise AI software can trade at much lower multiples when growth and margins disappoint. | Medium | SV026, SV027 |
| CV020 | Because Dataiku is private, illiquid, and economically opaque, it deserves a discount to the very highest public AI software multiples even if its category position is strong. | Medium | SV004, SV005, SV017 |
| CV021 | Conversely, Dataikus enterprise scale, governance positioning, and reference customer quality argue against valuing it at the lowest public AI software multiple tier. | Medium | SV003, SV006, SV029, SV030 |
| CV022 | The valuation debate is therefore not whether Dataiku deserves a premium at all, but where inside a very wide public comp band it belongs. | Medium | SV005, SV017 |
| CV023 | A reasonable bear-case valuation range is about $2.5B-$3.2B if growth slows, IPO timing slips, or hidden economics disappoint. | Medium | SV004, SV005, SV016, SV022, SV027 |
| CV024 | A reasonable base-case valuation range is about $3.5B-$4.5B if ARR scale is real and economics are decent but not elite. | Medium | SV003, SV005, SV017, SV018 |
| CV025 | A reasonable bull-case valuation range is about $5.0B-$6.5B if Dataiku can prove premium retention, strong margins, and credible IPO readiness. | Medium | SV003, SV004, SV006, SV017 |
| CV026 | The current disclosed $3.7B mark sits inside the base-case range, so the fairest present stance is fair rather than cheap or wildly stretched. | Medium | SV001, SV003 |
| CV027 | The evidence supports a track recommendation with medium confidence and a medium-high risk rating. | Medium | SV004, SV005, SV026 |
| CV028 | The missing economics—especially gross margin, NRR, cash burn, services mix, and concentration—prevent a stronger buy call despite the companys quality. | Medium | SV003, SV005, SV026 |
| CV029 | Dataikus strongest pro-valuation signals are scale, customer quality, and governance-first enterprise positioning. | Medium | SV003, SV006, SV029, SV030 |
| CV030 | Dataikus strongest anti-valuation signals are competitive intensity, implementation complexity, regulatory expansion, and economic opacity. | Medium | SV004, SV005, SV027, SV028 |
| CV031 | The scenario weighting used here is roughly 25% bull, 50% base, and 25% bear. | Medium | SV005, SV017 |
| CV032 | Price-sensitive upside from todays mark therefore depends on evidence that would move the company from the base case toward the bull case, not merely on continued category excitement. | Medium | SV004, SV024, SV025 |
| CV033 | If Dataiku can combine strong 2025 ARR momentum with credible IPO disclosures, the public-market path could unlock multiple expansion above the stale private mark. | Medium | SV003, SV004, SV006 |
| CV034 | If IPO timing slips or public-market investors demand cleaner economics, the valuation could compress below the last disclosed mark even without an outright business failure. | Medium | SV004, SV007, SV008, SV027 |
| CV035 | The fetched public comp set spans roughly 4.6x at the low end to 58.2x at the high end, proving that the right answer for Dataiku must be scenario-driven rather than a single-point multiple. | Medium | SV009, SV010, SV011, SV012, SV013, SV014, SV015, SV016, SV017, SV018, SV019, SV020, SV021, SV022 |
| CV036 | The final diligence asks should prioritize ARR recency, gross margin, NRR/churn, cash and burn, concentration, partner economics, and preference stack. | Medium | SV005, SV026 |
| CV037 | A thesis-break trigger would be materially weak retention or unexpectedly high customer concentration once private data is disclosed. | Medium | SV005, SV029 |
| CV038 | A second thesis-break trigger would be a major security or regulatory event that interrupts IPO timing or weakens trust with large enterprise customers. | Medium | SV004, SV027, SV029 |
| CV039 | A third thesis-break trigger would be an updated ARR bridge showing that growth has slowed meaningfully below what investors infer from the 2025 milestones. | Medium | SV002, SV003, SV005 |
| CV040 | Overall evidence quality is medium because the valuation analysis still relies on a stale private mark and public-comp proxies rather than current audited company disclosures. | Medium | SV001, SV005, SV024, SV026 |