Request Free Sample ×

Kindly complete the form below to receive a free sample of this Report

* Please use a valid business email

Leading companies partner with us for data-driven Insights

clients tt-cursor

Data Catalog Companies

ID: MRFR/ICT/4670-HCR
100 Pages
Nirmit Biswas
Last Updated: June 24, 2026

Data Catalog market enables organizations to discover, manage, and govern enterprise data assets with improved data quality, transparency, and compliance. Leading companies such as Informatica, Collibra, Alation, Microsoft, Google Cloud, and IBM are driving innovation through AI-powered metadata management, data governance, and analytics capabilities

Download PDF ×

We do not share your information with anyone. However, we may send you emails based on your report interest from time to time. You may contact us at any time to opt-out.

Data Catalog Market
Market Size
Forecast Period2026-2035
CAGR (2026-2035)21.14%
2025 Market SizeUSD 3.89 Billion
2035 Market SizeUSD 19.84 Billion
Key Players
Informatica
Collibra
Alation
Microsoft
Google Cloud
IBM
Opportunities
  • Data Mesh Architecture as a Catalyst
  • Emerging Market Digitalization
  • AI Model Governance and Data Lineage

Market Opening Overview

Why the Data Catalog Market Is Expanding at This Rate

The Data Catalog Market is compounding at 21.14% annually among the highest growth rates across the enterprise software landscape because three historically independent demand vectors have converged simultaneously. Per MRFR analysis, the market was valued at USD 3.89 Billion in 2025 and is projected to reach USD 19.84 Billion by 2035. The EU Data Act (effective September 2025) and U.S. Executive Order 14110 on AI governance have transformed catalog procurement from discretionary to compulsory in regulated industries: GDPR enforcement fines exceeded EUR 4.5 Billion cumulatively by end-2024, and DORA now requires financial institutions to maintain real-time data lineage for operational resilience reporting.

Simultaneously, the generative AI buildout has exposed a foundational problem enterprises cannot govern, trust, or audit AI models without knowing where the training data came from, which has made AI-powered data discovery and tagging a board-level priority rather than an IT programme. Cloud migration complexity completes the trifecta: as 85% of organizations adopt multi-cloud strategies, metadata silos proliferate, and only catalog infrastructure that spans AWS, Azure, and GCP can enforce coherent governance policy.

North America leads with approximately 44.9% market share (2025), anchored by hyperscaler ecosystems and Fortune 500 adoption; Asia-Pacific is the fastest-growing region at a 25.4% CAGR, driven by India’s Digital Personal Data Protection Act and China’s national data classification standards; Europe holds approximately 27%, compelled by DORA enforcement and the EU Data Spaces initiative. BFSI is the dominant end-user vertical at 26.4%, with Healthcare advancing at the fastest vertical growth rate of 24.1% CAGR.

What Structurally Separates Leaders from the Field

In the Data Catalog Market, the competitive moat is not feature breadth it is data graph depth and ecosystem lock-in. The leaders have built proprietary knowledge graphs that capture column-level lineage across heterogeneous stacks, behavioral usage patterns across thousands of analyst queries, and policy enforcement layers that survive infrastructure migrations.

This is architecture that takes years to replicate, not quarters. Informatica’s CLAIRE AI engine, trained on petabyte-scale metadata from thousands of enterprise deployments, classifies and tags assets with an accuracy that cold-start competitors cannot match. Collibra’s 1,000-customer governance workflow library, particularly its BFSI-specific approval workflows audited to Basel III standards, represents institutional knowledge embedded in product configuration.

Microsoft Purview and AWS Glue Data Catalog weaponize ecosystem bundling: Purview is included in E5 licensing, and Glue Catalog has zero marginal cost within AWS infrastructure spend a pricing architecture that pure-play vendors cannot match dollar-for-dollar. The vendors that will define the next decade are those who collapse the distinction between catalog, data quality, and AI governance into a single policy-enforcing fabric and who get there before platform consolidation renders standalone catalogs obsolete.

 Section 2: Top 10 Global Data Catalog Companies MRFR Rankings (2026)

All revenue figures validated from official company annual reports, SEC filings, or officially confirmed investor disclosures. Private company revenues marked ‘Undisclosed’ where no official published financials are available.

 

#

Company

HQ

Revenue (Validated)

Geo. Presence

Key Specialization

Notable Highlight

1

Informatica

Redwood City, CA, USA

USD ~1.67B FY2024 (est.) SEC 10-K FY2024, Guidance; formal annual filing pending

Americas, EMEA, APAC

AI-powered cloud data management, enterprise data catalog (CDGC), metadata governance, ETL lineage

Launched CLAIRE GPT (Oct 2024) as GenAI assistant for catalog; surpassed USD 1.68B Total ARR in Q3 2024 (SEC 10-Q)

2

Collibra

New York, USA / Brussels, Belgium

~EUR 149M / ~USD ~165M FY2024 Tracxn revenue estimate; private company, no published accounts

Americas, EMEA, APAC

Data Intelligence Cloud data governance, catalog, lineage, quality; strong in BFSI & regulated sectors

Named Leader in inaugural Gartner Magic Quadrant for Data Intelligence Platforms (Jan 2025, Collibra newsroom)

3

Alation

Redwood City, CA, USA

Undisclosed (private company)

Americas, EMEA, APAC, ANZ

Data intelligence platform behavioral catalog, AI-powered discovery, lineage; 650+ enterprise customers

Launched Alation Anywhere (Jun 2024) Slack-embedded catalog access for analyst workflow integration 

4

Microsoft (Purview)

Redmond, WA, USA

Not separately disclosed part of Azure/Intelligent Cloud segment (USD 105.4B group FY2024) Microsoft FY2024 Annual Report

Global (190+ countries)

Microsoft Purview Data Catalog Azure-native metadata management, M365 ecosystem integration, multi-cloud connectors

Expanded Purview with multi-cloud connectors for AWS S3 and GCP BigQuery (Mar 2025, Azure blog)

5

Google Cloud (Dataplex)

Mountain View, CA, USA

Not separately disclosed part of Google Cloud segment (USD 43.2B FY2024) Alphabet FY2024 Annual Report

Global

Dataplex / Data Catalog multi-cloud metadata management, BigQuery-native lineage, Vertex AI integration

Integrated Dataplex with Vertex AI for AI-powered data tagging within ML pipelines (Feb 2025, Google Cloud blog)

6

IBM (Watson Knowledge Catalog)

Armonk, NY, USA

Not separately disclosed IBM Software segment revenue USD 13.8B FY2024 IBM FY2024 Annual Report

Global (170+ countries)

Watson Knowledge Catalog hybrid cloud data governance, AI-powered discovery, IBM Cloud Pak for Data platform

Holds 7.6% market share in Metadata Management category (Mar 2024, PeerSpot); strong in IBM infrastructure estates

7

AWS (Glue Data Catalog)

Seattle, WA, USA

Not separately disclosed AWS total revenue USD 107.6B FY2024 Amazon FY2024 Annual Report 

Global

AWS Glue Data Catalog serverless metadata repository, deep S3/Redshift/Athena integration, data mesh support

Most widely deployed catalog in cloud-native data engineering; bundled with AWS Glue ETL at no separate catalog fee

8

Atlan

New York, NY, USA

Undisclosed (private company) raised USD 105M Series C at USD 750M valuation (May 2024, Atlan press release)

Americas, EMEA, APAC (19 regions)

Active Metadata Platform modern data stack-native (Snowflake, Databricks, dbt), developer-first governance, AI lineage

Series C at USD 750M valuation (May 2024); 7x revenue growth in 2 years; 400% enterprise sales growth Q1 2024

9

data.world

Austin, TX, USA

(private company)

Americas, EMEA

Enterprise data catalog knowledge graph-based discovery, open-data roots, data marketplace functionality

Supports semantic layer and knowledge graph metadata architecture; strong in US government and open data communities

10

SAP (Data Intelligence)

Walldorf, Germany

Not separately disclosed SAP total cloud revenue EUR 17.1B FY2024 SAP FY2024 Annual Report

Global (180+ countries)

SAP Data Intelligence ERP-native catalog, manufacturing & supply chain data governance, SAP BTP integration

Deepest catalog integration for SAP-centric enterprises; critical for organizations governing S/4HANA and ERP data estates

 

 Section 3: Detailed Company Profiles

1. Informatica | NYSE: INFA | Redwood City, CA, USA

Informatica’s strategic bet on a cloud-only, consumption-driven model abandoning perpetual licensing entirely in 2023 is the single most consequential product architecture decision in the Data Catalog Market this decade. Rather than defending a legacy install base, Informatica forced its enterprise customers through a cloud migration that simultaneously deepened product lock-in and accelerated ARR expansion.

Its Intelligent Data Management Cloud (IDMC) platform processes over 101 trillion cloud transactions monthly (Q3 2024, SEC 10-Q), generating behavioral metadata that continuously trains its CLAIRE AI engine a flywheel competitors without comparable production scale cannot replicate. 

The October 2024 launch of CLAIRE GPT converts catalog interaction from structured-form query to natural-language dialogue, compressing analyst onboarding from weeks to hours. With Total ARR of USD 1.68 Billion (Q3 2024) and Cloud Subscription NRR of 120%, Informatica’s renewal economics confirm that enterprises are expanding usage, not just renewing licenses. 

MRFR assesses that Informatica’s consumption pricing architecture where catalog usage revenue scales directly with enterprise AI and analytics workload growth structurally aligns vendor economics with the market’s fastest-growing demand driver.

 

2. Collibra | Private | New York, USA / Brussels, Belgium

Collibra has deliberately bet that the complexity of governance in regulated industries will pay off for specialization rather than breadth: their platform is constructed around structured approval workflows, policy documentation and the level of regulatory audit trail that generic catalogs don't reproduce.

More strategically significant is the acquisition by Collibra in 2025 of Raito, a data access management platform that extends governance from metadata description to real-time policy enforcement at the query level, although Gartner’s Leader status in the first-ever Data Intelligence Platforms Magic Quadrant (January 2025) confirms this positioning.

This is the bridge between catalog-as-registry to catalog-as-control-plane, addressing the very capability gap that enterprise security teams cite as the reason governance efforts stagnate. Collibra has a revenue of around EUR 149 million (FY2024 per Tracxn estimates; private company, no published accounts) and 1,000+ customers including major financial institutions. Its BFSI concentration is its deepest moat and largest concentration risk in case fintech platforms consolidate governance natively. 

MRFR believes that Collibra’s leadership journey hinges on its ability to productize its deep governance workflow for the mid-market before cloud-native companies commoditize the enterprise tier.

 

3. Alation | Private | Redwood City, CA, USA

Alation’s founding insight that catalog adoption fails when usage data is not embedded in the product itself gave it behavioral analytics as a native feature a decade before competitors recognized the differentiation. Its query log mining architecture, which surfaces the most trusted and frequently used tables directly in search results, reduces time-to-insight in ways that manually curated metadata glossaries cannot match at enterprise scale. The June 2024 launch of Alation Anywhere delivering catalog access as a browser extension and within Slack addresses the adoption friction that kills most catalog rollouts: analysts must leave their workflow to consult a separate tool.

By embedding catalog context inside the tools analysts already use, Alation collapses the compliance-versus-productivity tradeoff that has historically limited governance program reach. With 650+ enterprise customers and trusted by 40% of the Fortune 100 per company disclosures, Alation’s enterprise penetration is deep, but its challenge is monetizing the generative AI wave without becoming a feature within a hyperscaler’s broader governance bundle. 

MRFR assesses that Alation’s behavioral lineage architecture is a durable technical differentiation but only if it can anchor AI governance workflows before Microsoft Purview’s M365 bundling erodes mid-market renewal rates.

 

4. Microsoft (Purview) | NASDAQ: MSFT | Redmond, WA, USA

Microsoft Purview’s competitive advantage in the Data Catalog Market is neither its catalog feature depth nor its AI classification accuracy it is the fact that it is already deployed. For enterprises standardized on Azure and Microsoft 365, Purview is a licensing activation, not a procurement decision, which compresses the sales cycle from six months to six weeks.

The March 2025 addition of native multi-cloud connectors for AWS S3 and Google BigQuery is strategically significant: it converts Purview from an Azure-only registry into a credible cross-cloud governance layer, directly competing with Informatica and Collibra in enterprise accounts that run heterogeneous cloud estates.

Microsoft’s integration of Purview with Microsoft Fabric its unified analytics platform means catalog, data quality, and BI metadata are governed within a single control plane, a convergence architecture that analyst firms identify as the end-state for enterprise data governance. The constraint is portability: organizations that adopt Purview as their primary catalog incur structural switching costs that make multi-cloud strategy increasingly theoretical. 

MRFR assesses that Purview will capture the largest share of net-new catalog deployments among Azure-first enterprises through 2028, not through product superiority but through procurement inertia.

 

5. Google Cloud (Dataplex) | NASDAQ: GOOGL | Mountain View, CA, USA

Google’s entry into the enterprise catalog market via Dataplex reflects a fundamentally different architecture thesis from its peers: rather than building a catalog that connects to data, Google built a data governance layer that is native to the data lake. Dataplex governs data in place scanning BigQuery, Cloud Storage, and connected sources without requiring data movement or ETL re-engineering which eliminates the most common reason catalog projects stall in cloud-native organizations.

The February 2025 integration of Dataplex with Vertex AI creates a closed loop between data governance and model training: ML engineers can tag, validate, and trace training data provenance directly within the pipeline that produces AI models, rather than through a separate governance workflow.

This positions Google at the intersection of catalog and AI model governance the fastest-growing segment of the Data Catalog Market through 2035. Google Cloud’s constraint is enterprise relationship depth: outside of BigQuery-native organizations, Dataplex connector coverage for non-GCP sources lags Informatica and Collibra. 

MRFR assesses that Dataplex’s AI lineage architecture makes it the default catalog choice for organizations building GenAI pipelines on GCP, but its multi-cloud narrative requires continued investment to challenge the incumbents in heterogeneous enterprise accounts.

 

6. IBM (Watson Knowledge Catalog) | NYSE: IBM | Armonk, NY, USA

IBM’s catalog strategy is inseparable from its Cloud Pak for Data platform thesis: Watson Knowledge Catalog is most valuable as the governance layer within an IBM-native data and AI stack, not as a standalone procurement. For enterprises running Db2, DataStage, and Watson ML at scale, this integration depth creates genuine switching costs column-level lineage that traces from raw Db2 tables through DataStage transformations to Watson model outputs is a compliance artifact that no competitor can replicate without rebuilding the entire pipeline on IBM infrastructure.

IBM’s 7.6% market share in the Metadata Management category (March 2024, PeerSpot) reflects this installed-base loyalty. The challenge is the modernization cycle: IBM’s catalog interface and deployment experience lag cloud-native competitors by a generation, and enterprises initiating new catalog programs are more likely to evaluate Informatica or Atlan than IBM unless they are already committed to Cloud Pak. 

MRFR assesses that IBM’s catalog position is structurally defensive sticky within its installed base but unable to lead net-new enterprise procurement outside of IBM-primary IT organizations.

 

7. AWS (Glue Data Catalog) | NASDAQ: AMZN | Seattle, WA, USA

AWS Glue Data Catalog has achieved ubiquity in cloud-native data engineering not through sales motion but through infrastructure gravity: it is the default metadata store for every Glue ETL job, Athena query, and Lake Formation governance policy, meaning AWS data engineers encounter it before they evaluate alternatives.

This default-state deployment model has made Glue Data Catalog the most widely used catalog in the Data Catalog Market by deployment count, even though it generates no separate catalog revenue line it is priced as a free component within AWS Glue consumption spend. 

The strategic implication is that AWS is not competing to win catalog procurement decisions; it is competing to make catalog procurement decisions unnecessary for organizations already inside the AWS ecosystem. For enterprises with multi-cloud or on-premise footprints, Glue Data Catalog’s limited connector breadth beyond native AWS services remains a structural ceiling. 

MRFR assesses that AWS’s catalog strategy is a retention mechanism for the data engineering workload, not a standalone enterprise governance play but within that scope, its scale is unmatched.

 

8. Atlan | Private | New York, NY, USA

Atlan’s USD 750 million Series C valuation (May 2024) at a company stage where most governance platform peers were valued below USD 500 million reflects a specific investor conviction: that the modern data stack has created a generation of data engineers who refuse to use catalog tools that require IT deployment or vendor-managed professional services.

Atlan’s developer-first architecture API-first metadata APIs, dbt and Airbyte native connectors, Slack-embedded workflows is purpose-built for organizations running Snowflake, Databricks, and dbt as their core data infrastructure. Its 7x revenue growth in two years (per Atlan press release, May 2024) and 400% enterprise sales growth in Q1 2024 confirm that this architectural bet is being validated by enterprise procurement at scale.

The competitive risk is platform compression: if Snowflake’s Unity Catalog and Databricks’ Unity Catalog continue to embed native governance, Atlan’s connector-layer model may be disintermediated by the platforms it integrates with. 

MRFR assesses that Atlan’s window of independent platform leadership is the next 24–36 months, after which consolidation pressure from warehouse-native governance layers will require either a platform pivot or an M&A exit.

 

9. data.world | Private | Austin, TX, USA

data.world built its catalog architecture on a knowledge graph foundation at a time when the industry consensus was relational metadata repositories a bet that is now looking prescient as enterprise AI requires semantic understanding of data relationships, not just schema documentation. Its knowledge graph model enables entity resolution across heterogeneous sources (linking a ‘customer_id’ in Salesforce, a ‘cust_id’ in Snowflake, and a ‘CustomerKey’ in SQL Server as the same semantic entity) without requiring manual mapping a capability critical for organizations building AI training datasets that span multiple systems data.

The world’s strength in the U.S. government and open data communities reflects its early positioning as a data marketplace and collaboration platform, a heritage that has translated into experience with federated catalog architectures across jurisdictional boundaries. As a private company with no published financials, its revenue scale remains undisclosed. 

MRFR assesses that data.world’s knowledge graph architecture is a technically distinctive foundation for AI governance use cases, but its market share in commercial enterprise accounts is limited by brand awareness and sales capacity relative to VC-backed competitors.

 

10. SAP (Data Intelligence) | NYSE: SAP | Walldorf, Germany

SAP Data Intelligence occupies a structurally different competitive position from every other vendor in this ranking: it governs data that originates within SAP’s own application ecosystem S/4HANA, BW/4HANA, Ariba, SuccessFactors where no external catalog has the schema context, business object semantics, or change-event integration that SAP can deliver natively.

For the approximately 400,000 SAP enterprise customers globally, Data Intelligence is the only catalog that can trace lineage from raw SAP business transactions through BTP integration flows to downstream BI assets without requiring custom connector development. SAP’s cloud revenue of EUR 17.1 Billion (FY2024, SAP Annual Report) and its ongoing RISE with SAP migration program moving customers from on-premise ERP to cloud creates a structural pipeline for Data Intelligence adoption as part of cloud transformation initiatives. The ceiling is ecosystem boundary: outside the SAP estate, Data Intelligence has no competitive differentiation against Informatica, Collibra, or Atlan.

MRFR assesses that SAP Data Intelligence is the dominant catalog choice within SAP-centric enterprises, but its total addressable market is bounded by SAP’s customer footprint rather than the broader enterprise software universe.

 M&A Activity Tracker

Key verified transactions shaping the Data Catalog Market consolidation landscape (2022–2025):

 

Year

Acquirer

Target

Deal Value

Strategic Objective

2025

Collibra

Raito (Data Access Management)

Extend governance platform to include fine-grained data access controls closing the gap between catalog metadata and enforcement, enabling policy-as-code for AI and cloud data assets

2025

Collibra

Deasy Labs (AI Workflow Security)

Expand catalog coverage to unstructured data and AI workflows; critical for enterprises managing LLM training data provenance under EU AI Act compliance frameworks

2024

Informatica

Pending Salesforce acquisition (announced)

Not disclosed at time of writing see Informatica/Salesforce IR

Salesforce intent to acquire Informatica signals platform consolidation: embedding best-of-breed AI data management into the CRM/data cloud stack, compressing the standalone catalog market

2023

Collibra

Sifflet (Data Observability, France)

Integrate real-time data quality monitoring into governance workflows pre-empting Atlan and Monte Carlo by building observability natively into the Data Intelligence Cloud

2022

Alation

Secoda (integration partnership / bundling)

N/A

Broaden discovery ecosystem; address modern data stack coverage gaps with lighter-weight catalog entry points for SME and mid-market segments

 

Key Trend: M&A in the Data Catalog Market is driven by two structural forces: pure-play governance vendors acquiring data observability and access-control capabilities to build end-to-end data intelligence platforms (Collibra’s Raito and Sifflet acquisitions), and hyperscaler/CRM platforms acquiring catalog leaders to embed governance into broader cloud ecosystems (Salesforce-Informatica). The former consolidates the governance stack vertically; the latter threatens to eliminate standalone catalog procurement as a category.

 R&D Investment & Innovation Signals

Leading vendors are investing in AI-native catalog architectures, AI governance lineage, and federated deployment models:

 

•        Informatica’s CLAIRE GPT (launched October 2024) embeds generative AI directly into the catalog interface enabling natural-language data discovery, auto-documentation, and policy suggestion converting the catalog from a governance archive into an active AI agent. With 101+ trillion monthly cloud transactions processed, CLAIRE’s training dataset scale creates a classification accuracy advantage that cold-start AI features cannot replicate (Informatica IR, Q3 2024).

•        Collibra’s acquisition of Raito (June 2025) signals a strategic pivot from catalog-as-registry to catalog-as-enforcement: by embedding fine-grained data access controls directly into governance workflows, Collibra is building the policy-as-code layer that regulators under DORA and the EU AI Act increasingly require from financial institutions managing AI model training data.

•        Microsoft Purview’s March 2025 multi-cloud connector expansion adding native AWS S3 and Google BigQuery metadata ingestion is a competitive response to the argument that Purview only governs Azure data. If adopted at scale, it repositions Purview as a credible multi-cloud governance layer and compresses the addressable market for standalone catalog vendors in Azure-primary enterprises.

•        Google Cloud’s Dataplex-Vertex AI integration (February 2025) creates the first catalog-to-model lineage loop within a hyperscaler platform: ML engineers can now trace training data provenance from raw source through catalog classification to model training run, producing the audit trail required under EU AI Act Article 13. This positions Google at the center of the fastest-growing regulatory demand driver in the Data Catalog Market through 2030.

•        Atlan’s active metadata architecture continuously reading warehouses, pipelines, and BI tools to reconstruct an enterprise data graph in real time is the engineering approach that Gartner identifies as the next evolution beyond passive catalog repositories. Its developer-first deployment model (zero-day production readiness without professional services) is compressing enterprise time-to-value from six months to under four weeks, a capability gap relative to legacy catalog implementations that is driving competitive displacement in modern-stack organizations.

•        Federated catalog architectures are emerging as the response to cross-border data sovereignty regulations including China’s PIPL, the EU’s adequacy decisions, and India’s DPDP Act. Vendors building jurisdiction-aware metadata registries with residency-compliant lineage tracking will capture a premium market segment as multinational enterprises face conflicting regulatory requirements that single-region catalog deployments cannot satisfy.

•        AI model governance is opening a new catalog use case that did not exist before 2023: tracing the provenance of training data from raw sources through transformation pipelines to model outputs. Vendors who embed lineage-to-model tracing natively rather than requiring custom connector development will define the premium tier of the Data Catalog Market through the EU AI Act enforcement period beginning in 2026.