What once made a business hard to compete with may not anymore. As AI makes old advantages easier to copy, companies are betting on proprietary data to hold the line. The logic is appealing: If an organization has exclusive data and competitors don't, AI can't easily replicate it — an assumption now shaping product strategy, merger and acquisition rationale, and valuations.
But every company has proprietary data, even if it's just information on its own operations. What counts is whether that data holds up as a durable advantage against AI-driven disruption.
Take educational platform Chegg, which built its business around millions of expert-answered homework questions. Its massive library looked like a powerful defensive data moat — until generative AI arrived. Suddenly, customers could get the same answers without needing access to Chegg’s proprietary library.
Chegg’s data didn’t disappear. Its moat did, which prompts the question investors should be asking of every valuation based on proprietary data: Does it hold up in an AI-enabled market?
How to build a durable data moat in the age of AI
Static data decays. A good data moat protects an advantage, but a great one compounds it. More customers, interactions, and workflow ownership should make that advantage progressively harder to replicate — not just make the dataset bigger.
How that advantage becomes durable depends on the type of data and what makes it difficult to replicate.
Exclusive data creates a moat through scarcity
Every company brings some exclusive non-public data to the market. It’s information that is measured, created, collected, or controlled through the company’s privileged access to customers or segments of the market. That could mean exclusive licensing, aggregated customer insights, or data that’s created via research and analysis.
The moat is the barrier to access. The harder that information is for others to obtain or recreate, the stronger the advantage, the more effective the moat.
Think credit bureau data, CoStar’s property records, and CARFAX vehicle history records. CARFAX has built a broad network of data sources, including service and repair records, that underpins its vehicle history reports. These assets remain defensible because generative AI can't recreate what it can't access. Licensing, trusted relationships, proprietary collection, and direct creation don’t disappear when AI enters the picture. If anything, AI makes what's behind that barrier even more valuable.
Customer data creates a moat through workflow ownership
Customer information lives inside a platform and becomes essential to everyday business operations. Examples include patient records, ERP financials, HR platforms, and CRM histories.
The moat is where the data lives. Once software becomes the sole source of truth, embedded in workflows, reporting, compliance, and daily operations, it is difficult to replace — and everything downstream depends on it. But custody is not control. Customers may resist attempts to restrict how their data can be accessed by AI, putting pressure on platforms that treat access itself as part of the moat.
When data is combined across customers, it can become even more valuable by revealing industry benchmarks, patterns, and intelligence that no single dataset could surface alone. Turning system-of-record data into that kind of insight requires permission, architecture, and governance to pool, anonymize, or benchmark information responsibly. Workday, for example, uses de-identified data from thousands of customers to provide benchmarks across areas such as workforce composition, turnover, and compensation, while Toast uses restaurant performance data to provide benchmarking and insights to its customers.
For investors and operators, owning the information is only the start. As we explored in our analysis of AI’s impact on services, true advantage comes from turning it into results others cannot deliver.
Usage data creates a moat through learning at scale
Usage data captures what happens when people interact with a product: what they do, how they use it, and where problems occur.
In this case, the moat comes from learning at scale. More users generate more data, which can reveal insights that make the product better. Language app Duolingo, for example, learns from nearly a billion daily exercises to personalize lessons while payments platform Stripe uses payment activity across its network to improve fraud detection.
But more data does not automatically create a stronger moat. It must reveal something competitors cannot easily learn — and improve the product at the same time.
How to pressure-test an AI data-moat claim
Classifying data is the easy part. The real work is determining whether the advantage will last. For investors, three questions cut through the noise.
1. Is the exclusive information valuable enough?
Would the data materially improve a product, decision, prediction, or customer outcome? A lot of companies sit on data that checks the “proprietary” box but does little to change the product, decision, prediction, or customer outcome. Think basic customer demographics, routine transaction histories, or behavioral data that confirms what the market already knows. If the data does not change what the company builds, prices, or predicts, it is not a moat. It is just inventory.
2. Is it durably proprietary?
A data moat can lose its edge in two ways. The obvious one is a competitor catching up. The less obvious one is AI changing whether the data matters at all.
AI can erode a moat without ever gaining access to the underlying data. If it can infer, synthesize, or recreate the same value from other sources, competitors may no longer need the proprietary asset. Chegg is a case in point: its library remained proprietary, but AI could deliver the outcome users wanted without it.
3. Can the company operationalize it?
Valuable, proprietary information isn’t an advantage until a company can put it to work to improve products, decisions, or outcomes.
Insufficient instrumentation, legal or contractual restrictions, and lack of customer trust or feedback can all stand in the way. Without a differentiated outcome, there’s no moat.
How to strengthen a data moat
Identifying the data is only the beginning. Each type requires a different strategy to strengthen the moat and protect its value.
1. Exclusive non-public information: Widen the gap
Exclusive data is strongest when competitors cannot find, buy, scrape, or infer it. For example, credit bureaus protect their advantage through collection networks, history, and regulatory roles that are difficult to replicate.
To strengthen it: Focus investment on making the underlying information harder, not just more expensive, for competitors to recreate. Look for advantages that compound with time, scale, or access rather than advantages that competitors can overcome with enough capital.
2. Customer data: Turn custody into advantage
This type of data is strongest when customers cannot easily operate without it. The greater opportunity, however, is using that vantage point to reveal patterns no single customer could see alone.
To strengthen it: Own more of the workflow, not just the data. The more critical processes that run through your platform, the harder that position is to displace. Then use it to connect decisions, actions, and outcomes in ways point solutions cannot.
3. Usage data: Prove the flywheel
The best usage data creates a compounding advantage. More usage generates more insight, which improves the product, attracts more usage, and makes the next improvement possible. Over time, that experience becomes harder for competitors to catch up to.
To strengthen it: Build around what competitors can’t see. Focus on usage that reveals unique customer behavior, then turn those insights into product improvements competitors can’t easily match.
What data moats mean for operators
Classifying data is the diagnosis. Strategy comes next.
If the moat holds, strengthen it. Turn customer information into intelligence that can be benchmarked or productized — with the permission and trust to do it — or prove that more usage makes the product better, not just bigger. If the moat is weak, build elsewhere. Go deeper into the workflow, strengthen distribution, or secure partnerships competitors can’t copy.
Either way, do not take the answer on faith. Plenty of companies believe they have a moat but haven’t checked. Does the company actually have the rights to use the data as claimed? Is there evidence that the data is driving value?
The path that doesn't work is standing still, hoping the data holds.
The changing implications of data for M&A
For acquirers, data diligence should follow the same framework. A shaky data moat doesn’t just weaken the investment thesis. It can weaken the valuation, too.
Don’t let “proprietary data” sit on the diligence checklist unchallenged. Ask the hard question: What protects this business if the data stops being an edge?
Advantage should also show up in the numbers, through pricing power, retention, or new revenue opportunities. A thin data position backed by strong workflow ownership or customer trust can still be a good bet. Without either, the company may be underwriting a moat that no longer exists.