Data Gravity and Why It Dictates Your Topology

Data gravity is the observation that as a dataset grows large enough, it becomes cheaper and faster to bring compute and applications to the data than to move the data to them — and past a certain scale, that gravitational pull stops being a minor operational consideration and starts dictating an organisation’s entire topology, whether or not anyone ever made that a deliberate architectural decision.

I’ve written about the open lakehouse pattern and about multi-cloud as insurance or tax generally. Data gravity is the underlying physical and economic reality that makes both of those decisions harder than they initially appear — because a genuinely large dataset doesn’t move cheaply or quickly, regardless of how architecturally elegant a migration or a multi-cloud strategy looks on a whiteboard.

Why gravity is a genuine constraint, not just friction

Moving a large dataset between clouds or regions costs real money in egress fees, real time in transfer duration, and carries real operational risk during the migration window — costs that scale with data volume in a way that makes a small dataset’s migration trivial and a genuinely large dataset’s migration a multi-month, carefully-planned project with its own dedicated budget. Past a certain scale, the economically rational choice is rarely to move the data to wherever a new application needs to run — it’s to bring the application to wherever the data already, expensively, sits.

Free · 4 minutes

Do you know what could take the business down — and have you priced it?

Fourteen questions on concentration, third-party dependence, resilience, and incident readiness — the exposures a board is accountable for whether or not it can see them. Banded finding on screen, full sheet by email.

Where this quietly determines architecture nobody explicitly decided

An organisation’s core dataset, accumulated in one cloud provider over years, exerts a genuine gravitational pull on every subsequent architectural decision — a new analytics workload, a new machine-learning pipeline, a new application feature all get built, by default, in the same cloud the data already sits in, not because anyone evaluated and chose that provider for the new workload specifically, but because the cost and complexity of moving the data elsewhere dominates any other consideration. This is precisely how an organisation can end up single-cloud in practice despite nominally evaluating multi-cloud options for every new project — the data’s gravity has already decided the outcome before the evaluation genuinely began.

Why this matters for decisions made years before the gravity is felt

The genuine lesson isn’t that data gravity is avoidable — for any dataset large enough to matter, it isn’t. It’s that the gravitational pull should be treated as a deliberate, foreseeable consequence of an early platform decision, not a surprise discovered years later when a genuinely compelling reason to use a different provider arrives and the migration cost turns out to be prohibitive. The vendor-lock-in discipline I’ve written about generally applies with unusual force here, because data gravity is a lock-in mechanism that doesn’t require any contractual terms at all — pure physics and economics accomplish the same effect a restrictive contract would, and it accumulates silently, growing stronger with every additional volume of data committed to the same location.

What the gravity actually cost one firm

A logistics analytics firm, several years into accumulating a large operational dataset on one cloud provider, found a genuinely compelling reason to evaluate a competitor’s specialised analytics offering — meaningfully better performance for their specific workload, at a lower cost. The evaluation stalled entirely once the data-egress cost and transfer time for moving the core dataset were actually calculated: the migration would have cost more than a year of the competitor’s projected savings, purely in data movement, before accounting for the operational risk of the transfer itself. The firm stayed on its original provider, not because it was still the best technical fit, but because the data’s own gravity had made leaving economically irrational years after the original platform choice was made.

What actually accounts for this deliberately

Treating the initial choice of where core data will live as a genuinely weighty, long-horizon decision — considerably more consequential than most individual application or tooling choices, precisely because of the gravity it will eventually exert on every subsequent decision. And, where genuine multi-provider flexibility matters enough to justify the cost, building the open, portable data-layer discipline I’ve written about for the lakehouse pattern generally specifically to reduce the gravity’s pull before it fully sets in, rather than after.

Assessing how much gravitational pull a current core dataset already exerts on an organisation’s architecture, and what it would cost to preserve or restore genuine topology flexibility, is exactly the kind of architecture review a technology control assessment is built to run.

The original platform choice, years earlier, had been entirely reasonable given what was known at the time. Nobody had been negligent. The gravity had simply accumulated, quietly, the way it always does, until it was the deciding factor in a decision nobody thought they were making back then.

This is precisely the lock-in risk AI-ready data architecture is built to avoid at the foundation layer.

Getting the underlying platform choice right the first time is considerably cheaper than fighting gravity later — see Data Warehouse & Analytics Platform Strategy.

Recognising this pattern early is the only real defence against it — not avoiding the platform choice, which is unavoidable, but understanding from the start how much weight that early choice will eventually carry.

Free interactive tool

Website compliance checklist

What your site has to do, based on what it actually does

Answer as much or as little as you like — the list builds as you go. Nothing is stored against your name and no email is required.

Can you trust the architecture you have?

Architecture diagrams rarely show the reality of how systems actually operate. An independent review establishes what is really there.