Data model before tool selection. Say it to most vendors and watch the sentence not compute — their entire pitch depends on the tool being the decision. Say it to most consultancies and watch them agree in principle and then open a platform comparison deck anyway. Nobody in the fractional-CTO market I compete in says this and means it as the starting point rather than a caveat.
So let me mean it.
The failure this prevents
Here is the pattern I am called in to unwind, in some variation, most of the time. An organisation selects a platform — a CRM, a GRC tool, a data warehouse, an ERP — because a competitor uses it, the demo was persuasive, or someone senior had heard of it. The platform goes in. Some months later, someone needs an answer the platform cannot produce: which customers are in a specific regulatory category, which contracts carry which obligation, which asset supports which critical function. The platform is not wrong. It is faithfully reflecting a data model that was never designed to carry that question, because the question was never asked before the platform was bought.
The fix people reach for is a second tool — a reporting layer, an integration, a spreadsheet that becomes load-bearing and terrifying in equal measure. The actual problem was never addressed, because the actual problem sits one layer below the tool: what does “customer” mean here, and is it the same thing in every system that uses the word? In most organisations it is not. Marketing’s customer, finance’s customer, and the regulator’s definition of the same word are three different objects wearing one label, and no amount of tooling reconciles that. Only the model does.
What a data model actually has to do
A data model, in the sense I mean it, is not a schema diagram. It is the answer to a specific set of questions asked in the right order: what are the things this business needs to reason about — customers, assets, contracts, obligations, providers, incidents — and what has to be true about each one for the business to answer the questions a regulator, an auditor, or a board will eventually ask. Get that right and almost any competent platform can carry it. Get it wrong and the best platform on the market will faithfully automate the wrong answer, at scale, forever.
This is why I lead every engagement with the model rather than the market. Not because tools don’t matter — they do, and a bad tool on a good model is still a real cost — but because the model is the thing that has to be right first, and it is the thing almost nobody checks first. Regulation, covered on the previous page in this chain, tells you what has to be provable. The data model is the only thing that can actually carry that proof. The tool, covered next, is what you point at the model once it’s sound.
What this looks like in a real estate
A shipping operator under the IMO’s cyber-risk requirements needs to answer, for any vessel, what systems are on board, which are network-connected, and what the exposure looks like if one is compromised mid-voyage. The natural instinct is to buy a fleet-management platform and assume the answer falls out. It usually doesn’t, because “vessel” in the classification society’s world, “vessel” in the charter system, and “vessel” in the crewing platform are rarely the same record with the same identifiers — and a platform built on three disconnected definitions of its own core entity cannot produce one clean answer about it, however good its dashboards look. The fix is not a better dashboard. It’s making “vessel” mean one thing, consistently, before any tool is asked to report on it. Once that’s true, most competent fleet-management systems can carry the answer without difficulty.
Why this is unusually hard to fake
Data-model discipline is the kind of work that shows up only in its absence. A good model is invisible — questions get answered, audits go smoothly, the organisation stops noticing the plumbing. A bad model announces itself constantly, in duplicated records, contradictory reports, and the particular exhaustion of teams who have learned to distrust their own systems. Because the work is invisible when done well, it is chronically underinvested, and because it is unglamorous compared to a platform launch, it rarely wins a budget fight against a vendor with a better demo.
That is, honestly, the opportunity in doing it properly. Every regulation I write about — DORA’s register, CPS 230’s critical-operations mapping, the evidence chain underneath both — reduces, eventually, to the same test: can your data answer the question, correctly, on the regulator’s timeline, without a human reconstructing it from memory. A sound data model is not a nice-to-have underneath compliance. It is the only thing that makes compliance possible on demand rather than as a periodic ordeal.
If a recent audit, board question, or regulatory request took longer to answer than it should have, the model is almost always where the delay actually lived. A technology control assessment finds exactly where. See the full argument on How I Work, or the evidence chain this feeds in detail at evidence and assurance.