Get the data model right and most of your technology problems were never technology problems.
Every failed ERP. Every dashboard nobody trusts. Every AI pilot that died quietly between the demo and production. Every migration that ran double the budget and delivered half the promise. Pull the thread on any of them and you arrive at the same place, which is almost never where the post-mortem looked.
The model was wrong. The data didn’t mean one thing. The tool was asked to carry a distinction the model couldn’t hold, and it couldn’t, so people built spreadsheets, and the spreadsheets became the real system, and the expensive platform became an ornament.
Taxonomy is the discipline of fixing that at the root. Not the tool. The root.
What taxonomy actually is, once you strip the jargon
Taxonomy is how your organisation decides what things are. What counts as a customer. What makes two parts the same part. Which document is a matter and which is correspondence. Whether these three records are one entity wearing three names. It sounds abstract until you notice that every consequential technology question is really this question in disguise.
Ask a firm how many customers it has and watch what happens. If the answer takes a meeting, the problem is not the CRM. It’s that “customer” means one thing to sales, another to finance, a third to the regulator, and nothing consistent to the database — and no software will reconcile a distinction the organisation has never actually made.
Information architecture is taxonomy made operational: the structure that carries those decisions through your systems so that a machine, an auditor, or an AI can rely on them. Get it right and everything downstream gets easier, cheaper and more defensible. Get it wrong and you will keep buying tools to paper over a gap that no tool can close, because the gap is upstream of every tool.
Why this is the front door, not a footnote
Most consultancies treat data as a workstream — something you tidy up after the platform decision, in a governance phase nobody funds properly. That ordering is the mistake, and it’s the mistake I’m called in to unwind.
The model comes first. It comes before the tool, because the tool serves the model. It comes before the regulation work, because you cannot evidence an obligation your model cannot express. It comes before the AI, because retrieval quality, agent reliability and every hallucination you’re worried about trace back to whether the underlying content was classified with enough discipline for a machine to trust it.
That’s why this sits beside the framework, not under it. The framework tells you whether you can see, govern, run and change your estate. Taxonomy is what makes the answers to those questions true rather than aspirational.
Where taxonomy earns its keep
The discipline is general; the value is specific. A few of the places it pays for itself many times over:
Product and parts data. Manufacturers and distributors drowning in duplicate SKUs, unmanaged supersessions, and attributes that mean different things in different catalogues. This is not a PIM problem. The PIM is fine. The classification model you loaded into it was the problem, and a better PIM will hold the same confusion more efficiently.
Matter and document classification. Legal and professional-services firms buying document management systems to organise content whose underlying taxonomy nobody ever designed. The DMS vendor sells the shelves. Nobody sells the decision about what goes where and why, which is the part that determines whether anyone can find anything.
Regulatory evidence. The obligation-to-control-to-proof chain is a taxonomy problem before it is a compliance problem. If your controls aren’t classified against the obligations they satisfy, you cannot generate evidence — you can only assemble it, by hand, under time pressure, which is the thing every supervisor has learned to distrust.
AI retrieval and agents. Everyone is writing about RAG and nobody is connecting retrieval quality to classification discipline. An AI is only as reliable as the structure of what it retrieves from. Feed it a taxonomy that four teams defined four ways and you get a confident machine producing defensible-looking nonsense.
Standards crosswalks. UNSPSC, ETIM, eCl@ss, ISO 15926 — the industrial and scientific classification standards that let systems, suppliers and regulators speak the same language. Mapping between them is unglamorous, low-volume, and exactly the kind of thing that quietly determines whether an integration works or festers.
How I approach it
I don’t arrive with a template taxonomy and bend your business to fit it. That’s how you get a classification model that’s theoretically elegant and operationally ignored — the taxonomy equivalent of a beautiful process diagram nobody follows.
The work runs discovery, model, migrate, govern. Understand how the organisation actually distinguishes things today, including the informal distinctions living in people’s heads. Design a model that carries the distinctions that matter — the ones tied to a decision, a regulation, or a pound of value — and deliberately ignores the ones that don’t. Migrate to it without breaking what works. And govern it, because a taxonomy with no owner drifts back into the confusion you paid to escape, usually within a year.
It is slower than buying a tool. It is also the only thing that makes the tool worth buying.
- You know the data model is the problem. → Start a conversation
- You want the argument in full. → Get the data right first
- You want to see how bad it is. → Assessments