BI should be a query, not a project

When your taxonomy is right, business intelligence falls out of the data naturally.

Someone asks a question. The system answers it. The answer is consistent, trusted, and repeatable. The dashboard is a view, not an assembly exercise.

Most organisations do not have this. They have something that looks like it from a distance — reports, dashboards, a BI tool with a logo on it — but requires significant effort to produce a number anyone will stand behind. That effort is not analytics. It is compensation for a structural problem that was never addressed.

Free · 4 minutes

When two of your systems disagree, do you know which one to believe?

Fourteen questions on ownership, lineage, and quality — the difference between a number on a dashboard and a number you could defend. Banded finding on screen, full sheet by email.

How it starts

Nobody defines taxonomy on day one because there is nothing complex enough to require it.

A single product line, one market, one system. The data model is simple because the business is simple. Things are named informally, structured loosely, owned by whoever built them. It works. There is no visible cost to the informality.

Then the business grows. A new product line. A second market. An acquisition. A new platform bolted on because the old one could not do what was needed. Each addition is made in the context of the immediate requirement, not the existing data model — because there was never really a data model, only a working assumption that had not yet been tested.

The assumption breaks quietly. It does not announce itself.

The overlap problem

The damage is not usually that data is missing. It is that data exists in too many places, defined differently in each.

A product appears in the catalogue, in the ERP, in the CMS, and in the reporting layer. Each system holds its own version. Each was added at a different time, by a different team, for a different purpose. None of them was designated the canonical source.

When the versions diverge — and they always diverge — there is no rule for which one wins. The query returns different answers depending on where you ask the question. The report from the sales team does not match the report from finance. Both are technically correct, given their source. Neither can be trusted as the answer.

This is not a data problem. It is a taxonomy problem. There are no first class citizens. Nothing is the authoritative definition. Everything is a negotiation.

What gets hired to fix it

Organisations at this point often bring in data science capability.

That is the wrong tool for the problem. Data science applied to a broken taxonomy does not fix the taxonomy. It produces sophisticated outputs from an unreliable foundation. The models are impressive. The numbers remain untrustworthy.

More critically, it is not repeatable. One analyst runs the numbers one way. Another runs the same report next quarter and gets a different result. Both followed a reasonable method. The business cannot tell which answer is right because the underlying data does not have a definitive version.

Decisions get made on instinct. Or the report gets run again, slightly differently, until it produces a number that feels right. Neither is analytics.

What clean taxonomy actually enables

When first class citizens are defined — one canonical source for each entity, one owner, one authoritative definition — the downstream systems reference rather than replicate.

A product is defined once. The catalogue, ERP, and reporting layer all point to that definition. When it changes, everything that depends on it reflects the change. There is one answer to every question about that product, because there is one place the question can be asked.

BI becomes a query. The dashboard is configured, not constructed. KPIs are calculated consistently because the inputs are consistent. Filtering works because the categories are real. Search returns trusted results because the schema is coherent.

Analytics is the work of asking good questions, not reconciling competing answers.

The cost of deferring it

Taxonomy defined late costs more than taxonomy defined early. Not because the concepts are harder, but because by the time the problem is visible the mess is load-bearing.

Data structures that were never designed are now depended upon by three systems and a reporting layer. Unpicking them is not a schema exercise. It is a migration, a reconciliation, and a political conversation about whose version of the data is correct.

Every bolt-on added without a data model conversation made this harder. The business paid nothing visible at the time. It is paying now, in every report that takes a week and cannot be verified.

A few rules

Define your first class citizens before you need them. What are the canonical entities in your business — product, customer, order, contract? Name them, own them, and make everything else a reference to them.

Overlap in taxonomy is not flexibility. It is deferred conflict. Two definitions of the same thing will diverge. Plan for one definition.

If a business question requires an analyst to answer it, ask why. Standard KPIs should be queries, not projects. If they are not, the data model is the problem.

Bolt-ons require a data model conversation. Before a new system is integrated, establish how its entities map to your canonical ones. Not after.

Repeatable reporting is an architectural outcome, not a tooling one. The BI platform is not the problem. The schema underneath it is.


Richard King is CTO at Sixteen Pillars — technical leadership, architecture, and governance for organisations that need it without the overhead of a full-time hire.

Most technology problems are not technology problems. They are control problems.

The systems exist. The investment has been made. The question is whether leadership can understand, direct, evidence, and sustain what those systems produce. Find out where control exists — and where it only appears to.

Leave a comment