Boards keep asking “how good is the model,” and it’s the wrong question for almost everything currently being sold as an AI agent. The model is maybe a third of what determines whether an agent is safe to deploy. The other two-thirds is the harness — and almost nobody outside engineering has a mental model for what that word even means.
An AI agent is not a smarter chatbot. It’s a model wrapped in a harness — the scaffolding of tools, memory, permissions, and control logic that lets the model take actions in the world rather than just produce text a human reads and acts on. That distinction matters enormously for governance, because the model and the harness fail in completely different ways, and a board evaluating only the model’s capability is evaluating roughly a third of the actual risk.
What the harness actually is
The harness is everything sitting between the raw model and the outcome. It’s the set of tools the agent is permitted to call — send an email, query a database, execute code, place an order — and the specific scope of each permission. It’s the memory architecture deciding what the agent remembers across steps and how that memory can be manipulated. It’s the control logic determining when the agent stops and asks a human, versus proceeding autonomously. And it’s the monitoring layer, or its absence, that determines whether anyone would even notice if the agent did something wrong.
Free · 4 minutes
Would you survive contact with a determined attacker — or an auditor?
Fourteen questions on access, patching, detection, and recovery — the basics that prevent most real incidents, and the ones most often assumed rather than verified. Banded finding on screen, full sheet by email.
Two agents built on the identical underlying model can carry wildly different risk profiles purely because of harness design. A model with read-only access to a customer database, no ability to send communications, and a hard stop before any action affecting a live record is a fundamentally different risk than the same model given write access, an email-sending tool, and full autonomy to act without checking in — even though “the model” being evaluated is identical in both cases.
Why this reframes the governance question
The board question that actually discriminates isn’t “which model are we using” — it’s “what can this agent actually do, unsupervised, and what happens if it does the wrong thing.” That’s a harness question, not a model question, and it’s answerable in concrete terms: what tools does the agent have access to, what’s the blast radius if each tool is misused, what triggers a human checkpoint, and what’s the audit trail if something goes wrong. A vendor pitching “our agent uses the latest model” is answering a question that matters far less than the one they’re not being asked.
This also means agent governance can’t be a single organisation-wide policy the way a broad AI usage policy might be. Each agent deployment needs its own harness assessment, because the same underlying model deployed with a narrow, monitored harness in one context and a broad, autonomous harness in another carries genuinely different risk in each — evaluated by the same question every time: what can this specific configuration actually do, and what happens when it’s wrong.
A useful discipline for any board sponsoring agentic work: ask for the harness diagram before the model name. Which tools, which permissions, which checkpoints, which logs — in that order. The model name is the least informative fact in the room.
None of this requires distrusting agentic AI wholesale. It requires asking the question that actually locates the risk, which sits overwhelmingly in the scaffolding around the model, not in the model’s raw intelligence — a distinction most current AI governance conversations still haven’t caught up to.
The same model, two different verdicts
A fund administrator deployed the same underlying model in two configurations. The first: an agent with read-only access to fund documents, summarising investor correspondence for a human reviewer who approves every outgoing communication before it sends. The second, piloted by a different team without central coordination: an agent with the same model, given a mailbox tool and instructed to draft and send routine investor updates autonomously, with human review only on exception. The first configuration’s worst-case failure is a bad summary someone catches before it matters. The second’s worst-case failure is an incorrect statement about fund performance reaching an investor with nobody’s name on the decision to send it. Evaluating “the model” tells you nothing about which of these you’re actually looking at — only the harness does, and the fund administrator’s governance process, until it was pointed out, had no mechanism for distinguishing the two.
Two configurations, one model, two very different risk registers — and only one of the two questions actually being asked in most governance conversations today.
The harness question gets harder once agents start calling other agents — see delegation-chain accountability for what happens to responsibility three hops into a chain.
Assessing an agent deployment properly means assessing the harness — tool access, permission scope, human checkpoints, audit trail — not the model card. A technology control assessment can run exactly that assessment against a specific agentic deployment before it goes live.
Free interactive tool
Interactive deadline calculator
Check which regulations apply to you and when
Regulation across the EU, UK, US and Asia-Pacific has moved considerably in the past eighteen months, and several headline dates have shifted more than once. Twelve questions, about three minutes.
Results are shown on screen — no email required. A dated summary is available to download, and can be sent on if that's more useful. What we do with your answers.
Governance is what happens when nobody is watching.
Policies are easy. Consistent decision-making is harder. Understand where governance exists and where it has quietly become assumed.
Full Governance by Sixteen Pillars
Govern your business. Prove your compliance.
A board assurance cockpit for EU-regulated financial firms — tamper-evident, hash-chained proof of governance across DORA, GDPR, NIS2, ISO 27001, the EU AI Act and MiCA. In development.
See what's coming