You cannot govern what you cannot see, and most organisations deploying agents cannot see what their agents are actually doing. Traditional monitoring tells you whether a system is up and how fast it responds. An agent needs something more: you have to be able to observe what it decided and why, audit that record after the fact, and reverse what it did when it gets something wrong. Observability, auditability and reversibility are the operational trio that separates an agent you can run responsibly from one you are simply hoping behaves — and they have to be designed in, because none of them comes for free.
Why service monitoring is not enough
Monitoring a conventional service asks “is it working?” Monitoring an agent has to ask “what is it doing, and should it be?” An agent makes decisions and takes actions, so uptime and latency tell you almost nothing about the risk. The questions that matter are different: what did the agent decide, what did it act on, what reasoning or inputs led there, and is that within bounds? A firm that monitors its agents like services knows they are running but has no idea whether they are running amok — which is exactly the blind spot that turns a small agent error into a discovered-too-late incident.
The operational trio
- Observability is being able to see, in real time and in detail, what an agent is doing — the decisions, the actions, the inputs and the reasoning where available. This is what lets you notice misbehaviour while it is happening rather than after the consequences land.
- Auditability is being able to reconstruct, after the fact, what an agent did and why — a durable, trustworthy record. This is what lets you answer a regulator, investigate an incident, or defend a decision. An agent whose actions cannot be reconstructed is one you cannot account for.
- Reversibility is being able to undo what an agent did when it was wrong. Because agents act, their mistakes are actions taken, and the ability to roll back — or the design choice to make consequential actions reversible or checkpointed — is what caps the damage.
Designing for all three
- Instrument decisions, not just operations. Capture what the agent decided and acted on, with enough context to understand why, as a first-class output of the system.
- Make the audit trail durable and complete. The record has to survive and has to cover the decisions that matter, or auditability is a promise you cannot keep.
- Prefer reversible and checkpointed actions. Where an agent takes consequential action, design so it can be undone or requires a checkpoint, so a mistake is recoverable rather than final.
- Wire observability to intervention. Seeing misbehaviour only helps if someone or something can act on it; connect observation to the ability to pause or stop the agent.
Observability, auditability and reversibility are to agents what monitoring, logging and backups are to conventional systems — the unglamorous operational foundations that make the difference between a capability you run with confidence and one that is an incident waiting to be discovered. Firms that build them in can deploy agents into real work and answer for what they do; firms that skip them are trusting systems they cannot see, cannot reconstruct, and cannot undo.
Free · 4 minutes
Would you survive contact with a determined attacker — or an auditor?
Fourteen questions on access, patching, detection, and recovery — the basics that prevent most real incidents, and the ones most often assumed rather than verified. Banded finding on screen, full sheet by email.
Who this is for
This reading is for:
- CTOs and SREs operating AI agents in production
- Risk and compliance leads who need to answer for what agents did
- Boards asking how the firm would know if an agent misbehaved
- Teams whose monitoring was built for services, not decision-makers
Sixteen Pillars helps firms build the observability, auditability and reversibility that let them deploy agents into real work and actually answer for what those agents do. Pricing is published at /pricing/. If this is live for your organisation and you would like an independent reading, the place to start is a conversation.
Sixteen Pillars is a technology governance consultancy based in Cyprus. Engagements run remote across the EU, UK, and Middle East, with on-site time where the engagement requires it.
Free interactive tool
Interactive deadline calculator
Check which regulations apply to you and when
Regulation across the EU, UK, US and Asia-Pacific has moved considerably in the past eighteen months, and several headline dates have shifted more than once. Twelve questions, about three minutes.
Results are shown on screen — no email required. A dated summary is available to download, and can be sent on if that's more useful. What we do with your answers.
Governance is what happens when nobody is watching.
Policies are easy. Consistent decision-making is harder. Understand where governance exists and where it has quietly become assumed.
Full Governance by Sixteen Pillars
Govern your business. Prove your compliance.
A board assurance cockpit for EU-regulated financial firms — tamper-evident, hash-chained proof of governance across DORA, GDPR, NIS2, ISO 27001, the EU AI Act and MiCA. In development.
See what's coming