Your AI can follow every rule — and still do the wrong thing.
We spend a lot of time asking whether an AI system followed policy.
I think that's necessary.
I also think it's dangerously incomplete.
Imagine this:
An AI agent receives evidence E.
Policy version 7.3 says that under E, action X is permitted.
The agent proposes X.
The governance gate correctly evaluates the policy.
The correct authority approves it.
The tool executes exactly X.
Every control worked.
Every signature verifies.
Every policy check passes.
And six hours later we discover something uncomfortable:
Evidence E was wrong.
Nothing in the governance stack failed.
The system faithfully executed a bad premise.
That creates a distinction I think AI governance needs to make much more explicitly:
Policy compliance is not evidence validity.
And formal proof doesn't eliminate that boundary.
You can prove that the policy was internally consistent.
You can prove that the gate enforced it correctly.
You can prove who had authority.
You can prove exactly what was executed.
None of those prove that the evidence entering the decision represented reality.
So perhaps the interesting question isn't:
“Can we prove the AI followed the rules?”
It's:
“What happens when we can prove it followed the rules perfectly — and the premise was wrong?”
I think a serious governance architecture needs to preserve both.
Not just:
Evidence → Policy → Decision → Execution
but also what happens later:
Challenge → New evidence → Reevaluation → Authority response
without rewriting the original history.
Because sometimes there is no rogue agent.
No policy violation.
No unauthorized action.
No broken control.
Just a perfectly governed mistake.
How should a governance system respond when nothing violated policy — but reality proves the decision wrong?