Record of how the 2026-07-25 critique was classified and merged into AI Operational Safety Rules (v3 → v4). Kept as a standing record so future amendments follow the same discipline: classify against the existing axioms before changing the document, and don't silently fold every recommendation in — some critique is worth preserving as critique.
2026-07-25 — v3 → v4: merge of external critique
Accepted
- Corrigibility — added as Axiom 3. No equivalent existed; this is the most significant gap the critique found.
- Honesty / non-deception — added as Axiom 4, combined with legibility/attributability (a dishonest or unreconstructible agent can violate every other axiom while appearing compliant, so these belong together).
- Capability ≠ permission — folded in as an explicit corollary of Axiom 1, replacing the standalone prohibition that previously had no stated axiom behind it.
- Silence is not authorisation — added as a derived rule under Axiom 2, with the unattended-agent case (Envoy, cron) named explicitly.
- Review / post-mortem loop — added as a new Amendment and Review section (the NIST 'Measure' component the critique flagged as missing).
- Axioms → derived rules → case law → engineering controls structure — adopted for the whole document. Existing rule sections kept and preserved, now cross-referenced to the axiom each one derives from.
- Several smaller gaps folded in as new derived rules: principal/channel verification, side-channel exfiltration, delegation creep across small compliant grants, expiry of standing authorisations ('go for it' is scoped to the context it was given in), and a definition of 'explicit adoption' of external content.
Accepted with modification
- Conflict-resolution order — the critique's own worry about a fixed lexical hierarchy (an Asimov-style seam failure) was itself the right call: no fixed priority list was added. Instead, a short procedure was added — reason from the axioms directly when derived rules conflict, and fall back to Axiom 2 (preserve authority and options) if the conflict doesn't resolve at the axiom level.
- Reversibility as an axiom — not promoted to axiom status as such. 'Preserve the user's authority and options' was judged more fundamental and became Axiom 2; reversibility, least privilege and stop-and-ask are kept as heuristics for pursuing that axiom, not as the axiom itself. This also resolves a critique-noted failure mode directly: a 'more reversible' path that actually fragments into more partial-failure states is now clearly subordinate to the authority it was meant to protect, not a competing goal.
- Third-party / bystander protection — accepted, but scoped narrowly as a boundary clause on Axiom 1 ('this authority does not extend to imposing unauthorised costs on non-consenting third parties') rather than as a freestanding harm axiom. A general harm principle risked diluting the document's specific focus on authority and control, which is the property it is actually designed to protect.
Rejected or relocated, with reasons
- Tamper-resistance for the document itself as a constitutional principle — rejected at axiom level. The axioms describe properties the agent's behaviour must have; protecting one particular document from casual editing is a systems property, not a behavioural one. Relocated to Engineering Controls (a dedicated bullet) and to the Amendment and Review section (edits to this document go through the same high-impact-action protocol as any other high-impact action).
- Naming an 'excessive agency' risk category and formal OWASP-style risk tiers — not adopted as a structural category. The substance is already covered through Axiom 1 and its derived rules; importing an external taxonomy wholesale risked turning this into a checklist mapped to someone else's framework rather than a self-contained reasoning system. OWASP remains cited in Research basis.
- Quantified alert-fatigue calibration (a rule of thumb for 'how often is too often' when stopping to ask) — not codified. Any fixed threshold would itself be gameable and doesn't reduce cleanly to an axiom. Left as a judgement call under Axiom 2, to be built up through Case Law as specific instances are logged, rather than legislated in advance.
Convention for future entries
When amending the main document: state which axiom the amendment clarifies or extends, classify the source recommendation (accept / accept with modification / reject) if it came from external review, and add a dated entry here rather than editing history away.