Update verification.
A deployed agent that updates live sites, then judges its own work from the outside and reverses it when the site comes back worse.
Deployed agent — this configuration is running in production. Client identities are withheld.
Classified as an agent: it holds a defined responsibility and chooses, recommends or routes a permitted action from context.
No model anywhere in this entry. It is deterministic end to end — the behaviour comes from rules, arithmetic and explicit boundaries.
Objective
Apply pending updates to a running site without trusting that they worked, by proving the site still behaves as it did before and undoing everything when it does not.
Runtime
01
Plan
Resolves eligible updates, holding back those marked risky for the site unless explicitly included.
02
Capture
Audits the live site across desktop and mobile: status, console, page errors, failed requests, screenshots.
03
Back up
Exports the database and archives the exact plugin and core files it is about to touch.
04
Apply
Performs the update, then purges the cache so the next audit sees the real result.
05
Judge
Re-audits, compares against the capture, and rolls everything back unless the verdict is a clean pass.
Agent anatomy
Context
- Site inventory and per-site profile
- Pending plugin and core updates
- A per-site list of plugins marked risky
- The pre-update audit as the baseline of correct behaviour
Capabilities
- Determine which updates are eligible
- Audit a live site from a real browser
- Classify an outcome as pass, review or fail
- Restore database, plugins and core
- Report the verdict and its reasons
Activation
- Operator runs plan, audit or update
Authority
What can it change?
- Autonomy
- Escalating
- Horizon
- Operator-invoked
Read
- Site inventory and profiles
- Installed plugin and core versions
- The live site as a browser sees it
Write
- Plugin and core files, during an update
- Site database, during update or restore
- Backups, audits and reports
Conditional
- Update a plugin marked risky only when explicitly included
- Update a manually-managed site only on explicit approval
- Keep an update only when the post-audit verdict is a clean pass
Prohibited
- Mutate anything during plan or audit
- Update without a database and file backup taken first
- Keep an update whose verdict is fail or review
- Proceed when free disk space is below the configured floor
Escalate
- New error of any class after the update — rolls back
- Visual difference beyond threshold — rolls back rather than judging it cosmetic
- A failed rollback is reported as failed, never as recovered
Composition and verification
Made from explicit parts.
Composition
- Host orchestrator holding privileged operations
- Disposable browser container for auditing
- Deterministic verdict classifier
- Database, plugin and core backups
- Per-run reports and reasons
Verification
- Judged from outside the system, as a browser experiences the site
- Errors compared before against after, so pre-existing noise does not fail an update
- A visual-only breach is treated as suspect, not acceptable
- Rollback is mandatory — the flag that once made it optional is retained as a no-op
- Verdict classification covered by dependency-free unit tests
Evidence
Check it yourself.
Rollback is mandatory, not a flag someone can forget to pass.
A mechanism, shown in source
--rollback-on-fail) ;; # Backward-compatible: rollback is now mandatory.The flag that once made rollback optional is still accepted, and does nothing. Removing it would have broken existing callers silently; keeping it as a no-op means an old command line still runs, and still rolls back. Two unconditional triggers sit downstream — an error trap on any failing command, and the verdict check after the update — with no flag guarding either.
- Verdict tests
- 6 passing
- node --test verdict.test.mjs
- Runtime
- 62 ms
- Dependencies in the verdict logic
- none
WTB's own site estate, not a client's. The verdict logic decides pass, review or fail from before/after page state; it reads no site content and the site list lives in a separate config it never touches.
What this establishes: that the mechanism exists and behaves as described, in source you can read and by figures you can reproduce. It does not by itself show the mechanism operating in production, and it is not an outcome observed by anyone other than us.
Next step
What you can do with this.
This configuration runs in production, client identity withheld. To discuss the same for your operation, start with the problem rather than the system: a first call is thirty minutes and costs nothing.
Related entries