← Registry
Registry / 011AgentProductionNo model

Update verification.

A deployed agent that updates live sites, then judges its own work from the outside and reverses it when the site comes back worse.

Deployed agent — this configuration is running in production. Client identities are withheld.

Classified as an agent: it holds a defined responsibility and chooses, recommends or routes a permitted action from context.

No model anywhere in this entry. It is deterministic end to end — the behaviour comes from rules, arithmetic and explicit boundaries.

Objective

Apply pending updates to a running site without trusting that they worked, by proving the site still behaves as it did before and undoing everything when it does not.

Runtime

  1. 01

    Plan

    Resolves eligible updates, holding back those marked risky for the site unless explicitly included.

  2. 02

    Capture

    Audits the live site across desktop and mobile: status, console, page errors, failed requests, screenshots.

  3. 03

    Back up

    Exports the database and archives the exact plugin and core files it is about to touch.

  4. 04

    Apply

    Performs the update, then purges the cache so the next audit sees the real result.

  5. 05

    Judge

    Re-audits, compares against the capture, and rolls everything back unless the verdict is a clean pass.

Agent anatomy

Context

  • Site inventory and per-site profile
  • Pending plugin and core updates
  • A per-site list of plugins marked risky
  • The pre-update audit as the baseline of correct behaviour

Capabilities

  • Determine which updates are eligible
  • Audit a live site from a real browser
  • Classify an outcome as pass, review or fail
  • Restore database, plugins and core
  • Report the verdict and its reasons

Activation

  • Operator runs plan, audit or update

Authority

What can it change?

Autonomy
Escalating
Horizon
Operator-invoked

Read

  • Site inventory and profiles
  • Installed plugin and core versions
  • The live site as a browser sees it

Write

  • Plugin and core files, during an update
  • Site database, during update or restore
  • Backups, audits and reports

Conditional

  • Update a plugin marked risky only when explicitly included
  • Update a manually-managed site only on explicit approval
  • Keep an update only when the post-audit verdict is a clean pass

Prohibited

  • Mutate anything during plan or audit
  • Update without a database and file backup taken first
  • Keep an update whose verdict is fail or review
  • Proceed when free disk space is below the configured floor

Escalate

  • New error of any class after the update — rolls back
  • Visual difference beyond threshold — rolls back rather than judging it cosmetic
  • A failed rollback is reported as failed, never as recovered

Composition and verification

Made from explicit parts.

Composition

  • Host orchestrator holding privileged operations
  • Disposable browser container for auditing
  • Deterministic verdict classifier
  • Database, plugin and core backups
  • Per-run reports and reasons

Verification

  • Judged from outside the system, as a browser experiences the site
  • Errors compared before against after, so pre-existing noise does not fail an update
  • A visual-only breach is treated as suspect, not acceptable
  • Rollback is mandatory — the flag that once made it optional is retained as a no-op
  • Verdict classification covered by dependency-free unit tests

Evidence

Check it yourself.

Rollback is mandatory, not a flag someone can forget to pass.

A mechanism, shown in source

tools/wtb-wp-plugin-updates:383
--rollback-on-fail) ;; # Backward-compatible: rollback is now mandatory.

The flag that once made rollback optional is still accepted, and does nothing. Removing it would have broken existing callers silently; keeping it as a no-op means an old command line still runs, and still rolls back. Two unconditional triggers sit downstream — an error trap on any failing command, and the verdict check after the update — with no flag guarding either.

Verdict tests
6 passing
node --test verdict.test.mjs
Runtime
62 ms
Dependencies in the verdict logic
none

WTB's own site estate, not a client's. The verdict logic decides pass, review or fail from before/after page state; it reads no site content and the site list lives in a separate config it never touches.

What this establishes: that the mechanism exists and behaves as described, in source you can read and by figures you can reproduce. It does not by itself show the mechanism operating in production, and it is not an outcome observed by anyone other than us.

Next step

What you can do with this.

This configuration runs in production, client identity withheld. To discuss the same for your operation, start with the problem rather than the system: a first call is thirty minutes and costs nothing.

Start with the operating problem

Tell us what is not working.

A first call is thirty minutes and costs nothing. We say honestly whether we are the right people. You would be talking to the people at We Think Beautiful in Brussels — the ones who do the work, not an account layer above them.