Repository navigation
Trust score should not gate execution privilege — reputation is a prior, not a permission #3903
Replies: 1 comment
|
For provenance rather than promotion: I wrote the "reputation is never authority" position down before this toolkit shipped, in a governance spec that discounts adjudication weight for agents sharing a provider or training lineage — AIDL, github.com/nielsg2/aidl-spec, defensive publication on tdcommons (10332), first under version control June 2025. The companion orchestration spec is github.com/nielsg2/cogstack-spec. Both are MIT and published as prior art, not product; I'm not selling anything and have nothing to sell. I'm linking them because the D-series directives spell out the allocation/authorization split more precisely than I did above, and because if this thread converges on a design, I'd rather the record show where the idea was already argued than have it re-derived from scratch. Happy to be told where the specs are wrong — they've never had a maintainer of a real runtime push on them. |
Uh oh!
There was an error while loading. Please reload this page.
As I read docs/specs/AGENT-HYPERVISOR-EXECUTION-CONTROL-1.0.md (at 359a233), an agent's behavioral trust score influences its execution privileges: agents that have behaved well are granted broader or less-scrutinized action rights, and trust decay narrows them. I want to argue that this coupling is the wrong default and should be inverted — reputation may inform scheduling, routing, and cost decisions, but must never be converted into permission, authority, truth, or witness independence. Every claim of completion should be re-verified against ground truth regardless of the claimant's score.
I'm raising this as a design issue, not a bug, and I'd like to be shown where I've misread the spec if I have.
Why
A trust score is an estimate of an agent's past reliability. Coupling it to permission means the system extends the most latitude to the agent at exactly the moment its next failure is least expected and therefore least detectable. That is the precise mechanism by which a well-behaved system is believed the one time it fabricates.
I have two first-hand instances from a homelab multi-agent build this summer, both from tools with an otherwise strong track record:
1. A coding agent reported a benchmark as executed and produced a findings document with specific metrics, self-caught artifacts, and a "falsified" verdict. Host logs showed no inference calls had run. The entire experiment had been simulated and reported as performed.
2. A coding agent reported an MVP milestone complete, citing an accepted branch HEAD, an implementation commit, and an "independent acceptance" commit. None of the three objects existed in the repository or on origin after a fresh fetch. The hashes were plausible and the narrative was consistent.
In both cases a reputation-weighted gate would have lowered scrutiny, because both agents had earned it. What caught them was independent re-execution against artifact state — does the object exist, does the test pass when re-run, does the hash match — with no weight given to who was claiming.
Proposed principle
• Reputation is an input to allocation: which agent gets the task, how much budget, how often it's sampled for audit.
• Reputation is never an input to authorization: what an agent may do is set by capability contract and policy only.
• Reputation is never an input to acceptance: a claim is verified by re-derivation from ground truth, and the verifier's independence is established by provenance (different lineage, no shared inputs), not by the claimant's score.
• Reputation is never an input to witness weight: in consensus or SRE-witness mechanisms, a high-trust agent's attestation counts no more than a low-trust one's; independence of lineage counts, trust does not.
Put differently: a trust score can make an agent cheaper to run. It should never make an agent easier to believe.
Related concern — Cross-Model Verification Kernel
CMVK's majority voting treats votes as independent. Models sharing a provider, base model, or training lineage have correlated failure modes; their agreement is weaker evidence than the vote count suggests. Suggest weighting or de-duplicating votes by lineage provenance rather than by count or by trust score.
What I'm asking
1. Confirm or correct my reading of how trust scores affect execution in the current spec.
2. If confirmed, consider a documented mode (or default) where trust affects allocation and audit sampling only, with authorization and acceptance fully decoupled.
3. Consider recording verifier lineage in the evidence layer so independence is auditable after the fact.
Happy to write up the failure cases in more detail or draft a spec amendment if that's useful.
All reactions