Skip to content
breferrariPublic

About

Durable, scoped memory for coding agents.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

 

History

66 Commits

Folders and files

Repository files navigation

Vestige

Durable, scoped memory for coding agents. Memories are markdown, and reach is declared per memory — so a lesson learned in one repository reaches the repositories it applies to, and no others.

Works in Claude Code as a plugin, and in Codex CLI over the same MCP server.

Status: early. Benchmarked thoroughly, run in anger not at all. See Maturity.

Why reach is the whole design

A shared memory pool with no notion of scope has one failure mode, and it gets worse as it grows: everything is visible to everything. Ask about retries in the payments service and you get retry lessons from five other codebases, ranked above your own.

Vestige makes reach a property of the memory:

scope who sees it where it is stored
project only the repositories it names in the repo, if it names only that repo
platform any caller sharing a platform the global store
general everything the global store

Two rules are enforced rather than trusted, and they matter more than they look:

  • Reach is narrowed, never widened. Claiming general while naming specific projects is downgraded to those projects — by the memory's own admission — and the original claim is kept as claimed_scope so the downgrade can be audited. Without this, a filter is defeated by anything that over-claims: at a 24% over-claim rate, narrowing holds the caller's view at 73 documents and the scores do not move, while the same rate of claims that cannot be narrowed grows the view to 100 documents and costs 13 points of in-project recall and 18 of transfer (found@5 0.705 to 0.579, and 0.772 to 0.614).
  • A memory that would reach nobody is refused, not widened. Granting the widest reach because the narrowest could not be determined is the opposite of what narrowing is for.

What it does

  • remember — record a lesson that transfers. The routing test is exactly that: would this help someone working in a different repository?
  • recall — everything this project may see, ranked by specificity then recency.
  • search — semantic search inside what it may see. The reach filter runs first and ranking second, so nothing outside your reach can appear however well it matches.
  • explain — why each memory was shown or withheld, with what was claimed beside what was recorded. Every failure in this layer otherwise presents identically as "no results", and this is what tells an empty store apart from a reach mismatch.
  • memory_status — where the stores are and what is visible.

Plus the parts that make memory actually happen rather than sit unused: a protocol injected once per session, a gate that asks you to search before delegating discovery, a capture skill that decides what qualifies, and an audit skill that keeps the store lean.

Install

/plugin marketplace add breferrari/vestige
/plugin install vestige

Node 22+. qmd is required and is installed and kept current for you — it is not an optional accelerator: the reach filter alone, ordered by specificity and recency, puts the right memory first 4.4% of the time and in the top five 22%; with semantic ranking inside the same filtered view those become 45% and 93%.

For Codex, see codex/.

Storage is configuration, not policy

Declare stores in .vestige/config.json — in the repo, or in ~/.vestige/:

{
  "stores": [
    { "name": "project",  "kind": "repo",     "path": ".vestige/memories", "accepts": ["project"] },
    { "name": "team",     "kind": "external", "url": "git@example.com:acme/memories.git",
      "path": ".vestige/.team", "accepts": ["platform", "general"] },
    { "name": "personal", "kind": "local",    "path": "memories", "accepts": ["*"] }
  ]
}

An external store is a separate git repository, sparse-cloned into the workspace and kept out of your project's history — for teams who keep memory out of product repos on purpose. Add vestige-shared to sync it.

You never choose a location by hand: the narrowed scope picks the store, so a memory's reach and its location cannot disagree.

What it will refuse to publish

Memories bound for a shared store are scanned before anything is staged. Credentials, private keys, UUIDs, absolute home paths, email addresses, internal hostnames and private IPs are quarantined, not published — per file, so one bad memory never holds up the clean ones beside it.

Stated plainly, because a security claim without its limits is worse than none: this is a deny-list over shapes. It does not catch a secret spaced out character by character, or one described in prose. Write the lesson, not the evidence.

Prior art

Vestige is not a from-scratch design and does not pretend to be. It takes a distribution model from one system and a reach model from another, and most of what is new exists because combining them exposed a gap that only appears when both are present.

  • The shared-pool distribution model, the discovery gate, the capture and audit skills, and the deletion-review default come from mcs-cli/memory and mcs-cli/shared-memories by Bruno Guidolim.
  • The facet model, the write contract and the scope narrowing come from obsidian-mind — core/lib/om/ is its code, vendored rather than reimplemented.
  • What is new here is chiefly that reach computes storage, the filter-before-rank retrieval over a per-caller view, the write-time content gate, and a host-agnostic core that also runs on Codex.

PROVENANCE.md says which is which, component by component, so nobody has to guess — including the people whose work it builds on.

Maturity

Honest about what is and is not established:

  • Measured on a corpus built to match a real store — 183 memories, median 502 words, gated against a working vault's own profile before any score is taken:

    • What reach buys, and what it costs. Two questions, measured on the same corpus: a project finding its own memory, and a project finding one another repository wrote that declares it applies to them.

      own memory another project's top slots spent on another project's memory
      one index per project 0.984 — not in the index, at any k n/a
      one shared pool 0.475 0.597 130–147 of 183
      declared reach 0.710 0.772 0

      Against a shared pool it is better at both, and it is the only one of the three that reaches another repository's lesson without showing the caller memories that do not apply to it. Against a per-project index it trades: that configuration retrieves a project's own memories better than anything measured here, and cannot reach another repository's memory at any k — the dash is an absent capability rather than a low score, because the document is not in the index.

    • Within a project, the right memory reaches the top five 84–93% of the time and is first 33–53%, depending entirely on how the question is phrased. The three query registers are reported separately because averaging them describes none of them.

    • Zero hits from another project and zero junk, in every register and arm, against a shared pool's 130–147 of 183. It follows by construction — the filter runs before the engine, so an inapplicable memory is never a candidate — which is what makes it a guarantee rather than a ranking that usually behaves.

    • 80 of 80 planted secrets quarantined, zero contaminated blobs reaching git history, zero clean memories held back.

    • The behavioural layer verified inside live sessions rather than only in tests.

    Full method, every number, and what the measurements do not establish: RECORD.md. The harness is public at memory-stack-lab.

  • Established, and bad: it cannot decline. Asked something the store has no memory of, it returns five confident memories that are indistinguishable — on every axis a caller can observe — from a real answer. Structural, measured, unfixed. Every ranking figure above is therefore conditional on the question having an answer.

  • Not established: it has not been used in anger on real work over time. There is no episodic tier — tested, and deliberately not built. Consolidation proposes but never writes, by design. Defects in this codebase have been found by benchmarking and by cross-platform CI rather than by anything failing in use; assume there are more.

  • Measured at scale, and the honest version is narrower than it sounds: writing 3,600 memories across 180 projects costs a flat 3.8 ms each. That run scopes every memory to its own project, so the 20 documents a caller sees measures write cost rather than what reach does to a view. Counting the view properly — at this corpus's 31% org-wide sharing rate — a caller at 64 projects searches 411 of 1,280 documents. Declared reach gives a field roughly three times smaller than a shared pool, not a constant one, and composing that with the retrieval curve puts 64 projects well into the degraded region. See RECORD.md.

Documentation

RECORD.md how it was built, what was measured, why this shape. Start here to evaluate it
ARCHITECTURE.md how it works — the write path, the recall path, the sync path
PROVENANCE.md which component came from which prior system, and what is new

MIT.

About

Durable, scoped memory for coding agents.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages