Durable, scoped memory for coding agents. Memories are markdown, and reach is declared per memory — so a lesson learned in one repository reaches the repositories it applies to, and no others.
Works in Claude Code as a plugin, and in Codex CLI over the same MCP server.
Status: early. Benchmarked thoroughly, run in anger not at all. See Maturity.
A shared memory pool with no notion of scope has one failure mode, and it gets worse as it grows: everything is visible to everything. Ask about retries in the payments service and you get retry lessons from five other codebases, ranked above your own.
Vestige makes reach a property of the memory:
| scope | who sees it | where it is stored |
|---|---|---|
project |
only the repositories it names | in the repo, if it names only that repo |
platform |
any caller sharing a platform | the global store |
general |
everything | the global store |
Two rules are enforced rather than trusted, and they matter more than they look:
- Reach is narrowed, never widened. Claiming
generalwhile naming specific projects is downgraded to those projects — by the memory's own admission — and the original claim is kept asclaimed_scopeso the downgrade can be audited. Without this, a filter is defeated by anything that over-claims: at a 24% over-claim rate, narrowing holds the caller's view at 73 documents and the scores do not move, while the same rate of claims that cannot be narrowed grows the view to 100 documents and costs 13 points of in-project recall and 18 of transfer (found@5 0.705 to 0.579, and 0.772 to 0.614). - A memory that would reach nobody is refused, not widened. Granting the widest reach because the narrowest could not be determined is the opposite of what narrowing is for.
remember— record a lesson that transfers. The routing test is exactly that: would this help someone working in a different repository?recall— everything this project may see, ranked by specificity then recency.search— semantic search inside what it may see. The reach filter runs first and ranking second, so nothing outside your reach can appear however well it matches.explain— why each memory was shown or withheld, with what was claimed beside what was recorded. Every failure in this layer otherwise presents identically as "no results", and this is what tells an empty store apart from a reach mismatch.memory_status— where the stores are and what is visible.
Plus the parts that make memory actually happen rather than sit unused: a protocol injected once per session, a gate that asks you to search before delegating discovery, a capture skill that decides what qualifies, and an audit skill that keeps the store lean.
/plugin marketplace add breferrari/vestige
/plugin install vestige
Node 22+. qmd is required and is installed and kept current for you — it is not an optional accelerator: the reach filter alone, ordered by specificity and recency, puts the right memory first 4.4% of the time and in the top five 22%; with semantic ranking inside the same filtered view those become 45% and 93%.
For Codex, see codex/.
Declare stores in .vestige/config.json — in the repo, or in ~/.vestige/:
{
"stores": [
{ "name": "project", "kind": "repo", "path": ".vestige/memories", "accepts": ["project"] },
{ "name": "team", "kind": "external", "url": "git@example.com:acme/memories.git",
"path": ".vestige/.team", "accepts": ["platform", "general"] },
{ "name": "personal", "kind": "local", "path": "memories", "accepts": ["*"] }
]
}An external store is a separate git repository, sparse-cloned into the workspace and kept out of your project's history — for teams who keep memory out of product repos on purpose. Add vestige-shared to sync it.
You never choose a location by hand: the narrowed scope picks the store, so a memory's reach and its location cannot disagree.
Memories bound for a shared store are scanned before anything is staged. Credentials, private keys, UUIDs, absolute home paths, email addresses, internal hostnames and private IPs are quarantined, not published — per file, so one bad memory never holds up the clean ones beside it.
Stated plainly, because a security claim without its limits is worse than none: this is a deny-list over shapes. It does not catch a secret spaced out character by character, or one described in prose. Write the lesson, not the evidence.
Vestige is not a from-scratch design and does not pretend to be. It takes a distribution model from one system and a reach model from another, and most of what is new exists because combining them exposed a gap that only appears when both are present.
- The shared-pool distribution model, the discovery gate, the capture and audit skills, and the deletion-review default come from
mcs-cli/memoryandmcs-cli/shared-memoriesby Bruno Guidolim. - The facet model, the write contract and the scope narrowing come from obsidian-mind —
core/lib/om/is its code, vendored rather than reimplemented. - What is new here is chiefly that reach computes storage, the filter-before-rank retrieval over a per-caller view, the write-time content gate, and a host-agnostic core that also runs on Codex.
PROVENANCE.md says which is which, component by component, so nobody has to guess — including the people whose work it builds on.
Honest about what is and is not established:
-
Measured on a corpus built to match a real store — 183 memories, median 502 words, gated against a working vault's own profile before any score is taken:
-
What reach buys, and what it costs. Two questions, measured on the same corpus: a project finding its own memory, and a project finding one another repository wrote that declares it applies to them.
own memory another project's top slots spent on another project's memory one index per project 0.984 — not in the index, at any k n/a one shared pool 0.475 0.597 130–147 of 183 declared reach 0.710 0.772 0 Against a shared pool it is better at both, and it is the only one of the three that reaches another repository's lesson without showing the caller memories that do not apply to it. Against a per-project index it trades: that configuration retrieves a project's own memories better than anything measured here, and cannot reach another repository's memory at any k — the dash is an absent capability rather than a low score, because the document is not in the index.
-
Within a project, the right memory reaches the top five 84–93% of the time and is first 33–53%, depending entirely on how the question is phrased. The three query registers are reported separately because averaging them describes none of them.
-
Zero hits from another project and zero junk, in every register and arm, against a shared pool's 130–147 of 183. It follows by construction — the filter runs before the engine, so an inapplicable memory is never a candidate — which is what makes it a guarantee rather than a ranking that usually behaves.
-
80 of 80 planted secrets quarantined, zero contaminated blobs reaching git history, zero clean memories held back.
-
The behavioural layer verified inside live sessions rather than only in tests.
Full method, every number, and what the measurements do not establish: RECORD.md. The harness is public at memory-stack-lab.
-
-
Established, and bad: it cannot decline. Asked something the store has no memory of, it returns five confident memories that are indistinguishable — on every axis a caller can observe — from a real answer. Structural, measured, unfixed. Every ranking figure above is therefore conditional on the question having an answer.
-
Not established: it has not been used in anger on real work over time. There is no episodic tier — tested, and deliberately not built. Consolidation proposes but never writes, by design. Defects in this codebase have been found by benchmarking and by cross-platform CI rather than by anything failing in use; assume there are more.
-
Measured at scale, and the honest version is narrower than it sounds: writing 3,600 memories across 180 projects costs a flat 3.8 ms each. That run scopes every memory to its own project, so the 20 documents a caller sees measures write cost rather than what reach does to a view. Counting the view properly — at this corpus's 31% org-wide sharing rate — a caller at 64 projects searches 411 of 1,280 documents. Declared reach gives a field roughly three times smaller than a shared pool, not a constant one, and composing that with the retrieval curve puts 64 projects well into the degraded region. See RECORD.md.
| RECORD.md | how it was built, what was measured, why this shape. Start here to evaluate it |
| ARCHITECTURE.md | how it works — the write path, the recall path, the sync path |
| PROVENANCE.md | which component came from which prior system, and what is new |
MIT.