Status: Accepted · Builds on: memory-v2.md · Plan: ../plans/agent-memory.md
A host such as OpenHuman calls memory at fixed points of every agent turn:
- before the model runs, to inject context;
- after the model replies, to log the reply;
- when a session starts or resumes;
- when the prompt is truncated (compaction);
- in the background, to distil beliefs.
The v2 contract offers the primitives (store, fetch, recall, namespaces) but no shared shape for these moments. Each host would invent its own scope tree and its own latency rules, and moving to another engine would mean rewriting every one of those moments.
The host also needs a company brain: documents (PDF, markdown, Notion, GitHub) shared by every agent, kept apart by source type, and carrying no agent id.
- One standard layout for the brain, agent conversations and learnings, on top of the existing namespace model.
- One read primitive. A holistic recall across several scopes, rendered
as one token-budgeted block.
context.md, session start, pre-turn and compaction are all presets of it. - A fast live turn. The turn logs and recalls without waiting for indexing and without running a model.
- Explicit background work. Belief builds are values the host schedules, never hidden threads.
- Engine-agnostic. Everything is written against
MemoryEngine, and the conformance suite pins down every new engine obligation. Swapping engines changes only the engine the host constructs.
- Running a model inside this library. The host owns generation.
- Spawning tasks or owning a runtime.
- Thread storage beyond memory.
tinyagents-sessionowns session threads.
The layout sits below one root node. That root is core: the root namespace
by default, or a host node such as team:acme.
| Plan scope | Namespace | Kind |
|---|---|---|
core |
root | everything (holistic) |
core/brain/<source> |
root/source:<source> |
documents |
core/conversations/<agent> |
root/agent:<agent> |
conversations |
core/learnings |
root (shared) and any node (built beliefs) | learnings |
sourceis a new namespace segment kind (SegmentKind::Source). CortexDB acceptssourceas a scope type, so a brain source is a real scope:app:tinymemory/source:files/app:documents.BrainSourcenames the connector a document came from:files(local files of any format),web,notion,github, or any other id (a connected app's slug,gmail).pdfandmarkdownname the per-format nodes used beforefiles. A large source may be split intoproject:<collection>nodes below it (MemoryLayout::brain_collection); forgetting or rebuilding the source covers them.- A brain document carries no agent id. Its namespace is always its source's node, whatever metadata the caller passes.
- Each turn is stored as its own one-turn conversation item at the agent's
node. The item carries
thread_id,turns {first,last},agent_id, and theconversationsource keyed by thread. A retried turn with the same input is therefore a replay.
MemoryEngine::store_with(item, WriteOptions { wait })WaitFor::Visibleisstore.WaitFor::Acceptedmay return once the engine has durably accepted the item.- The default implementation serves both as
store.
MemoryEngine::store_many_with(items, WriteOptions { wait })is the same for a batch:WaitFor::Visibleisstore_many,WaitFor::Acceptedmay return once every item is durably accepted. The default serves both asstore_many.MemoryEngine::consolidate(ConsolidateRequest { reach, kinds })asks for a belief build and returns as soon as the job is taken.- The default refuses with
Unsupported. - The receipt status is one of
Started,ScheduledorCompleted, and the receipt names any job handles, the number of scopes covered and, for a completed build that reports it, the beliefsbuilt.
- The default refuses with
EngineDescriptor::consolidationdeclares how the engine consolidates:None,OnDemand,ScheduledorAutomatic(it rebuilds beliefs on its own after writes, and an explicitconsolidatestill builds at once).MemoryEngine::beliefs(BeliefsRequest { reach, query, limit })reads the beliefs an engine built and keeps apart from its stored items. With a query they are ranked for it; without one, the most confident come first, then the newest.- Each is a
Learninghit taggedBELIEF_TAG("belief") at the node of its sources. - A belief is not a stored item: it cannot be listed, fetched or forgotten by id. Forgetting its sources removes it.
- The default holds none. An engine whose beliefs are ordinary learning items (the reference engine) keeps it.
- Each is a
FetchRequest::beliefs(default0) asks a fetch for up to that many beliefs from what it read, returned inFetchPage::beliefs(first page only). An engine that keeps no beliefs apart returns none.- Conformance adds two checks:
store_with: a visible store is listed on return and replays; an accepted store answers with the item's own id.consolidate: a malformed request is refused, and the answer matches what the descriptor promises.beliefsrefuses a zero limit, and every belief it returns is a tagged learning within the reach asked for.
ReferenceEngineconsolidates on demand and deterministically: oneFactper document or conversation, taggedconsolidated.
Adding SegmentKind::Source and the descriptor field is a breaking change to
exhaustive matches and struct literals, so it ships in a major release.
HolisticRecall { query, sections, budget_tokens, title, exclude_ids, exclude_thread } produces a ContextPack { markdown, tokens, refs, sections, skipped, engine }.
- Each section is
ScopeSection { heading, filter, limit, query }, filled in one of three ways:Fetch: ranked retrieval, with no model.Answer: a synthesised answer, optionally falling back to fetch.Latest: newest first, then most confident, then latest turn.
- Reads run concurrently. A failing or empty section is reported in
skippedand never fails the pack. The only error is an invalid request. - Deduplication. An item is listed once, in its first (highest-priority) section. An answer citing an item does not hide it.
- Beliefs are learnings. When a pack has a learnings section (a fetched
or latest section that admits learnings), it gathers the engine's beliefs
into it:
- with a query, every fetched section asks its fetch for beliefs too
(
FetchRequest::beliefs), so they come from reads the pack makes anyway; - without one, the learnings section lists them (
beliefswith no query); - what all sections returned is merged, each belief once, and interleaved with the stored learnings rank by rank, stored learnings first.
- The section's filter applies to beliefs too.
- A failed belief read leaves the section to its stored learnings.
- Answered sections ask for none, because the engine's answer draws on its beliefs itself.
- So on CortexDB, what a build produced reaches every pack's Learnings
section: in
pre_turn,start_session, compaction andcontext.md.
- with a query, every fetched section asks its fetch for beliefs too
(
- Exclusions:
exclude_idsdrops named items, such as the turn just logged.exclude_thread { thread_id, from_turn }drops the turns still in the prompt.
- Dates. A conversation bullet whose turn carries a time
(
meta.observed_at, set fromPreTurn::atandPostTurn::at) is led by it,[YYYY-MM-DD HH:MM]. Hosts should pass the time the message was sent: without it, a superseded value and its correction look alike. - Budget. The block fits
budget_tokensat four characters per token. Bullets are trimmed from the last section first, then the last answer shortens. context.mdis this with fixed sections: one answered section per brief, then the latest learnings, under# Context, with frontmatter. Its output is unchanged.
| Call | Writes | Reads |
|---|---|---|
start_session { thread_id?, focus? } |
— | the resumed thread's latest turns, then the standard sections |
pre_turn { thread_id, turn_index, user_text, in_prompt_from, at? } |
the user turn, Accepted, run concurrently with the read |
the standard sections fetched for user_text, without this turn or the prompt's window |
post_turn { thread_id, turn_index, assistant_text, tool_calls, at? } |
the reply, Accepted |
— |
recall_for_compaction { thread_id, dropped, focus? } |
— | an answered summary of the thread (falling back to fetch), then the standard sections for the focus or the dropped turns' gist |
recall(query) |
— | the pre-turn read without logging |
run_background(job) |
the job's | — |
- Standard sections, in priority order: Learnings (the whole tree), one
section per core scope (core-scopes.md; none by default),
Brain (all documents, reading at most
BRAIN_SCOPES_PER_TURN(4) scopes: those the query names, then the most recently written), this agent's history, and team conversations (every agent; none in a layout that pools conversations, where they are the history's node). A zero limit inRecallPolicyleaves a section out. pre_turnnever fails on an engine error. A failed log is reported inTurnContext::log_errorand the pack is still returned.post_turnreports belief builds. It returns aBuildBeliefsjob for the agent's conversations everyRecallPolicy::build_beliefs_everyturns, counted asturn_index + 1, unless the engine declaresAutomatic; then it returns none, andhistory_buildstill asks for one.
Brainstores documents:ingestandingest_with(WaitFor)store aBrainDocumentat its source's node. The source kind defaults from theBrainSource.ingest_manybatches byMAX_STORE_MANY.- Each ingest returns the
BuildBeliefsjob for its source scope, andingest_manyone per source touched; on an engine that declaresAutomaticneither returns any (Ingested::jobisNone,jobsis empty), andBrain::buildasks for one explicitly.
Brain::searchfetches within one source or across the whole brain, andBrain::forgeterases one source.BackgroundJobisBuildBeliefs { request }orIngestBrain { documents }. It is serializable so a host can queue it.BackgroundRunner::runmapsUnsupportedtoJobOutcome::Skipped, so the same host code runs on any engine.
brain::brain_document(converter, raw, source?, meta)turns a file into aBrainDocument. Its source is the one the caller names, or the one the detected format implies.- CortexDB, Direct wire:
store_with(Accepted)writes without?wait=indexedand skips the visibility waits.consolidateposts{ "scope": … }tov1/beliefs/buildonce per scope that is held in reach and admitted. The server builds within the request, so the receipt isCompletedwith the beliefsbuilt; an answer that names a job instead makes itStartedwith the handles.- It declares
Automaticwhen its endpoint is CortexDB's managed API, which rebuilds beliefs on its own after writes, andOnDemandanywhere else.EngineSettings::consolidationoverrides that.
- CortexDB, TinyHumans wire: declares
Scheduledand sends nothing.
- The live turn (
pre_turn,post_turn) never waits for indexing and never runs a model. - A pack never contains the turn being logged, nor any turn of its thread at
or after
in_prompt_from. - A brain document never carries an agent id and always lives at its source's node.
- Siblings stay invisible except through the holistic reach. The layout reads the subtree of its root deliberately, because the brain and the learnings are shared.
- No lifecycle call spawns work. Every slow step is a returned
BackgroundJob.
- The conformance suite passes against the reference engine and both
CortexDB doubles. Fault injections for
store_withandconsolidatetrip their checks. - The existing
context.mdtests pass unchanged after the rebase onto holistic recall. - The same agent loop passes over both CortexDB doubles. On the Direct wire,
turns log without
wait=indexedand builds hitv1/beliefs/buildper scope. cargo run -p tinymemory-tools --example agent_loopshows the lifecycle offline.examples/cortex_agentandtests/live_cortex_lifecycle.rsrun against the harness.
- Whether the hosted TinyHumans backend will expose a belief-build route.
When it does, only the hosted descriptor and
cortex/engine/consolidate.rschange.