Repository navigation
Measure hosted memory's per-turn cost and stability: ingest paths, pack budget, prompt cache, PII, soak #7023
Description
Activity
- addedtaskWork item that is not primarily a bug or a feature.Work item that is not primarily a bug or a feature.priority: p2Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.Soon. Real but survivable — a rough edge, a gap, a thing that will bite later.
on Oct 6, 2026 Live ingest paths and re-ingest on production, 2026-10-06
The real
openhuman-corefrommain, headless in a throwawayHOME, bound to hosted memory (tinyhumans, status ok) with a memory-scoped test key. Every write was forgotten afterwards; a follow-up listing confirmed 0 stray documents. Core logs contained 0 key occurrences.Path Result Re-ingest / re-sync Learning ( memory_learn)✅ stored ✅ the same text again returns the same id Brain text ( memory_brain_ingesttext)✅ filed under markdown✅ the second call returns replayed: true, same idBrain file: markdown ✅ listed in 15 s, text extracted n/a Brain file: PDF ❌ refused: "the native converter does not handle pdf" → fixed in #7034 (office converter: PDF/DOCX/PPTX/XLSX) n/a Brain file: image (PNG) ❌ refused: no converter for images. Open; CortexDB has server-side image extraction ( CORTEX_IMAGE_*) as a possible routen/a Folder source ⚠️ both files read, second write lost toHTTP 502 Bad Gatewayfrommemory/events~65 s in (see tinyhumansai/cortexdb-saas#14)✅ no duplicates File source ✅ 1 document ✅ no duplicates Link source ✅ 1 document (64 s) ✅ no duplicates RSS source ✅ a 3-item text feed: all read, 2 visible within 150 s, the third still syncing (hosted latency). An image-only feed (xkcd) correctly stores nothing: "has no text to store" n/a Pinned by tests in #7034 (in-memory reference engine; no behaviour change needed):
- Prompt-cache stability: turn two's request reuses all of turn one's messages as the cached prefix, and turn one's pack never reaches turn two. On hoisting models (DeepSeek, native Anthropic), only the previous tail message drops out of the cache.
- Pack budget: with 1,500 learnings, a turn pack stays within
budget_tokensand repeats nothing.
Bug found and fixed upstream: a turn resumed after compaction pasted
start_session's pack besidepre_turn's. That is two budgets, with every learning injected twice (19 repeated lines in a test). The fix is tinyhumansai/tinymemory#206 (PreTurn::resumed, one pack); the OpenHuman side follows once it is pinned.Still open here: cost per turn, PII-redaction recall, soak, conversation-turn and v1-import ingest live. Most need recall latency fixed first (cortexdb-saas#14; likely tinyhumansai/tinymemory#204).
Triage: valid — Current review confirms this is a concrete, in-scope engineering issue or request: “Measure hosted memory's per-turn cost and stability: ingest paths, pack budget, prompt cache, PII, soak”. No complete resolution is evident in current main.
- addedtriage: validReviewed and confirmed as a valid actionable issueReviewed and confirmed as a valid actionable issue
on Oct 9, 2026 Triage: needs opinion — “Measure hosted memory's per-turn cost and stability: ingest paths, pack budget, prompt cache, PII, soak” needs a maintainer decision on product scope/priority or fresh reproduction evidence before its status can be settled. Please advise whether to pursue, narrow, or close it.
- addedtriage: needs-opinionValid issue awaiting a maintainer product or priority decisionValid issue awaiting a maintainer product or priority decisionand removedtriage: needs-opinionValid issue awaiting a maintainer product or priority decisionValid issue awaiting a maintainer product or priority decision
on Oct 9, 2026 Triage correction: valid — Current review confirms “Measure hosted memory's per-turn cost and stability: ingest paths, pack budget, prompt cache, PII, soak” is an actionable, in-scope gap; no complete resolution is evident in current main.
- addedsource: teamIssue opened by a repository memberIssue opened by a repository member
on Oct 9, 2026
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsTodo
Summary
Measure what hosted memory costs each turn, and prove it stays stable: every ingest path, re-ingest without duplicates, the per-turn pack's budget and prompt-cache behaviour, the effect of PII redaction on recall, and a soak run. Record the numbers on this issue.
Problem / Context
Split out of #6718 (Phase 3). Memory v2 (#6949, #6993) changed the shape of several of these criteria:
context.mdis gone. Each turn now recalls a token-budgeted memory pack ([memory.recall]).MemoryPackMiddlewareadds it to every model request of the turn withpush_ephemeral_instruction, as a request-only tail, and it is never persisted. On models that hoist system turns (DeepSeek, native Anthropic) it rides the tail user or tool message instead (openhuman: mid-turn system messages (stored-results list, validation and no-progress nudges) reset DeepSeek's prompt cache to the static prefix #6962). The design keeps the cached prefix intact, but nobody has checked that from real request bodies. This is the cost risk Stabilize the CortexDB-based memory engine in OpenHuman: validate hosted recall, cut over from tinycortex, then prove imports, context, efficiency and production readiness #6718 raised, the same class as Agent answers the previous message: thread history missing from the model request (clean install, 0.63.33) #6655, and it bears on Close the measured benchmark performance gaps: cache hit 80%→92%, cold start 5.6s→<1s, no CPU/RAM regression #7007's 92% cache-hit target.WaitFor::Visible. fix(memory): say beliefs are still building instead of showing an empty memory (#6718) #7014 covers what the UI says while beliefs are still being built.memory::guard::ScrubbingEngine. Redaction changes what gets stored, so its cost to recall needs measuring, not assuming. Stabilize the CortexDB-based memory engine in OpenHuman: validate hosted recall, cut over from tinycortex, then prove imports, context, efficiency and production readiness #6718 sawpii_redactions=8on a single insert.Scope
Acceptance criteria
memory_brain_ingest, including PDF and image conversion), each synced source kind (folder, file, link, GitHub, RSS, Composio), conversation turns, learnings, and the v1 import (Free memory migration from local TinyCortex storage to hosted CortexDB for current subscribers #7005 / fix(memory): stop a v1 import on credits or outage instead of skipping every item #7011).[memory.recall]'s token budget. The measured size is recorded.Related
docs/specs/memory-v2.md