feat(catalog): ARPA Skillware Mnemonic Matrix & Empirical Benchmark Suite - #10
Merged
rosspeili merged 2 commits intoSep 19, 2026
Merged
Conversation
… memories, and empirical benchmark - Persona skillware_operator: Invariant system prefix enforcing four-stage execution discipline (resolve, draft, preview, confirm), candidate disambiguation, and calm error de-escalation - 4 Curated skillware/ memories: interactive_slot_filling (work), entity_disambiguation (work), irreversible_action_crucible (incident), runtime_outage_and_grace (relational) - Lineage skillware_execution_mastery: 4-epoch progression across slot gathering, registry disambiguation, financial custody, and outage resilience - Flagship Integration Guide: docs/integrations/skillware_matrix.md with Motor Cortex vs Epistemic Cortex architecture, benchmark results, and code recipes - Simulation Suite: scripts/simulate_skillware_matrix.py evaluating 5 multi-turn scenarios across Directives, Brief, Full Matrix, and JIT Chunks - Tests & Validation: tests/test_skillware_bench.py and test_taxonomy_and_chunks.py with 112 passing tests, flake8, black, and repo integrity gate
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Integrates the Skillware Mnemonic Matrix, resolving the context dilemma when equipping autonomous agents with ARPA Skillware (v0.5.5):
instructions.mdper skill (6,000–15,000+ tokens in multi-skill chains), inflating costs and diluting model attention.confirmed: true), truncate zero-padded IDs (08472910->8472910), and crash on HTTP 429 rate limits.By pairing Skillware's lean
briefsignatures with MnemoLink's cached persona prefix and targeted JIT memory chunks, agents achieve 100% safety compliance while cutting prompt bloat by 50% to 84%.Key Changes
skillware_operator): Invariant prefix enforcing the four-stage discipline (resolve, draft, preview, confirm), candidate disambiguation, and calm error handling.skillware/):interactive_slot_filling(work): Conversational slot gathering, preview rendering, and confirmation gates.entity_disambiguation(work): Candidate formatting and leading-zero preservation for registry identifiers.irreversible_action_crucible(incident): The $48.2k burn postmortem enforcing EIP-55 checksums, slippage caps, and two-phase simulation.runtime_outage_and_grace(relational): Calmly shields raw HTTP 429/503 stack traces with polite backoff and user reassurance.skillware_execution_mastery): 4-epoch progression linking slot-filling through registry handling, financial custody, and outage resilience.docs/integrations/skillware_matrix.md): Architecture overview (Motor vs. Epistemic Cortex), recipe code samples, and full scenario economics.scripts/simulate_skillware_matrix.py): 5-scenario evaluation suite testing Directives, Brief, Full Matrix, and JIT Chunks.Empirical Benchmark Results
Prompt caching on the invariant persona prefix cuts multi-turn inference costs by over 95%.
Verification & Standards Gate
pytest tests/: 112 passed in 5.00s (100% offline CI isolation, zero network requests).flake8 mnemolink tests scripts: 0 errors, 0 warnings (≤100 char limit).black --check mnemolink tests scripts: Clean across all 30 files.python scripts/verify_repo.py: Zero-emoji gate passed; 43/43 markdown link paths resolved; all manifests/cards validated against Pydantic models.