Skip to content

feat(catalog): ARPA Skillware Mnemonic Matrix & Empirical Benchmark Suite - #10

Merged
rosspeili merged 2 commits into
ARPAHLS:mainfrom
rosspeili:feat/skillware-mnemonic-matrix-and-bench
Sep 19, 2026
Merged

rosspeili merged 2 commits into
ARPAHLS:mainfrom
rosspeili:feat/skillware-mnemonic-matrix-and-bench

Conversation

@rosspeili

Copy link
Copy Markdown
Contributor

Overview

Integrates the Skillware Mnemonic Matrix, resolving the context dilemma when equipping autonomous agents with ARPA Skillware (v0.5.5):

  • Directives Mode: Concatenates full instructions.md per skill (6,000–15,000+ tokens in multi-skill chains), inflating costs and diluting model attention.
  • Brief Mode: Strips edge-case safeguards (~25 tokens/skill), causing models to omit confirmation gates (confirmed: true), truncate zero-padded IDs (08472910 -> 8472910), and crash on HTTP 429 rate limits.

By pairing Skillware's lean brief signatures with MnemoLink's cached persona prefix and targeted JIT memory chunks, agents achieve 100% safety compliance while cutting prompt bloat by 50% to 84%.


Key Changes

  1. Persona (skillware_operator): Invariant prefix enforcing the four-stage discipline (resolve, draft, preview, confirm), candidate disambiguation, and calm error handling.
  2. 4 Curated Memories (skillware/):
    • interactive_slot_filling (work): Conversational slot gathering, preview rendering, and confirmation gates.
    • entity_disambiguation (work): Candidate formatting and leading-zero preservation for registry identifiers.
    • irreversible_action_crucible (incident): The $48.2k burn postmortem enforcing EIP-55 checksums, slippage caps, and two-phase simulation.
    • runtime_outage_and_grace (relational): Calmly shields raw HTTP 429/503 stack traces with polite backoff and user reassurance.
  3. Lineage (skillware_execution_mastery): 4-epoch progression linking slot-filling through registry handling, financial custody, and outage resilience.
  4. Flagship Integration Guide (docs/integrations/skillware_matrix.md): Architecture overview (Motor vs. Epistemic Cortex), recipe code samples, and full scenario economics.
  5. Empirical Benchmark (scripts/simulate_skillware_matrix.py): 5-scenario evaluation suite testing Directives, Brief, Full Matrix, and JIT Chunks.

Empirical Benchmark Results

Scenario Directives Brief JIT Chunks Bloat Cut Safety Score Cost / 1k Turns (Dir) Cost / 1k Turns (JIT)
1: Gmail Slot-Filling 1,775 tok 26 tok 569 tok 67.9% 100% $4.61 $0.29
2: Registry Disambiguation 3,509 tok 32 tok 562 tok 84.0% 100% $9.11 $0.29
3: DeFi Slippage & Custody 1,116 tok 27 tok 552 tok 50.5% 100% $2.90 $0.29
4: Upstream 429 Outage Grace 1,116 tok 27 tok 538 tok 51.8% 100% $2.90 $0.28
5: 3-Skill Chain (Mail+Reg+Opt) 6,005 tok 86 tok 1,312 tok 78.2% 100% $15.58 $0.68

Prompt caching on the invariant persona prefix cuts multi-turn inference costs by over 95%.


Verification & Standards Gate

  • pytest tests/: 112 passed in 5.00s (100% offline CI isolation, zero network requests).
  • flake8 mnemolink tests scripts: 0 errors, 0 warnings (≤100 char limit).
  • black --check mnemolink tests scripts: Clean across all 30 files.
  • python scripts/verify_repo.py: Zero-emoji gate passed; 43/43 markdown link paths resolved; all manifests/cards validated against Pydantic models.

… memories, and empirical benchmark

- Persona skillware_operator: Invariant system prefix enforcing four-stage execution discipline (resolve, draft, preview, confirm), candidate disambiguation, and calm error de-escalation
- 4 Curated skillware/ memories: interactive_slot_filling (work), entity_disambiguation (work), irreversible_action_crucible (incident), runtime_outage_and_grace (relational)
- Lineage skillware_execution_mastery: 4-epoch progression across slot gathering, registry disambiguation, financial custody, and outage resilience
- Flagship Integration Guide: docs/integrations/skillware_matrix.md with Motor Cortex vs Epistemic Cortex architecture, benchmark results, and code recipes
- Simulation Suite: scripts/simulate_skillware_matrix.py evaluating 5 multi-turn scenarios across Directives, Brief, Full Matrix, and JIT Chunks
- Tests & Validation: tests/test_skillware_bench.py and test_taxonomy_and_chunks.py with 112 passing tests, flake8, black, and repo integrity gate
@rosspeili
rosspeili merged commit 0a10582 into ARPAHLS:main Sep 19, 2026
5 checks passed
@rosspeili
rosspeili deleted the feat/skillware-mnemonic-matrix-and-bench branch September 19, 2026 18:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant