Skip to content

Replace Memento or IExtensionState for InstallTrackerSingleton with Custom Atomic State #2794

Description

@nagilson

Replace Memento-backed install tracking with transactional extension-owned storage

Summary

Replace direct use of VS Code Memento for install-tracker state with an extension-owned storage abstraction that supports fresh reads and atomic, durable read-modify-write transactions across VS Code extension hosts.

This work should be delivered in small, independently reviewable phases. The first phase must introduce only an abstraction over the existing behavior. A separate design phase must select and document the persistent format before the production backend changes.

Related issues:

Problem

InstallTrackerSingleton currently stores install records and live-session records under separate keys in ExtensionContext.globalState:

  • installed
  • dotnet.returnedInstallDirectories

These look like independent keys through the Memento API, but VS Code persists and broadcasts an extension's Memento as one state object. Memento.update() optimistically changes the extension host's cached copy, sends the whole state object to VS Code's storage service, and completes without providing a revision, compare-and-swap operation, fresh persisted read, or operation-correlated change event. A later whole-state notification can replace the extension host's cache with a snapshot created before another key was changed.

The mutexes in this repository serialize our extension code, but they do not serialize VS Code's internal replacement of the Memento cache. Consequently, acquiring installedLk and then calling globalState.get('installed') does not guarantee a fresh value from the persisted source of truth.

Issue #2792 demonstrates the impact on Windows and macOS x64:

  1. Automatic update successfully deletes outdated runtime directories.
  2. The first outdated install is successfully removed from installed and is absent in subsequent tracker reads.
  3. Cleanup of dead entries in dotnet.returnedInstallDirectories causes an older whole-Memento snapshot to become visible later.
  4. A subsequent uninstall reads that snapshot and persists one physically deleted runtime as installed.
  5. .NET Install Tool: Uninstall .NET then displays the stale record.

This is not an lsof or file-deletion failure. The attached logs report successful uninstalls, and the stale record appears only in tracker state.

Goals

  • Provide linearizable read-modify-write operations for install-tracker state while the repository's cross-process mutex is held.
  • Read the current source of truth after acquiring the mutex rather than relying on an extension-host cache.
  • Preserve install ownership and live-session behavior across VS Code windows and extension hosts.
  • Survive process termination without leaving a partially written state as authoritative.
  • Recover from an incomplete write or malformed latest record using a previous verified state.
  • Keep startup and common install lookup latency small.
  • Migrate existing users without reinstalling or orphaning managed .NET installations.
  • Keep machine-specific state from flowing to another machine, profile, remote host, container, or WSL environment.
  • Retain a testable storage interface so unit tests do not depend directly on VS Code.

Non-goals

  • Building a general database for arbitrary extension features.
  • Synchronizing managed installations between machines.
  • Using Settings Sync to copy installation or session records.
  • Depending on private VS Code storage APIs or monkey-patching ExtensionMemento.
  • Treating sleeps, repeated Memento.get() calls, or polling as a correctness mechanism.

Constraints

  • The state is machine- and extension-host-specific. Existing globalState.setKeysForSync([]) behavior must remain effective during migration.
  • Multiple VS Code windows may run the extension concurrently.
  • Local, remote, container, and WSL extension hosts may share a UI but must not accidentally share machine-local install records.
  • Writes must remain inside the extension-owned globalStorageUri directory.
  • Temporary files must be created in the same directory as their destination so atomic replacement does not cross file systems.
  • The implementation must work on Windows, macOS, and Linux and account for each platform's rename/replace behavior.
  • State files are data only. No stored value may be interpreted as executable code or an arbitrary path to write outside extension storage.

Proposed Phases

Each phase should normally be a separate commit or PR-sized commit series. Do not combine the abstraction, storage design, backend implementation, migration, and production cutover into one change.

Phase 1: Introduce a storage abstraction without changing behavior

  • Replace the IExtensionState = Memento type alias with an interface owned by this repository that preserves the required Memento-compatible surface:
    • get(key)
    • get(key, defaultValue)
    • update(key, value)
    • keys() where required
  • Add a production adapter that delegates directly to the existing VS Code Memento.
  • Pass one adapter instance through the extension and acquisition contexts instead of passing raw globalState throughout the library.
  • Keep the existing mock implementation behind the same interface.
  • Do not change persistence behavior, key names, migration, or tracker logic in this phase.
  • Compile the library, runtime extension, SDK extension, and sample.
  • Run the existing library unit and extension functional tests and confirm no behavioral regressions.

The abstraction should support substituting a file-backed implementation later without changing tracker consumers again. Avoid exposing VS Code-specific APIs that the future backend cannot implement honestly.

Phase 2: Select and document the filesystem design

  • Add a dedicated design document under Documentation/.
  • Prototype and compare at least the two designs below.
  • Record the selected consistency, durability, recovery, retention, migration, and performance guarantees.
  • Define the on-disk schema and versioning policy before implementing the production backend.

Option A: Append-only ledger / write-ahead log

Each update(key, value) appends a versioned transaction record instead of rewriting all state.

Questions the design must answer:

  • What constitutes a complete committed record after a crash?
  • Is the ledger itself the write-ahead log, or are intent and commit records separate?
  • Are checksums, lengths, operation IDs, and monotonically increasing revisions required?
  • How are truncated or corrupt tail records detected and ignored?
  • How are operations replayed and re-verified?
  • When is a compacted snapshot written, and how is compaction made atomic?
  • How large can the ledger become before startup or lookup performance degrades?
  • Can independent-key records preserve an invariant spanning installs and sessions, or must related changes share one transaction?

Potential benefits:

  • Small writes.
  • Complete operation history for recovery and diagnostics.
  • A corrupt or incomplete tail can be discarded without losing earlier committed operations.

Potential costs:

  • Replay and compaction complexity.
  • Harder schema evolution.
  • Append durability and multi-record transaction boundaries require careful platform-specific handling.

Option B: Full-state transaction snapshots

Each transaction writes the complete new tracker state. A minimal schema could be:

{
  "schemaVersion": 2,
  "revision": 42,
  "installed": [],
  "sessions": {}
}

Possible layouts include:

install-tracker-v2.json
install-tracker-v2.previous.json

or immutable generations:

install-tracker-v2/
  CURRENT
  state-0000000041-<operation-id>.json
  state-0000000042-<operation-id>.json

A transaction would:

  1. Acquire installedLk.
  2. Read and validate the latest committed state fresh from disk.
  3. Apply the mutation in memory.
  4. Write the complete new state to a uniquely named temporary file in the same directory.
  5. Flush and close the temporary file according to the selected durability guarantee.
  6. Atomically publish the new state using rename/replace or a committed generation pointer.
  7. Retain at least one previously verified state for recovery.
  8. Release installedLk.

Questions the design must answer:

  • Is a single replaced file sufficiently portable, or should immutable generations plus CURRENT be used?
  • How is CURRENT recovered if it is missing, stale, or corrupt?
  • Should readers scan for the highest valid revision as a fallback?
  • How many previous generations, or how much history, should be retained?
  • Is retention count-based, age-based, or both?
  • Are file and parent-directory flushes required for the promised durability level?
  • What is the maximum expected state size and measured read/write latency?

Potential benefits:

  • Fast reads of one complete state.
  • Simple invariant validation.
  • Straightforward schema versioning and rollback to the previous valid generation.

Potential costs:

  • More bytes written per update.
  • Atomic replacement and durability behavior must be validated on every supported platform.

Design evaluation criteria

The design document should compare the options using measured or explicitly reasoned results for:

  • Atomicity across keys and related invariants.
  • Cross-process freshness after acquiring installedLk.
  • Crash behavior before, during, and after publication.
  • Recovery from truncated, malformed, or checksum-invalid data.
  • Startup latency with realistic and worst-case state sizes.
  • Write latency and frequency during acquisition, update, and shutdown.
  • Compaction and retention complexity.
  • Schema migration and rollback behavior.
  • Testability and observability.
  • Windows, macOS, Linux, remote, WSL, container, and browser constraints.

Phase 3: Implement the selected backend behind the abstraction

  • Implement the selected filesystem store without making it the production default.
  • Keep all file access under globalStorageUri.
  • Require callers to hold the appropriate cross-process mutex for transactional mutations, or have the store acquire it consistently in one ownership layer.
  • Ensure a transaction reads state only after acquiring the mutex.
  • Validate schema, revision, and record invariants before returning state.
  • Add fault injection for write, flush, close, rename, replace, cleanup, and malformed-input failures.
  • Add telemetry/logging that identifies backend, schema version, recovery path, and operation ID without logging sensitive paths or state contents unnecessarily.
  • Keep the Memento adapter as the production implementation during this phase.

Phase 4: Audit and validate the backend independently

  • Conduct an independent code review focused on data loss, stale reads, lock ownership, deadlocks, path traversal, symlinks/reparse points, and crash recovery.
  • Run differential tests against an in-memory reference model.
  • Test abrupt termination at every transaction boundary.
  • Test concurrent writers from separate Node processes using the real IPC mutex.
  • Test malformed current state with a valid previous generation.
  • Test retention/compaction while readers and writers contend.
  • Measure startup and mutation performance with realistic and oversized state.
  • Do not replace the production Memento adapter until this audit is complete.

Phase 5: Wire the tracker to the new backend behind a controlled switch

  • Replace tracker use of the Memento adapter with the new backend in tests and an opt-in development configuration first.
  • Route install records and session records through the same transactional store when they participate in the same consistency domain.
  • Remove direct globalState.update() calls for migrated tracker keys.
  • Make getExistingInstalls() a read unless it is explicitly performing a migration or repair transaction.
  • Preserve a rollback path to the Memento backend during pre-release validation.
  • Do not enable the new backend by default until migration is implemented and tested.

Phase 6: Exercise the reported and cross-environment scenarios

  • Add an automated reproduction for [NETE2ESDK] Only the latest version should be remain in Uninstall .NET list after the older runtimes are removed by Update of Installs #2792 using a real VS Code Memento to demonstrate the old failure.
  • Run the same scenario against the new store and confirm no stale install record returns.
  • Cover one and multiple outdated runtime records.
  • Cover runtime and ASP.NET Core install groups.
  • Cover dead prior sessions and a currently live session.
  • Cover two VS Code windows updating tracker state concurrently.
  • Cover graceful shutdown, forced termination, and restart during a transaction.
  • Test Windows x64, macOS Intel, macOS ARM64, Ubuntu, and RHEL.
  • Test local Windows versus WSL and confirm neither reads the other's machine-local tracker state.
  • Test local versus SSH/dev-container extension hosts.
  • Test multiple VS Code profiles and user-data directories.
  • Verify state is not synchronized to another machine through Settings Sync.
  • Evaluate vscode.dev and github.dev: verify whether this extension can run there, identify the available filesystem implementation, and either support the backend or document and test the intentional unsupported behavior.
  • Attach sanitized logs and a concise state-transition timeline to the implementing PR or [NETE2ESDK] Only the latest version should be remain in Uninstall .NET list after the older runtimes are removed by Update of Installs #2792.
  • Preserve useful test artifacts for failures without collecting credentials, tokens, or unrelated user state.

The existing unit-test MockExtensionContext mutates an in-memory object synchronously and cannot reproduce VS Code's Memento replacement behavior. At least one functional test must run in a real extension host.

Phase 7: Add best-effort migration and enable production cutover

  • In InstallTrackerSingleton, detect existing legacy keys when the new store has not yet been initialized.
  • Acquire installedLk before migration.
  • Read and normalize the existing installed and dotnet.returnedInstallDirectories values once.
  • Validate local install records against expected extension-owned locations and existing executables before importing them.
  • Import both keys into one new schema transaction.
  • Mark migration complete in the new store, not by relying solely on another Memento update.
  • Make migration idempotent so interruption can safely retry.
  • If the new store already has a valid committed state, never overwrite it from stale Memento data.
  • If migration fails, leave the old keys untouched and continue with an explicit, observable fallback rather than silently losing records.
  • Initially leave old Memento keys in place but dormant. Define a later cleanup version only after rollback compatibility is no longer needed.
  • Enable the new backend by default only after migration and rollback tests pass.

Migration must account for legacy string records, installKey records, malformed records, missing owner arrays, local versus global installs, architectures, and install modes. It must not cause an elevation prompt or reinstall solely because state moved between backends.

Acceptance Criteria

  • After acquiring installedLk, the tracker reads the latest committed state from the authoritative store.
  • Install and session mutations cannot restore an older committed install list.
  • Two extension hosts cannot both commit from the same stale base without one observing the other's committed state first.
  • A process crash cannot make a partially written transaction authoritative.
  • The latest malformed state falls back to a previous valid state with an observable diagnostic.
  • [NETE2ESDK] Only the latest version should be remain in Uninstall .NET list after the older runtimes are removed by Update of Installs #2792 no longer reproduces on Windows or macOS Intel.
  • Existing users retain all valid managed install records after migration.
  • Machine-specific records do not flow between physical machines, Windows and WSL, or local and remote extension hosts.
  • Existing compile, unit, functional, lint, and packaging validation passes.
  • The selected filesystem format and its guarantees are documented.
  • The implementing PR includes the reproduction method, relevant sanitized logs, and test artifacts or links to retained CI artifacts.

Open Questions

  • Should the first production change move only session state to files, or move the complete tracker state at once?
  • Does the selected backend own mutex acquisition, or must callers prove they already hold it?
  • Is one composite state transaction required for every operation, or are some keys safe to store independently?
  • What durability guarantee is required: process-crash consistency, OS-crash consistency, or power-loss durability?
  • Which atomic replace strategy is reliable across all supported Node.js and operating-system versions?
  • What retention policy balances recovery value, privacy, disk usage, and startup cost?
  • Should physically missing local installs be removed as a repair transaction during reads, or only by an explicit reconciliation operation?
  • What behavior is required when extension storage is read-only or unavailable?
  • How long must rollback to the old Memento backend remain supported?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P1bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions