Skip to content

perf(store): store embeddings as float32 blobs and cap the wal file size - #523

Open
arabold wants to merge 2 commits into
mainfrom
perf/509-reduce-store-size
Open

arabold wants to merge 2 commits into
mainfrom
perf/509-reduce-store-size

Conversation

@arabold

@arabold arabold commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

What

Three changes that make the database smaller, prompted by #509:

  1. Embeddings are stored as float32 blobs instead of JSON text. documents.embedding held each vector as JSON, on top of the binary copy in documents_vec. New writes use a Float32Array buffer. Migration 018 converts existing rows with vec_f32. Read queries no longer select the column, since nothing in the code used it.
  2. The full-text update trigger only fires when the indexed columns change. These are content, metadata and page_id. Before, it fired on every update, so every embedding write deleted and re-inserted that chunk's full-text entry.
  3. The WAL file is trimmed after checkpoints. journal_size_limit = 64 MB makes SQLite cut the WAL back once a checkpoint resets it. Before, it kept its peak size until the last connection closed, which matches the 6.7 GB WAL in feat: Storage management in Web Dashboard, reclaimable space visibility, and post-scrape WAL cleanup #509.

Size of the effect

  • Embeddings: a 1536-dim vector takes about 20–32 KB as JSON and 6 KB as a blob. A scratch store with 30k embedded chunks went from 1,119 MB to 416 MB after conversion and VACUUM. This likely also explains most of feat: Storage management in Web Dashboard, reclaimable space visibility, and post-scrape WAL cleanup #509's 27 GB of free space: 301k chunks with JSON embeddings, later cleared by a model change, leave about that much behind.
  • WAL: in a test where a long read blocked checkpoints during a big write, the WAL reached 199 MB and stayed there. With the limit it went back to 64 MB after the next checkpoint. Normal use stays around 4 MB, so the limit rarely does anything.
  • Trigger: inserting chunks with embeddings is about 20% faster. A model change (UPDATE documents SET embedding = NULL) no longer rewrites the whole full-text index in one transaction.

Tested on a real database

A local store with 8 libraries and 46,535 chunks, all embedded with openai:text-embedding-3-small. The pre-migration state was rebuilt byte for byte: the OpenAI client returns float32 values as a plain array, so the old JSON is JSON.stringify(Array.from(Float32Array)).

File documents table Vector index Full-text index
Before 1,884 MB 1,430 MB 295 MB 156 MB
After this PR 842 MB (−55%) 389 MB 295 MB 156 MB
After this PR + #522 814 MB (−57%) 389 MB 295 MB 127 MB
  • Upgrade time: DocumentStore.initialize() (migration 018 plus the VACUUM) took 21 s.
  • Lossless: all 46,535 converted embeddings are byte-identical to the vectors written by the new code path.
  • Integrity: quick_check and the FTS5 integrity check pass.
  • Vector index: one row per embedded chunk. In a sample of 300, each stored embedding equals its indexed vector, and each chunk's nearest neighbour is itself at distance 0.
  • Search: a findByContent search through DocumentStore returns results.

Migration notes

  • The file is numbered 018 because feat!: library display names and overwrite protection #517 adds 017-add-library-display-name.sql.

  • The migration is safe to run again: it only converts rows that are still JSON, and it drops each trigger before creating it.

  • The migration first narrows the full-text trigger, so the conversion doesn't rewrite the whole full-text index.

  • It drops the vector triggers during the conversion, because documents_vec already holds the same vectors. Then it recreates them to pass the blob through; sqlite-vec accepts both blobs and JSON.

  • vec_f32 throws on malformed JSON, which would fail the migration. Every stored value already went through sqlite-vec's parser when it was written, via the triggers and the backfill in 011, so I didn't add a guard. A full guard made the conversion 2.5× slower.

  • First start after upgrading rewrites the documents table and then runs the usual post-migration VACUUM. The conversion took about 7 s per 30k embedded chunks locally. Large stores like the one in feat: Storage management in Web Dashboard, reclaimable space visibility, and post-scrape WAL cleanup #509 should expect a few minutes.

  • I checked that Buffer.from(Float32Array) produces the same bytes as vec_f32 on the JSON, so new and converted rows are identical.

Tests

  • DocumentStore.test.ts: embeddings written by addDocuments are 6,144-byte blobs equal to the vector in documents_vec. The test fails with the old JSON write path.
  • applyMigrations.test.ts:
    • a JSON embedding is converted to the same float32 values and is still found by KNN search
    • a content update still reindexes full text
    • the WAL size limit is set
  • Lint, typecheck and the full suite pass, except two E2E tests that start the server and timed out under full-suite load: telemetry-e2e in one run, the mcp-http-e2e heartbeat in another. Both pass on their own (9/9). The migration does nothing on the fresh databases these tests create.

Docs: updated the vector storage section and migration list in docs/concepts/data-storage.md, and added the WAL limit to the database-migrations spec.

Refs #509

Embeddings were written to documents.embedding as JSON text, 3-5x the size
of the binary vector and on top of the binary copy in documents_vec. They
are now written as float32 blobs, and migration 017 converts existing rows
with vec_f32. sqlite-vec accepts both forms, so the vector triggers pass
the blob through.

The full-text update trigger fired on every column, so each embedding
write rewrote that chunk's full-text entry. It now fires only for
content, metadata and page_id, the columns the index reads. The
migration narrows it before converting, so the conversion does not
rewrite the whole full-text index.

journal_size_limit is set to 64 MB so SQLite trims the WAL back after a
checkpoint instead of keeping its peak size until the last connection
closes.

Refs #509
Copilot AI balanced review requested due to automatic review settings October 4, 2026 23:24

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The storage, migration, trigger, configuration, test, and documentation changes are coherent with no unresolved correctness issues.

Review effort: Balanced
Findings: None

What changed in this PR

Optimizes SQLite storage by reducing embedding size, avoiding unnecessary FTS updates, and limiting retained WAL growth.

Changes:

  • Stores embeddings as float32 blobs and migrates existing JSON vectors.
  • Narrows FTS trigger updates while preserving vector synchronization.
  • Applies a 64 MiB WAL retention limit with tests and documentation.
File Description
src/​store/​types.ts Removes unused embedding data from retrieved chunk types.
src/​store/​DocumentStore.ts Writes binary embeddings and omits them from retrieval queries.
src/​store/​DocumentStore.test.ts Verifies binary embedding storage and vector-index parity.
src/​store/​assembly/​strategies/​HierarchicalAssemblyStrategy.test.ts Updates fixtures for the revised chunk type.
src/​store/​applyMigrations.ts Configures the WAL size limit.
src/​store/​applyMigrations.test.ts Tests migration conversion, FTS behavior, and WAL configuration.
db/​migrations/​017-store-embeddings-as-blobs.sql Converts embeddings and narrows maintenance triggers.
docs/​concepts/​data-storage.md Documents migrations and vector storage.
openspec/​specs/​database-migrations/​spec.md Specifies WAL retention behavior.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants