Repository navigation
Conversation
Embeddings were written to documents.embedding as JSON text, 3-5x the size of the binary vector and on top of the binary copy in documents_vec. They are now written as float32 blobs, and migration 017 converts existing rows with vec_f32. sqlite-vec accepts both forms, so the vector triggers pass the blob through. The full-text update trigger fired on every column, so each embedding write rewrote that chunk's full-text entry. It now fires only for content, metadata and page_id, the columns the index reads. The migration narrows it before converting, so the conversion does not rewrite the whole full-text index. journal_size_limit is set to 64 MB so SQLite trims the WAL back after a checkpoint instead of keeping its peak size until the last connection closes. Refs #509
There was a problem hiding this comment.
Copilot review overview
🟢 Approval recommended
The storage, migration, trigger, configuration, test, and documentation changes are coherent with no unresolved correctness issues.
Review effort: Balanced
Findings: None
What changed in this PR
Optimizes SQLite storage by reducing embedding size, avoiding unnecessary FTS updates, and limiting retained WAL growth.
Changes:
- Stores embeddings as float32 blobs and migrates existing JSON vectors.
- Narrows FTS trigger updates while preserving vector synchronization.
- Applies a 64 MiB WAL retention limit with tests and documentation.
| File | Description |
|---|---|
src/store/types.ts |
Removes unused embedding data from retrieved chunk types. |
src/store/DocumentStore.ts |
Writes binary embeddings and omits them from retrieval queries. |
src/store/DocumentStore.test.ts |
Verifies binary embedding storage and vector-index parity. |
src/store/assembly/strategies/HierarchicalAssemblyStrategy.test.ts |
Updates fixtures for the revised chunk type. |
src/store/applyMigrations.ts |
Configures the WAL size limit. |
src/store/applyMigrations.test.ts |
Tests migration conversion, FTS behavior, and WAL configuration. |
db/migrations/017-store-embeddings-as-blobs.sql |
Converts embeddings and narrows maintenance triggers. |
docs/concepts/data-storage.md |
Documents migrations and vector storage. |
openspec/specs/database-migrations/spec.md |
Specifies WAL retention behavior. |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
…play-name migration
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Three changes that make the database smaller, prompted by #509:
documents.embeddingheld each vector as JSON, on top of the binary copy indocuments_vec. New writes use aFloat32Arraybuffer. Migration 018 converts existing rows withvec_f32. Read queries no longer select the column, since nothing in the code used it.content,metadataandpage_id. Before, it fired on every update, so every embedding write deleted and re-inserted that chunk's full-text entry.journal_size_limit = 64 MBmakes SQLite cut the WAL back once a checkpoint resets it. Before, it kept its peak size until the last connection closed, which matches the 6.7 GB WAL in feat: Storage management in Web Dashboard, reclaimable space visibility, and post-scrape WAL cleanup #509.Size of the effect
UPDATE documents SET embedding = NULL) no longer rewrites the whole full-text index in one transaction.Tested on a real database
A local store with 8 libraries and 46,535 chunks, all embedded with
openai:text-embedding-3-small. The pre-migration state was rebuilt byte for byte: the OpenAI client returns float32 values as a plain array, so the old JSON isJSON.stringify(Array.from(Float32Array)).documentstableDocumentStore.initialize()(migration 018 plus the VACUUM) took 21 s.quick_checkand the FTS5 integrity check pass.findByContentsearch throughDocumentStorereturns results.Migration notes
The file is numbered 018 because feat!: library display names and overwrite protection #517 adds
017-add-library-display-name.sql.The migration is safe to run again: it only converts rows that are still JSON, and it drops each trigger before creating it.
The migration first narrows the full-text trigger, so the conversion doesn't rewrite the whole full-text index.
It drops the vector triggers during the conversion, because
documents_vecalready holds the same vectors. Then it recreates them to pass the blob through; sqlite-vec accepts both blobs and JSON.vec_f32throws on malformed JSON, which would fail the migration. Every stored value already went through sqlite-vec's parser when it was written, via the triggers and the backfill in 011, so I didn't add a guard. A full guard made the conversion 2.5× slower.First start after upgrading rewrites the
documentstable and then runs the usual post-migration VACUUM. The conversion took about 7 s per 30k embedded chunks locally. Large stores like the one in feat: Storage management in Web Dashboard, reclaimable space visibility, and post-scrape WAL cleanup #509 should expect a few minutes.I checked that
Buffer.from(Float32Array)produces the same bytes asvec_f32on the JSON, so new and converted rows are identical.Tests
DocumentStore.test.ts: embeddings written byaddDocumentsare 6,144-byte blobs equal to the vector indocuments_vec. The test fails with the old JSON write path.applyMigrations.test.ts:telemetry-e2ein one run, themcp-http-e2eheartbeat in another. Both pass on their own (9/9). The migration does nothing on the fresh databases these tests create.Docs: updated the vector storage section and migration list in
docs/concepts/data-storage.md, and added the WAL limit to thedatabase-migrationsspec.Refs #509