Skip to content

Cache (1/5): Add content-addressed local inference caches - #121

Merged
ErlisLushtaku merged 2 commits into
mainfrom
cache-on-118/01-local-store
Sep 22, 2026
Merged

ErlisLushtaku merged 2 commits into
mainfrom
cache-on-118/01-local-store

Conversation

@ErlisLushtaku

@ErlisLushtaku ErlisLushtaku commented Sep 9, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Adds content-addressed local SQLite caches for model completions and judge outputs.

Cache layout:

{store_root}/completions/{task}/{provider}/{model}/{descriptor_hash}/
├── metadata.json
└── completions.db

{store_root}/judgements/{task}/{provider}/{model}/{descriptor_hash}/
├── metadata.json
└── judgements.db

metadata.json stores the validated model descriptor. The databases use these schemas:

completions(
  input_hash TEXT PRIMARY KEY,
  input_text TEXT NOT NULL,
  completion TEXT NOT NULL,
  benchmark TEXT NOT NULL,
  instruction_id TEXT NOT NULL,
  model TEXT NOT NULL,
  pushed_by TEXT NOT NULL,
  pushed_at TEXT NOT NULL,
  run_id TEXT NOT NULL
)

judgements(
  input_hash TEXT PRIMARY KEY,
  judge_input TEXT NOT NULL,
  judge_completion TEXT NOT NULL,
  benchmark TEXT NOT NULL,
  instruction_id TEXT NOT NULL,
  model_a TEXT NOT NULL,
  model_b TEXT NOT NULL,
  judge TEXT NOT NULL,
  top_logprobs TEXT,
  orientation TEXT,
  pushed_by TEXT NOT NULL,
  pushed_at TEXT NOT NULL,
  run_id TEXT NOT NULL
)

top_logprobs is new relative to the earlier draft so judgements can restore the first-token logprobs used by #118 parsers.

@ErlisLushtaku ErlisLushtaku changed the title Cache: WIP (1/6) Add content-addressed local inference caches Cache: WIP (1/5) Add content-addressed local inference caches Sep 9, 2026
@kargibora
kargibora force-pushed the rebase/inference-usage-on-meta-eval branch from 58d6791 to d1ca639 Compare September 10, 2026 09:55
@ErlisLushtaku ErlisLushtaku changed the title Cache: WIP (1/5) Add content-addressed local inference caches Cache (1/5): Add content-addressed local inference caches Sep 15, 2026
@ErlisLushtaku
ErlisLushtaku marked this pull request as ready for review September 15, 2026 07:40
@ErlisLushtaku
ErlisLushtaku changed the base branch from rebase/inference-usage-on-meta-eval to main September 15, 2026 08:02
params: list[Any],
) -> pd.DataFrame:
if input_hashes is not None:
if not input_hashes:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why there are two checks here for input_hashes?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

None means no input-hash filter, while [] means no rows, so the checks preserve different query semantics. I added this to the tests.

Comment thread judgearena/cache_sqlite.py Outdated
str(row["benchmark"]),
str(row["instruction_id"]),
str(row["model_a"]),
str(row["model_b"]),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Note: For sample-wise this can be None. As we have discussed it better to make this more general

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yes, made model_b nullable now.

@kargibora
kargibora force-pushed the cache-on-118/01-local-store branch from b1156b8 to 52a4620 Compare September 15, 2026 12:54
@kargibora

Copy link
Copy Markdown
Member

I think it is a great cache design but I have some concerns apart from what I wrote above. So what we are planning is to have a seperate .db for

   return (
        Path(store_root)
        / kind
        / quote(task, safe="")
        / quote(provider, safe="")
        / quote(model, safe="")
    )

Although this makes sense in general, I feel like it is too fragmented. For each task/provider/model we are creating a database. However instead we can create this for kind/task only.

Recommendations / Questions

1) Database Boundary

As discussed above we have to think about the database design.

We can store provider, model and descriptor as SQL data and just assume cache is stored at

 {store_root}/{kind}/{task}.db

2) Metadata

Perhaps we can also store descriptors seperately

 CREATE TABLE descriptors (
       descriptor_id TEXT PRIMARY KEY,
       provider TEXT NOT NULL,
       model TEXT NOT NULL,
       descriptor_json TEXT NOT NULL
   );
   CREATE TABLE completions (
       descriptor_id TEXT NOT NULL,
       input_hash TEXT NOT NULL,
       input_text TEXT NOT NULL,
       completion TEXT NOT NULL,
       PRIMARY KEY (descriptor_id, input_hash)
   );

but this is not necessary.

3)

We do

   # cache_sqlite.py:207-210
   "INSERT OR REPLACE INTO completions VALUES (...)"

3)

Perhaps we can also put a versioning in cache if we want to apply migrations in the future.

Comment thread judgearena/cache_sqlite.py Outdated
with sqlite3.connect(temporary_db) as conn:
conn.execute(self.schema)
rows.to_sql(self.table, conn, if_exists="append", index=False)
os.replace(temporary_db, self.db_path)

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We might need some handling and protection for the cache. Assuming we can start two jobs at the same time, while reading it from

frames.append(pd.read_sql(f"SELECT * FROM {self.table}", conn))

if another process appends something to this table while

os.replace(temporary_db, self.db_path)

executed, than information will be lost. Also

 pd.read_sql(...)
   pd.concat(...)
   .sort_values(...)
   .drop_duplicates(...)
   rows.to_sql(...)

these operations load both databases and rewrites every row. If we have a huge cache this is expensive. We can use

   ATTACH DATABASE ? AS incoming; 

   INSERT INTO completions (...)
   SELECT ...
   FROM incoming.completions
   ON CONFLICT(descriptor_id, input_hash) DO UPDATE ...;

However this is not required and current is also enough (unless we use VERY large cache)

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, merge_from now uses an in-place transaction with ATTACH and ON CONFLICT, so it no longer rewrites or replaces the live database.

@ErlisLushtaku

ErlisLushtaku commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator Author

@kargibora I kept one database per role/task/provider/model/descriptor because each folder is also the inspection and synchronization unit, and separate files reduce writer contention. I'm not sure if this granularity has any downsides. Since each database already represents one descriptor, metadata.json can be beside it rather than repeated in a table. Schema changes use the schema_version field in the descriptor, so if we change the schema the cache will not collide.

Let me know if you agree with this.

@ErlisLushtaku
ErlisLushtaku merged commit a8d6232 into main Sep 22, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants