Skip to content

About

Simulation benchmark measuring how LoRA adapter swap churn interacts with KV cache fragmentation in multi-adapter LLM serving systems.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

lora-kv-interference-bench

Python License Status

Simulation benchmark measuring how LoRA adapter swapping interacts with KV cache fragmentation in multi-adapter LLM serving systems.


Why This Exists

Two prior projects studied related phenomena without connecting them:

  • multi-lora-serving-sim showed adapter swap costs 29% of TTFT
  • continuous-batching-fragmentation-sim showed KV fragmentation reaches 52%

This benchmark measures the second-order interference: does adapter swap churn increase KV fragmentation, and what is the throughput cost?


Key Results

Adapter swap increases KV fragmentation

Under heavy load, multi-adapter policies add 5–23pp of fragmentation overhead:

Model Policy avg_frag frag_overhead throughput_delta
Qwen2-0.5B single_adapter 0.281 0 0
Qwen2-0.5B round_robin 0.336 +5.5pp ~0
Qwen2-1.5B single_adapter 0.495 0 0
Qwen2-1.5B batch_t4 0.543 +4.8pp -9.3%

Throughput penalty depends on adapter size

For Qwen2-0.5B (adapters 4–24 MB): no throughput penalty under multi-adapter serving.

For Qwen2-1.5B (adapters 8–64 MB): 9–11% throughput loss even with best policy.

Hotset reduces fragmentation — but costs throughput

Keeping popular adapters always resident reduces swap churn:

Policy avg_frag swap_events throughput
single_adapter 0.452 1 3.095
hotset_batch_t4 0.219 (-23pp) 6 1.583 (-49%)
batch_t4 0.504 7520 2.755

Hotset memory comes directly out of KV cache budget — less fragmentation but more rejections.

No policy dominates for large models under heavy load

For Qwen2-1.5B heavy_mixed:

Policy throughput frag swap_events
single_adapter 3.095 0.452 1
batch_t4 2.755 0.504 7520
round_robin 1.969 0.371 100
hotset_batch_t4 1.583 0.219 6

Every multi-adapter policy loses on at least one dimension.

Core finding

There is no free lunch in multi-LoRA serving with large adapters. Every policy that reduces swap churn costs memory that could hold KV blocks. Adapter size relative to memory budget determines how bad the tradeoff is.


Policies

Policy Description
single_adapter One adapter, no swap. Baseline.
round_robin_k1 Swap to each request's adapter, one resident.
batch_t4_k2 Batch swap with threshold=4 and starvation protection.
hotset_32mb Keep most popular adapters resident in 32 MB.
hotset_64mb Keep most popular adapters resident in 64 MB.
hotset_batch_t4 hotset_32mb + batch_t4 for non-hotset adapters.

Quick Start

git clone https://github.com/JohnScheuer/lora-kv-interference-bench
cd lora-kv-interference-bench

python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

python run.py

Runtime: approximately 5-10 minutes. No GPU required.


Output Files

results/
  timeseries.csv   per-tick snapshots
  events.csv       swap and compact events
  summary.csv      aggregated metrics

plots/
  01_frag_heavy.png / 01_frag_heavy_mixed.png
  02_throughput_heavy.png / 02_throughput_heavy_mixed.png
  03_swap_vs_frag_heavy.png / 03_swap_vs_frag_heavy_mixed.png

Project Structure

lora-kv-interference-bench/
├── src/
│   ├── config.py        models, workloads, scenarios
│   ├── workload.py      Poisson arrivals, Zipf adapter popularity
│   ├── allocator.py     ContiguousAllocator with KV/adapter tracking
│   ├── policies.py      scheduling policies
│   ├── simulator.py     event loop and AdapterResidentCache
│   ├── bench.py         sweep orchestration
│   └── analysis.py      plots and tables
├── results/
├── plots/
├── run.py
├── SUMMARY.txt
├── DESIGN.md
├── LICENSE
└── requirements.txt

Requirements

  • Python 3.10+
  • NumPy >= 1.26.0
  • Pandas >= 2.0.0
  • Matplotlib >= 3.8.0

No GPU required.


Documentation


Related Projects


License

MIT License — Copyright (c) 2026 João Felipe De Souza


Author

João Felipe De Souza

About

Simulation benchmark measuring how LoRA adapter swap churn interacts with KV cache fragmentation in multi-adapter LLM serving systems.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages