Skip to content

About

Serving cost model for 2026 LLM inference stacks, integrating empirical data from 30+ portfolio benchmarks into a single auditable model.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

serving-cost-model-v2

Python License Status

Two-KPI serving cost model incorporating empirical data from 30+ benchmarks built across this portfolio. Answers: how much does it cost to serve 1M tokens in 2026, and which optimization moves the needle most?


Why This Exists

The original inference-cost-calculator (project 22) was built with data from 5 benchmarks. Since then, 30+ additional projects measured disaggregation, KV quantization, tiering, prefetch, fragmentation, mixed precision, eviction, and system prompt caching in detail.

This project synthesizes all of that into a single, auditable cost model.


Key Results

Cost per 1M output tokens — Qwen2-0.5B, medium_chat

Scenario USD/1M out Cost reduction
01_baseline $3.95 —
02_+compaction $2.40 -39.1%
03_+sys_cache $2.39 -39.5%
04_+int8_kv $1.40 -64.5%
05_+mixed_prec $1.31 -66.9%
06_+eviction $1.05 -73.5%
07_+prefetch $1.05 -73.5%
08_+disagg $0.095 -97.6%

Optimization ranking by marginal impact

Rank Optimization medium_chat tool_call
1 compaction 39.1% 39.1%
2 int8_kv 24.9% 23.3%
3 disaggregation 24.1% 22.6%
4 structured eviction 6.6% 6.1%
5 sys_prompt_cache 0.4% 4.3% ← workload-sensitive
6 mixed_precision 2.5% 2.3%
7 prefetch ~0.0% ~0.0%

Why two KPIs matter

Workload $/1M output $/1M request prefill share
medium_chat (0.5B) $0.095 $18.2 2.2%
tool_call (0.5B) $0.114 $2.7 26.4%

Tool calls are cheap per request but expensive per output token. System prompt cache is 10x more impactful in prefill-heavy workloads.


Empirical Sources

Optimization Source benchmark Key measurement
Fragmentation tax continuous-batching-fragmentation-sim 44% throughput loss without compaction
System prompt cache system-prompt-cache-bench hit_rate=0.80
INT8 KV kv-cache-quantization-bench 50% reduction, near-zero quality
Mixed precision mixed-precision-kv-policy 54.2% reduction, delta_nll~0
Structured eviction attention-sink-eviction-policy 62% reduction, low KL
Prefetch kv-cache-prefetch-bench 80% stall reduction
Disaggregation disaggregated-prefill-decode-sim 10.76x throughput gain
TTFT interference prefill-decode-interference-bench 1.25x / 1.50x inflation

Quick Start

git clone https://github.com/JohnScheuer/serving-cost-model-v2
cd serving-cost-model-v2

python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt

python run.py

Runtime: under 5 seconds. No GPU required.


Output Files

results/
  scenarios.csv     incremental scenario results with both KPIs
  sensitivity.csv   sensitivity analysis for the full stack

plots/
  01_waterfall_per_output_*.png
  02_waterfall_per_request_*.png
  03_cost_per_request_by_workload_*.png
  04_decomposition_*.png
  05_sensitivity_*.png

Project Structure

serving-cost-model-v2/
├── src/
│   ├── config.py      models, workloads, scenarios, empirical constants
│   ├── model.py       cost computation with decomposition
│   ├── bench.py       scenario and sensitivity sweep
│   └── analysis.py    plots and comparison tables
├── results/
├── plots/
├── run.py
├── SUMMARY.txt
├── DESIGN.md
├── LICENSE
└── requirements.txt

Documentation


Related Projects


License

MIT License — Copyright (c) 2026 João Felipe De Souza


Author

João Felipe De Souza

About

Serving cost model for 2026 LLM inference stacks, integrating empirical data from 30+ portfolio benchmarks into a single auditable model.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages