Two-KPI serving cost model incorporating empirical data from 30+ benchmarks built across this portfolio. Answers: how much does it cost to serve 1M tokens in 2026, and which optimization moves the needle most?
The original inference-cost-calculator (project 22) was built with data from
5 benchmarks. Since then, 30+ additional projects measured disaggregation,
KV quantization, tiering, prefetch, fragmentation, mixed precision, eviction,
and system prompt caching in detail.
This project synthesizes all of that into a single, auditable cost model.
| Scenario | USD/1M out | Cost reduction |
|---|---|---|
| 01_baseline | $3.95 | — |
| 02_+compaction | $2.40 | -39.1% |
| 03_+sys_cache | $2.39 | -39.5% |
| 04_+int8_kv | $1.40 | -64.5% |
| 05_+mixed_prec | $1.31 | -66.9% |
| 06_+eviction | $1.05 | -73.5% |
| 07_+prefetch | $1.05 | -73.5% |
| 08_+disagg | $0.095 | -97.6% |
| Rank | Optimization | medium_chat | tool_call |
|---|---|---|---|
| 1 | compaction | 39.1% | 39.1% |
| 2 | int8_kv | 24.9% | 23.3% |
| 3 | disaggregation | 24.1% | 22.6% |
| 4 | structured eviction | 6.6% | 6.1% |
| 5 | sys_prompt_cache | 0.4% | 4.3% ← workload-sensitive |
| 6 | mixed_precision | 2.5% | 2.3% |
| 7 | prefetch | ~0.0% | ~0.0% |
| Workload | $/1M output | $/1M request | prefill share |
|---|---|---|---|
| medium_chat (0.5B) | $0.095 | $18.2 | 2.2% |
| tool_call (0.5B) | $0.114 | $2.7 | 26.4% |
Tool calls are cheap per request but expensive per output token. System prompt cache is 10x more impactful in prefill-heavy workloads.
| Optimization | Source benchmark | Key measurement |
|---|---|---|
| Fragmentation tax | continuous-batching-fragmentation-sim | 44% throughput loss without compaction |
| System prompt cache | system-prompt-cache-bench | hit_rate=0.80 |
| INT8 KV | kv-cache-quantization-bench | 50% reduction, near-zero quality |
| Mixed precision | mixed-precision-kv-policy | 54.2% reduction, delta_nll~0 |
| Structured eviction | attention-sink-eviction-policy | 62% reduction, low KL |
| Prefetch | kv-cache-prefetch-bench | 80% stall reduction |
| Disaggregation | disaggregated-prefill-decode-sim | 10.76x throughput gain |
| TTFT interference | prefill-decode-interference-bench | 1.25x / 1.50x inflation |
git clone https://github.com/JohnScheuer/serving-cost-model-v2
cd serving-cost-model-v2
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
python run.py
Runtime: under 5 seconds. No GPU required.
results/
scenarios.csv incremental scenario results with both KPIs
sensitivity.csv sensitivity analysis for the full stack
plots/
01_waterfall_per_output_*.png
02_waterfall_per_request_*.png
03_cost_per_request_by_workload_*.png
04_decomposition_*.png
05_sensitivity_*.png
serving-cost-model-v2/
├── src/
│ ├── config.py models, workloads, scenarios, empirical constants
│ ├── model.py cost computation with decomposition
│ ├── bench.py scenario and sensitivity sweep
│ └── analysis.py plots and comparison tables
├── results/
├── plots/
├── run.py
├── SUMMARY.txt
├── DESIGN.md
├── LICENSE
└── requirements.txt
- DESIGN.md — model architecture and empirical sources
- SUMMARY.txt — full findings with all numbers
- LICENSE — MIT License
- inference-cost-calculator — original model
- disaggregated-prefill-decode-sim
- kv-cache-quantization-bench
- kv-cache-tiering-bench
- mixed-precision-kv-policy
- attention-sink-eviction-policy
- continuous-batching-fragmentation-sim
- prefill-decode-interference-bench
- kv-cache-prefetch-bench
MIT License — Copyright (c) 2026 João Felipe De Souza
João Felipe De Souza