Developed by Mohammad Ibrahim Saleem
Cost-aware LLM request routing: send each query to the cheapest model that can actually handle it, then prove the savings with real dollar-and-percentage measurements instead of vendor benchmark numbers. Includes a live chat UI, a Model Advisor, and a full architecture/methodology write-up — all as one site.
Powered by RouteLLM (Apache-2.0, UC Berkeley/LMSYS + Anyscale) as the routing engine — vendored in full under engine/ (see engine/NOTICE for attribution).
Built as a portfolio-grade demonstration of cost-aware model routing (an AI Engineer JD pattern: route across model tiers, track cost-per-resolved-task, prove savings vs. an always-expensive baseline), and secondarily to validate whether the approach is worth proposing internally at a company evaluating LLM API spend. See routellm-full-context.md for the full background and decision log, and routellm-setup-plan.md for the phased setup plan this project followed.
app/ is a small FastAPI + vanilla-JS site with four pages, all served from one process:
/— home/landing page: what SwitchLane does, headline stats, links to the other pages./chat— the live chat demo (below)./advisor— the Model Advisor (below)./about— full methodology, architecture, validation results, and setup guide as a real page, not a modal.
Routes every message through the routing classifier in real time and shows, per message and cumulatively for the session:
- which model tier handled it (weak/strong) and which model,
- tokens used and latency,
- % saved vs. always using the strong model — computed from the actual token usage priced at strong-model rates, so it's an honest per-message proxy rather than a second live call that would double spend,
- a session-wide savings percentage and weak/strong routing split.
Percentages, not raw dollars, are the headline number here — at these token volumes the real cost difference per message is a fraction of a cent, so a dollar figure undersells what's actually a 90%+ cost reduction on most queries.
Run it:
cd engine
python -m venv venv
venv\Scripts\activate # Windows; source venv/bin/activate on Linux/Mac
pip install -e .
pip install fastapi uvicorn tiktokenCreate engine/.env (gitignored) with your key:
OPENAI_API_KEY=sk-...
Then from the repo root:
python app/server.pyOpen http://127.0.0.1:8010. A mode switcher (Max savings / Balanced / Max quality) lets you pick between the three pre-calibrated thresholds from the validation run (20% / 30% / 50% strong-model traffic).
At /advisor — a tool for developers deciding which model to use for their own LLM call, not a live chat. Describe what the call does, paste a real example input (and optionally expected output) and your priority (cost-sensitive / balanced / quality-critical), and it recommends one of three models:
- runs the example through the routing classifier for a real quality signal (how much a top-tier model would actually help on this specific input),
- counts actual tokens with
tiktokento compute real per-call cost acrossgpt-5.6-luna/gpt-5.6-terra/gpt-5.6-sol(and a monthly projection if you give a call volume), - asks the strong model to reason about task-type fit using that real data, returning a grounded recommendation plus an alternative for a different priority.
Nothing here is guessed blind — the recommendation is backed by the same classifier and pricing math the rest of SwitchLane uses.
Full writeup: engine/RESULTS.md.
| Config | Total cost (15 prompts) | Pass rate | Cost per resolved task |
|---|---|---|---|
All-weak (gpt-5.6-luna) |
$0.000375 | 15/15 | $0.000025 |
All-strong (gpt-5.6-terra) |
$0.003028 | 15/15 | $0.000202 |
Router (mf, threshold 0.15609) |
$0.002135 | 15/15 | $0.000142 |
Router beat all-strong by 29.5% cost reduction at identical pass rate. Also documented: an honest caveat that this particular 15-prompt set was easy enough for the weak model alone to solve everything too, so it doesn't yet prove the router's quality value on its own — see RESULTS.md for what a stronger follow-up test would need.
routellm-full-context.md/routellm-setup-plan.md— background, decisions, and the phased setup plan this project followed.app/— the site (server.pyFastAPI backend;static/home.html,static/chat.html,static/advisor.html,static/about.html,static/shared.css).engine/— the full RouteLLM codebase (unmodified, Apache-2.0; seeengine/NOTICE), plus the validation scripts and results built on top of it:benchmark_prompts.py— the 15 hand-built easy/medium/hard test prompts with expected answers.run_pass_openai.py— runs the all-weak / all-strong passes via a plain OpenAI client call (noroutellm/torch import needed).run_pass_router.py— runs the router pass via RouteLLM'sController(mfrouter).run_benchmark.py— single-process version of all three passes (kept for reference; superseded by the split scripts above after a local out-of-memory issue loadingtorch— see RESULTS.md).smoke_test.py— the original quickstart smoke test.pass_weak.json/pass_strong.json/pass_router.json— raw per-prompt results (response, tokens, cost, latency) from the validation run.RESULTS.md— the full validation writeup.
From engine/ (after the install steps above):
python run_pass_openai.py gpt-5.6-luna pass_weak.json all-weak
python run_pass_openai.py gpt-5.6-terra pass_strong.json all-strong
python run_pass_router.pyAdjust the model IDs and per-MTok rates in each script to whatever strong/weak model pair and current pricing you're validating against — the ones here (gpt-5.6-terra / gpt-5.6-luna) were the specific pair available for this test.
Note on install extras: RouteLLM's pyproject.toml bundles sglang (GPU model-serving, for locally-hosted benchmark generation) into its eval extra. That's unnecessary when routing between two API-hosted models — pip install -e . (core only) plus fastapi uvicorn (for the chat UI) or matplotlib pandarallel tiktoken (for RouteLLM's own eval scripts) covers everything here without pulling in several GB of GPU-oriented dependencies.
requirements.txt and render.yaml at the repo root make this a connect-the-repo deploy:
- On Render: New → Blueprint, connect this repo. It reads
render.yamland creates a web service with the right build (pip install -r requirements.txt) and start (python app/server.py) commands. - Set the
OPENAI_API_KEYenv var in Render's dashboard (left blank inrender.yamlon purpose — never commit real keys). - Deploy.
Before you do, know the real constraint: loading the routing classifier (PyTorch + its checkpoint) uses 500MB+ of RAM at startup. Render's free tier (512MB) will very likely crash on boot — budget for at least their smallest plan with real headroom, not the free tier. render.yaml defaults to plan: starter; adjust to whatever Render currently offers with enough memory.
Other things worth knowing before deploying:
- The classifier's checkpoint downloads from Hugging Face on cold start (no persistent disk configured) — adds to every fresh boot's startup time.
- Session stats (the chat page's running savings ticker) live in an in-memory dict, per process — fine for a single instance, but won't stay consistent if you ever scale to multiple instances.
- No Blueprint access? Create the Web Service manually instead: connect the repo, set Build Command to
pip install -r requirements.txt, Start Command topython app/server.py, and addOPENAI_API_KEYas an environment variable.
engine/ vendors RouteLLM, Apache-2.0 licensed by its original authors (UC Berkeley/LMSYS + Anyscale) — unmodified except for the containing directory name; see engine/LICENSE and engine/NOTICE. Everything else in this repo (the site, validation scripts, results, context/planning docs) is original work.
Built by Mohammad Ibrahim Saleem — github.com/ibrahimsaleem.