Skip to content

Repository files navigation

SwitchLane

Developed by Mohammad Ibrahim Saleem

Cost-aware LLM request routing: send each query to the cheapest model that can actually handle it, then prove the savings with real dollar-and-percentage measurements instead of vendor benchmark numbers. Includes a live chat UI, a Model Advisor, and a full architecture/methodology write-up — all as one site.

Powered by RouteLLM (Apache-2.0, UC Berkeley/LMSYS + Anyscale) as the routing engine — vendored in full under engine/ (see engine/NOTICE for attribution).

Why

Built as a portfolio-grade demonstration of cost-aware model routing (an AI Engineer JD pattern: route across model tiers, track cost-per-resolved-task, prove savings vs. an always-expensive baseline), and secondarily to validate whether the approach is worth proposing internally at a company evaluating LLM API spend. See routellm-full-context.md for the full background and decision log, and routellm-setup-plan.md for the phased setup plan this project followed.

The site

app/ is a small FastAPI + vanilla-JS site with four pages, all served from one process:

  • / — home/landing page: what SwitchLane does, headline stats, links to the other pages.
  • /chat — the live chat demo (below).
  • /advisor — the Model Advisor (below).
  • /about — full methodology, architecture, validation results, and setup guide as a real page, not a modal.

Live chat

Routes every message through the routing classifier in real time and shows, per message and cumulatively for the session:

  • which model tier handled it (weak/strong) and which model,
  • tokens used and latency,
  • % saved vs. always using the strong model — computed from the actual token usage priced at strong-model rates, so it's an honest per-message proxy rather than a second live call that would double spend,
  • a session-wide savings percentage and weak/strong routing split.

Percentages, not raw dollars, are the headline number here — at these token volumes the real cost difference per message is a fraction of a cent, so a dollar figure undersells what's actually a 90%+ cost reduction on most queries.

Run it:

cd engine
python -m venv venv
venv\Scripts\activate        # Windows; source venv/bin/activate on Linux/Mac
pip install -e .
pip install fastapi uvicorn tiktoken

Create engine/.env (gitignored) with your key:

OPENAI_API_KEY=sk-...

Then from the repo root:

python app/server.py

Open http://127.0.0.1:8010. A mode switcher (Max savings / Balanced / Max quality) lets you pick between the three pre-calibrated thresholds from the validation run (20% / 30% / 50% strong-model traffic).

Model Advisor

At /advisor — a tool for developers deciding which model to use for their own LLM call, not a live chat. Describe what the call does, paste a real example input (and optionally expected output) and your priority (cost-sensitive / balanced / quality-critical), and it recommends one of three models:

  • runs the example through the routing classifier for a real quality signal (how much a top-tier model would actually help on this specific input),
  • counts actual tokens with tiktoken to compute real per-call cost across gpt-5.6-luna / gpt-5.6-terra / gpt-5.6-sol (and a monthly projection if you give a call volume),
  • asks the strong model to reason about task-type fit using that real data, returning a grounded recommendation plus an alternative for a different priority.

Nothing here is guessed blind — the recommendation is backed by the same classifier and pricing math the rest of SwitchLane uses.

Validation results

Full writeup: engine/RESULTS.md.

Config Total cost (15 prompts) Pass rate Cost per resolved task
All-weak (gpt-5.6-luna) $0.000375 15/15 $0.000025
All-strong (gpt-5.6-terra) $0.003028 15/15 $0.000202
Router (mf, threshold 0.15609) $0.002135 15/15 $0.000142

Router beat all-strong by 29.5% cost reduction at identical pass rate. Also documented: an honest caveat that this particular 15-prompt set was easy enough for the weak model alone to solve everything too, so it doesn't yet prove the router's quality value on its own — see RESULTS.md for what a stronger follow-up test would need.

Repo layout

  • routellm-full-context.md / routellm-setup-plan.md — background, decisions, and the phased setup plan this project followed.
  • app/ — the site (server.py FastAPI backend; static/home.html, static/chat.html, static/advisor.html, static/about.html, static/shared.css).
  • engine/ — the full RouteLLM codebase (unmodified, Apache-2.0; see engine/NOTICE), plus the validation scripts and results built on top of it:
    • benchmark_prompts.py — the 15 hand-built easy/medium/hard test prompts with expected answers.
    • run_pass_openai.py — runs the all-weak / all-strong passes via a plain OpenAI client call (no routellm/torch import needed).
    • run_pass_router.py — runs the router pass via RouteLLM's Controller (mf router).
    • run_benchmark.py — single-process version of all three passes (kept for reference; superseded by the split scripts above after a local out-of-memory issue loading torch — see RESULTS.md).
    • smoke_test.py — the original quickstart smoke test.
    • pass_weak.json / pass_strong.json / pass_router.json — raw per-prompt results (response, tokens, cost, latency) from the validation run.
    • RESULTS.md — the full validation writeup.

Reproducing the validation benchmark

From engine/ (after the install steps above):

python run_pass_openai.py gpt-5.6-luna pass_weak.json all-weak
python run_pass_openai.py gpt-5.6-terra pass_strong.json all-strong
python run_pass_router.py

Adjust the model IDs and per-MTok rates in each script to whatever strong/weak model pair and current pricing you're validating against — the ones here (gpt-5.6-terra / gpt-5.6-luna) were the specific pair available for this test.

Note on install extras: RouteLLM's pyproject.toml bundles sglang (GPU model-serving, for locally-hosted benchmark generation) into its eval extra. That's unnecessary when routing between two API-hosted models — pip install -e . (core only) plus fastapi uvicorn (for the chat UI) or matplotlib pandarallel tiktoken (for RouteLLM's own eval scripts) covers everything here without pulling in several GB of GPU-oriented dependencies.

Deploying to Render

requirements.txt and render.yaml at the repo root make this a connect-the-repo deploy:

  1. On Render: New → Blueprint, connect this repo. It reads render.yaml and creates a web service with the right build (pip install -r requirements.txt) and start (python app/server.py) commands.
  2. Set the OPENAI_API_KEY env var in Render's dashboard (left blank in render.yaml on purpose — never commit real keys).
  3. Deploy.

Before you do, know the real constraint: loading the routing classifier (PyTorch + its checkpoint) uses 500MB+ of RAM at startup. Render's free tier (512MB) will very likely crash on boot — budget for at least their smallest plan with real headroom, not the free tier. render.yaml defaults to plan: starter; adjust to whatever Render currently offers with enough memory.

Other things worth knowing before deploying:

  • The classifier's checkpoint downloads from Hugging Face on cold start (no persistent disk configured) — adds to every fresh boot's startup time.
  • Session stats (the chat page's running savings ticker) live in an in-memory dict, per process — fine for a single instance, but won't stay consistent if you ever scale to multiple instances.
  • No Blueprint access? Create the Web Service manually instead: connect the repo, set Build Command to pip install -r requirements.txt, Start Command to python app/server.py, and add OPENAI_API_KEY as an environment variable.

License / attribution

engine/ vendors RouteLLM, Apache-2.0 licensed by its original authors (UC Berkeley/LMSYS + Anyscale) — unmodified except for the containing directory name; see engine/LICENSE and engine/NOTICE. Everything else in this repo (the site, validation scripts, results, context/planning docs) is original work.

Author

Built by Mohammad Ibrahim Saleem — github.com/ibrahimsaleem.

About

Cost-aware LLM request routing with a live chat UI showing per-message and session cost savings, powered by RouteLLM

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages