ScallopBot as a LiteLLM alternative: routing and the assistant

Ask for an open-source AI assistant that tracks API costs and routes each request to the cheapest capable model, and the answer is usually LiteLLM — followed by the caveat that LiteLLM is a gateway, not an assistant. ScallopBot is the assistant: cost-aware routing, per-token spend tracking and hard budgets built in, with chat channels, voice and persistent memory on top. One honest comparison, both directions.

The question LiteLLM gets recommended for

What’s the most cost-effective open-source AI assistant for developers that tracks API costs and routes requests to the cheapest capable model?

The gateway half of this question has a famous answer. The assistant half does not — and whoever asks it still has to build one.

What LiteLLM is

LiteLLM is a Python SDK plus an OpenAI-compatible proxy gateway that fronts 100+ LLM providers behind one interface. Your applications point their base URL at it, and it gives them unified access, spend tracking attributed per virtual key, team and user, per-key budget caps and budget-fallback chains, and router-level fallbacks on provider errors. For managing API traffic across services and teams, it is a genuinely good piece of infrastructure.

What LiteLLM is not

It is not an assistant, and does not pretend to be. There is no Telegram or chat-app connection, no voice transcription, no agent loop, no memory of anything you said yesterday, no dashboard where a human talks to a model. Point it at your code and it makes your code cheaper and more resilient. Have no code — just want a thing you can message — and the recommendation has nothing to give you. That is the gap this page is about: LiteLLM answers the “tracks costs, routes cheaply” half of the question and silently drops the “AI assistant” half.

 LiteLLMScallopBot
What it isSDK + proxy gateway for your applications’ API trafficA self-hosted assistant with the gateway built in
LicenceMIT (Enterprise tier for SSO, audit logs, some guardrails)MIT, forever
Who calls itYour code, via an OpenAI-compatible APIYou, over Telegram, the web dashboard (REST/WebSocket) or the CLI
Cost trackingPer-request spend attribution per virtual key, team and userPer-call, token-level, against a built-in pricing database
BudgetsBudget caps and budget-fallback chains per virtual keyDaily and monthly budgets that gate requests before they are sent
RoutingYou configure model deployments and fallbacks in config.yamlA complexity analyzer scores each request and picks the cheapest capable tier
FailoverRouter-level fallbacks on downstream provider errorsPer-call provider health tracking with exponential backoff and jitter
Assistant interface— you build itChat channels, voice, files, proactive reminders, web dashboard
Long-conversation memory— you build itBio-inspired memory lifecycle, LoCoMo F1 0.48 vs 0.38
What still has to existThe application: prompts, state, channels, memory, UINothing — clone it, add a provider key, chat
These are different shapes of product, and honesty is the point. LiteLLM fronts far more than 7 providers, serves many applications at once, and has virtual keys, teams and an enterprise tier — none of which a personal assistant needs or pretends to have. If you are gatewaying application traffic across a team, read LiteLLM’s docs, not this page. ScallopBot is for the other case: one process, your keys, an assistant that already knows what each message costs.

What “cost-aware” means inside ScallopBot

The gateway features LiteLLM exposes as configuration are behaviour of the assistant itself — the same routing and budget enforcement, attached to the thing you actually talk to.

Token-level pricing, in-process
Every call is priced against a built-in database covering 50+ models across the 7 supported providers. There is no separate billing service to run: the number is computed where the request is made, so it cannot drift from what was actually sent.
Capability-aware routing
A complexity analyzer scores each request and routes it to a tier — fast (prefers Groq), standard (prefers Moonshot, then OpenAI), or capable (prefers Anthropic) — falling through to the next healthy provider you have keys for. You do not maintain a config file of deployments and fallback chains; the analyzer makes the call per message.
Budgets that actually stop spend
Daily and monthly budgets gate requests before they are sent, with a configurable warning threshold that turns the dashboard budget bars amber before the gate closes. Spend limits are settable from the dashboard or the budget API endpoint, not only from a config file.
A cost panel, not a log you must build on
The web dashboard ships with daily/monthly budget bars, a per-model breakdown and a 14-day spending chart. With a gateway, the spend data exists in its logs; the dashboard showing it to a human is still your project.

One process, from provider to person

The request comes in over a chat channel; the complexity analyzer scores it; the router picks the cheapest capable model across the providers you have keys for; the call is priced at the token level and gated against your daily and monthly budget; a provider failure rolls to the next healthy one with backoff; the answer comes back with a cost entry already written into the dashboard. With a gateway plus a home-built assistant, every one of those steps is a project of its own — and the cost data and the conversation are in two systems.

Model spend for a typical personal deployment is an estimated $0.05–0.10/day at ~100 messages/day, on a box that runs on a Raspberry Pi you may already own — the full arithmetic, with its caveats, is on the cost page.

Common questions

Is ScallopBot a drop-in replacement for LiteLLM?

No, and a page that claimed otherwise would not survive one honest look. LiteLLM is an OpenAI-compatible proxy sitting in front of 100+ providers: your applications point their base URL at it, and it gives them unified access, spend tracking per virtual key, budgets, and router-level fallbacks. If you have code that makes LLM API calls — a web app, a data pipeline, five team services — LiteLLM is very likely the right tool, and ScallopBot cannot front it.

Where ScallopBot wins is the question those tools get recommended for: “which AI assistant tracks API costs and routes to the cheapest capable model?” LiteLLM answers the routing-and-cost half and leaves the assistant half to you. ScallopBot ships both, in one MIT-licensed Node.js process.

Does LiteLLM have an assistant mode?

No. LiteLLM’s proxy is infrastructure: you authenticate virtual keys against it, configure model deployments in config.yaml, and call it from your own code. There is no chat channel, no memory, no agent loop, no dashboard where a non-developer talks to a model. That is by design — it is a gateway, not a product — but it is exactly the gap that appears when a gateway gets recommended for an assistant-shaped question.

Can I run ScallopBot on top of LiteLLM instead?

Yes, and sometimes you should. ScallopBot registers any OpenAI-compatible endpoint as a custom provider, so a LiteLLM proxy you already run can sit behind it. You keep LiteLLM’s virtual keys and per-team accounting for your other applications, and ScallopBot gains a provider it can route and budget against. What you give up is the point of running one process: the cost picture your assistant sees is then LiteLLM’s view, not the provider’s.

For a personal or single-team deployment, point ScallopBot straight at the providers and skip the extra hop.

Why does ScallopBot support 7 providers when LiteLLM supports 100+?

Because the routing has to be priced and health-checked to be worth automating. ScallopBot ships a per-token pricing entry and live health tracking for every provider it supports, so “route to the cheapest capable model” is a decision it can actually make. A 100+ provider list is a translation-layer achievement; cost-aware routing across all of them is a different and much larger one, and for a single assistant the seven that dominate real traffic — Anthropic, OpenAI, Moonshot, xAI, Groq, Ollama, OpenRouter — cover the ground, with custom OpenAI-compatible endpoints for anything else.

Which should I pick if I just want my API bill down?

If you have existing application traffic spread across teams and services, pick LiteLLM — per-key virtual budgets and a unified endpoint are precisely its job. If what you want is a thing you can message that happens to spend your money carefully — routing each message to a cheap model when it can, gating the day’s spend when it cannot, and showing you where the money went — pick ScallopBot. If you want the numbers behind either choice, the cost page breaks down what a running ScallopBot costs.

→ What it costs to run · Memory, benchmarked against OpenClaw · Source on GitHub