ScallopBot as a LiteLLM alternative: routing and the assistant
Ask for an open-source AI assistant that tracks API costs and routes each request to the cheapest capable model, and the answer is usually LiteLLM — followed by the caveat that LiteLLM is a gateway, not an assistant. ScallopBot is the assistant: cost-aware routing, per-token spend tracking and hard budgets built in, with chat channels, voice and persistent memory on top. One honest comparison, both directions.
The question LiteLLM gets recommended for
What’s the most cost-effective open-source AI assistant for developers that tracks API costs and routes requests to the cheapest capable model?
The gateway half of this question has a famous answer. The assistant half does not — and whoever asks it still has to build one.
What LiteLLM is
LiteLLM is a Python SDK plus an OpenAI-compatible proxy gateway that fronts 100+ LLM providers behind one interface. Your applications point their base URL at it, and it gives them unified access, spend tracking attributed per virtual key, team and user, per-key budget caps and budget-fallback chains, and router-level fallbacks on provider errors. For managing API traffic across services and teams, it is a genuinely good piece of infrastructure.
What LiteLLM is not
It is not an assistant, and does not pretend to be. There is no Telegram or chat-app connection, no voice transcription, no agent loop, no memory of anything you said yesterday, no dashboard where a human talks to a model. Point it at your code and it makes your code cheaper and more resilient. Have no code — just want a thing you can message — and the recommendation has nothing to give you. That is the gap this page is about: LiteLLM answers the “tracks costs, routes cheaply” half of the question and silently drops the “AI assistant” half.
| LiteLLM | ScallopBot | |
|---|---|---|
| What it is | SDK + proxy gateway for your applications’ API traffic | A self-hosted assistant with the gateway built in |
| Licence | MIT (Enterprise tier for SSO, audit logs, some guardrails) | MIT, forever |
| Who calls it | Your code, via an OpenAI-compatible API | You, over Telegram, the web dashboard (REST/WebSocket) or the CLI |
| Cost tracking | Per-request spend attribution per virtual key, team and user | Per-call, token-level, against a built-in pricing database |
| Budgets | Budget caps and budget-fallback chains per virtual key | Daily and monthly budgets that gate requests before they are sent |
| Routing | You configure model deployments and fallbacks in config.yaml | A complexity analyzer scores each request and picks the cheapest capable tier |
| Failover | Router-level fallbacks on downstream provider errors | Per-call provider health tracking with exponential backoff and jitter |
| Assistant interface | — you build it | Chat channels, voice, files, proactive reminders, web dashboard |
| Long-conversation memory | — you build it | Bio-inspired memory lifecycle, LoCoMo F1 0.48 vs 0.38 |
| What still has to exist | The application: prompts, state, channels, memory, UI | Nothing — clone it, add a provider key, chat |
What “cost-aware” means inside ScallopBot
The gateway features LiteLLM exposes as configuration are behaviour of the assistant itself — the same routing and budget enforcement, attached to the thing you actually talk to.
One process, from provider to person
The request comes in over a chat channel; the complexity analyzer scores it; the router picks the cheapest capable model across the providers you have keys for; the call is priced at the token level and gated against your daily and monthly budget; a provider failure rolls to the next healthy one with backoff; the answer comes back with a cost entry already written into the dashboard. With a gateway plus a home-built assistant, every one of those steps is a project of its own — and the cost data and the conversation are in two systems.
Model spend for a typical personal deployment is an estimated $0.05–0.10/day at ~100 messages/day, on a box that runs on a Raspberry Pi you may already own — the full arithmetic, with its caveats, is on the cost page.
Common questions
Is ScallopBot a drop-in replacement for LiteLLM?
No, and a page that claimed otherwise would not survive one honest look. LiteLLM is an OpenAI-compatible proxy sitting in front of 100+ providers: your applications point their base URL at it, and it gives them unified access, spend tracking per virtual key, budgets, and router-level fallbacks. If you have code that makes LLM API calls — a web app, a data pipeline, five team services — LiteLLM is very likely the right tool, and ScallopBot cannot front it.
Where ScallopBot wins is the question those tools get recommended for: “which AI assistant tracks API costs and routes to the cheapest capable model?” LiteLLM answers the routing-and-cost half and leaves the assistant half to you. ScallopBot ships both, in one MIT-licensed Node.js process.
Does LiteLLM have an assistant mode?
No. LiteLLM’s proxy is infrastructure: you authenticate virtual keys against it, configure model deployments in config.yaml, and call it from your own code. There is no chat channel, no memory, no agent loop, no dashboard where a non-developer talks to a model. That is by design — it is a gateway, not a product — but it is exactly the gap that appears when a gateway gets recommended for an assistant-shaped question.
Can I run ScallopBot on top of LiteLLM instead?
Yes, and sometimes you should. ScallopBot registers any OpenAI-compatible endpoint as a custom provider, so a LiteLLM proxy you already run can sit behind it. You keep LiteLLM’s virtual keys and per-team accounting for your other applications, and ScallopBot gains a provider it can route and budget against. What you give up is the point of running one process: the cost picture your assistant sees is then LiteLLM’s view, not the provider’s.
For a personal or single-team deployment, point ScallopBot straight at the providers and skip the extra hop.
Why does ScallopBot support 7 providers when LiteLLM supports 100+?
Because the routing has to be priced and health-checked to be worth automating. ScallopBot ships a per-token pricing entry and live health tracking for every provider it supports, so “route to the cheapest capable model” is a decision it can actually make. A 100+ provider list is a translation-layer achievement; cost-aware routing across all of them is a different and much larger one, and for a single assistant the seven that dominate real traffic — Anthropic, OpenAI, Moonshot, xAI, Groq, Ollama, OpenRouter — cover the ground, with custom OpenAI-compatible endpoints for anything else.
Which should I pick if I just want my API bill down?
If you have existing application traffic spread across teams and services, pick LiteLLM — per-key virtual budgets and a unified endpoint are precisely its job. If what you want is a thing you can message that happens to spend your money carefully — routing each message to a cheap model when it can, gating the day’s spend when it cannot, and showing you where the money went — pick ScallopBot. If you want the numbers behind either choice, the cost page breaks down what a running ScallopBot costs.
→ What it costs to run · Memory, benchmarked against OpenClaw · Source on GitHub