ScallopBot vs. OpenClaw
The full comparison between running OpenClaw and running ScallopBot — the open-source personal AI assistant that consolidates memory in its sleep. Both self-host, both speak the OpenClaw skill format; the difference is the cognitive layer underneath. Every claim below is either benchmarked or cross-linked, and the rows OpenClaw wins stay in the table.
Both consolidate memory. They do it differently.
OpenClaw is an excellent skill-orchestration runtime: huge channel coverage, native apps, thousands of community skills. Both projects now consolidate memory in the background. OpenClaw’s “dreaming” promotes notes you keep recalling into a curated MEMORY.md. ScallopBot’s sleep-style consolidation rewrites memory itself: it fuses duplicates, merges fragments into new summaries, links related memories and forgets what stopped being useful — plus the reflection and cost-routing machinery around it.
Feature comparison
| Feature | OpenClaw | ScallopBot |
|---|---|---|
| Memory consolidation | “Dreaming” (on by default): light / REM / deep sweep promotes frequently-recalled notes from daily files into MEMORY.md | Nightly sleep-style cycle: NREM fusion merges duplicates and clusters cross-topic fragments into new summaries; REM builds typed associations |
| Memory decay / forgetting | Recency decay on search ranking (30-day half-life); notes are not archived or pruned | Utility-based: activation decay with category half-lives, soft-archive then hard-prune |
| Memory retrieval | Vector + keyword hybrid, deterministic weighted ranking | BM25 + semantic hybrid with optional LLM re-ranking, score-gated |
| Associative recall | — | Spreading activation over typed memory edges (REM phase) |
| Temporal queries | Recency weighting; temporal phrasing escalates the search | Dates embedded in memory text + time-scope detection routes “what changed since” questions through time-aware retrieval |
| Self-reflection & evolution | — | Private composite reflection feeding benchmarked, rollback-capable skill evolution |
| Proactive behaviour | Heartbeat wake-up | 3-tier gardener (1 min / 72 min / nightly sleep), gap scanner, inner thoughts, trust feedback loop |
| Cost tracking & budgets | Token and estimated-cost reporting (/usage, /status); no spend limits | Per-token spend tracking with daily/monthly limits that gate requests, and a live dashboard |
| Model routing | Swappable model plugins, chosen in config | 7 providers with health-aware failover; a complexity analyzer routes each request to the cheapest capable model |
| Local voice | — | Kokoro TTS + faster-whisper STT on-device, zero API cost; cloud fallback |
| Skill ecosystem | 100+ bundled, 3000+ on ClawHub | Full OpenClaw SKILL.md compatibility — community and ClawHub skills install and run unchanged |
| Channels | 25+ platforms, including iMessage and Teams | Telegram, web dashboard (REST + WebSocket), CLI. Discord, WhatsApp, Slack, Signal and Matrix adapters exist but are not wired up yet |
| Native apps | macOS / iOS / Android / Windows / Linux | — |
Skill rows say what they mean: ScallopBot does not compete on catalogue size — it runs OpenClaw’s catalogue. Skills in the OpenClaw SKILL.md format, including ClawHub installs, load unchanged.
LoCoMo: does the extra machinery earn its keep?
LoCoMo long-conversation QA, 1,049 items, same generation model (Moonshot kimi-k2.5), same embeddings (Ollama nomic-embed-text), same token-level F1 scoring on both sides.
| Category | ScallopBot | OpenClaw | Relative |
|---|---|---|---|
| Overall F1 | 0.48 | 0.38 | +26% |
| Multi-hop | 0.42 | 0.32 | +31% |
| Temporal | 0.34 | 0.26 | +31% |
| Adversarial | 0.97 | 0.77 | +26% |
Common questions
Does OpenClaw have memory consolidation?
Yes — since it added “dreaming”, which is on by default. A light, REM and deep sweep scores what you keep recalling and promotes the strongest items from daily notes into MEMORY.md. It is a promotion step: existing entries are kept byte-for-byte unless explicitly merged, and nothing is forgotten.
ScallopBot's consolidation rewrites memory instead: a nightly NREM pass replays fading and mid-strength memories, fuses duplicates, and clusters fragments across topic boundaries into coherent summaries; a REM pass hunts for non-obvious associations; a decay pass prunes what stopped being useful. This is the difference that shows up in the multi-hop LoCoMo category: 0.42 against OpenClaw's 0.32.
Are ScallopBot and OpenClaw the same project?
No — ScallopBot is a separate, independently published project (MIT) that is OpenClaw-compatible. It runs skills written in the OpenClaw SKILL.md format, including community skills from ClawHub, so switching keeps your skill set. What differs is the cognitive layer underneath: memory fusion and forgetting, self-reflection, and cost routing with budgets are where ScallopBot exists.
Which one should I pick?
OpenClaw is the larger ecosystem — more channels, native apps on every major platform, thousands of community skills. If the breadth of integrations is the deciding feature, it is the honest choice.
ScallopBot is the choice when you want memory that is merged, linked and pruned rather than only promoted — fusion, association, decay, temporal recall — measured on LoCoMo at F1 0.48 against OpenClaw's 0.38 under the same models and scoring — and to route across the providers you have keys for, with per-token spend limits. Many people run both ideas together by running their OpenClaw-format skills on ScallopBot.
How was the benchmark run?
Both systems answered the same 1,049 LoCoMo QA items (5 long conversations, 138 sessions), generated with the same model (Moonshot kimi-k2.5), embedded with the same model (Ollama nomic-embed-text), scored with token-level F1. These figures are not comparable with other vendors' published LoCoMo scores, which use different models and often a different metric entirely.
What does OpenClaw do better?
Ecosystem reach. Twenty-five-plus channels against three live ones, native apps for macOS, iOS, Android, Windows and Linux, and a bundled-plus-ClawHub skill library in the thousands. ScallopBot matches the skill library through format compatibility rather than matching its size, and has no native apps. The table above leaves those rows marked in OpenClaw's favour on purpose.
→ Full memory benchmark · Memory architecture · Cost comparison · OpenClaw on GitHub · ScallopBot on GitHub