ScallopBot vs. OpenClaw

The full comparison between running OpenClaw and running ScallopBot — the open-source personal AI assistant that consolidates memory in its sleep. Both self-host, both speak the OpenClaw skill format; the difference is the cognitive layer underneath. Every claim below is either benchmarked or cross-linked, and the rows OpenClaw wins stay in the table.

Both consolidate memory. They do it differently.

OpenClaw is an excellent skill-orchestration runtime: huge channel coverage, native apps, thousands of community skills. Both projects now consolidate memory in the background. OpenClaw’s “dreaming” promotes notes you keep recalling into a curated MEMORY.md. ScallopBot’s sleep-style consolidation rewrites memory itself: it fuses duplicates, merges fragments into new summaries, links related memories and forgets what stopped being useful — plus the reflection and cost-routing machinery around it.

NREM consolidation
Every night ScallopBot replays fading and mid-strength memories the way slow-wave sleep replays a day: duplicates are fused, fragments of the same story are clustered across topic boundaries into a coherent summary, and recurring facts are strengthened. OpenClaw’s dreaming promotes and annotates notes but keeps the originals as written; ScallopBot merges them into fewer, cleaner memories.
REM association
A second, high-noise pass wanders the memory graph looking for links that literal retrieval would never surface, and an LLM judge keeps only associations that are novel, plausible and useful. They are stored as typed edges, so recall can follow them later. This is where multi-hop questions gain: 0.42 against OpenClaw’s 0.32 on LoCoMo.
Retrieval that declines
Recall is BM25 + embeddings merged and optionally re-ranked by an LLM, then score-gated: if nothing clears the bar, nothing is injected. Adversarial questions — the ones with no answer — score 0.97 against 0.77 because the system says “I don’t know” instead of confabulating.
Cost as a first-class feature
Per-token spend tracking, daily and monthly budget limits that stop requests before they are sent, and automatic routing of each request to the cheapest capable model across the providers you have keys for, with health-aware failover. Estimated $0.05–0.10/day of model spend at a hundred messages a day — see the cost breakdown.
OpenClaw ships fast. The OpenClaw column reflects its public README and docs as of October 2026. If a row goes stale because OpenClaw added the feature, that is a bug in this page — open an issue or a PR against the repo and it will be corrected. Comparisons that quietly outlive their facts help nobody.

Feature comparison

FeatureOpenClawScallopBot
Memory consolidation“Dreaming” (on by default): light / REM / deep sweep promotes frequently-recalled notes from daily files into MEMORY.mdNightly sleep-style cycle: NREM fusion merges duplicates and clusters cross-topic fragments into new summaries; REM builds typed associations
Memory decay / forgettingRecency decay on search ranking (30-day half-life); notes are not archived or prunedUtility-based: activation decay with category half-lives, soft-archive then hard-prune
Memory retrievalVector + keyword hybrid, deterministic weighted rankingBM25 + semantic hybrid with optional LLM re-ranking, score-gated
Associative recall—Spreading activation over typed memory edges (REM phase)
Temporal queriesRecency weighting; temporal phrasing escalates the searchDates embedded in memory text + time-scope detection routes “what changed since” questions through time-aware retrieval
Self-reflection & evolution—Private composite reflection feeding benchmarked, rollback-capable skill evolution
Proactive behaviourHeartbeat wake-up3-tier gardener (1 min / 72 min / nightly sleep), gap scanner, inner thoughts, trust feedback loop
Cost tracking & budgetsToken and estimated-cost reporting (/usage, /status); no spend limitsPer-token spend tracking with daily/monthly limits that gate requests, and a live dashboard
Model routingSwappable model plugins, chosen in config7 providers with health-aware failover; a complexity analyzer routes each request to the cheapest capable model
Local voice—Kokoro TTS + faster-whisper STT on-device, zero API cost; cloud fallback
Skill ecosystem100+ bundled, 3000+ on ClawHubFull OpenClaw SKILL.md compatibility — community and ClawHub skills install and run unchanged
Channels25+ platforms, including iMessage and TeamsTelegram, web dashboard (REST + WebSocket), CLI. Discord, WhatsApp, Slack, Signal and Matrix adapters exist but are not wired up yet
Native appsmacOS / iOS / Android / Windows / Linux—

Skill rows say what they mean: ScallopBot does not compete on catalogue size — it runs OpenClaw’s catalogue. Skills in the OpenClaw SKILL.md format, including ClawHub installs, load unchanged.

LoCoMo: does the extra machinery earn its keep?

LoCoMo long-conversation QA, 1,049 items, same generation model (Moonshot kimi-k2.5), same embeddings (Ollama nomic-embed-text), same token-level F1 scoring on both sides.

CategoryScallopBotOpenClawRelative
Overall F10.480.38+26%
Multi-hop0.420.32+31%
Temporal0.340.26+31%
Adversarial0.970.77+26%
These numbers are not comparable with other vendors’ published LoCoMo scores. Mem0, Zep and others report under different models, retrieval budgets, and often a different metric (LLM-as-judge accuracy, not token F1). The only comparison in this table is the controlled one, and its OpenClaw arm is the configuration measured when the benchmark was published — OpenClaw’s memory has changed since, so a re-run may move these numbers. Methodology and per-category detail: the memory benchmark page and the memory architecture page.

Common questions

Does OpenClaw have memory consolidation?

Yes — since it added “dreaming”, which is on by default. A light, REM and deep sweep scores what you keep recalling and promotes the strongest items from daily notes into MEMORY.md. It is a promotion step: existing entries are kept byte-for-byte unless explicitly merged, and nothing is forgotten.

ScallopBot's consolidation rewrites memory instead: a nightly NREM pass replays fading and mid-strength memories, fuses duplicates, and clusters fragments across topic boundaries into coherent summaries; a REM pass hunts for non-obvious associations; a decay pass prunes what stopped being useful. This is the difference that shows up in the multi-hop LoCoMo category: 0.42 against OpenClaw's 0.32.

Are ScallopBot and OpenClaw the same project?

No — ScallopBot is a separate, independently published project (MIT) that is OpenClaw-compatible. It runs skills written in the OpenClaw SKILL.md format, including community skills from ClawHub, so switching keeps your skill set. What differs is the cognitive layer underneath: memory fusion and forgetting, self-reflection, and cost routing with budgets are where ScallopBot exists.

Which one should I pick?

OpenClaw is the larger ecosystem — more channels, native apps on every major platform, thousands of community skills. If the breadth of integrations is the deciding feature, it is the honest choice.

ScallopBot is the choice when you want memory that is merged, linked and pruned rather than only promoted — fusion, association, decay, temporal recall — measured on LoCoMo at F1 0.48 against OpenClaw's 0.38 under the same models and scoring — and to route across the providers you have keys for, with per-token spend limits. Many people run both ideas together by running their OpenClaw-format skills on ScallopBot.

How was the benchmark run?

Both systems answered the same 1,049 LoCoMo QA items (5 long conversations, 138 sessions), generated with the same model (Moonshot kimi-k2.5), embedded with the same model (Ollama nomic-embed-text), scored with token-level F1. These figures are not comparable with other vendors' published LoCoMo scores, which use different models and often a different metric entirely.

What does OpenClaw do better?

Ecosystem reach. Twenty-five-plus channels against three live ones, native apps for macOS, iOS, Android, Windows and Linux, and a bundled-plus-ClawHub skill library in the thousands. ScallopBot matches the skill library through format compatibility rather than matching its size, and has no native apps. The table above leaves those rows marked in OpenClaw's favour on purpose.

→ Full memory benchmark · Memory architecture · Cost comparison · OpenClaw on GitHub · ScallopBot on GitHub