/slm-recall
Search and retrieve facts, decisions, and past context from SuperLocalMemory. Use when the user asks to recall, find, search, or "what did we decide/say about X". Triggers multi-channel semantic retrieval with reranking; always call before storing anything new.
$ npx -y skills add qualixar/superlocalmemory --skill slm-recall --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/slm-recall
Context preview
The summary Claude sees to decide when to auto-load this skill.
Search and retrieve facts, decisions, and past context from SuperLocalMemory. Use when the user asks to recall, find, search, or "what did we decide/say about X". Triggers multi-channel semantic retrieval with reranking; always call before storing anything new.
SKILL.md
slm-recall.SKILL.mdname: slm-recall
description: Search and retrieve facts, decisions, and past context from SuperLocalMemory. Use when the user asks to recall, find, search, or "what did we decide/say about X". Triggers multi-channel semantic retrieval with reranking; always call before storing anything new.
when_to_use: |
- "What did we decide about X?"
- "Recall anything about Y"
- "Do we have context on the Z feature?"
- "Find stored information about authentication / the database / error handling"
- "Search for what I said about Y"
- Automatically before any non-trivial task, to surface prior context
allowed-tools: recall, search, fetch, list_recent, Bash
slm-recall — Search & Retrieve Memory
Retrieve stored facts, decisions, and past context from SuperLocalMemory using multi-channel retrieval. The golden rule: **recall before you remember**.
---
When to use recall vs search vs fetch vs list_recent
| Situation | Tool | |-----------|------| | Conceptual or paraphrase query ("what did we agree on for auth?") | `recall` — full multi-channel retrieval + rerank | | Exact keyword match needed ("find facts containing BM25") | `search` — FTS5 BM25 only, lower latency | | You have a specific `fact_id` from a prior result | `fetch` — exact lookup, full detail | | Browse newest entries without a query | `list_recent` |
Use `recall` as the default. `search` is a fallback for zero-result recall on a known exact term. `fetch` is for when you already know the ID.
---
Recall-before-remember discipline
Before storing anything new, always call `recall` first. If a near-duplicate fact already exists, call `update_memory(fact_id, content)` to refine it rather than creating a duplicate. Duplicates degrade retrieval quality for every future session.
---
MCP-first workflow
1. Standard recall
recall(
query="authentication strategy decision",
limit=20, # default 20; reduce to 5 for quick pre-task checks
session_id="<sid>", # pass the session_id returned by session_init
fast=False, # default False; True enables faster, reduced-channel retrieval
)
Real response shape (`--json` equivalent):
{
"success": true,
"results": [
{
"fact_id": "f8a2bc91",
"content": "Decided to use JWT with 1h expiry for API auth (2026-06-10)",
"score": 0.87,
"confidence": 0.91,
"trust_score": 0.84,
"fact_type": "decision",
"channel_scores": {
"semantic": 0.88,
"lexical": 0.61,
"temporal": 0.72,
"structural": 0.55
}
}
],
"count": 1,
"query_type": "semantic",
"channel_weights": {
"semantic": 0.4,
"lexical": 0.2,
"temporal": 0.2,
"structural": 0.2
},
"retrieval_time_ms": 134,
"no_confident_match": false
}**Refine on low confidence.** `recall` returns confidence signals with every result. If `no_confident_match` is `true` (or `answer_confidence` is low / `abstained` is `true`), do NOT invent a memory — rewrite the query into 1–3 more specific sub-queries (split multi-hop questions; try entity names, synonyms, or broader phrasing) and call `recall` again before concluding nothing was found. A confident match → use it directly. SLM returns fast local results (~1–2s, no server-side LLM round on the hot path) and lets you, the calling model, drive this refinement.
2. Passing session_id
Pass the `session_id` returned by `session_init`. It threads engagement signals through to the ranker so each recall contributes to improving retrieval for your project over time. Omitting it degrades the learning loop — recall works correctly, but feedback is not attributed to the session.
3. Fast mode
Use `fast=True` for pre-tool-call checks where sub-second response matters. This enables a faster, reduced-channel mode. Core semantic and keyword channels always run; additional graph and contextual channels are skipped.
recall(query="rate limiting approach", limit=5, session_id="<sid>", fast=True)
4. Keyword fallback via search
When `recall` returns zero results on a specific term, try `search`:
search(query="BM25 indexing", limit=10, profile_id="")
`profile_id=""` uses the active profile. Response has `success`, `results`, and `count` but no `channel_scores` or `query_type`.
5. Pull full detail for a known fact
fetch(fact_ids="f8a2bc91,d4c1e203")
Returns the full record for each ID: `entities`, `lifecycle`, `access_count`, `importance`, `observation_date`, `referenced_date`. Use this when the recall summary (120-char truncation in `list_recent`) is not enough.
6. Browse recent memories
list_recent(limit=20, profile_id="")
Returns facts newest-first. Content is truncated to 120 chars. Use `fetch` once you have the `fact_id` for full content.
---
How multi-channel retrieval works
`recall` runs multiple candidate producers in parallel — semantic vector similarity, keyword matching, temporal recency weighting, and contextual graph channels — then fuses and reranks the combined results, with an optional entity-graph score enhancement. The `channel_weights` field in the response shows how each channel contributed for that query. Weights adapt over time based on engagement signals attributed via `session_id`.
To inspect per-channel scores for a real query against your own data:
slm trace "<query>" [--limit N] [--json]
No benchmark numbers are cited here; performance is workload-dependent.
---
CLI fallback (when MCP is unavailable)
# Multi-channel semantic recall
slm recall "<query>" [--limit N] [--fast] [--json]
# Opt into shared/global facts for one query (v3.6.15 — off by default)
slm recall "<query>" --include-global --include-shared
# Keyword/FTS5 search (alias: slm search)
slm search "<query>" [--limit N] [--json]
# Per-channel score breakdown
slm trace "<query>" [--limit N] [--json]
# Browse recent memories
slm list [--limit N] [--json]
Flags verified
Read more
name: slm-recall description: Search and retrieve facts, decisions, and past context from SuperLocalMemory. Use when the user asks to recall, find, search, or "what did we decide/say about X". Triggers multi-channel semantic retrieval with reranking; always call before storing anything new. when_to_use: | - "What did we decide about X?" - "Recall anything about Y" - "Do we have context on the Z feature?" - "Find stored information about authentication / the database / error handling" - "Search for what I said about Y" - Automatically before any non-trivial task, to surface prior context allowed-tools: recall, search, fetch, list_recent, Bash
slm-recall — Search & Retrieve Memory
Retrieve stored facts, decisions, and past context from SuperLocalMemory using multi-channel retrieval. The golden rule: **recall before you remember**.
---
When to use recall vs search vs fetch vs list_recent
| Situation | Tool | |-----------|------| | Conceptual or paraphrase query ("what did we agree on for auth?") | `recall` — full multi-channel retrieval + rerank | | Exact keyword match needed ("find facts containing BM25") | `search` — FTS5 BM25 only, lower latency | | You have a specific `fact_id` from a prior result | `fetch` — exact lookup, full detail | | Browse newest entries without a query | `list_recent` |
Use `recall` as the default. `search` is a fallback for zero-result recall on a known exact term. `fetch` is for when you already know the ID.
---
Recall-before-remember discipline
Before storing anything new, always call `recall` first. If a near-duplicate fact already exists, call `update_memory(fact_id, content)` to refine it rather than creating a duplicate. Duplicates degrade retrieval quality for every future session.
---
MCP-first workflow
1. Standard recall
recall( query="authentication strategy decision", limit=20, # default 20; reduce to 5 for quick pre-task checks session_id="<sid>", # pass the session_id returned by session_init fast=False, # default False; True enables faster, reduced-channel retrieval )
Real response shape (`--json` equivalent):
{
"success": true,
"results": [
{
"fact_id": "f8a2bc91",
"content": "Decided to use JWT with 1h expiry for API auth (2026-06-10)",
"score": 0.87,
"confidence": 0.91,
"trust_score": 0.84,
"fact_type": "decision",
"channel_scores": {
"semantic": 0.88,
"lexical": 0.61,
"temporal": 0.72,
"structural": 0.55
}
}
],
"count": 1,
"query_type": "semantic",
"channel_weights": {
"semantic": 0.4,
"lexical": 0.2,
"temporal": 0.2,
"structural": 0.2
},
"retrieval_time_ms": 134,
"no_confident_match": false
}**Refine on low confidence.** `recall` returns confidence signals with every result. If `no_confident_match` is `true` (or `answer_confidence` is low / `abstained` is `true`), do NOT invent a memory — rewrite the query into 1–3 more specific sub-queries (split multi-hop questions; try entity names, synonyms, or broader phrasing) and call `recall` again before concluding nothing was found. A confident match → use it directly. SLM returns fast local results (~1–2s, no server-side LLM round on the hot path) and lets you, the calling model, drive this refinement.
2. Passing session_id
Pass the `session_id` returned by `session_init`. It threads engagement signals through to the ranker so each recall contributes to improving retrieval for your project over time. Omitting it degrades the learning loop — recall works correctly, but feedback is not attributed to the session.
3. Fast mode
Use `fast=True` for pre-tool-call checks where sub-second response matters. This enables a faster, reduced-channel mode. Core semantic and keyword channels always run; additional graph and contextual channels are skipped.
recall(query="rate limiting approach", limit=5, session_id="<sid>", fast=True)
4. Keyword fallback via search
When `recall` returns zero results on a specific term, try `search`:
search(query="BM25 indexing", limit=10, profile_id="")
`profile_id=""` uses the active profile. Response has `success`, `results`, and `count` but no `channel_scores` or `query_type`.
5. Pull full detail for a known fact
fetch(fact_ids="f8a2bc91,d4c1e203")
Returns the full record for each ID: `entities`, `lifecycle`, `access_count`, `importance`, `observation_date`, `referenced_date`. Use this when the recall summary (120-char truncation in `list_recent`) is not enough.
6. Browse recent memories
list_recent(limit=20, profile_id="")
Returns facts newest-first. Content is truncated to 120 chars. Use `fetch` once you have the `fact_id` for full content.
---
How multi-channel retrieval works
`recall` runs multiple candidate producers in parallel — semantic vector similarity, keyword matching, temporal recency weighting, and contextual graph channels — then fuses and reranks the combined results, with an optional entity-graph score enhancement. The `channel_weights` field in the response shows how each channel contributed for that query. Weights adapt over time based on engagement signals attributed via `session_id`.
To inspect per-channel scores for a real query against your own data:
slm trace "<query>" [--limit N] [--json]
No benchmark numbers are cited here; performance is workload-dependent.
---
CLI fallback (when MCP is unavailable)
# Multi-channel semantic recall slm recall "<query>" [--limit N] [--fast] [--json] # Opt into shared/global facts for one query (v3.6.15 — off by default) slm recall "<query>" --include-global --include-shared # Keyword/FTS5 search (alias: slm search) slm search "<query>" [--limit N] [--json] # Per-channel score breakdown slm trace "<query>" [--limit N] [--json] # Browse recent memories slm list [--limit N] [--json]
Flags verified
World's first local-only AI memory to break 74% retrieval and 60% zero-LLM on LoCoMo. No cloud, no APIs, no data leaves your machine. Additionally, mode C (LLM/Cloud) - 87.7% LoCoMo. Research-backed. arXiv: 2603.14588
Repo: qualixar/superlocalmemory
Other skills on superlocalmemory.
- /slm-cache
KV cache for repeated reads — call slm_cache_get(key) first; on a miss do the expensive operation then slm_cache_set(key, value, ttl_seconds) to store it; on a hit use the returned value directly; always fail-open (hit:false on any error, never raises); saves tokens when the
Open skill - /slm-compress
Compress large text, tool output, or transcripts to reduce context-window usage while keeping the full 1M window intact — call slm_compress(content, mode, reversible, ttl_seconds) to shrink content; if the result is lossy a ccr_id is returned so you can call slm_retrieve(ccr_id)
Open skill - /slm-governance
Enterprise compliance and governed workspace behavior for SuperLocalMemory. Covers role-based access (admin/member/viewer), retention policies, audit trail, GDPR data export/erase, and how agents must behave when operating under workspace governance. Requires power MCP profile
Open skill - /slm-graph
Index and query a codebase as a structural graph — build the code graph, trace blast radius of a change, find callers/callees/inheritors, semantic code search by meaning, assemble PR review context, and detect what changed since last index. Use when the user asks how code
Open skill - /slm-loop
Run gate-verified bounded loops with SuperLocalMemory as the durable ledger. Use when a task has a checkable acceptance condition (tests, schema, lint, reconciliation) and you must iterate until an INDEPENDENT gate passes — never stopping just because the agent believes it is
Open skill - /slm-mesh
Cross-session peer coordination via the SLM mesh network. Lets multiple AI agent sessions on the same machine discover each other, send messages, share lightweight state, and lock files to avoid conflicts. Requires full, power, or mesh MCP profile. All 8 tools are MCP-only —
Open skill

