ai-output-validation
Validates, parses, and sanitizes AI-generated outputs before they reach end users or downstream systems. Structured output enforcement, schema validation, and…
Patterns for Retrieval-Augmented Generation (RAG) and agent memory systems. Retrieves only relevant context, prevents context bloat, and maintains coherent state across sessions.
$ npx -y skills add DevelopersGlobal/ai-agent-skills --skill rag-and-memory --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/rag-and-memoryContext preview
The summary Claude sees to decide when to auto-load this skill.
Patterns for Retrieval-Augmented Generation (RAG) and agent memory systems. Retrieves only relevant context, prevents context bloat, and maintains coherent state across sessions.
name: rag-and-memory description: Patterns for Retrieval-Augmented Generation (RAG) and agent memory systems. Retrieves only relevant context, prevents context bloat, and maintains coherent state across sessions. category: build applies-to: [claude, gemini, cursor, copilot, any] version: 1.0.0
RAG and memory systems are how AI agents work with knowledge that exceeds their context window. Done well: agents give accurate, grounded answers. Done poorly: context overflow, hallucination from stale retrieval, and performance degradation.
This skill covers the design principles and failure modes of RAG and memory architectures for production AI systems.
1. Identify what the agent needs to remember:
2. Match the memory type to the store:
| Memory Type | Recommended Store | |------------|------------------| | In-session facts | Context window (summarized) | | User preferences | Key-value store (Redis, DynamoDB) | | Document corpus | Vector database (Pinecone, Weaviate, pgvector) | | Long-term facts | Structured DB + caching |
**Verify:** Each type of information the agent needs has a defined storage mechanism.
3. **Chunking strategy**: Break documents into chunks at semantic boundaries (paragraphs, sections) — not arbitrary character counts. 4. **Embedding model**: Match the embedding model to your query type. Use the same model for indexing and retrieval. 5. **Retrieval**: Retrieve top-K most semantically similar chunks. K = 3–7 is usually optimal. 6. **Re-ranking**: After retrieval, re-rank by relevance using a cross-encoder. Top K becomes top 3–5 for the prompt. 7. **Context injection**: Inject retrieved chunks into the prompt with clear source citations.
**Verify:** Retrieved chunks are genuinely relevant to the query before injecting into context.
8. **Summarize, don't accumulate**: For long sessions, summarize previous turns rather than appending them indefinitely. 9. **Retrieve, don't pre-load**: Only load context relevant to the current query. Don't pre-load everything. 10. **Set context budgets**: Define maximum token allocations for: system prompt, retrieved context, conversation history, user message. 11. **Compress before injecting**: Summarize long retrieved documents to extract the relevant portion only.
**Verify:** Total prompt length is within model limits with buffer. Retrieved context is relevant to current query.
12. If retrieval returns no relevant results: say so — do not hallucinate an answer. 13. If retrieved documents are outdated: surface the document date to the user. 14. If confidence is low: present the retrieved source and let the user evaluate. 15. Design for "no relevant information found" as a first-class outcome.
**Verify:** System has defined behavior for failed/empty retrieval.
16. Track retrieval quality:
17. Track answer quality: Use RAGAS or similar evaluation framework. 18. Monitor: context length per query, retrieval latency, hallucination rate.
**Verify:** Baseline metrics established. Retrieval precision > 80%.
| Excuse | Rebuttal | |--------|----------| | "Let's just put everything in the context" | Context bloat degrades quality and costs money. Retrieve what's needed. | | "The model knows this from training" | Training knowledge is stale. Use RAG for current information. | | "Vector search is good enough without re-ranking" | Re-ranking improves precision significantly. It's a small cost for large quality gain. | | "We'll fix retrieval quality later" | Poor retrieval quality compounds into poor answer quality. Fix it now. |
AI agent skills for production grade applications
Validates, parses, and sanitizes AI-generated outputs before they reach end users or downstream systems. Structured output enforcement, schema validation, and…
Design stable, versioned, self-documenting APIs. Easy to use correctly, hard to use incorrectly. Apply Hyrum's Law from day one.
Automated quality gates from commit to production. Every merge to main is potentially shippable. No manual steps in the deployment path.
Get layered, context-aware explanations of unfamiliar code. Understand what it does, why it was written that way, and how to work with it safely.
Structured code review focusing on correctness, security, and maintainability. Correctness before style. Every reviewer comment must be actionable.
Load minimum necessary context into agent context windows. Prevents token bloat, reduces cost, and improves focus. Only load what the current task needs.