audio-generation
Guide to audio generation and understanding in MassGen. Covers text-to-speech, music, sound…
This skill provides semantic search capabilities using embedding-based similarity matching for code and text. Enables meaning-based search beyond keyword matching, with optional document parsing (PDF, DOCX, PPTX) support.
$ npx -y skills add massgen/massgen --skill semtools --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/semtoolsContext preview
The summary Claude sees to decide when to auto-load this skill.
This skill provides semantic search capabilities using embedding-based similarity matching for code and text. Enables meaning-based search beyond keyword matching, with optional document parsing (PDF, DOCX, PPTX) support.
name: semtools description: This skill provides semantic search capabilities using embedding-based similarity matching for code and text. Enables meaning-based search beyond keyword matching, with optional document parsing (PDF, DOCX, PPTX) support. license: MIT
Perform semantic (meaning-based) search across code and documents using embedding-based similarity matching.
The semtools skill provides access to Semtools, a high-performance Rust-based CLI for semantic search and document processing. Unlike traditional text search (ripgrep) which matches exact strings, or structural search (ast-grep) which matches syntax patterns, semtools understands **semantic meaning** through embeddings.
**Key capabilities:**
1. **Semantic Search**: Find code/text by meaning, not just keywords 2. **Workspace Management**: Index large codebases for fast repeated searches 3. **Document Parsing**: Convert PDFs, DOCX, PPTX to searchable text (requires API key)
Semtools excels at **discovery** - finding relevant code when you don't know the exact keywords, function names, or syntax patterns.
Use the semtools skill when you need meaning-based search:
**Semantic Code Discovery:**
**Documentation & Knowledge:**
**Use Cases:**
**Choose semtools over file-search (ripgrep/ast-grep) when:**
**Still use file-search when:**
Semtools provides three CLI commands you can use via `execute_command`:
**All commands work out-of-the-box** in your execution environment. Document parsing requires the LLAMA_CLOUD_API_KEY environment variable to be set.
Find files and code sections by semantic meaning:
# Basic semantic search search "authentication logic" src/ # Search with more context (5 lines before/after) search "error handling" --n-lines 5 src/ # Get more results (default: 3) search "database queries" --top-k 10 src/ # Control similarity threshold (0.0-1.0, lower = more lenient) search "API endpoints" --max-distance 0.4 src/
**Parameters:**
**Output format:**
Match 1 (similarity: 0.12)
File: src/auth/handlers.py
Lines: 42-47
----
def authenticate_user(username: str, password: str) -> Optional[User]:
"""Authenticate user credentials against database."""
user = get_user_by_username(username)
if user and verify_password(password, user.password_hash):
return user
return None
----
Match 2 (similarity: 0.18)
File: src/middleware/auth.py
...For large codebases, create workspaces to cache embeddings and enable fast repeated searches:
# Create/activate workspace workspace use my-project # Set workspace via environment variable export SEMTOOLS_WORKSPACE=my-project # Index files in workspace (workspace auto-detected from env var) search "query" src/ # Check workspace status workspace status # Clean up old workspaces workspace prune
**Benefits:**
**When to use workspaces:**
Convert documents to searchable markdown (requires LlamaParse API key):
# Parse PDFs to markdown parse research_papers/*.pdf # Parse Word documents parse reports/*.docx # Parse presentations parse slides/*.pptx # Parse and pipe to search parse docs/*.pdf | xargs search "neural networks"
**Supported formats:**
**Configuration:**
# Via environment variable
export LLAMA_CLOUD_API_KEY="llx-..."
# Via config file
cat > ~/.parse_config.json << EOF
{
"api_key": "llx-...",
"max_concurrent_requests": 10,
"timeout_seconds": 3600
}
EOF**Important:** Document parsing is **optional**. Semantic search works without it.
When you know what you're looking for conceptually but not by name:
# Step 1: Broad semantic search search "rate limiting implementation" src/ # Step 2: Review results, refine query sea
🚀 MassGen is an open-source multi-agent scaling system that runs in your terminal, autonomously orchestrating frontier models and agents to collaborate, reason, and produce high-quality results. | Join us on Discord: discord.massgen.ai
Guide to audio generation and understanding in MassGen. Covers text-to-speech, music, sound…
Complete guide for integrating a new LLM backend into MassGen. Use when adding a new provider…
Guide for creating evolving skills - detailed workflow plans that capture what you'll do,…
This skill should be used when agents need to search codebases for text patterns or…
Guide to image generation and editing in MassGen. Use when creating images, editing existing…
Guide for creating properly structured YAML configuration files for MassGen. This skill…