ai-infrastructure-hugg…
Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation,…
Qdrant vector database -- collection management, point operations, payload filtering, named vectors, quantization, recommendations, snapshots
$ npx -y skills add agents-inc/skills --skill api-vector-db-qdrant --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/api-vector-db-qdrantContext preview
The summary Claude sees to decide when to auto-load this skill.
Qdrant vector database -- collection management, point operations, payload filtering, named vectors, quantization, recommendations, snapshots
name: api-vector-db-qdrant description: Qdrant vector database -- collection management, point operations, payload filtering, named vectors, quantization, recommendations, snapshots
> **Quick Guide:** Use `@qdrant/js-client-rest` (v1.17.x) for high-performance vector search. Collections define vector dimensions and distance metrics upfront -- mismatches cause silent failures. Use `must`/`should`/`must_not` filter clauses with payload conditions (not Pinecone-style `$eq`/`$and`). Payload indexes are optional but critical for filter performance at scale -- create them explicitly with `createPayloadIndex()`. Named vectors let you store multiple embeddings per point (e.g., title + content). Quantization (scalar/binary/product) trades accuracy for memory and speed. The `query()` method is the universal search endpoint -- prefer it over the older `search()` method.
---
<critical_requirements>
> **All code must follow project conventions in CLAUDE.md** (kebab-case, named exports, import ordering, `import type`, named constants)
**(You MUST create payload indexes with `createPayloadIndex()` for any field used in filters -- unindexed fields cause full scans that degrade linearly with collection size)**
**(You MUST use `must`/`should`/`must_not` filter syntax -- Qdrant does NOT use `$eq`/`$and`/`$or` operators like Pinecone)**
**(You MUST match vector dimensions exactly between embedding model output and collection config -- dimension mismatches cause silent upsert failures or corrupt search results)**
**(You MUST set `wait: true` on writes when subsequent reads depend on the data -- Qdrant writes are asynchronous by default and may not be immediately visible)**
</critical_requirements>
---
**Additional resources:**
---
**Auto-detection:** Qdrant, QdrantClient, @qdrant/js-client-rest, createCollection, upsert, query, scroll, recommend, setPayload, createPayloadIndex, must, should, must_not, payload, named vectors, quantization, vector database, similarity search, semantic search, RAG retrieval, embedding search
**When to use:**
**Key patterns covered:**
**When NOT to use:**
---
<philosophy>
Qdrant is a **high-performance open-source vector database** built in Rust, designed for filtered similarity search at scale. The core principle: **store vectors with rich payloads, search by similarity, filter by payload conditions.**
**Core principles:**
1. **Payload is first-class** -- Unlike databases that treat metadata as secondary, Qdrant's payload system supports complex nested JSON, multiple data types, and granular indexing. Use payloads for filtering, not just annotation. 2. **Index what you filter** -- Payload indexes are not automatic. Create explicit indexes on fields used in filters via `createPayloadIndex()`. Without indexes, filters cause full collection scans. 3. **Named vectors for multi-modal** -- A single point can hold multiple named vectors (e.g., title embedding + content embedding). Search targets a specific named vector. This avoids duplicating payloads across collections. 4. **Quantization for scale** -- Scalar (4x compression), binary (32x), and product quantization trade accuracy for memory savings. Configure at collection or per-vector level. Use `always_ram: true` to keep quantized vectors in memory for speed. 5. **Writes are async by default** -- Upserts return before data is persisted to all replicas. Set `wait: true` when immediate consistency matters (e.g., read-after-write flows).
</philosophy>
---
<patterns>
Create a QdrantClient connected to a local instance or Qdrant Cloud. See [examples/core.md](examples/core.md) for full examples.
// Good Example
import { QdrantClient } from "@qdrant/js-client-rest";
function createQdrantClient(): QdrantClient {
const url = process.env.QDRANT_URL;
const apiKey = process.env.QDRANT_API_KEY;
if (!url) {
throw new Error("QDRANT_URL environment variable is required");
}
return new QdrantClient({ url, apiKey });
}
export { createQdrantClient };**Why good:** URL and API key from environment, validation before construction, n
The official skills marketplace for Agents Inc. 150+ skills covering everything from React and Prisma to Redis, ElevenLabs, and infrastructure tooling. Pick the skills that match your stack and install them via Claude Code. Need more control?
Repo: agents-inc/skills
Hugging Face Inference SDK patterns for TypeScript/Node.js — InferenceClient setup, chat completion, text generation, streaming, embeddings, image generation,…
LiteLLM proxy server setup, TypeScript client patterns via OpenAI SDK, model routing, fallbacks, load balancing, spend tracking, virtual keys, and production…
Serverless GPU compute platform for AI model deployment — web endpoints, GPU functions, model serving, and TypeScript client patterns
Local LLM inference with the Ollama JavaScript client -- chat, streaming, tool calling, vision, embeddings, structured output, model management, and…
Replicate SDK patterns for TypeScript/Node.js -- client setup, predictions, streaming, webhooks, file handling, model versioning, deployments, and training
Together AI SDK patterns for TypeScript — client setup, chat completions, streaming, structured output, function calling, embeddings, image generation,…