Skip to content
Development
Skill

/cloudflare-workers-ai

Cloudflare Workers AI for serverless GPU inference. Use for LLMs, text/image generation, embeddings, or encountering AI_ERROR, rate limits, token exceeded errors.

From plugin
secondsky-claude-skills
219183 skills42 agents62 commands2 MCP
Install
$ npx -y skills add secondsky/claude-skills --skill cloudflare-workers-ai --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/cloudflare-workers-ai

Context preview

The summary Claude sees to decide when to auto-load this skill.

Cloudflare Workers AI for serverless GPU inference. Use for LLMs, text/image generation, embeddings, or encountering AI_ERROR, rate limits, token exceeded errors.

SKILL.md

cloudflare-workers-ai.SKILL.md
name: cloudflare-workers-ai
description: "Cloudflare Workers AI for serverless GPU inference. Use for LLMs, text/image generation, embeddings, or encountering AI_ERROR, rate limits, token exceeded errors."

metadata:
  keywords:
    - workers ai
    - cloudflare ai
    - ai bindings
    - llm workers
    - "@cf/meta/llama"
    - workers ai models
    - ai inference
    - cloudflare llm
    - ai streaming
    - text generation ai
    - ai embeddings
    - image generation ai
    - workers ai rag
    - ai gateway
    - llama workers
    - flux image generation
    - stable diffusion workers
    - vision models ai
    - ai chat completion
    - AI_ERROR
    - rate limit ai
    - model not found
    - token limit exceeded
    - neurons exceeded
    - ai quota exceeded
    - streaming failed
    - model unavailable
    - workers ai hono
    - ai gateway workers
    - vercel ai sdk workers
    - openai compatible workers
    - workers ai vectorize

license: MIT

Cloudflare Workers AI - Complete Reference

Production-ready knowledge domain for building AI-powered applications with Cloudflare Workers AI.

**Status**: Production Ready ✅ **Last Updated**: 2025-11-21 **Dependencies**: cloudflare-worker-base (for Worker setup) **Latest Versions**: wrangler@4.81.0, @cloudflare/workers-types@4.20260408.0

---

Table of Contents

1. [Quick Start (5 minutes)](#quick-start-5-minutes) 2. [Workers AI API Reference](#workers-ai-api-reference) 3. [Model Selection Guide](#model-selection-guide) 4. [Common Patterns](#common-patterns) 5. [AI Gateway Integration](#ai-gateway-integration) 6. [Rate Limits & Pricing](#rate-limits--pricing) 7. [Production Checklist](#production-checklist)

---

Quick Start (5 minutes)

1. Add AI Binding

**wrangler.jsonc:**

{
  "ai": {
    "binding": "AI"
  }
}

2. Run Your First Model

export interface Env {
  AI: Ai;
}

export default {
  async fetch(request: Request, env: Env): Promise<Response> {
    const response = await env.AI.run('@cf/meta/llama-3.1-8b-instruct', {
      prompt: 'What is Cloudflare?',
    });

    return Response.json(response);
  },
};

3. Add Streaming (Recommended)

const stream = await env.AI.run('@cf/meta/llama-3.1-8b-instruct', {
  messages: [{ role: 'user', content: 'Tell me a story' }],
  stream: true, // Always use streaming for text generation!
});

return new Response(stream, {
  headers: { 'content-type': 'text/event-stream' },
});

**Why streaming?**

  • Prevents buffering large responses in memory
  • Faster time-to-first-token
  • Better user experience for long-form content
  • Avoids Worker timeout issues

---

Workers AI API Reference

Core API: `env.AI.run()`

const response = await env.AI.run(model, inputs, options?);

| Parameter | Type | Description | |-----------|------|-------------| | `model` | string | Model ID (e.g., `@cf/meta/llama-3.1-8b-instruct`) | | `inputs` | object | Model-specific inputs (see model type below) | | `options.gateway.id` | string | AI Gateway ID for caching/logging | | `options.gateway.skipCache` | boolean | Skip AI Gateway cache |

**Returns**: `Promise<ModelOutput>` (non-streaming) or `ReadableStream` (streaming)

Input Types by Model Category

| Category | Key Inputs | Output | |----------|------------|--------| | **Text Generation** | `messages[]`, `stream`, `max_tokens`, `temperature` | `{ response: string }` | | **Embeddings** | `text: string \| string[]` | `{ data: number[][], shape: number[] }` | | **Image Generation** | `prompt`, `num_steps`, `guidance` | Binary PNG | | **Vision** | `messages[].content[].image_url` | `{ response: string }` |

📄 **Full model details**: Load `references/models-catalog.md` for complete model list, parameters, and rate limits.

---

Model Selection Guide

Text Generation (LLMs)

| Model | Best For | Rate Limit | Size | |-------|----------|------------|------| | `@cf/meta/llama-3.1-8b-instruct` | General purpose, fast | 300/min | 8B | | `@cf/meta/llama-3.2-1b-instruct` | Ultra-fast, simple tasks | 300/min | 1B | | `@cf/qwen/qwen1.5-14b-chat-awq` | High quality, complex reasoning | 150/min | 14B | | `@cf/deepseek-ai/deepseek-r1-distill-qwen-32b` | Coding, technical content | 300/min | 32B | | `@hf/thebloke/mistral-7b-instruct-v0.1-awq` | Fast, efficient | 400/min | 7B |

Text Embeddings

| Model | Dimensions | Best For | Rate Limit | |-------|-----------|----------|------------| | `@cf/baai/bge-base-en-v1.5` | 768 | General purpose RAG | 3000/min | | `@cf/baai/bge-large-en-v1.5` | 1024 | High accuracy search | 1500/min | | `@cf/baai/bge-small-en-v1.5` | 384 | Fast, low storage | 3000/min |

Image Generation

| Model | Best For | Rate Limit | Speed | |-------|----------|------------|-------| | `@cf/black-forest-labs/flux-1-schnell` | High quality, photorealistic | 720/min | Fast | | `@cf/stabilityai/stable-diffusion-xl-base-1.0` | General purpose | 720/min | Medium | | `@cf/lykon/dreamshaper-8-lcm` | Artistic, stylized | 720/min | Fast |

Vision Models

| Model | Best For | Rate Limit | |-------|----------|------------| | `@cf/meta/llama-3.2-11b-vision-instruct` | Image understanding | 720/min | | `@cf/unum/uform-gen2-qwen-500m` | Fast image captioning | 720/min |

---

Common Patterns

Pattern 1: Chat with Streaming

app.post('/chat', async (c) => {
  const { messages } = await c.req.json<{ messages: Array<{ role: string; content: string }> }>();
  const stream = await c.env.AI.run('@cf/meta/llama-3.1-8b-instruct', { messages, stream: true });
  return new Response(stream, { headers: { 'content-type': 'text/event-stream' } });
});

Pattern 2: RAG (Retrieval Augmented Generation)

// 1. Generate embedding for query
const embeddings = await env.AI.run('@cf/baai/bge-base-en-v1.5', { text: [userQuery] });
// 2. Search Vectorize
const matches = await env.VECTORIZE.query(embeddings.data[0], { topK: 3 });
// 3. Build contex
Read more
Ships withsecondsky-claude-skills

145 production-ready skills for Claude Code CLI 🔌 Platform / Harness Support These plugins ship as Claude Code marketplace plugins (.claude-plugin/ manifests) and Codex CLI plugins (.codex-plugin/ manifests).

Get the whole plugin

Other skills on secondsky-claude-skills.