Skip to content
Automation
Command

/gen

Generate an image with ComfyUI from a text prompt

From plugin
comfyui-mcp
52211 skills4 agents11 commands
Install
$ npx -y skills add artokun/comfyui-mcp --agent claude-code

How it fires

How this command gets triggered: by you, by Claude, or both.

  • Fires itselfClaude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/gen

Context preview

What this command does when you run it.

Generate an image with ComfyUI from a text prompt

Command definition

gen.md
description: Generate an image with ComfyUI from a text prompt
argument-hint: A text description of the image to generate

/comfy-gen — Generate an Image

The user wants to generate an image using ComfyUI. Their prompt is provided as the argument to this command.

Instructions

1. **Parse the user's prompt.** The argument after `/comfy-gen` is the image description: $ARGUMENTS

If no argument was provided, ask the user what they'd like to generate.

2. **Check available models.** Use the `list_local_models` tool with `model_type: "checkpoints"` to see what checkpoints are installed.

3. **If no checkpoints are installed — acquire one automatically.**

Do NOT ask the user to download a model. Instead, find and download one yourself:

**Option A — CivitAI (preferred if CIVITAI_API_TOKEN is set):**

  • Use WebFetch to search CivitAI's REST API: `https://civitai.com/api/v1/models?query=SDXL&types=Checkpoint&sort=Most+Downloaded&limit=5`
  • Pick the top result that matches the user's needs (SDXL for general, Flux for quality, etc.)
  • Get the download URL from the model version: `https://civitai.com/api/download/models/{modelVersionId}?token={CIVITAI_API_TOKEN}`
  • Use `download_model` with `target_subfolder: "checkpoints"` to download it

**Option B — HuggingFace (default):**

  • Use `download_model` `action:"search"` to find a suitable checkpoint (e.g., query "SDXL" or "stable diffusion xl")
  • Find the direct `.safetensors` download URL from the model page (typically `https://huggingface.co/{repo}/resolve/main/{filename}`)
  • Use `download_model` with `target_subfolder: "checkpoints"` to download it

**Sensible defaults by category:**

  • General purpose: SDXL base (`stabilityai/stable-diffusion-xl-base-1.0`)
  • Fast generation: SDXL Turbo or Lightning
  • High quality: Flux.1 schnell or dev (if user has enough VRAM)
  • Anime: Anything-XL or similar

Tell the user what you're downloading and why. Large checkpoints can be 2-7 GB.

4. **Create the workflow.** Use the `create_workflow` tool with:

  • `template`: `"txt2img"`
  • `params`: Include `positive_prompt` from the user's input. Set `checkpoint` to the model filename. Use sensible defaults (1024x1024 for SDXL, 20 steps, cfg 8).

5. **Enqueue the workflow.** Pass the workflow JSON from step 3 to `enqueue_workflow(action="enqueue")`. This returns immediately with a `prompt_id`.

6. **Monitor progress in background.** Start a background task to track the job:

   Bash(run_in_background: true):
   node "${CLAUDE_PLUGIN_ROOT}/scripts/monitor-progress.mjs" <prompt_id>

This connects to ComfyUI's WebSocket and prints real-time step progress (e.g., `step 3/14 (21%)`), then reports completion with output filenames. You do NOT need to poll `queue` (action:"status") — the background task handles everything.

Continue the conversation while waiting. Check the background task output when notified it completed.

If the job fails, the monitor prints error details (node, message). Use `get_history(action="diagnose")` for the full traceback plus anything missing.

**Fallback**: If the background script is unavailable, use `queue` (action:"status") to poll until `done` is true.

7. **Show the result.** Once the background monitor reports completion, use `get_image (action:"list_outputs")` (limit 1) to find the newest image. Read it with the Read tool to display it to the user.

8. **Open the image.** Open the image so the user can see it immediately without navigating to the output folder. Use the Bash tool with the appropriate command for the OS:

  • **macOS**: `open /path/to/image.png`
  • **Linux**: `xdg-open /path/to/image.png`
  • **Windows**: `start "" "/path/to/image.png"`

The image will be saved to ComfyUI's output directory (check `get_system_stats` for the `--output-directory` arg, or default to `~/Documents/ComfyUI/output/`).

Model Selection Logic

When choosing a checkpoint, consider:

  • **User's prompt**: Photorealistic → SDXL or Juggernaut XL. Anime/illustration → Anything-XL. Abstract → Flux.
  • **Available VRAM**: Check `get_system_stats`. Flux needs ~12GB. SDXL needs ~6GB. SD 1.5 works on ~4GB.
  • **Speed vs quality**: For quick tests, prefer turbo/lightning models (4-8 steps). For quality, use full models (20+ steps).

CivitAI Integration

If the environment variable `CIVITAI_API_TOKEN` is available, prefer CivitAI for model discovery because it has:

  • A larger selection of fine-tuned models
  • Community ratings and reviews
  • Better categorization (photorealistic, anime, illustration, etc.)

CivitAI REST API endpoints:

  • **Search**: `GET https://civitai.com/api/v1/models?query={query}&types=Checkpoint&sort=Most+Downloaded&limit=5`
  • **Model details**: `GET https://civitai.com/api/v1/models/{modelId}`
  • **Download**: `GET https://civitai.com/api/download/models/{modelVersionId}?token={token}`

Always include the `token` query parameter when downloading.

Example

User: `/comfy-gen a beautiful sunset over mountains with golden light`

Steps:

  • List checkpoints → none found
  • Search HuggingFace for "SDXL" → find `stabilityai/stable-diffusion-xl-base-1.0`
  • Download `sd_xl_base_1.0.safetensors` to checkpoints folder
  • Create txt2img workflow with the prompt and downloaded model
  • Enqueue the workflow, poll for completion
  • Show the generated image

Notes

  • Always randomize the seed (let the template handle it) unless the user requests a specific seed
  • If the user specifies dimensions, aspect ratio, steps, or other parameters, pass them through to `create_workflow`
  • For negative prompts, use a sensible default like "blurry, low quality, deformed" unless the user specifies one
  • If ComfyUI is not reachable, tell the user to check that it's running
  • After downloading a model, it's immediately available — no restart needed
Read more
Ships withcomfyui-mcp

The local-first, agent-native control plane for ComfyUI — an MCP server + live sidebar agent that generates images, video and audio, authors and runs workflows, manages models and custom nodes, and edits your live ComfyUI graph in natural language.

Get the whole plugin, auto-invoked
Stats
522
Stars
0
Views
84
Forks
Active
Maintenance
TypeScript
Language
MIT
License
2h ago
Last commit
5mo ago
Created

Repo: artokun/comfyui-mcp