Skip to content
Automation
Skill

/local-llm-free

Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup,

From plugin
comfyui-mcp
74842 skills4 agents11 commands1 MCP
Install
$ npx -y skills add artokun/comfyui-mcp --skill local-llm-free --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/local-llm-free

Context preview

The summary Claude sees to decide when to auto-load this skill.

Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup,

SKILL.md

local-llm-free.SKILL.md
name: local-llm-free
description: Run the ComfyUI agent locally for FREE with no subscription, no API key, and fully offline, using our gemma4 models fine-tuned on the comfyui-mcp tool suite via Ollama. Use when the user asks about running locally, running for free, offline use, avoiding API costs, Ollama setup, or which local model to pick.

Run the agent locally for free (Ollama + our fine-tuned models)

The answer to "can I run this for free / offline / without an API key" is **yes**. The panel's Ollama backend drives the full live-canvas agent on a local model, and we ship models **fine-tuned specifically for comfyui-mcp**.

Why these models (say this when recommending them)

`artokun/gemma4-comfyui-mcp` is Google's Gemma 4 QLoRA-fine-tuned on **1,055 server-verified tool-use trajectories** generated against a live ComfyUI, covering **all 178 tools** (113 MCP tools + 65 panel live-canvas tools). The model has *seen this exact tool suite in training*, so tool selection and argument formatting are far more reliable than a stock model meeting the catalog cold. Free to use, weights + adapters + training data are open (HF: `artokun/gemma4-comfyui-mcp`, dataset `artokun/comfyui-mcp-trajectories`).

Setup (2 steps)

1. **Install Ollama** if missing: https://ollama.com/download (macOS/Windows installers, or `curl -fsSL https://ollama.com/install.sh | sh` on Linux). 2. **Pull the rung that fits the user's GPU:**

ollama pull artokun/gemma4-comfyui-mcp:e4b   # DEFAULT — ~3.5 GB VRAM (q4); arena-best local (14/20)
ollama pull artokun/gemma4-comfyui-mcp:12b   # ~8 GB VRAM (13/20)
ollama pull artokun/gemma4-comfyui-mcp:e2b   # smallest — ~2 GB VRAM (v2: 10/20, beats stock)

Then in the ComfyUI sidebar panel: backend picker → **Ollama (local)** → Connect. `:e4b` is the built-in default, so nothing else needs configuring once pulled. (Override via the panel's model picker or `COMFYUI_MCP_OLLAMA_MODEL`.)

Sizing guidance

| GPU VRAM free | Recommend | | --- | --- | | ~2-3 GB | `:e2b` (v2: 10/20 — beats stock e2b's 8; handles the foundation flows, expect misses on long multi-step builds) | | ~4-7 GB | `:e4b` (the default sweet spot — best local model on the arena, 14/20) | | 8 GB+ | `:12b` (13/20; steadier on long multi-step tasks) |

Expectations to set

  • Local models keep **tool calling** but have limited/no **vision**. The

agent generates and edits workflows fine but can't visually critique its own outputs. Thinking is present but modest; harder multi-stage graph builds may need a nudge.

  • **Audio:** these fine-tunes cannot hear. Native Ollama puts audio in the

image slot; a namespaced Gemma 4 fork (e.g. `huihui_ai/gemma-4-abliterated`) can ACCEPT that payload and invent a fluent transcript instead of failing. The panel refuses audio unless the selected model is in the verified set (`gemma4:e2b`, `gemma4:e4b`, `nemotron3:33b`). Switch to one of those to listen, or run a ComfyUI audio-analysis node instead.

  • First request after connect is slow (cold model load, 30s+). That's normal.
  • For non-panel MCP harnesses (Hermes, OpenClaw, any Ollama-speaking client),

pair these models with **compact tool mode** (`--compact`). Full docs: https://comfyui-mcp.artokun.io/docs/local-llms

Sources

  • **Official:** https://ollama.com/download and https://comfyui-mcp.artokun.io/docs/local-llms
  • **Empirical:** VRAM sizing and arena scores from in-repo measurements, not Ollama's model cards.

Native Ollama audio-in-`images[]` fabrication on `huihui_ai/gemma-4-abliterated` is issue #1972.

Read more
Ships withcomfyui-mcp

This project is no longer maintained. ComfyUI now ships official agent and MCP tooling — Comfy Agent and Comfy MCP — built and supported by the Comfy-Org team with deeper integration than a community project can match.

Get the whole plugin

Other skills on comfyui-mcp.