AI-powered image generation using Google Gemini or OpenAI (gpt-image-2), integrated with Claude Code.
> /plugin marketplace add guinacio/claude-image-gen> /plugin install media-pipeline@media-pipeline-marketplace
Repo: guinacio/claude-image-gen
What's inside
AI-powered image generation using Google Gemini or OpenAI (gpt-image-2), integrated with Claude Code.
background and outputFormat control, on OpenAI modelsThe plugin installs the core image-generation skill + CLI + MCP server in one step—no separate configuration needed. Specialized workflows are intentionally opt-in.
# Add the marketplace
/plugin marketplace add guinacio/claude-image-gen
# Install the plugin
/plugin install media-pipeline@media-pipeline-marketplace
Or install directly from GitHub:
/plugin install guinacio/claude-image-gen
Once installed:
Tip: Since the skill runs the CLI directly, you can disable the MCP server in Claude Code's MCP list to reduce startup overhead. The skill will continue to work without it.
For Claude Desktop users, install the pre-built extension:
media-pipeline.mcpb from Releases.mcpb fileFor developers who want to customize or build from source:
cd mcp-server
npm install
npm run bundle
cd mcp-server
GEMINI_API_KEY=your-api-key-here node build/cli.bundle.js \
--prompt "Landing page hero image for a fintech startup" \
--aspect-ratio "16:9"
The CLI routes to Gemini or OpenAI based on the model name and returns structured JSON on stdout. It does not require the MCP server layer. Any custom output path must still remain inside the configured output directory.
Option A: Using MCP server
claude mcp add --transport stdio media-pipeline \
--env GEMINI_API_KEY=your-api-key-here \
-- node /path/to/claude-image-gen/mcp-server/build/bundle.js
The -- separates Claude CLI flags from the server command.
Option B: Manual config
Add to your Claude Code config (~/.claude.json):
{
"mcpServers": {
"media-pipeline": {
"command": "node",
"args": ["/path/to/claude-image-gen/mcp-server/build/bundle.js"],
"env": {
"GEMINI_API_KEY": "${GEMINI_API_KEY}",
"GEMINI_DEFAULT_MODEL": "${GEMINI_DEFAULT_MODEL:-gemini-3-pro-image-preview}",
"OPENAI_API_KEY": "${OPENAI_API_KEY}",
"OPENAI_DEFAULT_MODEL": "${OPENAI_DEFAULT_MODEL:-gpt-image-2}",
"IMAGE_PROVIDER": "${IMAGE_PROVIDER:-gemini}",
"IMAGE_OUTPUT_DIR": "${IMAGE_OUTPUT_DIR:-~/generated-images}",
"GEMINI_REQUEST_TIMEOUT_MS": "${GEMINI_REQUEST_TIMEOUT_MS:-60000}",
"MEDIA_PIPELINE_LOG_LEVEL": "${MEDIA_PIPELINE_LOG_LEVEL:-info}"
}
}
}
}
The ${VAR:-default} syntax uses environment variables with fallback defaults.
If not using the plugin:
cp -r skills/image-generation ~/.claude/skills/
Specialized workflows live outside skills/, so the plugin does not discover
or install them automatically. The current character-reference workflow is for
dressing Blender character renders and preparing garments for image-to-3D tools.
From a cloned repository:
python -m pip install -r optional-workflows/character-reference-sheets/requirements.txt
cp -r optional-workflows/character-reference-sheets ~/.claude/skills/
Built a specialized workflow of your own? See Contributing a Specialized Workflow.
To create your own .mcpb extension for Claude Desktop:
cd mcp-server
npm install -g @anthropic-ai/mcpb
npm run pack:mcpb
This creates mcp-server/media-pipeline.mcpb using bundled runtime entry points for both the MCP server and the standalone CLI.
The packed extension is committed to the repo and served from the Releases page, so it must not fall behind source. pack:mcpb records a hash of every packed file in mcp-server/mcpb-contents.json, and npm test recomputes them — change the source without repacking and the suite fails. The archive itself is not byte-reproducible (zip stores timestamps), which is why the inputs are hashed rather than the .mcpb.
mcp-server/package.json is the single source of truth for the version. mcp-server/manifest.json, mcp-server/package-lock.json, .claude-plugin/plugin.json and .claude-plugin/marketplace.json are all derived from it — never edit their version fields by hand.
To bump:
cd mcp-server && npm version minor --no-git-tag-version
npm version updates package.json and the lockfile, then the version lifecycle script propagates it to the remaining manifests and stages them. npm run sync:version does the same propagation on its own, and npm run check:version reports drift without writing. Both npm test and npm run pack:mcpb fail on drift, so a release cannot ship with mismatched manifests.
Use create_asset to create a hero image for a tech startup website
The skill will proactively suggest image generation when:
| Variable | Required | Default | Description |
|---|---|---|---|
GEMINI_API_KEY | At least one of GEMINI_API_KEY / OPENAI_API_KEY | - | Your Gemini API key |
GEMINI_DEFAULT_MODEL | No | gemini-3-pro-image-preview | Default Gemini model to use |
OPENAI_API_KEY | At least one of GEMINI_API_KEY / OPENAI_API_KEY | - | Your OpenAI API key |
OPENAI_DEFAULT_MODEL | No | gpt-image-2 | Default OpenAI model to use |
IMAGE_PROVIDER | No | gemini | Provider (gemini or openai) used when a request omits model |
IMAGE_OUTPUT_DIR | No | ~/generated-images | Where to save images. Left unset, the MCP server writes to generated-images in your home directory — not the current directory, since a server launched by Claude Desktop/Code has an unpredictable working directory. Relative paths are resolved against that working directory. ~, $HOME, and %USERPROFILE% are expanded. The bundled plugin config (.mcp.json) sets ./generated-images explicitly, and the CLI defaults to the current directory instead (see --output-dir) |
GEMINI_REQUEST_TIMEOUT_MS | No | 60000 | Request timeout, applies to both Gemini and OpenAI requests |
MEDIA_PIPELINE_LOG_LEVEL | No | info | Stderr logging level |
The server is dual-provider: it routes each request to Google Gemini or OpenAI (gpt-image-2) automatically based on the model name — models starting with gpt-image or dall-e go to OpenAI, everything else goes to Gemini. When a request omits model entirely, IMAGE_PROVIDER selects which provider's default model is used.
Aspect ratios on OpenAI: OpenAI image models only support 1024x1024, 1536x1024, and 1024x1536. Requested aspect ratios are mapped to the nearest supported size, and if the mapping isn't exact the response includes a warning describing the substitution.
Reference images: supported on both providers — pass up to 5 reference image paths to guide generation.
mask, background, outputFormat: OpenAI models only — Gemini models reject them.
mask takes an absolute path to a PNG whose transparent areas are the region the model repaints; everything else is preserved from the base image. It requires referenceImages, and OpenAI documents the mask as needing an alpha channel and the same dimensions as the base image. Checked locally: the file really is a PNG, it is under the API's 4MB mask limit (masks are capped well below reference images, whose limit here is 20MB), and it is not fully opaque — an opaque mask marks nothing, so it is rejected before the request is sent. Left to the API: the dimension match, and whether a PNG carrying transparency in a tRNS chunk instead of an alpha channel is accepted — that case is sent with a warning rather than blocked.background (auto, transparent, opaque) chooses how the background is handled. transparent needs an alpha-capable outputFormat, so png is selected automatically when outputFormat is omitted. gpt-image-2 rejects transparent, and that combination is refused before the request is sent.outputFormat (png, jpeg, webp) sets the encoding of the returned image; jpeg cannot carry transparency.Apart from the gpt-image-2/transparent case above, these three options have not been verified against a live API — when a model refuses one, the API's own reason is returned.
Gemini models are fetched dynamically from the Gemini API at runtime; the CLI and MCP tool validate Gemini model choices against the current image-capable model list, and GEMINI_DEFAULT_MODEL is used when available. For OpenAI, gpt-image-2 is the default model (OPENAI_DEFAULT_MODEL); any gpt-image*/dall-e* model name routes to OpenAI.
| Ratio | Best For |
|---|---|
1:1 | Social media, thumbnails |
16:9 | Hero images, presentations |
9:16 | Mobile stories, vertical banners |
4:3 | Blog posts, general web |
3:2 | Photography-style images |
On OpenAI models, aspect ratios other than
1:1/3:2/2:3-equivalent are mapped to the nearest supported size (see Providers).
Use this formula for effective prompts:
[Style] [Subject] [Composition] [Context/Atmosphere]
Example:
Minimalist 3D illustration of abstract geometric shapes floating in space,
soft gradient background from deep purple to electric blue, subtle glow effects,
modern professional aesthetic, wide composition for website header
See skills/image-generation/references/prompt-crafting.md for advanced techniques.
CLI Mode (Default) - Used by the skill:
Claude → Skill → Bash → bundled CLI → Gemini API / OpenAI API
FAQ
media-pipeline is a Claude Code plugin with 1 hand-picked skill for content work, indexed on Flowy. Install it with the command on its page. It includes image-generation. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it