Skip to content
Automation
Skill

/inno-figure-gen

Generate/edit images with OpenAI gpt-image-2 by default, falling back to Gemini (gemini-3.1-flash-image-preview) when OPENAI_API_KEY is unset. Supports text-to-image + image-to-image; 1K/2K/4K; use --input-image for editing, --provider to force a provider, --model to override

From plugin
dr-claw
1.1k174 skills
Install
$ npx -y skills add OpenLAIR/dr-claw --skill inno-figure-gen --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/inno-figure-gen

Context preview

The summary Claude sees to decide when to auto-load this skill.

Generate/edit images with OpenAI gpt-image-2 by default, falling back to Gemini (gemini-3.1-flash-image-preview) when OPENAI_API_KEY is unset. Supports text-to-image + image-to-image; 1K/2K/4K; use --input-image for editing, --provider to force a provider, --model to override

SKILL.md

inno-figure-gen.SKILL.md
name: inno-figure-gen
description: >
  Generate/edit images with OpenAI gpt-image-2 by default, falling back to
  Gemini (gemini-3.1-flash-image-preview) when OPENAI_API_KEY is unset.
  Supports text-to-image + image-to-image; 1K/2K/4K; use --input-image for
  editing, --provider to force a provider, --model to override the model.

Image Generation & Editing (GPT Image default, Gemini fallback)

Generate new images or edit existing ones. The script picks a provider based on which API keys are available:

1. **OpenAI** `gpt-image-2` — used when `OPENAI_API_KEY` is set (default). 2. **Gemini** `gemini-3.1-flash-image-preview` — used when only `GEMINI_API_KEY` is set.

When both keys are set, OpenAI is picked by default; force Gemini with `--provider gemini`. If an auto-selected OpenAI call fails at runtime (quota, moderation, network), the script transparently falls back to Gemini when a Gemini key is available.

Usage

The script is at `scripts/generate_image.py` **relative to this skill's directory** (the directory containing this `SKILL.md`). Resolve the full path from the skill's location before running. Do not hardcode `~/.codex/...`, because the skill may be installed in a different location.

Keep the distinction clear:

  • The **script path** tells you where to find `generate_image.py`.
  • The **output path** is controlled by the current working directory plus `--filename`.
  • Run from the user's working directory so relative filenames save output there, not in the skill directory.

**Generate new image:**

uv run <this-skill-directory>/scripts/generate_image.py --prompt "your image description" --filename "output-name.png" [--resolution 1K|2K|4K] [--provider auto|openai|gemini] [--model MODEL] [--openai-api-key KEY | --gemini-api-key KEY]

**Edit existing image:**

uv run <this-skill-directory>/scripts/generate_image.py --prompt "editing instructions" --filename "output-name.png" --input-image "path/to/input.png" [--resolution 1K|2K|4K] [--provider auto|openai|gemini] [--model MODEL] [--openai-api-key KEY | --gemini-api-key KEY]

**Important:** Always run from the user's current working directory so images are saved where the user is working, not in the skill directory.

Default Workflow (draft → iterate → final)

Goal: fast iteration without burning time on 4K until the prompt is correct.

  • Draft (1K): quick feedback loop
  • `uv run <this-skill-directory>/scripts/generate_image.py --prompt "<draft prompt>" --filename "yyyy-mm-dd-hh-mm-ss-draft.png" --resolution 1K`
  • Iterate: adjust prompt in small diffs; keep filename new per run
  • If editing: keep the same `--input-image` for every iteration until you’re happy.
  • Final (4K): only when prompt is locked
  • `uv run <this-skill-directory>/scripts/generate_image.py --prompt "<final prompt>" --filename "yyyy-mm-dd-hh-mm-ss-final.png" --resolution 4K`

Resolution Options

The script accepts three resolution tiers (uppercase K required):

  • **1K** (default) - ~1024px
  • **2K** - ~2048px
  • **4K** - ~4096px (Gemini) / **3840×2160 landscape** under OpenAI

Map user requests to API parameters:

  • No mention of resolution → `1K`
  • "low resolution", "1080", "1080p", "1K" → `1K`
  • "2K", "2048", "normal", "medium resolution" → `2K`
  • "high resolution", "high-res", "hi-res", "4K", "ultra" → `4K`

**OpenAI 4K note:** `gpt-image-2` supports non-square 4K outputs within its size limits. `--resolution 4K` with OpenAI maps to `3840×2160`.

Provider & Model Selection

Two providers are available:

| Provider | Default model | When chosen | | -------- | ------------------------------------ | ------------------------------------------------------------ | | OpenAI | `gpt-image-2` | `--provider auto` (default) when `OPENAI_API_KEY` is set, or `--provider openai` | | Gemini | `gemini-3.1-flash-image-preview` | `--provider auto` when only `GEMINI_API_KEY` is set, or `--provider gemini`, or as runtime fallback from a failed auto-OpenAI call |

Override either default with `--model`. Note: the model name is provider-specific; passing a Gemini model name while the script falls back to Gemini automatically will not preserve a user-specified OpenAI model (each provider uses its own default during fallback).

Common model options:

  • **OpenAI:** `gpt-image-2` (default)
  • **Gemini:** `gemini-3.1-flash-image-preview` (default, fast), `gemini-3-pro-image-preview` (higher quality, slower)

API Keys

The script resolves provider-specific keys first, while preserving the original Gemini-only `--api-key` behavior:

1. Explicit, provider-specific flags: `--openai-api-key KEY`, `--gemini-api-key KEY` 2. Generic `--api-key KEY` — kept for backward compatibility with the original Gemini-only script. Under `auto`, it is treated as a Gemini key unless an OpenAI key is provided by `--openai-api-key` or `OPENAI_API_KEY`. Under explicit `--provider openai`, it is treated as an OpenAI key; under explicit `--provider gemini`, it is treated as a Gemini key. 3. Environment variables: `OPENAI_API_KEY`, `GEMINI_API_KEY`

When `--provider auto` is selected and OpenAI fails at runtime, fallback to Gemini requires `--gemini-api-key`, `--api-key`, or the `GEMINI_API_KEY` env var.

If no key is resolvable for the chosen provider, the script exits with a clear error message listing both ways to fix it.

Preflight + Common Failures (fast fixes)

  • Preflight:
  • `command -v uv` (must exist)
  • At least one of: `test -n "$OPENAI_API_KEY" -o -n "$GEMINI_API_KEY"` (or pass `--openai-api-key`, `--gemini-api-key`, or backward-compatible `--api-key`)
  • If editing: `test -f "path/to/input.png"`
  • Common failures:
  • `Error: No API key found...` → set `OPENAI_API_KEY` or `GEMINI_API_KEY`, or pass an explicit `--*-api-key` flag
  • `Error loading input image:` → wrong path / unreadable file; verify `--input-image` points to a real image
  • `[war
Read more
Ships withdr-claw

A Super AI Lab with massive AI Doctors as Assistants. Best IDE for Research via AI Power.

Get the whole plugin
Stats
1,091
Stars
119
Forks
Active
Maintenance
JavaScript
Language
6d ago
Last commit
6mo ago
Created

Repo: OpenLAIR/dr-claw

Other skills on dr-claw.