deck-to-pptx
Build a PowerPoint .pptx file with python-pptx, on this deployment, without the deck engine.…
Generate images via Nano Banana (Gemini 2.5/3.1 Flash Image) on OpenRouter. Use when the user asks to draw, illustrate, render or generate any kind of picture/diagram/scene.
$ npx -y skills add EverMind-AI/Raven --skill image-gen --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/image-genContext preview
The summary Claude sees to decide when to auto-load this skill.
Generate images via Nano Banana (Gemini 2.5/3.1 Flash Image) on OpenRouter. Use when the user asks to draw, illustrate, render or generate any kind of picture/diagram/scene.
name: image-gen
description: Generate images via Nano Banana (Gemini 2.5/3.1 Flash Image) on OpenRouter. Use when the user asks to draw, illustrate, render or generate any kind of picture/diagram/scene.
metadata: '{"raven": {"requires": {"env": ["OPENROUTER_API_KEY"]}}}'Generates one or more images from a text prompt (and optionally one or more input images) by calling Google's **Nano Banana** family on OpenRouter:
OpenRouter speaks the OpenAI-compatible chat-completions API for these models, with two extras:
1. The request must include `"modalities": ["image", "text"]` so the server knows to return image bytes, not just a description. 2. The response carries images in a top-level ``message.images`` array (NOT in ``content`` — that field still holds optional commentary text).
import os, base64, json, urllib.request
KEY = os.environ["OPENROUTER_API_KEY"]
MODEL = "google/gemini-2.5-flash-image" # or 3.1 for "Nano Banana 2"
def generate_image(prompt: str, out_path: str = "out.png") -> str:
req = urllib.request.Request(
"https://openrouter.ai/api/v1/chat/completions",
data=json.dumps({
"model": MODEL,
"messages": [{"role": "user", "content": prompt}],
"modalities": ["image", "text"],
}).encode(),
headers={
"Authorization": f"Bearer {KEY}",
"Content-Type": "application/json",
},
method="POST",
)
with urllib.request.urlopen(req, timeout=120) as resp:
data = json.loads(resp.read())
msg = data["choices"][0]["message"]
# ``message.images[i].image_url.url`` is a data URI:
# "data:image/png;base64,<base64-bytes>"
url = msg["images"][0]["image_url"]["url"]
b64 = url.split(",", 1)[1]
with open(out_path, "wb") as f:
f.write(base64.b64decode(b64))
text_note = msg.get("content") or ""
return f"Wrote {out_path} ({len(b64)//1024} KB). Model said: {text_note[:200]!r}"{
"model": "google/gemini-2.5-flash-image",
"messages": [
{"role": "user", "content": "A red circle on white background"}
],
"modalities": ["image", "text"]
}For **image input + image output** (edit / vary / extend), use the standard multipart `content` form OpenAI clients accept. The image part can be either a remote URL or a base64 data URI:
{
"role": "user",
"content": [
{"type": "text", "text": "Make this watercolor style"},
{"type": "image_url",
"image_url": {"url": "data:image/png;base64,iVBORw0KGgo..."}}
]
}{
"choices": [{
"message": {
"role": "assistant",
"content": "optional text commentary",
"images": [
{
"type": "image_url",
"image_url": {
"url": "data:image/png;base64,<bytes>"
}
}
]
}
}],
"usage": {
"prompt_tokens": 8,
"completion_tokens": 1295,
"total_tokens": 1303,
"cost": 0.0383,
"completion_tokens_details": {"image_tokens": 1290}
}
}~**$0.04 per 1024×1024 PNG** at the time of writing (1290 image tokens × $0.00003/tok via the v2.5 flash-image model). Burst test: 10 images ≈ **$0.40**. Budget gates accordingly when looping in agent code.
Trigger keywords / patterns you should recognize:
Don't use this skill for:
(matplotlib / plotly) rather than a generative image. Charts need accurate numbers; Nano Banana will hallucinate axes.
graphviz instead.
OpenRouter falls back to text-only without it.
Anthropic + some Google models from China-mainland IPs. Set ``HTTPS_PROXY`` to a proxy that exits via a non-blocked region.
the prompt; don't retry the same string.
instead of decoding. Always strip the ``"data:image/png;base64,"`` prefix and ``base64.b64decode`` the rest before writing.
One Surface, All Agents: Raven generates DAGs and orchestrates multiple specialized agents for complex tasks. Raven is the harness of harnesses, built for recursive self-improvement (RSI).
Build a PowerPoint .pptx file with python-pptx, on this deployment, without the deck engine.…
构建以内容实体的发布、发现、阅读、引用、修订与归档为核心的网站;适用于单篇出版物、小型静态站、出版或机构站、文档与知识库、目录与档案、CMS…
构建、重构、诊断和验证以玩家能动性为核心的二维游戏与趣味体验。用于二维规则/动作/益智游戏、互动玩具、叙事探索、节奏体验、生成体验和本地多人;按子型选择专业运行时与内容工具,分别验证机制、…
构建、改造或诊断由可执行模型驱动的交互解释器、计算器与仿真。适用于用户通过参数、状态、步骤或事件理解规律的任务;负责主型分路、reference model、独立…
构建、重做、诊断或提升最终由浏览器消费的高完成度视觉前端。用于 HTML/CSS/JS/TS、React/Vue/Svelte/Astro、SVG/Canvas/WebGL 或…
为浏览器内需要持续操作对象、推进多步状态并产生可检查结果的产品与专业工具设计、实现、重构或诊断完整界面。用于运营与审核工作台、任务流业务产品、创作与编辑工具、分析与管理工具、资产或文件处理…