agent-activity
Streams what the agent is doing into the room, as rows the desktop client renders in an **events drawer** above the composer (collapsed: avatar, pulsing dots,…
Generate and edit images using Gemini Flash Image, and generate videos using Veo. Supports text-to-image, image editing, text-to-video, and image-to-video.
$ npx -y skills add sonichi/sutando --skill image-generation --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/image-generationContext preview
The summary Claude sees to decide when to auto-load this skill.
Generate and edit images using Gemini Flash Image, and generate videos using Veo. Supports text-to-image, image editing, text-to-video, and image-to-video.
name: image-generation description: "Generate and edit images using Gemini Flash Image, and generate videos using Veo. Supports text-to-image, image editing, text-to-video, and image-to-video."
Generate images and videos using Gemini APIs.
# Text-to-image python3 "$SKILL_DIR/scripts/generate.py" --prompt "A futuristic city skyline at night" # Edit an existing image python3 "$SKILL_DIR/scripts/generate.py" --input photo.jpg --prompt "Add dramatic clouds" # Specify output path python3 "$SKILL_DIR/scripts/generate.py" --prompt "A cute robot mascot" --output mascot.png # Text-to-video python3 "$SKILL_DIR/scripts/generate.py" --video --prompt "A timelapse of a city at sunset" # Video with portrait aspect ratio python3 "$SKILL_DIR/scripts/generate.py" --video --prompt "Ocean waves" --aspect 9:16 # Image-to-video (animate a reference image) python3 "$SKILL_DIR/scripts/generate.py" --video --input scene.jpg --prompt "Animate this scene with gentle wind" # Specify output python3 "$SKILL_DIR/scripts/generate.py" --video --prompt "Dancing robot" --output robot.mp4
| Flag | Description | Default | |------|-------------|---------| | `--prompt` | Text prompt describing what to generate | (required) | | `--input` | Input image path(s) for editing/reference | None | | `--output` | Output file path | `generated-{timestamp}.png` or `.mp4` | | `--model` | Gemini model to use | `gemini-2.5-flash-image` / `veo-3.1-generate-preview` | | `--video` | Generate video instead of image | false | | `--aspect` | Video aspect ratio | `16:9` | | `--quality` | JPEG quality (1-100, images only) | 90 |
My AI Stand — Realtime by Day, Rewriting Itself by Night. Summon my AI superpower. Voice, vision, screen, meetings, calls when I'm engaged. Learns my patterns, ships its own code when I'm not. Runs across my Macs, interacts with people & their Stands.
Repo: sonichi/sutando
Streams what the agent is doing into the room, as rows the desktop client renders in an **events drawer** above the composer (collapsed: avatar, pulsing dots,…
Local Agent Registry — a standalone, dependency-free service that tracks running Claude Code (and other) agent instances. Agents self-register on startup and…
**Prefer the `ag2-space` MCP tools when they are connected and the room exposes them** — availability is per-room and per-actor, so check…
Deterministic final-answer normalizer — a last-step pass for any task that ends in a *precise* answer (a number, a short string, a comma-list). Applies the…
Transcribes audio files and voice notes to text via Gemini 2.5-flash. Integrates with Slack, Discord, and Telegram bridges so voice clips surface as readable…
Act back on the owner's Bee wearable — the TOOL half of the Bee integration (channels-vs-tools split). The Bee CHANNEL (ag2-sparrow's `sources/bee.py` watcher)…