gemini-api-dev
Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation,…
Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK. Includes workflows
$ npx -y skills add google-gemini/gemini-skills --skill gemini-omni-flash-api --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/gemini-omni-flash-apiContext preview
The summary Claude sees to decide when to auto-load this skill.
Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK. Includes workflows
name: gemini-omni-flash-api description: Use this skill for generative video editing, text-to-video, image-referenced video generation, first-frame-to-video, first-and-last-frame transitions, and video extensions using Gemini Omni 1.1 Flash (gemini-omni-1.1-flash) via the official google-genai SDK. Includes workflows for pre-processing/optimizing high-resolution or long source videos with ffmpeg, stripping audio for full sound regeneration, and handling turn-by-turn video editing and parallel execution.
This skill uses the Gemini Omni 1.1 Flash model (`gemini-omni-1.1-flash`) to perform text to video generation, image to video generation (first frame and last frame transitions), video extensions (up to 40s), and video editing.
> [!WARNING] > **Important Regional Restrictions**: Uploading videos to use for video edits or extensions is **NOT** available in the EEA, Switzerland, the United Kingdom, and some US states. If a video-to-video edit completes quickly with empty outputs (`total_output_tokens: 0` or no video content), it is likely due to this restriction.
1. **Text to video**: Generating videos from a text prompt. 2. **First frame to video**: Generating videos from a starting image (`--first-frame`). 3. **First and last frame transition**: Generating videos interpolating between a starting image and a final image (`--first-frame` and `--last-frame`; note: `--last-frame` must be used with `--first-frame`). 4. **Video extensions**: Extending existing videos by up to 10 seconds per turn, up to a total length of 40 seconds (`--extend` or `--previous-interaction-id`). 5. **Video editing and refinement**: Editing existing videos (maximum duration 10 seconds), applying stylistic changes, or performing inpainting/outpainting. 6. **Image and video referenced generation**: Using style, character, or object references from images or videos to guide video generation.
1. **Analyze request**: Determine the target task (e.g., first-frame-to-video, first-and-last-frame transition, video extension, reference-guided editing) and identify any input media assets. 2. **Run SDK scripts**:
3. **Retrieve and process output**: Outputs are saved to the local filesystem (e.g. `media/`). Report back the completed media path to the user.
pip install -U google-genai
export GEMINI_API_KEY="your-api-key"
Use the following Python scripts to upload media with the Files API, prepare input videos with ffmpeg, and generate video outputs using the Interactions API.
1. **[upload_file.py](scripts/upload_file.py)**: Uploads local media (images and videos) to the Files API and polls until `ACTIVE`. If uploading a video larger than 25MB, it prints an informative warning/tip highlighting that Gemini Omni Flash is optimized for editing 10s videos at 720p/24fps, and recommends pre-processing with `prep_video.py` first to speed up the upload.
./scripts/upload_file.py path/to/image.png
2. **[generate_video.py](scripts/video/generate_video.py)**: Performs end-to-end video generation and downloads the output video. It detects and uploads local media references (images or videos) before calling the Interactions API. Large video assets (>25MB) will trigger informative pre-processing recommendations without blocking the upload.
./scripts/video/generate_video.py "A close-up of a cat drinking tea" --output media/cat_tea.mp4
Gemini Omni 1.1 Flash natively supports four output resolutions across both landscape (`16:9`) and portrait (`9:16`) aspect ratios:
# High-definition (1080p)
./scripts/video/generate_video.py "A cinematic drone shot over misty mountains at sunrise" --resolution 1080p --output media/mountains_1080p.mp4
# Ultra-high-definition 4K (A library of skills for the Gemini API, SDK and model interactions.
Repo: google-gemini/gemini-skills
Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation,…
Use this skill when building real-time, bidirectional streaming applications with the Gemini Live API. Covers WebSocket-based audio/video/text streaming, voice…