ASR
Implement speech-to-text (ASR/automatic speech recognition) capabilities using the…
Implement specialized video understanding capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to analyze video content, understand motion and temporal sequences, extract information from video frames, describe video scenes, or perform video-based AI
$ npx -y skills add jjyaoao/helloagents --skill video-understand --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/video-understandContext preview
The summary Claude sees to decide when to auto-load this skill.
Implement specialized video understanding capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to analyze video content, understand motion and temporal sequences, extract information from video frames, describe video scenes, or perform video-based AI
name: video-understand description: Implement specialized video understanding capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to analyze video content, understand motion and temporal sequences, extract information from video frames, describe video scenes, or perform video-based AI analysis. Optimized for MP4, AVI, MOV, and other common video formats. license: MIT
This skill provides specialized video understanding functionality using the z-ai-web-dev-sdk package, enabling AI models to analyze, describe, and extract information from video content including motion, temporal sequences, and scene changes.
**Skill Location**: `{project_path}/skills/video-understand`
this skill is located at above path in your project.
**Reference Scripts**: Example test scripts are available in the `{Skill Location}/scripts/` directory for quick testing and reference. See `{Skill Location}/scripts/video-understand.ts` for a working example.
Video Understanding focuses specifically on video content analysis, providing capabilities for:
**IMPORTANT**: z-ai-web-dev-sdk MUST be used in backend code only. Never use it in client-side code.
The z-ai-web-dev-sdk package is already installed. Import it as shown in the examples below.
For quick video analysis tasks, you can use the z-ai CLI instead of writing code. This is ideal for simple video descriptions, testing, or automation.
# Analyze a video from URL z-ai vision --prompt "Summarize what happens in this video" --image "https://example.com/video.mp4" # Note: Use --image flag for video URLs as well z-ai vision -p "Describe the key events" -i "https://example.com/presentation.mp4"
# Analyze a local video file z-ai vision -p "What activities are shown in this video?" -i "./recording.mp4" # Save response to file z-ai vision -p "Provide a detailed summary" -i "./meeting.mp4" -o summary.json
# Complex scene understanding with thinking z-ai vision \ -p "Analyze this video and identify: 1) Main events, 2) People and their actions, 3) Timeline of key moments" \ -i "./event.mp4" \ --thinking \ -o analysis.json # Action detection z-ai vision \ -p "Identify all actions performed by people in this video" \ -i "./sports.mp4" \ --thinking
# Stream the video analysis z-ai vision -p "Describe this video content" -i "./video.mp4" --stream
**Use CLI for:**
**Use SDK for:**
For better performance and reliability with local videos, consider: 1. Uploading videos to a CDN and using URLs 2. For shorter videos, convert key frames to images for faster analysis 3. For long videos, consider chunking or sampling at intervals
import ZAI from 'z-ai-web-dev-sdk';
async function analyzeVideo(videoUrl, prompt) {
const zai = await ZAI.create();
const response = await zai.chat.completions.createVision({
messages: [
{
role: 'user',
content: [
{
type: 'text',
text: prompt
},
{
type: 'video_url',
video_url: {
url: videoUrl
}
}
]
}
],
thinking: { type: 'disabled' }
});
return response.choices[0]?.message?.content;
}
// Usage examples
const summary = await analyzeVideo(
'https://example.com/presentation.mp4',
'Summarize the key points presented in this video'
);
const actionDetection = await analyzeVideo(
'https://example.com/sports.mp4',
'Identify and describe all athletic actions performed in this video'
);import ZAI from 'z-ai-web-dev-sdk';
async function understandVideoScenes(videoUrl) {
const zai = await ZAI.create();
const prompt = `Analyze this video and provide:
1. Overall summary of the video content
2. Main scenes or segments (with approximate timestamps if possible)
3. Key people or characters and their roles
4. Important actions or events in chronological order
5. Setting and environment description
6. Overall mood or tone`;
const response = await zai.chat.completions.createVision({
messages: [
{
role: 'user',
content: [
{ type: 'text', text: prompt },
{ type: 'video_url', video_url: { url: videoUrl } }
]
}
],🤖 生产级多智能体框架 - 工具响应协议、上下文工程、会话持久化、子代理机制等16项核心能力 HelloAgents 是一个基于 OpenAI 原生 API 构建的生产级多智能体框架,集成了工具响应协议(ToolResponse)、上下文工程(HistoryManager/TokenCounter)、会话持久化(SessionStore)、子代理机制(TaskTool)、乐观锁(文件编辑)、熔断器(CircuitBreaker)、Skills 知识外化、TodoWrite 进度管理、DevLog
Implement speech-to-text (ASR/automatic speech recognition) capabilities using the…
Implement large language model (LLM) chat completions using the z-ai-web-dev-sdk. Use this…
Implement text-to-speech (TTS) capabilities using the z-ai-web-dev-sdk. Use this skill when…
Implement vision-based AI chat capabilities using the z-ai-web-dev-sdk. Use this skill when…
Comprehensive document creation, editing, and analysis with support for tracked changes,…