ASR
Implement speech-to-text (ASR/automatic speech recognition) capabilities using the…
Implement text-to-speech (TTS) capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to convert text into natural-sounding speech, create audio content, build voice-enabled applications, or generate spoken audio files. Supports multiple voices, adjustable
$ npx -y skills add jjyaoao/helloagents --skill TTS --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/TTSContext preview
The summary Claude sees to decide when to auto-load this skill.
Implement text-to-speech (TTS) capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to convert text into natural-sounding speech, create audio content, build voice-enabled applications, or generate spoken audio files. Supports multiple voices, adjustable
name: TTS description: Implement text-to-speech (TTS) capabilities using the z-ai-web-dev-sdk. Use this skill when the user needs to convert text into natural-sounding speech, create audio content, build voice-enabled applications, or generate spoken audio files. Supports multiple voices, adjustable speed, and various audio formats. license: MIT
This skill guides the implementation of text-to-speech (TTS) functionality using the z-ai-web-dev-sdk package, enabling conversion of text into natural-sounding speech audio.
**Skill Location**: `{project_path}/skills/TTS`
This skill is located at the above path in your project.
**Reference Scripts**: Example test scripts are available in the `{Skill Location}/scripts/` directory for quick testing and reference. See `{Skill Location}/scripts/tts.ts` for a working example.
Text-to-Speech allows you to build applications that generate spoken audio from text input, supporting various voices, speeds, and output formats for diverse use cases.
**IMPORTANT**: z-ai-web-dev-sdk MUST be used in backend code only. Never use it in client-side code.
Before implementing TTS functionality, be aware of these important limitations:
function splitTextIntoChunks(text, maxLength = 1000) {
const chunks = [];
const sentences = text.match(/[^.!?]+[.!?]+/g) || [text];
let currentChunk = '';
for (const sentence of sentences) {
if ((currentChunk + sentence).length <= maxLength) {
currentChunk += sentence;
} else {
if (currentChunk) chunks.push(currentChunk.trim());
currentChunk = sentence;
}
}
if (currentChunk) chunks.push(currentChunk.trim());
return chunks;
}The z-ai-web-dev-sdk package is already installed. Import it as shown in the examples below.
For simple text-to-speech conversions, you can use the z-ai CLI instead of writing code. This is ideal for quick audio generation, testing voices, or simple automation.
# Convert text to speech (default WAV format) z-ai tts --input "Hello, world" --output ./hello.wav # Using short options z-ai tts -i "Hello, world" -o ./hello.wav
# Use specific voice z-ai tts -i "Welcome to our service" -o ./welcome.wav --voice tongtong # Adjust speech speed (0.5-2.0) z-ai tts -i "This is faster speech" -o ./fast.wav --speed 1.5 # Slower speech z-ai tts -i "This is slower speech" -o ./slow.wav --speed 0.8
# MP3 format z-ai tts -i "Hello World" -o ./hello.mp3 --format mp3 # WAV format (default) z-ai tts -i "Hello World" -o ./hello.wav --format wav # PCM format z-ai tts -i "Hello World" -o ./hello.pcm --format pcm
# Stream audio generation z-ai tts -i "This is a longer text that will be streamed" -o ./stream.wav --stream
**Use CLI for:**
**Use SDK for:**
import ZAI from 'z-ai-web-dev-sdk';
import fs from 'fs';
async function textToSpeech(text, outputPath) {
const zai = await ZAI.create();
const response = await zai.audio.tts.create({
input: text,
voice: 'tongtong',
speed: 1.0,
response_format: 'wav',
stream: false
});
// Get array buffer from Response object
const arrayBuffer = await response.arrayBuffer();
const buffer = Buffer.from(new Uint8Array(arrayBuffer));
fs.writeFileSync(outputPath, buffer);
console.log(`Audio saved to ${outputPath}`);
return outputPath;
}
// Usage
await textToSpeech('Hello, world!', './output.wav');import ZAI from 'z-ai-web-dev-sdk';
import fs from 'fs';
async function generateWithVoice(text, voice, outputPath) {
const zai = await ZAI.create();
const response = await zai.audio.tts.create({
input: text,
voice: voice, // Available voices: tongtong, chuichui, xiaochen, jam, kazi, douji, luodo
speed: 1.0,
response_format: 'wav',
stream: false
});
// Get array buffer from Response object
const arrayBuffer = await response.arrayBuffer();
const buffer = Buffer.from(new Uint8Array(arrayBuffer));
fs.writeFileSync(outputPath, buffer);
return outputPath;
}
// Usage
await generateWithVoice('Welcome to our service', 'tongtong', './welcome.wav');import ZAI from '
🤖 生产级多智能体框架 - 工具响应协议、上下文工程、会话持久化、子代理机制等16项核心能力 HelloAgents 是一个基于 OpenAI 原生 API 构建的生产级多智能体框架,集成了工具响应协议(ToolResponse)、上下文工程(HistoryManager/TokenCounter)、会话持久化(SessionStore)、子代理机制(TaskTool)、乐观锁(文件编辑)、熔断器(CircuitBreaker)、Skills 知识外化、TodoWrite 进度管理、DevLog
Repo: jjyaoao/helloagents
Implement speech-to-text (ASR/automatic speech recognition) capabilities using the…
Implement large language model (LLM) chat completions using the z-ai-web-dev-sdk. Use this…
Implement vision-based AI chat capabilities using the z-ai-web-dev-sdk. Use this skill when…
Comprehensive document creation, editing, and analysis with support for tracked changes,…
Comprehensive Finance API integration skill for real-time and historical financial data…
Transform UI style requirements into production-ready frontend code with systematic design…