ai-image-editing
The AI image-editing router — inpainting/object removal, background removal, upscaling, outpainting, old-photo restoration, and retouch, routed task-first to…
The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline. Use when someone wants to edit in Descript, edit a podcast or interview, remove filler words/silences, clean up audio (Studio Sound), fix a
$ npx -y skills add social-media-skills/skills --skill descript --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/descriptContext preview
The summary Claude sees to decide when to auto-load this skill.
The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline. Use when someone wants to edit in Descript, edit a podcast or interview, remove filler words/silences, clean up audio (Studio Sound), fix a
name: descript description: >- The Descript craft skill — edit talk content (podcasts, interviews, talking-head video) by editing the transcript instead of the timeline. Use when someone wants to edit in Descript, edit a podcast or interview, remove filler words/silences, clean up audio (Studio Sound), fix a flubbed word without re-recording (Overdub), auto-cut between speakers, turn one recording into clips + show notes + chapters, or asks about Descript's plans, credits, or Underlord. Uses the WORDS framework plus the interview rule: concision yes, meaning-flips never. Reads the recording's content skill + brand-profile/voice-builder first. The agent plans the edit (API/MCP where connected); the HUMAN verifies by ear and approves; WoopSocial publishes the exports. Overdub is consent-verified own-voice-only; tiers/credits are verified in-app. Distinct from capcut (visual short-form), captions-and-clipping/opus-clip (clip selection at scale), ai-voiceover (dedicated TTS), and podcast-and-audiograms (the strategy). version: 1.0.0
The **talk-content editing tool skill** — write the edit in the transcript, Overdub with consent, refine the sound, dress the visuals, and ship the cuts. The agent plans (and can drive Underlord via API/MCP where connected); the **human verifies by ear and approves**; **WoopSocial publishes** the exports. (Ships with `tools/integrations/descript.md`.)
For dialogue-heavy content, editing the transcript beats scrubbing a timeline: delete the sentence, the clip disappears; move the paragraph, the footage follows — reviews report ~60–70% editing-time cuts for talk content. But the paradigm has two sharp edges the top 1% respect. **(1) The voice spine:** Overdub's consent-verified, **own-voice-only** design is the model, not an obstacle — it exists so nobody types words into someone else's mouth; and the craft truth is it shines on flubbed *words*, not paragraphs (long Overdub drifts synthetic — re-record those). **(2) The meaning spine:** text-editing makes it dangerously easy to rearrange a guest into saying something they didn't — **concision yes, meaning-flips never**, and the human owns the final cut of anyone else's words. Operationally: the **accuracy pass is mandatory** (transcript errors become wrong edits AND wrong captions), and since the Sept 2025 overhaul, the workflow must be **credit-aware** — media minutes count everything you import, and formerly-unlimited AI features are metered.
1. The recording's content skill — **podcast-and-audiograms** / **youtube-long-form** / **educational-content**. 2. **brand-profile** + **voice-builder** (written outputs) + **design-and-templates** (captions/layout).
(Depth: `references/the-words-framework.md`.)
(Underlord), restructure by moving paragraphs — decide in text, **verify by ear.**
re-record); vocabulary/credit limits; disclose synthetic speech where required.
level speakers; extreme noise is a re-record, not a rescue.
transcript, human-approved B-roll, Eye Contact used honestly; beat-sync/color route elsewhere.
show notes + chapters + a text post; route onward and publish via WoopSocial.
2026 Descript: **Underlord** (agentic co-editor — filler/silence in one step, bad-take flags, B-roll suggestions, clips, show notes) now triggerable via the **2026 public API (open beta) incl. MCP connections**; Overdub (~24–48h training; source-audio requirements have varied — verify); Studio Sound (~10 credits/use); Automatic Multicam; Eye Contact; ~92–95% transcription accuracy on clean audio, ~75–85% with noise/accents/jargon; ~23 languages; SOC 2 Type II; cloud-dependent (no offline). **Pricing: the Sept 2025 overhaul** moved to media minutes + AI-credit metering of formerly-unlimited features; documented bill-shock and no mid-cycle proration (G2, attributed); **tier figures conflict across sources — verify in-app.** Full detail: `references/descript-2026-reality.md`. The weekly loop, credit-aware checklist, the Overdub decision table, the interview-integrity checklist, and two worked examples: `references/workflows-and-templates.md`.
(exact human steps otherwise — no pretended automation); the **human verifies by ear** (pacing/tone/fairness don't live in text) and approves — the agent never fabricates "that cut sounds great." **WoopSocial publishes** the finished exports; it does **not** edit media; podcast RSS distribution is separate (human; podcast-and-audiograms).
real mouth; **AI-disclosure** for synthetic speech where required (EU AI Act; C2PA). **Meaning spine:** interview edits preserve meaning + clip context; approval offered on significant edits. **Never fabricate** tiers, credits, or metrics — verify in-app. (Full scope: `references/scope-and-connections.md`.)
**descript (this)** = text-based talk-content editing · **capcut** = beat-synced visual short-form (the hybrid: master here, style cuts there) · **captions-and-clipping / opus-clip** = clip selection
Give your AI agent the skills of a top-1% social media team. 106 of them, free.
The AI image-editing router — inpainting/object removal, background removal, upscaling, outpainting, old-photo restoration, and retouch, routed task-first to…
The AI music + sound-design skill for social -- original/licensed audio beds and sound design for Reels/TikToks/Shorts/videos. Use when someone needs…
Use to get a brand and its content CITED and RECOMMENDED by AI answer engines — the GEO (Generative Engine Optimization) / AI-search-visibility skill. Run when…
The model-agnostic AI-video router and brief — the counterpart to image-prompt. Use when someone asks "which AI video tool should I use," "make an AI video,"…
The AI narration / voiceover mini-skill (ElevenLabs-led). Use when someone wants an "AI voiceover," "narration," "text-to-speech for a video," "voice for my…
Social media analytics and reporting — read native platform data honestly and turn it into next actions. Use when someone wants to "check my analytics," "see…