Skip to content
Agent Orchestration
Skill

/omh-media-input

[omh] Audio, video, or screenshot to process: user-sent media - audio, video, YouTube links, screenshots, receipts, OCR, meeting recordings, transcripts, timestamps, and clip summaries, gated for source, permission, and hallucination risk. Use when the user says:

BOOST
From plugin
oh-my-hermes
3.2k145 skills
Install
$ npx -y skills add rlaope/oh-my-hermes --skill omh-media-input --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/omh-media-input

Context preview

The summary Claude sees to decide when to auto-load this skill.

[omh] Audio, video, or screenshot to process: user-sent media - audio, video, YouTube links, screenshots, receipts, OCR, meeting recordings, transcripts, timestamps, and clip summaries, gated for source, permission, and hallucination risk. Use when the user says:

SKILL.md

omh-media-input.SKILL.md
name: "omh-media-input"
description: "[omh] Audio, video, or screenshot to process: user-sent media - audio, video, YouTube links, screenshots, receipts, OCR, meeting recordings, transcripts, timestamps, and clip summaries, gated for source, permission, and hallucination risk. Use when the user says: media-input-operator, media input operator, media input, audio transcription, audio transcript, transcribe audio, transcribe this audio, meeting recording."
metadata:
  hermes:
    tags: [workflow, oh-my-hermes, media]
    category: media
    phase: media-input-task
    role: guide
    quality_tier: workflow-surface-gated

Media Input Operator

This is a Hermes-native `media-input-operator` workflow skill.

Why This Exists

`media-input-operator` exists so Hermes users can ask for this workflow in chat and get a structured, checkable answer instead of an improvised one.

Do Not Use When

  • The request is already handled by a narrower explicit skill with stronger evidence.
  • The user asks OMH to secretly run external platforms, connectors, schedulers, file exports, or runtime agents.
  • The only safe answer is to ask for missing authority, credentials, target, or observed evidence first.

Examples

Good example:

  • Prompt: media-input-operator transcribe this audio meeting and summarize action items with evidence and timestamp boundaries.
  • Expected behavior: Produce `prepare_media_input_card` with required context, wrapper actions, and not-evidence boundaries.
  • Why: The prompt names a real workflow surface that Hermes can orchestrate without hiding execution.

Bad example:

  • Prompt: media-input-operator invent a YouTube transcript and claim the timestamps are verified without media evidence.
  • Expected behavior: Report the missing observed evidence or authority instead of claiming the external step happened.
  • Why: Prepared OMH guidance is not platform, runtime, connector, file, memory, or delivery evidence.

Completion Checklist

  • Media type, source location, permission boundary, transcript availability, language, requested output, timestamp requirement, and stop condition are explicit.
  • Downloads, uploads, ASR, transcript extraction, speaker labels, copyrighted media access, and provider setup are gated or marked missing.
  • Transcript text, OCR output, screenshot text, receipt fields, timestamps, quotes, action items, and media-summary claims are reported only from observed media or supplied transcript/extraction evidence.

Recovery Notes

  • If the media or transcript is missing, ask for the smallest source, file, transcript, or provider result needed.
  • If the request is broad current-source research about a video topic, route to research or source-finder before summary.
  • If the user wants a PPT/PDF/report generated from the media summary, route to materials-package after media input evidence is clear.
  • If the request is about whether a live duplex voice connector keeps whole spoken turns, route to external-connector-readiness for a realtime_voice_trial_receipt/v1 rather than treating a supplied recording as that evidence.

Workflow Lane

  • Current lane: **Materials and visual summaries** (`design-orchestration`, `apple-design`, `design-quality-gate`, `award-bar-score`, `frontend`, `accessibility-audit`, `visual-qa`, `content-operator`, `+6 more`) - web, accessibility, visual QA, files, and packages.
  • If intent belongs to another lane, hand back to `oh-my-hermes` or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: `omh-routing/references/skill-common-rail.md`.

Use When

Use when Hermes should prepare or supervise audio/video transcript, YouTube/video summary, OCR, screenshot text extraction, receipt image parsing, or timestamped media extraction work without claiming media access, download, transcription, OCR output, or factual summary evidence.

Strong routing signals: `media-input-operator`, `media input operator`, `media input`, `audio transcription`, `audio transcript`, `transcribe audio`, `transcribe this audio`, `meeting recording`, `recording transcript`, `video transcript`, `youtube summary`, `youtube video`, `summarize youtube`, `summarize this youtube`, `video summary`, `summarize this video`, `ocr image`, `image ocr`, `photo ocr`, `picture ocr`, `graphic ocr`, `screenshot ocr`, `ocr this image`, `ocr receipt image`, `ocr this receipt image`, `receipt ocr`, `receipt image ocr`, `receipt text`, `receipt text from image`, `receipt fields`, `receipt fields from image`, `receipt image extraction`, `receipt image text`, `receipt image fields`, `parse receipt image`, `receipt image parse`, `receipt image into fields`, `image text extraction`, `extract text from image`, `extract text from this image`, `screenshot text extraction`, `extract text from screenshot`, `extract text from this screenshot`, `screenshot to text`, `timestamps`, `with timestamps`, `clip summary`, `podcast summary`, `webinar summary`, `오디오 전사`, `음성 전사`, `회의 녹음`, `녹음 요약`, `영상 요약`, `유튜브 요약`, `youtube 요약`, `이미지 ocr`, `이미지 OCR`, `이미지 텍스트 추출`, `이미지에서 텍스트 추출`, `영수증 ocr`, `영수증 OCR`, `영수증 이미지 ocr`, `영수증 이미지 OCR`, `스크린샷 텍스트 추출`, `스크린샷에서 텍스트 추출`, `타임스탬프`, `타임라인 요약`

Catalog Metadata

Category: `media` Phase: `media-input-task` Hermes role: `guide` Quality tier: `workflow-surface-gated` Reasoning demand: `standard`

Quality bar:

  • Name the user-facing workflow objective, required context, next action, and stop condition.
  • Separate prepared guidance from observed platform, runtime, connector, file, memory, or delivery evidence.
  • Expose missing tools, credentials, targets, or observations as user-visible gaps.

Handoff policy:

Keep this as Hermes-facing orchestration guidance first. Prepare executor, connector, gateway, or host-runtime handoff only when the user accepts that next step and observed evidence can be recorded.

Required inputs:

  • user request
  • target context
  • delivery or status expectation
  • known missing evidence

Expected

Read more
Ships withoh-my-hermes

English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.

Get the whole plugin
Stats
3,206
Stars
244
Forks
Active
Maintenance
Python
Language
MIT
License
4h ago
Last commit
4mo ago
Created
9h ago
Added

Repo: rlaope/oh-my-hermes

Other skills on oh-my-hermes.