Skip to content
Agent Orchestration
Skill

/omh-llm-app-dev

[omh] LLM-powered feature to build: LLM app development: prepare a build handoff for an LLM-powered feature with a pinned provider boundary, schema-first outputs, versioned prompt files, grounded retrieval, and an eval suite as a shipped deliverable. Use when the user says:

BOOST
From plugin
oh-my-hermes
3.3k145 skills
Install
$ npx -y skills add rlaope/oh-my-hermes --skill omh-llm-app-dev --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/omh-llm-app-dev

Context preview

The summary Claude sees to decide when to auto-load this skill.

[omh] LLM-powered feature to build: LLM app development: prepare a build handoff for an LLM-powered feature with a pinned provider boundary, schema-first outputs, versioned prompt files, grounded retrieval, and an eval suite as a shipped deliverable. Use when the user says:

SKILL.md

omh-llm-app-dev.SKILL.md
name: "omh-llm-app-dev"
description: "[omh] LLM-powered feature to build: LLM app development: prepare a build handoff for an LLM-powered feature with a pinned provider boundary, schema-first outputs, versioned prompt files, grounded retrieval, and an eval suite as a shipped deliverable. Use when the user says: llm-app-dev, llm app development, llm application development, build an llm app, build an llm feature, llm feature development, build a rag pipeline, rag pipeline."
metadata:
  hermes:
    tags: [workflow, oh-my-hermes, delivery]
    category: delivery
    phase: llm-app-dev
    role: operator
    quality_tier: delivery-gated

Llm App Dev

This is a Hermes-native `llm-app-dev` workflow skill.

Why This Exists

`llm-app-dev` exists because the failure modes of an LLM feature are not the failure modes of the code around it. A floating model alias, a prompt buried in a string literal, an output scraped out of prose with a regex, and a retrieval layer nobody measured all pass code review and all fail in production, and without a golden set nobody can tell whether the next prompt edit helped or hurt.

Do Not Use When

  • The subject is comparing executors or agent harnesses - Codex against Claude Code against Hermes coding - rather than evaluating the product's own model calls; use `agent-evaluation`.
  • An agent run is already stuck, looping, or drifting and needs diagnosis; use `agent-debug`.
  • The subject is the harness's own context window, prompt caching, or token budget rather than the application being built; use `context-budget-review`.
  • The request is a prompt-injection, secret-handling, or dependency risk gate on work that already exists; use `security-safety-review`.
  • The feature makes no model call - the LLM is only mentioned as the subject being discussed - so this is a direct answer, not a build handoff.
  • The ask is training the weights themselves on your own data -- SFT, DPO, RLVR, or a LoRA adapter; use `model-finetuning`.

Examples

Good example:

  • Prompt: $llm-app-dev we are adding an invoice-field extractor that calls a model per upload - set it up so we can change the prompt later without guessing.
  • Expected behavior: Name the rails, put the provider call behind one client module with a pinned model ID, declare the extraction schema and the repair path, lay the prompt out as a versioned file, and specify the golden set and validators that let the next prompt edit be compared against this baseline.
  • Why: The feature is a real model call whose output another system consumes, which is exactly where an unpinned model, an inline prompt, and a missing golden set become expensive later.

Bad example:

  • Prompt: $llm-app-dev the extractor is done - confirm the new prompt is better than the old one.
  • Expected behavior: Prepare the paired baseline-vs-candidate comparison and state that no result exists until the run is observed; report nothing about which prompt is better.
  • Why: Better is a claim about an observed run. Without one, the comparison is a design, and calling it a result is the false-green this workflow exists to prevent.

Completion Checklist

  • Every rail - provider boundary, structured output, prompt artifacts, retrieval grounding, evaluation - is either decided or explicitly deferred with a reason.
  • One client boundary owns the model ID, credentials, timeout, retry, and backoff, and no credential appears in source, prompts, tests, or examples.
  • The model ID is exact, and it is recorded next to any result meant to be compared.
  • Every model response is validated against a declared schema, with a bounded repair path and a loud failure - no prose scraping.
  • Prompts are files with a version identifier, and system rules, task instruction, and injected context are separated.
  • Untrusted retrieved or user-supplied content is fenced from the instruction channel and cannot change the task.
  • The eval deliverables - golden set, task-level validators, baseline-vs-candidate comparison - exist as committed artifacts, and retrieval is evaluated before generation when retrieval is in the path.
  • Token, latency, and cost figures come from an observed run or stay null; no design output is reported as an eval result, implementation, review, CI, or merge evidence.
  • If the feature communicates through a public board, the destination carries a public-audience label, each action class names its own authority and outbound data, the exact draft and its host-recorded approval reference travel with the request through compaction and executor handoff, and no publication is reported without an observed connector result.

Recovery Notes

  • If the exact model ID or provider is not decided yet, name the candidates and prepare the boundary against a config value rather than choosing one silently.
  • If no failing case can be stated, the golden set has no seed: collect the real failures first, because a golden set written from imagination measures the imagination.
  • If a response cannot be made to satisfy the schema after one bounded repair, treat that as a schema or prompt defect and record it as a golden-set case rather than loosening validation.
  • If retrieval quality was never measured, stop before scoring generation and route the retrieval evaluation first; a generation score on unmeasured retrieval is not attributable.
  • If the comparison run did not emit tokens or cost, leave those fields null and say the harness did not report them; never reconstruct them from pricing tables.
  • If a public-board send returned no confirmed outcome, do not retry: read the board back or resolve the receipt first, because a duplicate public post cannot be withdrawn the way a failed private write can be repeated.

Workflow Lane

  • Current lane: **Coding handoff** (`idea-to-deploy`, `llm-app-dev`, `cto-loop`, `deploy-and-monitor`, `code-review`, `build-failure-triage`, `verification-gate`, `security-safety-review`, `+28 more`) - coding owners, handoffs, review, CI, and m
Read more
Ships withoh-my-hermes

English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.

Get the whole plugin
Stats
3,264
Stars
244
Forks
Active
Maintenance
Python
Language
MIT
License
10h ago
Last commit
4mo ago
Created
4d ago
Added

Repo: rlaope/oh-my-hermes

Other skills on oh-my-hermes.