Skip to content
Agent Orchestration
Skill

/omh-model-finetuning

[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use

BOOST
From plugin
oh-my-hermes
3.3k145 skills
Install
$ npx -y skills add rlaope/oh-my-hermes --skill omh-model-finetuning --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/omh-model-finetuning

Context preview

The summary Claude sees to decide when to auto-load this skill.

[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use

SKILL.md

omh-model-finetuning.SKILL.md
name: "omh-model-finetuning"
description: "[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use when the user says: model-finetuning, model finetuning, model fine-tuning, fine-tune a model, fine-tune the model, fine tune a model, fine tune the model, fine-tune."
metadata:
  hermes:
    tags: [workflow, oh-my-hermes, planning]
    category: planning
    phase: model-finetuning
    role: planner
    quality_tier: baseline-comparison-gated

Model Finetuning

This is a Hermes-native `model-finetuning` workflow skill.

Why This Exists

`model-finetuning` exists because producing a model had no owner: `model-optimization` onboards a model into OMH, `inference-serving` serves one that exists, and `llm-app-dev` builds on top of one, while an SFT or DPO question reached `workflow-learning` with no baseline comparison and no way to answer that training is not needed.

First Steps

  • Ask what the untuned model and the best prompt score on the held-out examples before discussing any method.
  • Ask what shape the data has -- demonstrations, preference pairs, or answers a program can check.

Do Not Use When

  • The ask is onboarding a new model generation into OMH's routing, calibration, or pricing; use `model-optimization`.
  • The ask is serving an existing model behind an endpoint or benchmarking that endpoint; use `inference-serving`.
  • The ask is building an application on top of a hosted model -- RAG, structured output, prompt versions; use `llm-app-dev`.
  • The ask is learning from an OMH run, a missed route, or a skill improvement candidate; use `workflow-learning`.

Examples

Good example:

  • Prompt: should we fine-tune a model for our support replies or is a better prompt enough
  • Expected behavior: Ask for the held-out examples and the current prompt's score, try few-shot and retrieval against them, and return `do_not_finetune` if one closes the gap; otherwise pick SFT from the reply demonstrations and gate promotion on beating the untuned baseline.
  • Why: Most prompt-shaped gaps close without training, and training first hides that the cheaper fix was enough.

Bad example:

  • Prompt: the fine-tuned checkpoint got 0.82 on our eval, ship it
  • Expected behavior: Refuse to promote on a standalone score: run the same held-out eval on the untuned baseline and compare before promotion.
  • Why: A score with no baseline cannot show the training helped at all.

Completion Checklist

  • The fine-tune decision is stated, and `do_not_finetune` was considered first.
  • The method is chosen from the data's shape and its failure mode is named.
  • The held-out split was drawn before training and checked for overlap.
  • Promotion cites the same eval observed on the untuned baseline and the candidate.
  • OMH ran nothing, and every score cites observed output or is marked unverified.

Recovery Notes

  • If no held-out eval exists, building one is the first step; say so before any training plan.
  • If the candidate does not beat the baseline, keep the baseline serving and report the gap rather than retraining blindly.

Workflow Lane

  • Current lane: **Research and company ops** (`product-docs`, `source-finder`, `web-research`, `research`, `model-optimization`, `inference-serving`, `model-finetuning`, `research-brief`, `+20 more`) - research, signals, ops, and briefings.
  • If intent belongs to another lane, hand back to `oh-my-hermes` or name the adjacent workflow.
  • Shared product, routing, compatibility, and evidence rules: `omh-routing/references/skill-common-rail.md`.

Use When

Use when someone wants to fine-tune a model on their own data, or is deciding whether to: supervised fine-tuning (SFT), preference tuning (DPO), reinforcement learning from verifiable rewards (RLVR), or a LoRA adapter. The output is a decision on whether to train at all, the method chosen from the data available, a training data plan with a held-out split, a comparison against the untuned baseline, and a checkpoint promotion gate; OMH trains nothing and runs no eval.

Strong routing signals: `model-finetuning`, `model finetuning`, `model fine-tuning`, `fine-tune a model`, `fine-tune the model`, `fine tune a model`, `fine tune the model`, `fine-tune`, `fine tune`, `fine-tuning`, `fine tuning`, `fine-tuned model`, `fine-tuned checkpoint`, `finetune`, `finetuning`, `sft`, `supervised fine-tuning`, `dpo`, `direct preference optimization`, `rlvr`, `verifiable rewards`, `lora`, `qlora`, `lora adapter`, `preference data`, `untuned baseline`, `held-out eval`

Catalog Metadata

Category: `planning` Phase: `model-finetuning` Hermes role: `planner` Quality tier: `baseline-comparison-gated` Reasoning demand: `standard`

Quality bar:

  • Measure the untuned model and the cheaper fixes on the held-out eval before proposing any training.
  • Load `references/finetuning-method.md` for the decision ladder, the method table, the data checklist, and the promotion procedure instead of recalling them.
  • Choose the method from the data that exists, not from the method that is fashionable.
  • Compare every candidate against the untuned baseline on the same eval, never against its own previous run alone.
  • Keep prepared, trained, evaluated, and promoted as separate states for every checkpoint.

Handoff policy:

Keep the fine-tune decision, the method choice, the data plan, the baseline comparison, and the promotion gate in Hermes. Losses, eval scores, and comparisons are recorded only from executor, operator, or wrapper observed output; OMH never launches a training run, calls a model, or runs an eval.

Required inputs:

  • the task the model fails at, and the held-out examples that show the failure
  • what prompting, few-shot examples, or retrieval were already tried, and what they scored
  • the data available: d
Read more
Ships withoh-my-hermes

English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.

Get the whole plugin
Stats
3,264
Stars
244
Forks
Active
Maintenance
Python
Language
MIT
License
10h ago
Last commit
4mo ago
Created
4d ago
Added

Repo: rlaope/oh-my-hermes

Other skills on oh-my-hermes.