model-onboarding
Onboard a new model generation or sibling into oh-my-hermes: probe router recognition,…
[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use
$ npx -y skills add rlaope/oh-my-hermes --skill omh-model-finetuning --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/omh-model-finetuningContext preview
The summary Claude sees to decide when to auto-load this skill.
[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use
name: "omh-model-finetuning"
description: "[omh] Fine-tuning a model on your own data -- SFT, DPO, RLVR or a LoRA adapter: decide first whether prompting or retrieval already closes the gap, choose the method from the data you have, and promote a checkpoint only when it beats the untuned baseline on a held-out eval. Use when the user says: model-finetuning, model finetuning, model fine-tuning, fine-tune a model, fine-tune the model, fine tune a model, fine tune the model, fine-tune."
metadata:
hermes:
tags: [workflow, oh-my-hermes, planning]
category: planning
phase: model-finetuning
role: planner
quality_tier: baseline-comparison-gatedThis is a Hermes-native `model-finetuning` workflow skill.
`model-finetuning` exists because producing a model had no owner: `model-optimization` onboards a model into OMH, `inference-serving` serves one that exists, and `llm-app-dev` builds on top of one, while an SFT or DPO question reached `workflow-learning` with no baseline comparison and no way to answer that training is not needed.
Good example:
Bad example:
Use when someone wants to fine-tune a model on their own data, or is deciding whether to: supervised fine-tuning (SFT), preference tuning (DPO), reinforcement learning from verifiable rewards (RLVR), or a LoRA adapter. The output is a decision on whether to train at all, the method chosen from the data available, a training data plan with a held-out split, a comparison against the untuned baseline, and a checkpoint promotion gate; OMH trains nothing and runs no eval.
Strong routing signals: `model-finetuning`, `model finetuning`, `model fine-tuning`, `fine-tune a model`, `fine-tune the model`, `fine tune a model`, `fine tune the model`, `fine-tune`, `fine tune`, `fine-tuning`, `fine tuning`, `fine-tuned model`, `fine-tuned checkpoint`, `finetune`, `finetuning`, `sft`, `supervised fine-tuning`, `dpo`, `direct preference optimization`, `rlvr`, `verifiable rewards`, `lora`, `qlora`, `lora adapter`, `preference data`, `untuned baseline`, `held-out eval`
Category: `planning` Phase: `model-finetuning` Hermes role: `planner` Quality tier: `baseline-comparison-gated` Reasoning demand: `standard`
Quality bar:
Handoff policy:
Keep the fine-tune decision, the method choice, the data plan, the baseline comparison, and the promotion gate in Hermes. Losses, eval scores, and comparisons are recorded only from executor, operator, or wrapper observed output; OMH never launches a training run, calls a model, or runs an eval.
Required inputs:
English | 한국어 | 日本語 | 中文 Install once. Keep Hermes. Add a stronger operating layer. Planning, research, creation, coding handoffs, operations, and project memory with explicit evidence boundaries.
Repo: rlaope/oh-my-hermes
Onboard a new model generation or sibling into oh-my-hermes: probe router recognition,…
Review oh-my-hermes pull requests that have not been reviewed at their current head commit.…
Backfill labels across oh-my-hermes issues and pull requests. Run manually to sweep…
[omh] Screen-reader or keyboard accessibility gaps: prepare WCAG, keyboard, focus,…
[omh] Hermes badges unlocked and achievement progress: achievements observation: summarize…
[omh] Technical proposal facing adversarial scrutiny: independent perspectives attack a…