Skip to content
Development
Skill

/grpo-rlvr-training

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or

From plugin
wshobson-agents
40k183 skills137 agents93 commands
Install
$ npx -y skills add wshobson/agents --skill grpo-rlvr-training --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/grpo-rlvr-training

Context preview

The summary Claude sees to decide when to auto-load this skill.

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or

SKILL.md

grpo-rlvr-training.SKILL.md
name: grpo-rlvr-training
description: Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.

GRPO & RLVR Training

This skill assumes `finetuning-method-selection` already routed here because the target behavior has a verifiable pass/fail signal — not demonstrations (`lora-qlora-recipes`) or preference pairs (`preference-optimization`). What follows is when RL is the right tool, the reference recipe, the mandatory reward-inspection gate, and how to pick a GRPO variant when the base recipe misbehaves.

**Input:** a routing decision (RLVR via GRPO) plus a verifier (code executor, test suite, schema checker, or grader) for the target task. **Output format:** a validated GRPO config — the kwarg values in `references/grpo-memory.md` and the reward functions in `references/reward-functions.md`, not free-form advice — that `llm-finetuning-training-engineer` consumes directly.

When RL Applies

GRPO+RLVR only pays off when task success is **algorithmically checkable** — a unit test passes, a parser accepts the output, a tool call matches an expected schema, a math answer matches a ground truth. If grading the output requires human judgment or a subjective rubric, that's an eval-harness and judge-calibration problem first — see `eval-harness-first` — not a reason to skip straight to RL.

Before opening a GRPO run, confirm the model can **sometimes** succeed on the target task already. RL sharpens an existing capability by reweighting toward the samples that already work; it does not install a capability from zero.

  • **The model never succeeds, even at low

temperature across many samples:** the gap is format or task understanding, not policy refinement. Route back to SFT first (`lora-qlora-recipes`) and only return to this skill once the base success rate is nonzero.

  • **The model succeeds sometimes,

inconsistently:** this is the GRPO sweet spot — proceed to The Recipe below.

The standing rule for the whole plugin: **DPO for taste, GRPO for reasoning.** If the signal is a preference between two acceptable outputs, that's `preference-optimization`, not this skill.

The Recipe

The reference recipe is TRL's `GRPOTrainer` with vLLM-backed generation:

from trl import GRPOConfig, GRPOTrainer

grpo_args = GRPOConfig(
    output_dir="./outputs-grpo",
    use_vllm=True,
    vllm_mode="colocate",       # single GPU; "server" for multi-GPU
    num_generations=8,          # floor — fewer starves the group-relative baseline
    learning_rate=5e-7,         # settled range for GRPO
    beta=0.01,                  # KL coefficient vs the reference policy
    per_device_train_batch_size=8,
    gradient_accumulation_steps=4,
    bf16=True,
    logging_steps=10,
    seed=3407,
)

trainer = GRPOTrainer(
    model=SFT_CHECKPOINT,
    args=grpo_args,
    reward_funcs=[format_reward, correctness_reward],   # references/reward-functions.md
    train_dataset=prompts,       # prompt-only — GRPO generates its own completions
    processing_class=tokenizer,
)

trainer.train()
  • **`vllm_mode="colocate"`** runs generation and

training on the same GPU — the default for a single-GPU box.

  • **`vllm_mode="server"`** points at a separate

vLLM server process and is the multi-GPU path — generation and training don't compete for the same device.

  • **`num_generations` ≥ 8** is a floor, not a

suggestion: GRPO's advantage estimate is relative to the group mean, and fewer than 8 samples per prompt produces a noisy baseline.

  • **Reward is composite** — a format reward (did

the output parse / match the required structure) plus a correctness reward (did the answer verify). A well-formed-but-wrong answer and a malformed one should not score identically; correctness alone loses that signal.

  • **`learning_rate=5e-7`** and **`beta=0.01`** are

the settled starting point; deviate only after the base run is stable and reward-inspected (below).

Memory sizing for this recipe by target size class: `references/grpo-memory.md`.

The Inspection Rule

**Run the reward function against 50–100 sampled outputs and manually read the results before starting the actual training run.** This is a gate, not a one-time sanity check.

If the reward function's judgment disagrees with a human reading of that sample, fix the reward function first. Training against an uninspected reward, or tuning hyperparameters to compensate for one silently scoring the wrong thing, is how a run reward-hacks: the model optimizes cleanly toward the wrong target, and that doesn't surface as a training-loop bug.

This inspection is a Phase 1 gate input for `/finetune` — the same 50–100-sample read that catches a broken reward function here is what that command checks for before it lets a GRPO brief proceed.

Complete reward function implementations to inspect against — exact-match, schema-validation, unit-test-execution, a length-penalty wrapper, and a rubric-as-reward judge pattern: `references/reward-functions.md`.

Variant Selection

The base recipe above is the default. Reach for a variant only when a specific failure mode shows up, not preemptively:

| Failure mode | Variant | Why | |---|---|---| | Entropy collapse / degenerate long chain-of-thought | **DAPO** | Decouples clip bounds and relaxes the KL penalty that over-regularizes exploration on long reasoning traces | | Reward or output length trends up regardless of quality | **Dr.GRPO** | Removes GRPO's length-normalization bias so reward tracks correctness, not completion length | | Training a mixture-of-experts model | **GSPO** | Moves the importance-sampling ratio to the sequence level instead of per-token — per-token ratios are unstable on MoE routing, so GSPO is required here, not optional |

Start wi

Read more
Ships withwshobson-agents

Production-ready agentic workflow building blocks: 94 plugins, 202 agents, 183 skills, 105 commands — built for Claude Code and consumed natively by OpenAI Codex CLI, Cursor, OpenCode, the Antigravity CLI, GitHub Copilot, and Pi from a single Markdown source.

Get the whole plugin

Other skills on wshobson-agents.