Skip to content

/rigorous-evaluation

Use when evaluating or reporting model performance and choosing metrics, thresholds, and plots that FIT the problem type. Picks the right metrics per task (binary uses ROC-AUC and PR-AUC; multiclass uses macro F1 and a confusion matrix; detection uses mAP; segmentation uses Dice

From plugin
823 skills1 agents1 commands
shell
$ npx -y skills add mxslr/mlcraft --skill rigorous-evaluation --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/rigorous-evaluation
How auto-invocation works

Context preview

The summary Claude sees to decide when to auto-load this skill.

Use when evaluating or reporting model performance and choosing metrics, thresholds, and plots that FIT the problem type. Picks the right metrics per task (binary uses ROC-AUC and PR-AUC; multiclass uses macro F1 and a confusion matrix; detection uses mAP; segmentation uses Dice
Ships withmlcraft

A research-first AI/ML research-engineer workflow for Claude Code

Get the whole plugin, auto-invoked

Other skills on mlcraft.