Skip to content
Automation
Skill

/evaluating-machine-learning-models

Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.

From plugin
vibe-skills
2.7k200 skills8 agents3 commands
Install
$ npx -y skills add foryourhealth111-pixel/Vibe-Skills --skill evaluating-machine-learning-models --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/evaluating-machine-learning-models

Context preview

The summary Claude sees to decide when to auto-load this skill.

Evaluate trained machine learning models with the right metrics and comparison logic. Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.

SKILL.md

evaluating-machine-learning-models.SKILL.md
name: evaluating-machine-learning-models
description: |
  Evaluate trained machine learning models with the right metrics and comparison logic.
  Use for benchmark review, threshold selection, calibration, validation, and model comparison; not for feature engineering or leakage auditing.
allowed-tools: Read, Write, Edit, Grep, Glob, Bash(cmd:*)
version: 1.0.0
author: Jeremy Longshore <jeremy@intentsolutions.io>
license: MIT

Model Evaluation Suite

Use this skill when the model exists and the question is whether it is good enough.

Overview

This skill focuses on choosing and interpreting the right evaluation metrics for the problem, then comparing candidate models or thresholds.

When to Use This Skill

  • Comparing candidate models with consistent metrics
  • Reviewing precision/recall/F1/AUC, regression error, calibration, or ranking quality
  • Stress-testing validation strategy before deployment or publication

Not For / Boundaries

  • Building the training pipeline itself: use `scikit-learn` for classical modeling or `ml-pipeline-workflow` for end-to-end workflow ownership
  • Engineering features: use `preprocessing-data-with-automated-pipelines`
  • Checking train/test contamination: use `ml-data-leakage-guard`

Typical Outputs

  • Metric suite recommendations
  • Model comparison tables
  • Notes on threshold tradeoffs, calibration, and validation weaknesses

Related Skills

  • `scikit-learn` for class-level error breakdowns and confusion matrices
  • `scientific-reporting` when the evaluation must become a deliverable
Read more
Ships withvibe-skills

VibeSkills is a general-purpose Skill that automatically routes local Skills and intelligently orchestrates harness workflows.

Get the whole plugin

Other skills on vibe-skills.