Skip to content
Automation
Skill

/hugging-face-evaluation

Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.

From plugin
lihongwei-cn
5200 skills1 agent
Install
$ npx -y skills add LiHongwei-cn/lihongwei-cn --skill hugging-face-evaluation --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/hugging-face-evaluation

Context preview

The summary Claude sees to decide when to auto-load this skill.

Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.

SKILL.md

hugging-face-evaluation.SKILL.md
name: hugging-face-evaluation
description: Add and manage evaluation results in Hugging Face model cards. Supports extracting eval tables from README content, importing scores from Artificial Analysis API, and running custom model evaluations with vLLM/lighteval. Works with the model-index metadata format.
risk: unknown
source: community

Overview

This skill provides tools to add structured evaluation results to Hugging Face model cards. It supports multiple methods for adding evaluation data:

  • Extracting existing evaluation tables from README content
  • Importing benchmark scores from Artificial Analysis
  • Running custom model evaluations with vLLM or accelerate backends (lighteval/inspect-ai)

When to Use

  • You need to add structured evaluation results to a Hugging Face model card.
  • You want to import benchmark data or run custom evaluations with vLLM, lighteval, or inspect-ai.
  • You are preparing leaderboard-compatible `model-index` metadata for a model release.

Integration with HF Ecosystem

  • **Model Cards**: Updates model-index metadata for leaderboard integration
  • **Artificial Analysis**: Direct API integration for benchmark imports
  • **Papers with Code**: Compatible with their model-index specification
  • **Jobs**: Run evaluations directly on Hugging Face Jobs with `uv` integration
  • **vLLM**: Efficient GPU inference for custom model evaluation
  • **lighteval**: HuggingFace's evaluation library with vLLM/accelerate backends
  • **inspect-ai**: UK AI Safety Institute's evaluation framework

Version

1.3.0

Dependencies

Core Dependencies

  • huggingface_hub>=0.26.0
  • markdown-it-py>=3.0.0
  • python-dotenv>=1.2.1
  • pyyaml>=6.0.3
  • requests>=2.32.5
  • re (built-in)

Inference Provider Evaluation

  • inspect-ai>=0.3.0
  • inspect-evals
  • openai

vLLM Custom Model Evaluation (GPU required)

  • lighteval[accelerate,vllm]>=0.6.0
  • vllm>=0.4.0
  • torch>=2.0.0
  • transformers>=4.40.0
  • accelerate>=0.30.0

Note: vLLM dependencies are installed automatically via PEP 723 script headers when using `uv run`.

IMPORTANT: Using This Skill

⚠️ CRITICAL: Check for Existing PRs Before Creating New Ones

**Before creating ANY pull request with `--create-pr`, you MUST check for existing open PRs:**

uv run scripts/evaluation_manager.py get-prs --repo-id "username/model-name"

**If open PRs exist:** 1. **DO NOT create a new PR** - this creates duplicate work for maintainers 2. **Warn the user** that open PRs already exist 3. **Show the user** the existing PR URLs so they can review them 4. Only proceed if the user explicitly confirms they want to create another PR

This prevents spamming model repositories with duplicate evaluation PRs.

---

> **All paths are relative to the directory containing this SKILL.md file.** > Before running any script, first `cd` to that directory or use the full path.

**Use `--help` for the latest workflow guidance.** Works with plain Python or `uv run`:

uv run scripts/evaluation_manager.py --help
uv run scripts/evaluation_manager.py inspect-tables --help
uv run scripts/evaluation_manager.py extract-readme --help

Key workflow (matches CLI help):

1) `get-prs` → check for existing open PRs first 2) `inspect-tables` → find table numbers/columns 3) `extract-readme --table N` → prints YAML by default 4) add `--apply` (push) or `--create-pr` to write changes

Core Capabilities

1. Inspect and Extract Evaluation Tables from README

  • **Inspect Tables**: Use `inspect-tables` to see all tables in a README with structure, columns, and sample rows
  • **Parse Markdown Tables**: Accurate parsing using markdown-it-py (ignores code blocks and examples)
  • **Table Selection**: Use `--table N` to extract from a specific table (required when multiple tables exist)
  • **Format Detection**: Recognize common formats (benchmarks as rows, columns, or comparison tables with multiple models)
  • **Column Matching**: Automatically identify model columns/rows; prefer `--model-column-index` (index from inspect output). Use `--model-name-override` only with exact column header text.
  • **YAML Generation**: Convert selected table to model-index YAML format
  • **Task Typing**: `--task-type` sets the `task.type` field in model-index output (e.g., `text-generation`, `summarization`)

2. Import from Artificial Analysis

  • **API Integration**: Fetch benchmark scores directly from Artificial Analysis
  • **Automatic Formatting**: Convert API responses to model-index format
  • **Metadata Preservation**: Maintain source attribution and URLs
  • **PR Creation**: Automatically create pull requests with evaluation updates

3. Model-Index Management

  • **YAML Generation**: Create properly formatted model-index entries
  • **Merge Support**: Add evaluations to existing model cards without overwriting
  • **Validation**: Ensure compliance with Papers with Code specification
  • **Batch Operations**: Process multiple models efficiently

4. Run Evaluations on HF Jobs (Inference Providers)

  • **Inspect-AI Integration**: Run standard evaluations using the `inspect-ai` library
  • **UV Integration**: Seamlessly run Python scripts with ephemeral dependencies on HF infrastructure
  • **Zero-Config**: No Dockerfiles or Space management required
  • **Hardware Selection**: Configure CPU or GPU hardware for the evaluation job
  • **Secure Execution**: Handles API tokens safely via secrets passed through the CLI

5. Run Custom Model Evaluations with vLLM (NEW)

⚠️ **Important:** This approach is only possible on devices with `uv` installed and sufficient GPU memory. **Benefits:** No need to use `hf_jobs()` MCP tool, can run scripts directly in terminal **When to use:** User working in local device directly when GPU is available

Before running the script

  • check the script path
  • check uv is installed
  • check gpu is available with `nvidia-smi`

Running the script

uv run scripts/train_sft_example.py

Features

  • **vLLM Backend**: High-performance GPU inference (5-10x
Read more
Ships withlihongwei-cn

MUNDO - THE EMPEROR. Complete AI orchestration system with 1208 skills, 25 capability modules, self-evolving, collective consciousness. GitHub Actions 24/7 automation.

Get the whole plugin
Stats
5
Stars
1
Forks
Maintained
Maintenance
Python
Language
MIT
License
1mo ago
Last commit
4mo ago
Created

Repo: LiHongwei-cn/lihongwei-cn

Other skills on lihongwei-cn.

cheat-on-content
Skill

cheat-on-content

给所有想把"感觉"变成可校准预测的内容创作者。**方法论通用**——打分 → 盲预测 → T+3d 复盘 → 进化 rubric 的循环适用任何能被量化(播放 / 阅读 / 收听 / 点击)的内容。**rubric 是循环的内容,不是循环本身**——当前内置一份观点视频 rubric(参考博主 25+…

cheat-bump
Skill

cheat-bump

提议并执行 rubric 或 bucket 升级。两种模式:**完整 rubric bump**(最高风险动作,5 步强制 + 跨模型审核)和 **--bucket-only 轻量重校**(只换 bucket 边界,不动 rubric 公式)。**Phase 2 强制走 cheat-score-blind…

cheat-init
Skill

cheat-init

cheat-on-content 的首次 onboarding 与脚手架创建器。统一流程——所有用户都走相同 5 阶段闭环,唯一区别是"发过视频的人"会在 init 时多一步:抓取已有视频建立历史 context(用于后续 cheat-seed 给更贴合的选题、更准的…

cheat-migrate
Skill

cheat-migrate

把老用户的 .cheat-state.json 升级到当前 schema_version。读 migrations/registry.md 算迁移链,按顺序应用每一步迁移文件。幂等:跑两次结果一样。失败停在中间版本不前进。触发词:"迁移"/"升级 state"/"migrate"/"我的 state…

cheat-persona
Skill

cheat-persona

从复盘评论数据派生 / 刷新账号的受众画像,写入 audience.md。这是和 rubric 平行的第二个派生物——rubric 答"怎么打分",persona 答"谁在看"。cheat-seed 选题 / 写稿时读它。**audience.md 含实绩信号,cheat-score-blind…