Skip to content

/agent-evaluation

Evaluate LLM agents and tool-using workflows—task success, tool accuracy, latency/cost, safety, and regression suites. Use when shipping agent features, comparing prompts/models, or debugging agent failures.

shell
$ npx -y skills add charlieviettq/awesome-agent-skill --skill agent-evaluation --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.
  • You can call itInvoke it directly when you want it.
  • Slash command/agent-evaluation
How auto-invocation works

Context preview

The summary Claude sees to decide when to auto-load this skill.

Evaluate LLM agents and tool-using workflows—task success, tool accuracy, latency/cost, safety, and regression suites. Use when shipping agent features, comparing prompts/models, or debugging agent failures.
Ships withawesome-agent-skill

Curated skill pack for LLM agents in engineer and science workflow (Cursor & Claude ready).

Get the whole plugin, auto-invoked
Stats
22
Stars
0
Views
8
Forks
Active
Maintenance
Python
Language
MIT
License
15d ago
Last commit
2mo ago
Created

Repo: charlieviettq/awesome-agent-skill

Other skills on awesome-agent-skill.