llm-output-schema-cons…
Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object(). Use when writing…
Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Use after writing a judge prompt to verify it agrees with human judgment.
$ npx -y skills add growthxai/output --skill output-eval-validate-judge --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/output-eval-validate-judgeContext preview
The summary Claude sees to decide when to auto-load this skill.
Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Use after writing a judge prompt to verify it agrees with human judgment.
name: output-eval-validate-judge description: Validate LLM judges against human labels using TPR/TNR metrics and train/dev/test splits. Use after writing a judge prompt to verify it agrees with human judgment. allowed-tools: [Bash, Read, Write, Edit]
An LLM judge is only useful if it agrees with human judgment. This skill walks you through calibrating a judge against human-labeled data using True Positive Rate (TPR) and True Negative Rate (TNR) metrics. Do this **before** trusting any `judgeVerdict()`, `judgeScore()`, or `judgeLabel()` evaluator in your eval suite.
1. **A judge `.prompt` file** — Written following `output-eval-judge-prompt` 2. **~100 human-labeled traces** — With binary pass/fail labels for the failure mode this judge targets. Aim for ~50 pass and ~50 fail. Minimum: 20 pass and 20 fail. 3. **Labels stored in dataset YAML** — Each dataset has `ground_truth.evals.<evaluator_name>.verdict: pass` or `fail`
This process applies **only to LLM-based judges**. For code-based `Verdict.*` evaluators, write unit tests instead.
Split your labeled datasets into three groups:
| Split | % of Data | Purpose | Example (100 datasets) | |-------|-----------|---------|----------------------| | **Train** | 10-20% | Source of few-shot examples in the judge prompt | 15 datasets | | **Dev** | 40-45% | Iterate on judge prompt, measure TPR/TNR | 42 datasets | | **Test** | 40-45% | Final held-out measurement, run once | 43 datasets |
Use a naming convention or subdirectories to separate splits:
**Option A: Name prefixes**
tests/datasets/ ├── train_formal_pass_01.yml ├── train_casual_fail_01.yml ├── dev_technical_pass_01.yml ├── dev_ambiguous_fail_01.yml ├── test_simple_pass_01.yml ├── test_contradictory_fail_01.yml └── ...
**Option B: Subdirectories**
tests/datasets/
├── train/
│ ├── formal_pass_01.yml
│ └── casual_fail_01.yml
├── dev/
│ ├── technical_pass_01.yml
│ └── ambiguous_fail_01.yml
└── test/
├── simple_pass_01.yml
└── contradictory_fail_01.ymlExecute the eval workflow against only the dev-split datasets:
# Run with cached output on dev datasets npx output workflow test <workflowName> --cached \ --dataset dev_technical_pass_01,dev_ambiguous_fail_01,dev_formal_pass_02,...
Or if using subdirectories, list the dev dataset names:
npx output workflow test <workflowName> --cached \ --dataset $(ls tests/datasets/dev/ | sed 's/.yml//' | tr '\n' ',')
Save the output. You need the judge's verdict for each dataset to compare against ground truth.
Use `--json` to get machine-readable results:
npx output workflow test <workflowName> --cached --dataset <dev_datasets> --json
The output includes per-dataset, per-evaluator verdicts that you can compare against `ground_truth.evals.<evaluator_name>.verdict`.
For the evaluator you're validating, build a confusion matrix from the dev results.
Using "fail" as the positive class (what you're trying to detect):
| | Judge says Fail | Judge says Pass | |---|---|---| | **Human says Fail** | True Positive (TP) | False Negative (FN) | | **Human says Pass** | False Positive (FP) | True Negative (TN) |
**TPR (True Positive Rate)** = TP / (TP + FN)
**TNR (True Negative Rate)** = TN / (TN + FP)
Dev set results for `check_tone` evaluator (42 datasets):
| | Judge: Fail | Judge: Pass | |---|---|---| | **Human: Fail** | 18 (TP) | 3 (FN) | | **Human: Pass** | 2 (FP) | 19 (TN) |
Raw accuracy = (TP + TN) / total = (18 + 19) / 42 = 88.1%
This looks fine, but masks problems. If your dataset were 90% pass (class imbalance), a judge that always says "pass" would get 90% accuracy while catching zero failures (TPR = 0%). TPR and TNR measure what actually matters: catching failures and not crying wolf.
For every case where the judge disagrees with the human label, determine the root cause.
The judge said "pass" but the human said "fail." For each:
1. Read the trace and the judge's critique 2. Determine why the judge missed it:
The judge said "fail" but the human said "pass." For each:
1. Read the trace and the judge's critique 2. Determine why the judge flagged it:
The open-source TypeScript framework for building AI workflows and agents. Designed for Claude Code — describe what you want, Claude builds it, with all the best practices already in place. One framework.
Repo: growthxai/output
Zod schema constraints that Anthropic rejects or silently ignores when sent as structured-output tool definitions via aiSdk.Output.object(). Use when writing…
Guide to the providerOptions structure in .prompt files — decision tree for where an option goes, common mistakes, per-provider quick reference, and Anthropic…
Implement an Output SDK workflow from a plan document. Use when the user asks to build, implement, or code a workflow from an existing plan, or after…
View and edit encrypted credentials in an Output.ai project. Use when adding secrets, updating API keys, verifying credential values, or retrieving a specific…
Wire encrypted credentials to environment variables using the credential: convention. Use when setting up LLM provider keys (ANTHROPIC_API_KEY, OPENAI_API_KEY)…