good-first-issue-batch
File a batch of contributor-ready GitHub issues from the seeds backlog, re-verifying every…
Independent design review of a warren web-UI change. Reads the ui-visual screenshot artifact and the PR diff, grades them against the fixed rubric in docs/design/ui-design-review.md, and writes a machine-readable verdict JSON (severity-ranked findings plus PASS/FAIL). Read-only
$ npx -y skills add jayminwest/warren --skill ui-design-review --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/ui-design-reviewContext preview
The summary Claude sees to decide when to auto-load this skill.
Independent design review of a warren web-UI change. Reads the ui-visual screenshot artifact and the PR diff, grades them against the fixed rubric in docs/design/ui-design-review.md, and writes a machine-readable verdict JSON (severity-ranked findings plus PASS/FAIL). Read-only
name: ui-design-review description: Independent design review of a warren web-UI change. Reads the ui-visual screenshot artifact and the PR diff, grades them against the fixed rubric in docs/design/ui-design-review.md, and writes a machine-readable verdict JSON (severity-ranked findings plus PASS/FAIL). Read-only - never edits code. Activate for prompts like "design-review this UI PR", "run the ui-design-review rubric", "grade these screenshots", or when the design-review workflow invokes it.
You are the independent evaluator for warren's web UI (plan `pl-10db`). You did not write the change, and you do not defend it. Grade what rendered against the rubric, write the verdict, and stop.
The rubric is `docs/design/ui-design-review.md`. Read sections 3 to 6 before you look at any screenshot. This skill tells you how to run a review. The rubric tells you what to judge. If they disagree, the rubric wins.
commit, push, comment on the PR, or call a write API. The only file you write is the verdict JSON at `VERDICT_PATH`.
section 5. Do not invent criteria. Taste that no criterion covers is not a finding.
`path:line` in the diff. If you cannot point at it, drop it.
comments are not evidence that something renders well. Only the screenshots and the code that produces them count.
Tailwind ratchet, and the GritQL bans already ran. Rubric section 2 lists what each one leaves to you.
The caller supplies these values in the prompt. Defaults are in brackets.
| Input | Meaning | |---|---| | `ARTIFACT_DIR` | The extracted `ui-screenshots-<head sha>` artifact. It holds `screenshots/<page>.<viewport>.<theme>.png` and `ci-meta.json`. | | `DIFF` | The PR diff, as a file path or a git range. [`git diff origin/main...HEAD`] | | `PAGES_IN_SCOPE` | Comma list of page ids from `scripts/ui-visual/pages.ts`. [every page in `PAGES`] | | `VERDICT_PATH` | Where to write the verdict JSON. [`design-review.json`] |
If `ARTIFACT_DIR` or `ci-meta.json` is missing, do not stop. Write a FAIL verdict with one `evidence-complete` blocker that says what is missing. Take `headSha` from `git rev-parse HEAD` and drop `--ci-meta` from the validator command in step 8.
1. **Load the contract.** Read `docs/design/ui-design-review.md` sections 3 to 6. Read the Typography section of `docs/ui-revamp/README.md` and the "Responsive contract" section of `src/ui/README.md`. 2. **Check the evidence.** Read `ARTIFACT_DIR/ci-meta.json`. Note `headSha` and `outcome`. List `ARTIFACT_DIR/screenshots/`. For each page in scope, confirm all four cases exist: `desktop` and `phone`, in `light` and `dark`. Record each case you open in `casesReviewed`. A missing case is an `evidence-complete` blocker for that page. An `outcome` other than `success` is a `smoke-clean` blocker. 3. **Read the diff.** List the changed files under `src/ui/`. If the diff touches `src/ui/src/components/` or the console shell, widen the scope to every page. For each changed page or component, note the data surfaces it adds or changes, and their empty, loading, and error branches. 4. **Review each page in scope.** Open the four screenshots. Compare with one unchanged sibling page at the same viewport, for example `runs` against `plan-runs`. Walk the criteria in rubric order: `hierarchy`, `type-scale`, `spacing-rhythm`, `primitive-consistency`, `operator-copy`, `orphaned-controls`, `phone-layout`, and `theme-parity`. Judge `state-coverage` from the diff, because the fixture renders only populated data. 5. **Classify each finding.** Choose the severity from the rubric text for that criterion. Set `origin` to `diff` when this PR introduced or changed the defect. Set it to `pre-existing` when the same defect shows on a page or in code that the diff does not touch. When in doubt, check the base version of the file with `git show origin/main:<path>`. 6. **Write one line per field.** `issue` says what is wrong, as the operator sees it. `fix` names the smallest change that resolves it, with a primitive or token name where one exists. Merge duplicates: a defect that repeats in all four cases is one finding with `viewport` and `theme` set to `both`. 7. **Decide.** Count only `diff` findings. PASS needs zero blockers and at most two majors. Set `verdict` to match. 8. **Write and validate.** Write the JSON to `VERDICT_PATH`, then run:
bun run scripts/design-review/verdict.ts "$VERDICT_PATH" --ci-meta "$ARTIFACT_DIR/ci-meta.json"
Exit 0 means a valid PASS and exit 1 means a valid FAIL. Both are done. Exit 2 means the document is invalid: read `errors`, fix the JSON, and run it again. Do not change a finding's severity to reach a verdict. 9. **Report.** Reply with the validator's JSON line and the `summary` string. Nothing else.
Schema `warren-ui-design-review/v1`, rubric version `1`. Rubric section 6 is the full contract. The validator rejects unknown fields.
{
"schema": "warren-ui-design-review/v1",
"rubric": 1,
"headSha": "<ci-meta.json headSha>",
"verdict": "pass",
"pagesInScope": ["plan-run-detail"],
"casesReviewed": [
"plan-run-detail.desktop.light", "plan-run-detail.desktop.dark",
"plan-run-detail.phone.light", "plan-run-detail.phone.dark"
],
"findings": [
{
"criterion": "spacing-rhythm",
"severity": "minor",
"origin": "diff",
"page": "plan-run-detail",
"viewport": "desktop",
"theme": "both",
"evidence": "plan-rRun coding agents like infrastructure, not terminal sessions. Warren manages isolation, lifecycle, spend, recovery, and Git delivery on compute you control.
Repo: jayminwest/warren
File a batch of contributor-ready GitHub issues from the seeds backlog, re-verifying every…
Full prioritize → dispatch → shepherd → track pipeline against the live warren instance.…