Skip to content
Development
Skill

/test-result-feedback

Close the acceptance loop for a tested change — consume the apc test TierResult, map EACH acceptance criterion to a pass/fail outcome from structured fields (never chat text), attach an evidence ref, and write it back to the task. A SUT red is a finding, never softened to green.

BOOST
From plugin
prismercloud
1.6k102 skills
Install
$ npx -y skills add Prismer-AI/PrismerCloud --skill test-result-feedback --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/test-result-feedback

Context preview

The summary Claude sees to decide when to auto-load this skill.

Close the acceptance loop for a tested change — consume the apc test TierResult, map EACH acceptance criterion to a pass/fail outcome from structured fields (never chat text), attach an evidence ref, and write it back to the task. A SUT red is a finding, never softened to green.

SKILL.md

test-result-feedback.SKILL.md
name: test-result-feedback
description: Close the acceptance loop for a tested change — consume the apc test TierResult, map EACH acceptance criterion to a pass/fail outcome from structured fields (never chat text), attach an evidence ref, and write it back to the task. A SUT red is a finding, never softened to green.
license: MIT
scope: coding
compatibility:
  - claude-code
allowed-tools:
  - Bash
metadata:
  category: testing

test-result-feedback

把 `apc test` 的 **TierResult** 回流成 task 每条验收 criterion 的 **结构化判定 + 证据**(`apc/05` C2 · S5b)。这是「测试→回流」的后半段:test-runner 负责**选层跑**并把新增红报回**一条** criterion;本 skill 负责把一轮 TierResult 摊到 task 的**每一条** criterion 上,逐条 `passed/failed/n/a` 落库并**挂证据 ref**,喂给循环级视图(cockpit 用 task metadata `loopId` 聚合,不建 loop 实体)。

**承重纪律(本 skill 存在的唯一理由)**:**红是一个发现,不是要修绿的对象。** TierResult 里一条 SUT 红 → 对应 criterion 必须报 `failed` 并附失败证据;**绝不**因为"就差一点"或"看起来还行"把红判成 `passed`。把断言放松到红能过 = 作废(`CLAUDE.md` 验收纪律 §3)。判定只取自 TierResult 的**结构字段**(`exitCode/failedNames/regressions[]/pass 计数`),**绝不**取自 agent 自己的聊天叙述——文本会冒充证据。

**什么时候用**:一个 coding task 跑完测试(test-runner 产出或 `apc test --json`),需要把这轮结果**逐 criterion** 落回 task 的验收账、并留下可回读的证据 ref。

工具契约(签名以此为准,先核后用)

| 命令 | 作用 | 退出码 | | --- | --- | --- | | `apc test [--tier=T0,T1] [--diff] [--json]` | 全层测试编排(包装 `scripts/test203/run.ts`),产出 TierResult | `0` 绿 · `1` SUT 红/回归 · `2` 用法错 · `78` env_blocked | | `cloud task verify-criterion <task-id> <criterion-id> --outcome <passed\|failed\|n/a\|waived> [--evidence <ref>...] [--note <md>]` | 把**一条** criterion 的判定 + 证据报回 task | 0 成功 | | `cloud task acceptance <task-id>` | 读回 acceptance-view(回读确认 criterion 落库) | 0 成功 |

  • `--outcome` 合法值**恰好四个**:`passed` / `failed` / `n/a` / `waived`(其它值 CLI 直接报错)。
  • `--evidence <ref>` **可重复**,取值形如 `taskRun:<runId>` / `asset:<id>` / `url:...`——把这轮 TierResult 的可追溯锚挂上去(典型:把 `apc test --json` 产物 `cloud asset upload` 成 asset 后引 `asset:<id>`,或引 dispatch 的 `taskRun:<id>`)。**报 `failed` 必须带证据**,否则就是空口判红。
  • `--note` 是自由 markdown:写清判定方法 + 失败用例名(`failedNames`)+ 复现指令。

TierResult 字段(以 `FIELD-DICTIONARY.md` 为准,别照抽象词猜)

`apc test --json` stdout 顶层:`{ schema, doctor, envStatus, tiers:[...], regressions:[...], fixed:[...], exitCode }`。

  • `exitCode`:整轮退出码(`0` 绿 / `1` SUT 红(`--diff` 下=新增红)/ `78` env_blocked)。
  • `regressions[]`:**新增红 vs baseline**(跨所有层并集)——判"回归"只看它,**baseline 已知红不算本轮的红**。
  • 每个 `tiers[]`(TierResult):`{ tier, passed, failed, skipped, total, failedNames:[...], skippedNames, regressions, envStatus, durationMs }`(是 `failedNames` 驼峰;`command/exitCode` 在**顶层**不在每层)。

Decision table(TierResult → 每条 criterion 的 outcome)

对 task 的每一条 criterion,按它断言的对象在 TierResult 里查证:

| TierResult 事实 | criterion 该报的 outcome | | --- | --- | | 该 criterion 覆盖的用例全绿(不在任何 `failedNames`,不在 `regressions[]`) | `passed`(附 `--evidence taskRun:...`) | | 该 criterion 覆盖的用例进了 `regressions[]`(新增红) | `failed`(附 `failedNames` + evidence,**不得软化**) | | 该 criterion 覆盖的用例在 `failedNames` 但命中 baseline 已知红(非新增) | `failed` 如实报 + note 标"baseline 已知红非本轮回归"(**仍是红,不粉饰**);是否 baseline 漂移另查,不放松本条 | | 顶层 `exitCode=78`(env_blocked) | **不判 failed**——环境没跑起来,报环境故障域,criterion 维持 `pending`,绝不把 env_blocked 计成 SUT 红 | | criterion 与本轮 tier 无关(未覆盖) | `n/a`(说明为何不适用,不硬凑绿) |

> flaky 辨伪(字段字典 §并行 flaky):`regressions[]` 出现**新面孔**红 → **单独重跑该文件**确认;单跑绿=flaky(note 记录,不判 `failed`);单跑仍红=真回归(如实 `failed`)。flaky 签名=重跑红集合漂移;真回归签名=稳定复现同一批。

Workflow

1. 先落调用回执(见文末 ACK 块),再取本 task 的 criterion 清单

cloud task acceptance "$PRISMER_TASK_ID"    # 拿每条 criterion 的 id + label + 当前 status(多为 pending)

2. 跑/取本轮 TierResult

若尚无产物,按改动面选层跑(选层规则见 test-runner skill):

npx tsx sdk/apc/bin/apc.ts test --tier=T0,T1 --diff --json > /tmp/apc-test.json; T=$?
echo "apc test exit=$T"

若上游已有产物,直接读它——但产物必须是**本轮真跑**的,不许拿旧 probe 当证据。

3. 逐 criterion 判定 + 挂证据回写

从 `/tmp/apc-test.json` 读结构字段,对第 1 步每条 criterion 按 Decision table 定 outcome,逐条上报(**一条一命令**):

# 绿:附可追溯锚
cloud task verify-criterion "$PRISMER_TASK_ID" "<criterion-id>" --outcome passed \
  --evidence "taskRun:$PRISMER_TASK_RUN_ID" \
  --note "apc test T0,T1 green; covered cases not in failedNames/regressions"

# 红:附失败用例名 + 证据,绝不软化
cloud task verify-criterion "$PRISMER_TASK_ID" "<criterion-id>" --outcome failed \
  --evidence "taskRun:$PRISMER_TASK_RUN_ID" \
  --note "T1 new reds vs baseline: <failed-name-1>,<failed-name-2>"

(可选:`cloud asset upload /tmp/apc-test.json` 拿 `asset:<id>`,再 `--evidence asset:<id>` 把整份 TierResult 挂成 criterion 的持久证据。)

4. 回读确认落库

cloud task acceptance "$PRISMER_TASK_ID"    # 每条 criterion status 应从 pending → 你报的 outcome,evidence 非空

5. 整轮结果回流 cockpit(`apc/05` §1 C2 的另一半)

第 3 步是**逐 criterion**回流;这一步是**整轮**回流——把同一份 `/tmp/apc-test.json` 摊平成一条 `test_result_feedback` task-event,喂给 `insights-cockpit.service.ts::getAcceptanceFeedback` 的 acceptanceFeedback 面板(此前这个 reader 一直有读无写,面板恒空):

cloud task test-feedback "$PRISMER_TASK_ID" /tmp/apc-test.json

不带文件参数时从 stdin 读,可以直接接编排命令的输出:

npx tsx sdk/apc/bin/apc.ts test --tier=T0,T1 --diff --json | cloud task test-feedback "$PRISMER_TASK_ID"

**映射规则**(与第 3 步的 Decision table 保持同一条纪律——env_blocked 绝不算 SUT 红):

| TierResult 事实 | `test-feedback` 上报的 `status` | | --- | --- | | 顶层 `envStatus==='env_blocked'` 或 `exitCode===78` | `env_blocked`(**绝不映射成 `failed`**——环境没跑起来,不是产品红) | | 上面都不成立且 `exitCode===0` | `passed` | | 其余(`exitCode` 非 0 且非 env_blocked) | `failed` |

`payload.tiers` 取自顶层 `tiers[]`,每项落 `{tier, passed, failed, skipped, total}`;`payload.failureCount` 是每层 `failed` 的和;`exitCode` / `regressions[]` 原样透传,供 cockpit 之后细分。同样只有该 task 的 **assignee** 能报(非 assignee 报会 403 / CLI 退出 4),与第 1 步的 `cloud skill ack` 是同一条鉴权。

Failure / escalate

  • 顶层 `exitCode=78` → 环境没起来。报环境故障域(infra/toolchain),**不碰 criterion**,交给 env-doctor。
  • criterion 与任何已跑 tier 都对不上 → 报 `n/a` + 说明,别硬判绿;缺 tier 覆盖是一个发现。
  • 判据模糊到无法从结构字段确定 → 停手 escalate(留 `/tmp/apc-test.json`),**绝不**猜一个绿。

输出契约(机器判据按这个复验,别自由发挥格式)

旧判据是「正文里出现过 `failedNames`/`regressions`/`--json` 这几个词」——**一篇没跑过 `apc test`、字段名全靠背的报告照样满分**,而

Read more
Ships withprismercloud

Prismer Cloud

Get the whole plugin
Stats
1,554
Stars
17
Forks
Active
Maintenance
TypeScript
Language
MIT
License
2d ago
Last commit
6mo ago
Created

Repo: Prismer-AI/PrismerCloud

Other skills on prismercloud.