bug-reproduce
Turn a known bug into a tight, red-capable reproducer, then prove the reproducer locks that…
Close the acceptance loop for a tested change — consume the apc test TierResult, map EACH acceptance criterion to a pass/fail outcome from structured fields (never chat text), attach an evidence ref, and write it back to the task. A SUT red is a finding, never softened to green.
$ npx -y skills add Prismer-AI/PrismerCloud --skill test-result-feedback --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/test-result-feedbackContext preview
The summary Claude sees to decide when to auto-load this skill.
Close the acceptance loop for a tested change — consume the apc test TierResult, map EACH acceptance criterion to a pass/fail outcome from structured fields (never chat text), attach an evidence ref, and write it back to the task. A SUT red is a finding, never softened to green.
name: test-result-feedback description: Close the acceptance loop for a tested change — consume the apc test TierResult, map EACH acceptance criterion to a pass/fail outcome from structured fields (never chat text), attach an evidence ref, and write it back to the task. A SUT red is a finding, never softened to green. license: MIT scope: coding compatibility: - claude-code allowed-tools: - Bash metadata: category: testing
把 `apc test` 的 **TierResult** 回流成 task 每条验收 criterion 的 **结构化判定 + 证据**(`apc/05` C2 · S5b)。这是「测试→回流」的后半段:test-runner 负责**选层跑**并把新增红报回**一条** criterion;本 skill 负责把一轮 TierResult 摊到 task 的**每一条** criterion 上,逐条 `passed/failed/n/a` 落库并**挂证据 ref**,喂给循环级视图(cockpit 用 task metadata `loopId` 聚合,不建 loop 实体)。
**承重纪律(本 skill 存在的唯一理由)**:**红是一个发现,不是要修绿的对象。** TierResult 里一条 SUT 红 → 对应 criterion 必须报 `failed` 并附失败证据;**绝不**因为"就差一点"或"看起来还行"把红判成 `passed`。把断言放松到红能过 = 作废(`CLAUDE.md` 验收纪律 §3)。判定只取自 TierResult 的**结构字段**(`exitCode/failedNames/regressions[]/pass 计数`),**绝不**取自 agent 自己的聊天叙述——文本会冒充证据。
**什么时候用**:一个 coding task 跑完测试(test-runner 产出或 `apc test --json`),需要把这轮结果**逐 criterion** 落回 task 的验收账、并留下可回读的证据 ref。
| 命令 | 作用 | 退出码 | | --- | --- | --- | | `apc test [--tier=T0,T1] [--diff] [--json]` | 全层测试编排(包装 `scripts/test203/run.ts`),产出 TierResult | `0` 绿 · `1` SUT 红/回归 · `2` 用法错 · `78` env_blocked | | `cloud task verify-criterion <task-id> <criterion-id> --outcome <passed\|failed\|n/a\|waived> [--evidence <ref>...] [--note <md>]` | 把**一条** criterion 的判定 + 证据报回 task | 0 成功 | | `cloud task acceptance <task-id>` | 读回 acceptance-view(回读确认 criterion 落库) | 0 成功 |
`apc test --json` stdout 顶层:`{ schema, doctor, envStatus, tiers:[...], regressions:[...], fixed:[...], exitCode }`。
对 task 的每一条 criterion,按它断言的对象在 TierResult 里查证:
| TierResult 事实 | criterion 该报的 outcome | | --- | --- | | 该 criterion 覆盖的用例全绿(不在任何 `failedNames`,不在 `regressions[]`) | `passed`(附 `--evidence taskRun:...`) | | 该 criterion 覆盖的用例进了 `regressions[]`(新增红) | `failed`(附 `failedNames` + evidence,**不得软化**) | | 该 criterion 覆盖的用例在 `failedNames` 但命中 baseline 已知红(非新增) | `failed` 如实报 + note 标"baseline 已知红非本轮回归"(**仍是红,不粉饰**);是否 baseline 漂移另查,不放松本条 | | 顶层 `exitCode=78`(env_blocked) | **不判 failed**——环境没跑起来,报环境故障域,criterion 维持 `pending`,绝不把 env_blocked 计成 SUT 红 | | criterion 与本轮 tier 无关(未覆盖) | `n/a`(说明为何不适用,不硬凑绿) |
> flaky 辨伪(字段字典 §并行 flaky):`regressions[]` 出现**新面孔**红 → **单独重跑该文件**确认;单跑绿=flaky(note 记录,不判 `failed`);单跑仍红=真回归(如实 `failed`)。flaky 签名=重跑红集合漂移;真回归签名=稳定复现同一批。
cloud task acceptance "$PRISMER_TASK_ID" # 拿每条 criterion 的 id + label + 当前 status(多为 pending)
若尚无产物,按改动面选层跑(选层规则见 test-runner skill):
npx tsx sdk/apc/bin/apc.ts test --tier=T0,T1 --diff --json > /tmp/apc-test.json; T=$? echo "apc test exit=$T"
若上游已有产物,直接读它——但产物必须是**本轮真跑**的,不许拿旧 probe 当证据。
从 `/tmp/apc-test.json` 读结构字段,对第 1 步每条 criterion 按 Decision table 定 outcome,逐条上报(**一条一命令**):
# 绿:附可追溯锚 cloud task verify-criterion "$PRISMER_TASK_ID" "<criterion-id>" --outcome passed \ --evidence "taskRun:$PRISMER_TASK_RUN_ID" \ --note "apc test T0,T1 green; covered cases not in failedNames/regressions" # 红:附失败用例名 + 证据,绝不软化 cloud task verify-criterion "$PRISMER_TASK_ID" "<criterion-id>" --outcome failed \ --evidence "taskRun:$PRISMER_TASK_RUN_ID" \ --note "T1 new reds vs baseline: <failed-name-1>,<failed-name-2>"
(可选:`cloud asset upload /tmp/apc-test.json` 拿 `asset:<id>`,再 `--evidence asset:<id>` 把整份 TierResult 挂成 criterion 的持久证据。)
cloud task acceptance "$PRISMER_TASK_ID" # 每条 criterion status 应从 pending → 你报的 outcome,evidence 非空
第 3 步是**逐 criterion**回流;这一步是**整轮**回流——把同一份 `/tmp/apc-test.json` 摊平成一条 `test_result_feedback` task-event,喂给 `insights-cockpit.service.ts::getAcceptanceFeedback` 的 acceptanceFeedback 面板(此前这个 reader 一直有读无写,面板恒空):
cloud task test-feedback "$PRISMER_TASK_ID" /tmp/apc-test.json
不带文件参数时从 stdin 读,可以直接接编排命令的输出:
npx tsx sdk/apc/bin/apc.ts test --tier=T0,T1 --diff --json | cloud task test-feedback "$PRISMER_TASK_ID"
**映射规则**(与第 3 步的 Decision table 保持同一条纪律——env_blocked 绝不算 SUT 红):
| TierResult 事实 | `test-feedback` 上报的 `status` | | --- | --- | | 顶层 `envStatus==='env_blocked'` 或 `exitCode===78` | `env_blocked`(**绝不映射成 `failed`**——环境没跑起来,不是产品红) | | 上面都不成立且 `exitCode===0` | `passed` | | 其余(`exitCode` 非 0 且非 env_blocked) | `failed` |
`payload.tiers` 取自顶层 `tiers[]`,每项落 `{tier, passed, failed, skipped, total}`;`payload.failureCount` 是每层 `failed` 的和;`exitCode` / `regressions[]` 原样透传,供 cockpit 之后细分。同样只有该 task 的 **assignee** 能报(非 assignee 报会 403 / CLI 退出 4),与第 1 步的 `cloud skill ack` 是同一条鉴权。
旧判据是「正文里出现过 `failedNames`/`regressions`/`--json` 这几个词」——**一篇没跑过 `apc test`、字段名全靠背的报告照样满分**,而
Repo: Prismer-AI/PrismerCloud
Turn a known bug into a tight, red-capable reproducer, then prove the reproducer locks that…
Review a diff against its acceptance criteria in four segments (convention adherence, bug…
Five-dimension design audit (frontend UI/UX · server data-model & flow · endpoint spec ·…
Before merge, mechanize Documentation-First — derive the code delta from git diff, then…
Diagnose the local dev machine before any APC loop step — run apc env doctor, classify each…
Close out a local coding task on the bound daemon — stage, commit, branch, merge, push via…