Skip to content
Development
Skill

/ia-reflect

Session retrospective and skill audit. Use when asked to reflect, do a retrospective, review lessons learned, audit what went well or wrong, or review session effectiveness.

From plugin
whetstone
3333 skills19 agents38 commands1 MCP
Install
$ npx -y skills add iliaal/whetstone --skill ia-reflect --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/ia-reflect

Context preview

The summary Claude sees to decide when to auto-load this skill.

Session retrospective and skill audit. Use when asked to reflect, do a retrospective, review lessons learned, audit what went well or wrong, or review session effectiveness.

SKILL.md

ia-reflect.SKILL.md
name: ia-reflect
class: tool
description: >-
  Session retrospective and skill audit. Use when asked to reflect, do a
  retrospective, review lessons learned, audit what went well or wrong, or
  review session effectiveness.

Reflect

Success Criteria

  • Every mistake/friction point cites the specific moment and its impact
  • Improvements are actionable and prioritized (cap defined in step 4)
  • Each skill audit proposes measurable changes (not vague suggestions)
  • Memory persistence follows existing authorization, or the user selects concrete proposed items before any write
  • If review activity occurred, review-trap candidates are reported; persist only with authorization, or explicitly report no candidates

Process

1. Session Review

Scan the full conversation. For each finding, cite the specific exchange (quote or paraphrase) and its impact.

| Category | Signal | |----------|--------| | **Mistakes** | Wrong outputs, incorrect assumptions, hallucinated facts | | **Friction** | Repeated clarifications, verbose responses, misread intent | | **Wasted effort** | Work discarded, wrong approaches tried first | | **Wins** | Approaches worth repeating, smooth interactions |

Skip one-time typos, external tool failures, and issues outside agent control.

2. Review Activity Scan (if applicable)

Collect candidates in the response. A retrospective alone does not authorize memory writes or skill edits; apply only changes already authorized by the user or approved in steps 4 and 5.

If the session included PR or MR review activity in either direction, run this scan before moving on. Skip only if no reviews happened.

**Inbound (my code was reviewed):** For each review comment received:

  • Did I accept it? If yes, what pattern did the reviewer catch that I missed? Is it a recurring blind spot? Propose a one-line memory candidate for step 4.
  • Did I push back? If I was right and the reviewer was wrong, nothing to capture. If I was wrong and had to retract mid-thread, capture what I learned.

**Outbound (I reviewed someone else's code):** For each comment I authored:

  • Was it accepted? Nothing to capture -- good call.
  • Was it rejected with a valid counter? That's a review trap. Capture the pattern: what heuristic did I apply that produced a wrong comment?

"No harvestable items" is a valid outcome -- say so explicitly. Don't let the step quietly drop off.

3. Operational Learnings

Before listing improvements, scan the session for operational insights worth preserving. Apply the 5-minute filter: would knowing this save 5+ minutes in a future session? If yes, include it. Examples: a project-specific quirk, a project command that failed for a project-specific reason, an approach that worked better than expected.

Exclude harness-level noise — "File has not been read yet", token-limit truncations, bash-quoting slips, and other tooling artifacts. Those aren't project learnings; capture the *project's* behavior, not the agent's mechanics.

Also scan for **information-access gaps**: points where the session stalled or guessed because the agent lacked read access to something a human would have checked — dev-server logs, a third-party dashboard, a staging database, CI output. Distinct from the harness noise excluded above: a one-off tooling hiccup isn't reusable, but a standing access gap is, since granting access pays off in every future session. Each gap is an improvement candidate ("grant readonly access to X" or "pipe X into a file the agent can read"), often higher-leverage than a prompt tweak.

4. Improvements

Numbered list of **concrete improvements**, ranked by impact. Each item: one sentence, imperative, actionable. Cap at 10 items: if more surface, the bottom items are noise -- drop them rather than batching or splitting.

For items not already authorized for persistence, present the concrete candidates and ask which to remember. Use the active harness's supported approval interface, or ask directly in chat. Do not ask again for items the user already authorized.

Save authorized items in the project's configured memory location using the active harness's file-editing tool and memory format. In Claude Code, inspect `~/.claude/projects/<project-slug>/memory/` and its MEMORY.md index; use the configured project slug rather than inventing one.

Before writing, grep the existing memory directory for the item's key terms. On a near-duplicate, update that file instead of adding a second. On a direct contradiction with an entry already on file ("use tabs" when "use spaces" is recorded), do not blind-append — surface both and let the user choose merge, replace, or keep-both. Silent duplicate and contradiction accumulation is the main way a curated memory index rots.

5. Skill Audit (if skills were used)

For each skill invoked during the session:

**A. Self-check gate** -- If the skill lacks success criteria + verification loop:

  • Propose `## Success Criteria` at top (3-5 measurable checks)
  • Propose `## Self-Check` at bottom: "Verify all success criteria are met before presenting output. If not, iterate (max 5 times)."

**B. Token efficiency** -- Flag: redundant phrasing, mergeable sections, oversized examples, "Claude already knows this" content, inert frontmatter metadata.

**C. Other** -- Missing edge cases, vague directives (rewrite as measurable criteria or remove), naked negations (add "do Y instead" or remove).

**D. Guidance mismatch** -- fires when a skill was invoked and its advice turned out wrong, stale, or inapplicable *here*. A, B, and C all judge a skill standing alone; this one anchors the finding to the line that actually misfired. Record four fields, all required:

  • the **verbatim excerpt** from SKILL.md or its reference that produced the wrong behavior
  • the **project context** that made it not apply (language, runner, framework version, house convention)
  • **what happened** when it was followed
  • **what was done instead**

A skill invoked with no mi

Read more
Ships withwhetstone

A Claude Code plugin that makes AI coding agents follow engineering discipline. Plan before coding. Verify before claiming done. Find root cause before patching. Review before merge. Skills activate based on file type and task signals, not manual toggling.

Get the whole plugin

Other skills on whetstone.