background-watch-hook
Use `vibe watch` to run a managed Harness waiter that returns to the same conversation later. Best for reviews, CI, files, logs, and other…
Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch messages — using real run evidence. Use when an Agent misbehaves (stalls, over-asks,
$ npx -y skills add avibe-bot/avibe --skill agent-prompt-audit --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/agent-prompt-auditContext preview
The summary Claude sees to decide when to auto-load this skill.
Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch messages — using real run evidence. Use when an Agent misbehaves (stalls, over-asks,
name: agent-prompt-audit slug: agent-prompt-audit description: Audit and improve the prompt surface of Avibe Agents across backends (Claude, Codex/GPT, OpenCode) — global and project rules, Agent system prompts, Skills, delegation briefs, and Task and Watch messages — using real run evidence. Use when an Agent misbehaves (stalls, over-asks, over-reaches, ignores or over-applies a rule), after a model or backend change, or when the user asks to review, clean up, or tighten prompts. version: 0.1.0
An Agent's behavior comes from every piece of text that reaches it over its lifecycle, not just its system prompt. The audit's job is to find which text causes the behavior the user sees, on the backend and model that actually ran it, and to propose the smallest change that fixes it. Judge each instruction by what it does to behavior, not by its length: sometimes the fix is adding a missing reason or exit, and a clean surface is a valid result.
Deliver a report of findings — each with its evidence, confidence, and a concrete proposed change — and apply changes only when asked.
**Evidence over reading.** What Agents actually did beats what the text seems to say. Start from real runs and user corrections ("you stopped", "why didn't you report", "don't ask me that"), trace each symptom to the line that caused it or the missing line that would have prevented it, and use `git blame` to learn what incident a rule was written for and whether it still happens. A finding without evidence or documented model behavior is a flag, not a fix.
**Context is kept; constraints must earn their place.** Facts only the author knows — environment, contracts, ownership, quality bar, the reason behind a rule — are what prompts are for. Behavioral constraints are what go stale. Keep exact scripts where one sequence is safe (destructive commands, auth, merge gates, key custody), prohibitions against failures that still reproduce, and the scope bounds that make autonomy safe.
**Say intent and reason, not pressure or method.** Caps, `MUST/NEVER`, and emphasis without a reason make current models rigid; step scripts for judgment work and strategy coaching usually do worse than the model's own plan; fixed formats, word caps, and "don't narrate" rules produce silence or starved answers. Fossils — named-model workarounds, incident numbers as authority, "now/no longer" phrasing, one session's stumble made permanent — should become the current rule they stand for.
**Every stop needs an exit.** Agents run across turns, wake on callbacks, and hand work to each other, so the costliest defects are lifecycle gaps: a "stop/wait" with no statement of what the turn produces instead, asking permission for reversible in-scope steps, continuing without bounds after repeated failure, waiting with no durable waiter or expiry meaning, briefs missing the goal or report target, callbacks that say "done" without the result, and Task/Watch messages that restate rules on every fire.
**One home per rule, at the layer whose timing fits.** Always-loaded and recurring text has the most leverage and deserves the most scrutiny. Duplicates that disagree force the Agent to guess; keep the mechanism in one place and a principle or pointer elsewhere. Agreeing fallbacks are fine. Long procedures belong in on-demand Skills, not always-loaded rules.
**Shared text runs on every backend.** GPT/Codex tend to follow a bare prohibition or stop literally, so they need scope and exit conditions; strong Claude models tend to over-reach, so they need scope bounds and a definition of done; tool names and native mechanics dangle on other backends. Take model-specific behavior from the vendor's current docs, and lower confidence when you cannot reach them.
**A removal is a hypothesis.** For contested changes, compare behavior before and after with a scratch run on the target that produced the failure, and read the transcript rather than asking the model whether it needs the rule.
Verify against the current machine; these are starting points. Each backend also reads its own native configuration — config directories moved by environment variables (`CLAUDE_CONFIG_DIR`, `CODEX_HOME`, OpenCode's config path), and native subagent definitions such as `.claude/agents/`, `.codex/agents/`, or OpenCode agents — so resolve what the target backend actually loads rather than assuming default paths.
| Layer | Where | How it changes | | --- | --- | --- | | Avibe runtime prompt | `vibe debug prompt export --format json` lists every source; `vibe debug prompt export --format json --context-file <file>` renders a composition from the inputs you supply (backend, Agent instructions, Skill directory, context), so it approximates the target only as well as those inputs match (history in the Avibe repo `core/prompts/`, if checked out) | Proposal to the Avibe repository | | Global rules and native backend config | `~/.claude/CLAUDE.md`, `~/.codex/AGENTS.md`, Codex `developer_instructions` in `$CODEX_HOME/config.toml` (default `~/.codex`), OpenCode `instructions` in global or project `opencode.json[c]`, … | Edit the source if the file is generated or imports others | | Project rules | nearest `AGENTS.md` / `CLAUDE.md` chain | The repository's own delivery process | | Agent system prompt, model, effort | `vibe agent show <name> --json` | `vibe agent update <name> --system-prompt-file <file>` | | Skills | user skill dirs (follow symlinks), Avibe `skills/`, project `.agents/skills/` | The directory's owner | | Task and Watch messages (re-sent every fire) | `vibe task list` / `vibe watch list` for ids, then `vibe task show <id>` / `vibe watch show <id>` for the full text | `vibe task update`, `vibe watch update` | | User preferences (read on demand) | `~/.avibe/state/user_preferences.md`; inspect only the reported user's part, and only when the transcript shows it was read | The user | | Delegation briefs and callbacks
The local-first Agent OS — your AI partner lives on your own machine. Drive the official Claude Code, Codex & OpenCode from your browser or any chat app.
Repo: avibe-bot/avibe
Use `vibe watch` to run a managed Harness waiter that returns to the same conversation later. Best for reviews, CI, files, logs, and other…
Use Avibe Harness for durable Agent delegation, Sessions, scheduled Tasks, Watches, Runs, queues, and work that must continue beyond the current turn.
Use Avibe Vault for API keys, tokens, passwords, protected credentials, authenticated HTTP requests, or digest signing without exposing secret values to the…
Safely inspect and modify local Avibe configuration, routing, runtime settings, watches, scheduled tasks, Avibe Cloud remote access, and operational state.…
Build, inspect, update, restore, or share Avibe Show Pages for visual explanations, diagrams, reports, or interactive prototypes. Covers the page workspace and…