Skip to content
Development
Skill

/verify-omnigent

Drive Omnigent the way a user does and prove a behavior with recorded evidence, using an isolated server, runner, and mock model plus a feature map of every user entry point. Load before reproducing a user-facing bug, before claiming a fix works, or when reviewing whether a

BOOST
From plugin
omnigent
10k21 skills
Install
$ npx -y skills add omnigent-ai/omnigent --skill verify-omnigent --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/verify-omnigent

Context preview

The summary Claude sees to decide when to auto-load this skill.

Drive Omnigent the way a user does and prove a behavior with recorded evidence, using an isolated server, runner, and mock model plus a feature map of every user entry point. Load before reproducing a user-facing bug, before claiming a fix works, or when reviewing whether a

SKILL.md

verify-omnigent.SKILL.md
name: verify-omnigent
description: Drive Omnigent the way a user does and prove a behavior with recorded evidence, using an isolated server, runner, and mock model plus a feature map of every user entry point. Load before reproducing a user-facing bug, before claiming a fix works, or when reviewing whether a change covered every surface (session vs. new-session composer, terminal strip vs. agent terminal, all twelve native harnesses). Covers the web UI, native harness terminals, and the CLI.

Verify Omnigent

Use this skill to see a behavior happen in the real app and to prove a change, not to reason about it from code. It has two parts:

  • **An isolated instance.** [`scripts/verify-env`](scripts/verify-env) wraps

`python -m dev.repro_env`: a server, runner, and mock model server on private ports, with their own config, data, Claude, and Codex directories. It never touches `~/.omnigent`, a running host daemon, or another developer server.

  • **A feature map.** [feature map](../../README.md) lists each user-facing

feature's entry points, the tests that drive them, and the traps. A fix is verified only when every entry point listed for its feature has proof.

Run all commands from the repository root. Put `scripts/` on your path or call `feature-map/skills/verify-omnigent/scripts/verify-env` directly.

Launch

1. Install dependencies, Chromium, and build the web UI for the checkout:

   uv sync --frozen --extra all --group test
   uv run --no-sync playwright install --with-deps chromium
   pnpm install --frozen-lockfile --filter web && pnpm --filter web run build

2. Start an instance and load its paths:

   verify-env start          # waits until the runner is online, about 10 seconds
   eval "$(verify-env paths)"

`start` writes `VERIFY_ROOT` to `.omnigent/verify/current`, so later shells find the same instance. The instance stops itself after its lease (default 90 minutes; `--lease SECONDS`, 60 to 21600).

Ready means the server answers, the runner reports online, and the mock model server answers. `start` fails with the environment's error and log path otherwise.

Ready does not check browser launch. Use the Chromium version installed by this checkout's Playwright package; if the environment provides it through `PLAYWRIGHT_BROWSERS_PATH`, keep that path available to the test process. `--ui-skip-build` reuses the existing bundle, so rebuild after frontend edits. Native-terminal journeys also need the relevant CLI and terminal prerequisites (such as tmux); a browser installation alone does not provide them.

In CI, the repro workflow already runs this environment. Use `python -m dev.repro_env exec -- ...` there instead of starting another.

Doctor

Run `verify-env doctor` before the first drive, after any failed drive, and whenever something looks off. It checks that the instance is ready, that its supervisor is alive, that the runner and model server answer, and it warns when the checkout has moved since launch. A warning about a moved checkout means the instance runs old code: stop it and start a new one.

Drive

1. Open the matching file in [feature map](../../README.md) and list every entry point for the behavior in question. 2. For each entry point, run the named test through the instance, recording on:

   verify-env run -- python -m pytest <tests/...py::test_name> \
     --ui-skip-build --video=on --screenshot=on \
     --output="$VERIFY_EVIDENCE/<feature>"

Tests under `tests/browser_ui/` need no instance: `uv run pytest <test> --browser-ui-skip-build --video=on --output=...`. Tests the feature file marks "own environment" also run with plain `uv run pytest`. Server/transport/component tests use plain pytest without browser recording or build flags. 3. For an entry point with no test, hand-drive it with Playwright against `OMNIGENT_REPRO_SERVER_URL` inside `verify-env run -- python <script>`, and script model replies with the mock helpers named in the feature map README. 4. To reproduce a bug, run the same journey on the unfixed code first and keep that evidence; then run it again on the fix.

A test that fails before reaching the user action has a setup failure, not a reproduction. Check its browser error and server/runner logs. When a test starts its own server or runner from inside an agent, check whether it inherited the parent's `OMNIGENT_RUNNER_*`, `RUNNER_SERVER_URL`, or `OMNIGENT_PROCESS_LOG_FILE`. Use an isolated child environment for that test; do not change the controlling agent's environment. Record setup failures, skipped variants, and unavailable credentials separately from behavior results.

Evidence

Everything under `$VERIFY_EVIDENCE` is the proof, one directory per feature. Instance logs, the database, and the mock model's recorded requests stay in `$VERIFY_ENV`. For each claim, record the checkout commit, installed harness versions, feature file, entry point ID, command, and resulting artifact. Keep source review, component tests, real process checks, browser drives, and live provider checks distinct. A map contract pass verifies references and structure; it is not evidence that the journeys ran. The proof standards are in the [feature map README](../../README.md#proof-and-coverage). Mock runs prove Omnigent's integration with Claude and Codex, not a live vendor model.

Cleanup

`verify-env stop` stops only this instance's supervisor, which stops its own server, runner, and model processes and saves the model's request log. It never deletes `$VERIFY_ROOT`, so the evidence survives; remove old roots under `.omnigent/verify/` or `/tmp/verify-omnigent.*` yourself once the evidence is attached. Never kill Omnigent processes by name: a developer's own server and host daemon may be running on the same machine.

Helpers

`scripts/verify-env` is the only helper. Its subcommands are `start [--lease SECONDS]`, `doctor`, `run -- COMMAND...`, `paths`, an

Read more
Ships withomnigent

Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.

Get the whole plugin

Other skills on omnigent.