An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experiments, inspect evidence, and compare outcomes. Turn task evidence into actionable team and organization insights.
> /plugin marketplace add QoderAI/better-harness> /plugin install better-harness@better-harness
Repo: QoderAI/better-harness
What's inside
Analyze and improve your coding workflow with: Claude Code, Codex Desktop, Codex CLI, Qoder Desktop/CLI, Cursor, or GitHub Copilot CLI.
Choose the host you already use to get its exact installation, verification, invocation, and report-output steps. Better Harness does not use one universal entrypoint across every host.
This README shows inline setup for the most common hosts. Additional supported hosts (Qwen Code, Pi, Kimi Code, WorkBuddy, and Grok) keep their steps and boundaries in the installation guide and the public Host Adapter Matrix; see More adapters. README placement is a display choice, not a support-level claim.
Better Harness scopes behavior claims to relevant Task Episodes and the surrounding project mechanisms. Qoder and Cursor produce host-native Canvas reports; Claude Code, Codex, Qwen Code, GitHub Copilot, and Kimi Code produce self-contained HTML with paired Markdown. Missing or partial evidence remains explicit. See the Host Adapter Matrix for current coverage and output differences.
The report keeps missing evidence explicit and turns supported gaps into prioritized findings with an impact, expected output, scoped repair, and acceptance checks.
For delivery tracing, the interactive Harness Inspector follows product intent through agent activity, sessions, files, and commits in a read-only workspace, keeping evidence strength and limitations visible:
After you have comparable reports over time, the history view shows how the five Agent Work Loop dimensions move:
The static final frame summarizes historical Harness reports. It shows recorded trends, not causal proof of improvement. See how the demo was recorded.
AI coding agents change code fast, but the workflow around them is often the weak point:
Reviewing only the final diff misses these system-level problems. Better Harness analyzes the workflow around the diff: it gathers project evidence (and session evidence where supported), evaluates five connected dimensions, and turns concrete gaps into prioritized findings — each tied to its evidence, expected outcome, repair boundary, and validation route, so a team can improve one issue at a time.
Better Harness uses a feedforward-and-feedback loop that combines guidance available before work starts with signals available after the agent acts:
AGENTS.md, specs, Skills, and acceptance criteria
steer the agent before it acts.Across that loop, it evaluates five parts of delivery — the Agent Work Loop:
| Dimension | The question it answers | Backed by |
|---|---|---|
| Task Understanding | Does the agent know the goal and what "done" means? | Rules, AGENTS.md, specs, DESIGN.md |
| Controlled Execution | Is the work on supported, repeatable paths? | Skills, commands, MCP tools, sandbox boundaries |
| Change Validation | Is there evidence the change actually works? | Tests, lint, Hooks, observable diagnostics |
| Reliable Delivery | Does AI speed bypass quality checks or acceptance? | Human review, approvals, CI/CD, recovery paths |
| Learning Capture | Does the next task benefit from this one? | Loop Discovery, reusable SDLC Skills, Memory |
Running /better-harness establishes a task-bounded baseline and, depending on
the host, produces a visual report, a Markdown report, or both. The report
combines the five-part overview, prioritized findings, detected agent assets,
and an evidence brief. Each finding includes a repair action that drafts a
scoped fix plan for review.
Better Harness is deliberately honest: unobserved behavior stays explicit instead of becoming an unsupported score or claim. Passing a current check proves that the intervention was exercised; only a comparable later result can prove that the loop improved.
Better Harness opens three connected layers, not only a slash-command prompt:
/better-harness workflow, evidence
collectors, analyzers, renderers, and thin
host adapters.The three layers share the same boundary: configured assets can establish that a mechanism exists, but only linked task evidence can establish that it was used or improved an outcome.
The architecture keeps the three evidence domains independent until unified analysis by the lead agent. Every result retains a visible evidence source, owner, and validation route.
Installation differs by coding agent. Install Better Harness separately for each host, except that Qoder CLI can use the version bundled with Qoder Desktop. After installing or updating a plugin, start a new session or task so the host reloads its plugin inventory.
Register this repository as a Claude Code marketplace:
/plugin marketplace add QoderAI/better-harness
Then install Better Harness:
/plugin install better-harness@better-harness
Verify discovery from the shell:
claude plugin details better-harness@better-harness
The details should include Skills (1) better-harness. Then start a new Claude
session in the repository you want to analyze and run the report prompt:
/better-harness analyze this project's AI coding workflow and generate an evidence-backed report
Claude Code defaults to a self-contained report.html with paired report.md
and findings.json under the repository's .claude/better-harness report root.
Ask for inline or no-files output to keep the result in chat only. Workspace-
matching local Claude sessions are included when available; missing evidence
stays explicit rather than being inferred.
@better-harness analyze this project's AI coding workflow and generate an evidence-backed report
Use https://github.com/QoderAI/better-harness.git with Git ref main.

Add the repository source:
codex plugin marketplace add \
'https://github.com/QoderAI/better-harness.git' \
--ref main
Then inspect and install Better Harness:
codex plugin list --marketplace better-harness
codex plugin add better-harness@better-harness
Start a new Codex task in the repository you want to analyze and run the report prompt:
$better-harness:better-harness analyze this project's AI coding workflow and generate an evidence-backed report
Use the repository URL with marketplace add, not a raw marketplace.json
URL. Current Codex builds use plugin add and --marketplace; examples that
use plugin install or --source target a different CLI contract.
Better Harness is built into the Qoder desktop app, so no Marketplace or local plugin installation is required there. Choose either entry point:
From a session: Open the repository you want to analyze, start a new session, and run the report prompt:
/better-harness analyze this project's AI coding workflow and generate an evidence-backed report
From Quest (Qoder 1.18.0+): Open Quest, then select Better Harness (Beta) from the left sidebar.
Showing a partial view of a very large repo.
FAQ
better-harness is a Claude Code plugin with 4 hand-picked skills for development work, indexed on Flowy. Install it with the command on its page. It includes memory-recap, generate-harness-dsl, better-harness. Its skills do not fire on their own yet. Request auto-invocation to have Flowy route them as you prompt. Free and open source.
Is this plugin yours?
Claim it with GitHubSubmit a pluginPromote it