Skip to content
Testing
Skill

/create-adapter

Scaffold a new Harbor benchmark adapter by running `harbor adapter init` and then guide implementation using the Adapters Agent Guide as the authoritative spec.

From plugin
lhtb
5285 skills
Install
$ npx -y skills add zli12321/LHTB --skill create-adapter --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/create-adapter

Context preview

The summary Claude sees to decide when to auto-load this skill.

Scaffold a new Harbor benchmark adapter by running `harbor adapter init` and then guide implementation using the Adapters Agent Guide as the authoritative spec.

SKILL.md

create-adapter.SKILL.md
name: create-adapter
description: Scaffold a new Harbor benchmark adapter by running `harbor adapter init` and then guide implementation using the Adapters Agent Guide as the authoritative spec.

Create Adapter

Bootstrap a new benchmark adapter in the Harbor repository. This skill scaffolds the adapter directory with `harbor adapter init` and then defers to the adapter tutorial for every implementation decision.

Authoritative reference

The adapter tutorial is the authoritative specification for this skill. Read it in full before taking any action beyond scaffolding:

docs/content/docs/datasets/adapters.mdx

That path is relative to the Harbor repo root (a skill prerequisite — see below). The tutorial contains:

  • Required directory structures for the generated tasks and the adapter code package.
  • Step-by-step process (steps 1-9) covering benchmark analysis, oracle verification, parity experiments, dataset registration, and submission.
  • Schemas for `task.toml`, `parity_experiment.json`, `adapter_metadata.json`, and `dataset.toml`.
  • Parity matching criterion, pre-flight checklist, and debug playbook.
  • README format rules (machine-parsed; deviations break automation).

Do not substitute prior knowledge for the contents of that file. Treat it as the contract.

Prerequisites

  • Harbor CLI is installed and available on `PATH` (`harbor --version` succeeds).
  • Working directory is the Harbor repository root.
  • Harbor checkout is current: run `git fetch origin && git status` and pull `main` if behind. Stale checkouts miss recent adapter and agent fixes and are a common source of spurious parity failures later on.

Workflow

1. Read the tutorial

Before any filesystem or CLI action, Read `docs/content/docs/datasets/adapters.mdx` in full. Pay particular attention to:

  • **Required Directory Structures** — the contract for generated task and adapter layouts.
  • **Step 1. Understand the Original Benchmark** — what to identify upstream before coding.
  • **Step 8. Register the Dataset → Naming rules** — the `name` field and `<org>/<task>` format requirements.

2. Gather benchmark context from the user

Collect the following before scaffolding. If the user has not provided an item, ask before proceeding — these inputs shape the scaffold and the tutorial steps that follow.

| Field | Why it matters | |-------|---------------| | Adapter name | Lowercase, hyphen-separated. Must match the benchmark's common identifier (e.g., `swe-bench`, `aider-polyglot`). Becomes the directory under `adapters/` and, after underscore conversion, the Python package name. | | Human-readable name | Passed via `--name`; appears in the generated README. | | Upstream repo URL | Needed for tutorial Step 1 (benchmark analysis) and for the `original_parity_repo` field later. | | Oracle solutions available? | If the benchmark ships reference solutions, use them. If not, oracle solutions must be LLM-generated (tutorial Step 3 → "Benchmarks without oracle solutions"). | | Agent scenario | Which of Step 4's scenarios applies: (1) existing compatible agent, (2) fork + add LLM agent, or (3) custom agent. This shapes the parity plan. |

3. Create the adapter branch

Per tutorial Step 2, work on a dedicated branch:

git checkout -b <adapter-name>-adapter

4. Run the scaffold

Prefer the non-interactive form when both names are known:

harbor adapter init <adapter-name> --name "<Human-Readable Name>"

Otherwise run `harbor adapter init` interactively and let the CLI prompt.

Expected output: a new directory at `adapters/<adapter-name>/` containing `pyproject.toml`, `README.md`, `src/<adapter_name>/`, and the template task files under `src/<adapter_name>/task-template/`. Verify the directory exists before continuing.

5. Hand off to the tutorial

Continue from "Step 1. Understand the Original Benchmark" in the tutorial. Do not invent structure, field names, or workflow beyond what the guide specifies.

**High-priority gotchas** (each is documented in the tutorial, but these are the most common adapter-build failures — surface them proactively as you work through the steps):

  • **Every generated `task.toml` must contain a `name` field under `[task]`.** `main.py` is responsible for deriving a sanitized, unique, registry-safe name for every task. Tasks without a `name` cannot be registered. See the tutorial's "Naming rules" table.
  • **Task names must be stable across adapter runs.** Unstable names churn registry digests on republish. If upstream lacks stable identifiers, mint a deterministic scheme (e.g., `{dataset}-1`, `{dataset}-2`) from a reproducible sort.
  • **`version = "1.0"` in `task.toml` is the schema version — leave it alone.** Dataset versions are publish-time tags requested in the PR description, not a field in `task.toml` or `dataset.toml`.
  • **`main.py` must support `--output-dir`, `--limit`, `--overwrite`, and `--task-ids`.** These flags are required for reproducible runs and task-level debugging.
  • **The generated `README.md` is parsed by downstream automation.** Fill in every section exactly as the template defines; put extra context in the **Notes** section or in the `notes` fields of `parity_experiment.json` / `adapter_metadata.json`. Do not add, rename, reorder, or remove sections.
  • **Do not run parity experiments unilaterally.** Tutorial Step 4 requires team coordination on agents, models, and number of runs before incurring API costs. Complete sanity checks first, and execute full runs symmetrically on both sides.

Reference adapters by scenario

When implementation questions come up, point at an existing adapter that matches the benchmark's shape:

| Scenario | Example adapter | When to use | |----------|----------------|-------------| | Compatible agent already exists | `adapters/adebench/` | Upstream already supports Claude-Code / Codex / OpenHands / Gemini-CLI | | Fork upstream + add LLM agent | `adapters/evoeval/` | LLM-based benchmark with

Read more
Ships withlhtb

Long-Horizon Terminal-Bench is a 46-task benchmark for measuring how well LLM agents sustain useful work in a containerized terminal over hundreds of steps.

Get the whole plugin
Stats
529
Stars
34
Forks
Active
Maintenance
Python
Language
Apache-2.0
License
4d ago
Last commit
1mo ago
Created

Repo: zli12321/LHTB