swe-marathon-five-arm
Migrate historical SWE-Marathon agent configurations to the shared Codex runtime.
Use for `/loopx-pr-review` or evidence-backed PR queue review. Run `loopx pr-review` first, execute the capability-owned review plan for each selected exact head, then publish full bilingual PR reviews (complete Chinese five-block review plus one concise English verdict) that
$ npx -y skills add loopx-project/loopx --skill loopx-pr-review --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/loopx-pr-reviewContext preview
The summary Claude sees to decide when to auto-load this skill.
Use for `/loopx-pr-review` or evidence-backed PR queue review. Run `loopx pr-review` first, execute the capability-owned review plan for each selected exact head, then publish full bilingual PR reviews (complete Chinese five-block review plus one concise English verdict) that
name: loopx-pr-review description: Use for `/loopx-pr-review` or evidence-backed PR queue review. Run `loopx pr-review` first, execute the capability-owned review plan for each selected exact head, then publish full bilingual PR reviews (complete Chinese five-block review plus one concise English verdict) that match the verified findings. Use `loopx-pr-merge` for approval or merge actions.
This skill is a thin host adapter. The built-in `pull-request-review` capability owns review depth, evidence requirements, completeness, and verdict policy through the CLI packet. Do not copy those rules into this skill or replace them with a host-specific checklist.
Use this skill for `/loopx-pr-review`, explicit PR reviews, or review queues by state or time window. Route approval, merge, self-merge, and admin bypass to `loopx-pr-merge` (optional repo-kept workflow, not installed by default) after the evidence review is complete; it never replaces this skill's exact-head gate.
For named PRs, resolve heads and run repeatable `--target-exact-head NUMBER@HEAD_OID`. An omitted `--state` keeps ordinary queue discovery open-only while exact targets remain lifecycle-neutral; use explicit `--state merged|all` only for deliberate history or post-merge audit.
Translate only explicit filters:
When omitted, the CLI resolves `pull_request_review` from the standard machine capability editor; an absent namespace keeps the default `other-developers-first`. Words such as `today`, `open`, or `merged` are filters, not permission to return a table only; stats-only output needs an explicit opt-out such as `只统计` or `stats only`.
Save the full first JSON packet before printing a compact projection. Keep all paths named by `agent_response_contract.required_packet_fields_to_preserve`:
Do not pipe the only copy through `jq`. When an exhaustive queue request has `result_completeness.complete=false`, rerun with its `recommended_limit` before reviewing; `limit_scope=exact_targets` is already complete for the named targets.
Read `review_execution_contract.policy_revision` from the packet; require the result's `review_policy_revision` to equal it, never a literal this file pins. If missing or unequal, do not publish APPROVE; a conservative REQUEST_CHANGES is allowed only when it names the incompatible-policy gap. Do not retain expired temporary worktree overrides. Report incompatible policy rather than downgrading the review.
Read the PR's existing comments and cited documents first: a maintainer comment names the contract the change is judged against.
Receive pending current-session requests through `review_execution_contract.decision_procedure.establish_goal` before generic queue selection, including other agents sharing a GitHub account; task intake and claimed action authority are separate judgments. Resolve the selected requests with repeatable `--target-exact-head NUMBER@HEAD_OID`. Follow `scheduling_policy` and its ranked actionable `review_sequence`; explicit current-request PR selection may override ordering only, never `pull_requests[].review_action_kind` or exact-head idempotency. Generic `re-review`, `重新review`, and `复审` wording selects the named PR; it is not a force-refresh token. Todo/monitor prose may not select work. When `review_action_kind` is null, the row stays in `pull_requests` inventory but must not appear in `review_sequence`; its `review_plan` and `review_template` are null and `evidence_commands` is empty. Do one compact exact-head conclusion readback and report the existing verdict or bounded invalid/missing reason. Run a fresh audit only when the user explicitly requests fresh evidence despite that result, or supplies a concrete new concern; regenerate with `--fresh-audit-exact-head NUMBER@HEAD_OID`, then execute the complete current plan and never inherit the earlier approval. For every actionable PR:
1. Record the packet's exact head, then follow `review_execution_contract.decision_procedure`: the current goal and `problem_context` delivery judgment first, including on re-review, then `evidence_commands` and repository-native validation. 2. Fill `review_plan.result_template`; preserve missing evidence as `unverified` and never infer `verified` from metadata or CI. Execute its repository-reuse, default-off, authority and real-path counterfactuals rather than repeating them as prose. Fill `result.reviewer` per `review_execution_contract.reviewer_declaration`, open the body with its `body_marker` line, and read `problem_context.spec_basis`'s specification before the diff. 3. Apply `completion_gate` literally: save final Markdown in `review_body`, then check evidence and that exact body. Follow capability-owned floors and scope counterfactuals; prose cannot replace missing execution:
loopx --format json pr-review --check-result review-result.
A control plane with a durable state kernel for long-horizon agents and teams. Keep work moving and improving across sessions, with less human attention.
Repo: loopx-project/loopx
Migrate historical SWE-Marathon agent configurations to the shared Codex runtime.
在 Terminal-Bench 4.0 上做 codex harness 五臂对照(裸 codex / 原生 /goal / LoopX 三模式)。复用 SWE-Marathon…
Inspect authorized LoopX Goals, Todos and deliveries to explain progress, identify owner…
Use when acting as the operator or post-run analyst of a LoopX-managed benchmark experiment…
Qualify the exact final diff for a LoopX-managed goal. Use when goal policy enables…
Use when a connected LoopX project is asked to read, remember, record, index, register, or…