paperclip-api
Use when managing Paperclip AI agent companies - creating tasks, managing agents, approving…
Use when choosing between engines or models for an agent pipeline and the answer must come from measurement on your own data, not from marketing pages: speech recognition for meeting recordings, the model that turns a transcript into notes, the critic model that checks it. Runs
$ npx -y skills add serejaris/personal-corp-os --skill benchmark --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/benchmarkContext preview
The summary Claude sees to decide when to auto-load this skill.
Use when choosing between engines or models for an agent pipeline and the answer must come from measurement on your own data, not from marketing pages: speech recognition for meeting recordings, the model that turns a transcript into notes, the critic model that checks it. Runs
name: benchmark description: >- Use when choosing between engines or models for an agent pipeline and the answer must come from measurement on your own data, not from marketing pages: speech recognition for meeting recordings, the model that turns a transcript into notes, the critic model that checks it. Runs every variant in an isolated container, measures time, cost per hour of input (USD and RUB, with price source and date) and quality (WER/CER and course terms for ASR; code checks plus a judge of a different model for LLM steps), writes results.json/csv and one static analytics page. Triggers on "бенчмарк", "сравни движки", "сколько стоит час записи", "какую модель взять для пайплайна", "benchmark the pipeline", "compare ASR engines".
Measure a pipeline on your own input before you pick an engine. One task file lists the variants; the scripts run them one by one in an isolated container and write numbers you can defend: time, cost per hour of input with a dated price source, quality against a reference.
Shape of a pipeline this skill knows:
`input (recording) → ASR engine → transcript → author model (notes, chapters) → code checks → judge model`
Each arrow is a slot. A variant fills one slot and keeps the rest fixed.
Not for: load testing, latency SLOs of a live service, model evals on public datasets.
1. **Isolation.** Run in a dedicated container (LXC/VM), not on a laptop and not next to production: the timing is clean and the keys live only there. Egress: provider APIs; private network closed except the services the task needs (for example a self-hosted ASR). 2. **Keys only from env.** The container gets a root-only env file; scripts read variable names from the task and never print values. Error bodies are redacted (some providers echo the key back). 3. **Same harness for every model.** LLM steps run through `claude -p` (Claude Code headless) against Anthropic-compatible endpoints; only the model changes. Tools, prompt, permissions stay the same. 4. **Judge is never the author's model.** `llm.py` refuses such a variant. 5. **No answer leaks.** The author must not see the published result of the same input (a «form reference» that is the answer itself). Give it a sibling of the same kind. 6. **Prices carry source and date.** `prices.json`: every item has `source` and `checked`; FX from the central bank of the day. A number without a source does not go into the report. 7. **What was not run is a row too.** Missing key, no balance, provider closed to new clients: write the reason and the list price, do not drop the row. 8. **Private input stays private.** Transcripts and recordings go to a private repo; the analytics page shows only aggregate numbers.
1. **Task file** (`task.example.json`): input audio, language, terms list, reference transcript, ASR variants, LLM variants with judges, `not_run` rows with reasons. 2. **Container.** Create it by your infra rules (the section «Container» below). Install `ffmpeg`, Python venv with `httpx jiwer faster-whisper`, Node and `@anthropic-ai/claude-code`. Put keys into `/etc/bench.env` (root 600) through a pipe from your secret store. 3. **ASR:** `python3 scripts/asr.py --task task.json --out runs/<name>` (engines run sequentially; ElevenLabs credits are read before and after). 4. **LLM:** write an adapter for your pipeline (contract in `scripts/llm.py`), then `python3 scripts/llm.py --task task.json --out runs/<name> --stage author` on the container and `--stage judge` where the judge's key lives (a subscription login on another machine is fine: the judge only reads files). 5. **Score:** `python3 scripts/score.py --task task.json --out runs/<name> --prices prices.json`. 6. **Page:** `python3 scripts/report.py --results runs/<name>/results.json --out report.html --notes notes.md`. 7. **Method note** next to the results: input, reference origin, what was not run and why, known biases.
| Slot | Engine key | Tested | Needs | |---|---|---|---| | ASR | `elevenlabs` (Scribe, optional keyterms) | yes | `ELEVENLABS_API_KEY` | | ASR | `openai_compat` (`/v1/audio/transcriptions`: self-hosted GigaAM, speaches, cloud) | yes | server URL, optional key | | ASR | `faster_whisper` (CPU, int8) | yes | model download once | | ASR | `openrouter_audio` (chat with `input_audio`, chunked) | request path only | `OPENROUTER_API_KEY` with balance | | LLM | `zai` (GLM via `api.z.ai/api/anthropic`) | yes | `ZAI_API_KEY` | | LLM | `anthropic-local` (the login on this machine) | yes | Claude Code logged in | | LLM | `openrouter`, `deepseek`, `anthropic-api`, `anthropic-oauth` | not yet | key with balance |
Yandex SpeechKit and SaluteSpeech have no adapter yet: add one to `ENGINES` in `asr.py` (input: audio path, output: text, raw response, usage) when you have a key.
ASR (against the reference, after lower case, `ё→е`, punctuation and speaker tags removed):
LLM: code checks before and after the fix round, fix rounds, judges' verdicts with critical / major / minor counts, tokens, turns, wall time.
Cost per hour of input: list price (`api_price`), tokens × list price (`tokens`), provider-reported cost (`reported`), or the CPU-seconds share of your server's month (`server_s
Manager uses your local Manager Config for repositories and boards; the distributed skill contains no personal workspace logs.
Use when managing Paperclip AI agent companies - creating tasks, managing agents, approving…
Orchestrate iterative visual style searches with branch prompts, decision graphs, feedback…
Робот агента Personal Corp. Агент делает себе голову робота из набора деталей, хранит её в…
Use when user asks for Claude Code usage stats, weekly analytics, project activity summary,…
Use when needing strategic project analysis from multiple independent expert perspectives.…
Use when creating or refactoring CLAUDE.md files - enforces best practices for size,…