coordinate-external-ag…
Coordinate independently operated external agents through durable handoffs. Use when work crosses hosts, sessions, accounts, services, queues, boards, pull…
Route language-model training, data, evaluation, checkpoint, representation, steering, ablation, jailbreak, tool-calling, Apple-runtime, and benchmark requests. Use when the primary workflow or Socket owner is unclear.
$ npx -y skills add gaelic-ghost/socket --skill choose-model-lab-workflow --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/choose-model-lab-workflowContext preview
The summary Claude sees to decide when to auto-load this skill.
Route language-model training, data, evaluation, checkpoint, representation, steering, ablation, jailbreak, tool-calling, Apple-runtime, and benchmark requests. Use when the primary workflow or Socket owner is unclear.
name: choose-model-lab-workflow description: Route language-model training, data, evaluation, checkpoint, representation, steering, ablation, jailbreak, tool-calling, Apple-runtime, and benchmark requests. Use when the primary workflow or Socket owner is unclear.
Select one primary workflow, name any supporting workflows, and make the evidence boundary explicit before work begins.
| Requested outcome | Primary skill | | --- | --- | | Define a hypothesis, controls, budget, and artifacts | `design-model-experiment` | | Curate, transform, split, or document examples | `prepare-language-model-dataset` | | Run SFT, LoRA, QLoRA, or a full parameter update | `fine-tune-language-model` | | Measure capability, behavior, quality, or safety | `evaluate-language-model` | | Decide which checkpoint is better and why | `compare-model-checkpoints` | | Choose Core AI, Core ML, MLX, ExecuTorch, or Foundation Models | `choose-apple-model-runtime` | | Locate or test internal representations | `research-model-representations` | | Apply activation or weight-space behavior steering | `steer-language-model-behavior` | | Remove or suppress a refusal direction | `ablate-refusal-representations` | | Measure jailbreak or prompt-injection robustness | `evaluate-jailbreak-resilience` | | Measure tool selection, arguments, execution, or recovery | `evaluate-tool-calling-model` | | Compare latency, memory, energy, throughput, or artifact size | `benchmark-model-runtime` | | Run preference optimization such as DPO/ORPO | Keep the experiment and eval here; use the supported TRL workflow through `fine-tune-language-model` until a stable standalone skill is earned | | Pretrain or continue pretraining a foundation model | Do not collapse it into fine-tuning; define the distributed/corpus contract and treat `train-language-model` as a deferred skill candidate | | Merge adapters or quantize/package an artifact | Use `compare-model-checkpoints` around the exact transformation; use the project-native tool and evaluate the deployable output | | Evaluate an agent skill, plugin, or host harness rather than a model protocol | Hand off to `agent-engineering-skills` and `agent-portability-skills` |
State:
1. the primary skill; 2. supporting skills in execution order; 3. the controlled variable; 4. the artifact or metric that proves completion; 5. any paid compute, data-access, or deployment authorization required.
Do not silently turn research planning into a paid run, model publication, or production-system test.
Stuff for Agents on macOS Promo audio: Socket Codex Marketplace Promo
Coordinate independently operated external agents through durable handoffs. Use when work crosses hosts, sessions, accounts, services, queues, boards, pull…
Assign worktree, branch, write, validation, integration, and cleanup ownership before parallel repository work. Use when a worker will inspect or modify…
Design framework-neutral agent and automation workflows before implementation. Use when choosing between Codex app automations, codex exec, Codex subagents,…
Design evaluation workflows for agent, skill, prompt, and automation behavior before implementation. Use when choosing eval cases, graders, thresholds,…
Design safe n8n workflows with deterministic routing, credentials, idempotency, recovery, local-model checks, drafts, and exact approval gates.
Coordinate bounded worker tasks with a launch envelope, report-back, escalation, and synthesis contract. Use before spawning, resuming, steering, cancelling,…