analyze_current
Read and understand the current baseline implementation. Extract all relevant information about the existing approach without modifying anything, and record…
Define the comparison metrics and extract baseline values from the current implementation. Record them as a structured JSON entry so downstream phases and final evaluation can read them directly.
$ npx -y skills add Upsonic/Upsonic --skill benchmark --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/benchmarkContext preview
The summary Claude sees to decide when to auto-load this skill.
Define the comparison metrics and extract baseline values from the current implementation. Record them as a structured JSON entry so downstream phases and final evaluation can read them directly.
Define the comparison metrics and extract baseline values from the current implementation. Record them as a structured JSON entry so downstream phases and final evaluation can read them directly.
Phase 3 — after both current analysis and research analysis are complete.
| Parameter | Type | Description | |-----------|------|-------------| | experiment_path | path | `experiments/{research_name}/` |
1. **Define comparison metrics:**
2. **Extract baseline values:**
3. **Append a Phase 3 entry to `{experiment_path}/log.json`** under `phases`:
{
"name": "Phase 3: Benchmark",
"completed_at": "2026-04-17T10:45:00Z",
"metrics": [
{
"name": "accuracy",
"description": "Fraction of correctly classified samples.",
"higher_is_better": true,
"baseline": 0.8726,
"needs_computation": false
},
{
"name": "f1",
"description": "F1 score (binary, positive class).",
"higher_is_better": true,
"baseline": 0.7277,
"needs_computation": false
},
{
"name": "roc_auc",
"description": "Area under the ROC curve.",
"higher_is_better": true,
"baseline": 0.9274,
"needs_computation": false
},
{
"name": "training_time_seconds",
"description": "Wall-clock training time.",
"higher_is_better": false,
"baseline": null,
"needs_computation": true
}
],
"notes": "training_time_seconds must be added to both notebooks for a fair comparison."
}Do not overwrite earlier entries; append to the `phases` array.
Read and understand the current baseline implementation. Extract all relevant information about the existing approach without modifying anything, and record…
Compare baseline and new implementation results. Produce the machine-readable final report `result.json`, update `experiments.json`, and append a row to…
Set up and manage the experiment folder structure. This is Phase 0 — it runs before any analysis begins. All bookkeeping files are JSON (never markdown).
Create a new Jupyter notebook implementing the method from the research paper, using the same data as the baseline. Record implementation details and measured…
Maintain a **machine-readable** progress file so dashboards, CLIs, and notebooks can poll the experiment's state at any time. The file is a JSON document —…