analyze_current
Read and understand the current baseline implementation. Extract all relevant information about the existing approach without modifying anything, and record…
Compare baseline and new implementation results. Produce the machine-readable final report `result.json`, update `experiments.json`, and append a row to `comparison.json`.
$ npx -y skills add Upsonic/Upsonic --skill evaluate --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/evaluateContext preview
The summary Claude sees to decide when to auto-load this skill.
Compare baseline and new implementation results. Produce the machine-readable final report `result.json`, update `experiments.json`, and append a row to `comparison.json`.
Compare baseline and new implementation results. Produce the machine-readable final report `result.json`, update `experiments.json`, and append a row to `comparison.json`.
Phase 5 — after the new implementation is complete and metrics are collected.
| Parameter | Type | Description | |-----------|------|-------------| | experiment_path | path | `experiments/{research_name}/` | | research_name | string | Name of this experiment |
1. **Collect all metrics** from `log.json` (Phase 3 baseline entry + Phase 4 new method entry).
2. **Determine verdict:**
3. **Write `{experiment_path}/result.json`** in the exact schema below. Always valid JSON; never leave fields undefined — use `null` for unknown values.
{
"name": "{research_name}",
"verdict": "BETTER",
"summary": "2-3 paragraphs explaining what the new method does, how it fundamentally differs from the baseline, and what trade-offs it makes.",
"explanation": "2-3 sentences explaining WHY this verdict was reached. Reference specific metrics and their differences. Be concrete — mention numbers, not vague statements.",
"comparison": {
"metrics": [
{
"name": "accuracy",
"current": 0.853,
"new": 0.872,
"diff": 0.019,
"diff_display": "+0.019",
"unit": null,
"higher_is_better": true,
"better": "new"
},
{
"name": "training_time_seconds",
"current": 2.0,
"new": 45.0,
"diff": 43.0,
"diff_display": "+43.0",
"unit": "seconds",
"higher_is_better": false,
"better": "current"
}
]
},
"file_locations": {
"current_notebook": "experiments/{research_name}/current.ipynb",
"current_data": "experiments/{research_name}/current_data/",
"new_notebook": "experiments/{research_name}/new.ipynb",
"research_source": "experiments/{research_name}/research.pdf",
"experiment_log": "experiments/{research_name}/log.json"
}
}4. **Update `experiments/experiments.json`:**
5. **Update `experiments/comparison.json`:**
{
"name": "{research_name}",
"date": "YYYY-MM-DD",
"baseline": "{baseline_model}",
"new_method": "{new_method}",
"key_metric": {"name": "accuracy", "baseline": 0.853, "new": 0.872},
"verdict": "BETTER"
}6. **Update `{experiment_path}/log.json`** — append a Phase 5 entry:
{
"name": "Phase 5: Evaluate",
"completed_at": "2026-04-17T11:40:00Z",
"verdict": "BETTER",
"key_change": "accuracy +0.019 (new > current)",
"files_written": ["result.json", "experiments.json", "comparison.json"]
}Read and understand the current baseline implementation. Extract all relevant information about the existing approach without modifying anything, and record…
Define the comparison metrics and extract baseline values from the current implementation. Record them as a structured JSON entry so downstream phases and…
Set up and manage the experiment folder structure. This is Phase 0 — it runs before any analysis begins. All bookkeeping files are JSON (never markdown).
Create a new Jupyter notebook implementing the method from the research paper, using the same data as the baseline. Record implementation details and measured…
Maintain a **machine-readable** progress file so dashboards, CLIs, and notebooks can poll the experiment's state at any time. The file is a JSON document —…