dr-claw
Dr. Claw skill for OpenClaw project discovery, idea intake, waiting-session triage, structured session control, event-driven notifications, and mobile…
Periodically check WandB metrics during training to catch problems early (NaN, loss divergence, idle GPUs). Avoids wasting GPU hours on broken runs. Use when training is running and you want automated health checks.
$ npx -y skills add OpenLAIR/dr-claw --skill aris-training-check --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/aris-training-checkContext preview
The summary Claude sees to decide when to auto-load this skill.
Periodically check WandB metrics during training to catch problems early (NaN, loss divergence, idle GPUs). Avoids wasting GPU hours on broken runs. Use when training is running and you want automated health checks.
name: aris-training-check description: Periodically check WandB metrics during training to catch problems early (NaN, loss divergence, idle GPUs). Avoids wasting GPU hours on broken runs. Use when training is running and you want automated health checks. argument-hint: "[wandb-run-path]" allowed-tools: Bash(*), Read, Grep, Glob, Write, Edit, mcp__codex__codex, mcp__codex__codex-reply license: MIT metadata: author: wanshuiyin/ARIS version: "1.0.0"
Periodically read WandB metrics during training to catch problems early. Do not wait until training finishes to discover it was a waste of GPU time.
import wandb
api = wandb.Api()
run = api.run("<entity>/<project>/<run_id>")
history = run.history()If WandB is unreachable (API error, network issue), fall back to reading the log file directly via SSH:
ssh server "tail -100 /path/to/training.log"
Check these signals:
| Signal | Judgment | Action | |--------|----------|--------| | NaN/Inf in loss | **Clearly bad** | Stop training, investigate | | Loss diverging (increasing for >N steps) | **Clearly bad** | Stop training, investigate | | Eval metrics significantly worse than baseline | **Clearly bad** | Stop training, investigate | | Loss decreasing, metrics improving | **Clearly fine** | Continue, increase check interval | | Loss flat but not diverging | **Unsure** | → Step 3 (Codex judgment) | | Metrics noisy, can't tell trend | **Unsure** | → Step 3 (Codex judgment) | | Slightly worse than baseline but still early | **Unsure** | → Step 3 (Codex judgment) |
Only escalate to Codex when the signal is ambiguous. For clearly good or clearly bad signals, act directly.
mcp__codex__codex:
config: {"model_reasoning_effort": "high"}
prompt: |
TRAINING HEALTH CHECK — need your judgment on ambiguous metrics.
Run: <entity>/<project>/<run_id>
Current epoch/step: X / Y total
Training loss (last 10 checkpoints): [values]
Eval metrics (last 3 evals): [values]
Baseline reference: [numbers from paper/reproduction]
What I'm unsure about: [specific concern]
Please respond with exactly one of:
- STOP: clearly problematic, should kill training
- CONTINUE: looks fine, check again next interval
- WAIT: not enough data to judge, check again sooner| Decision | Action | |----------|--------| | **Stop** | Kill the training session. Save the WandB run URL, key metrics, and reason for stopping. Log to project notes for debugging. | | **Continue** | Do nothing. Will be invoked again at next interval (increase interval if consistently healthy). | | **Wait** | Do nothing but keep the current short interval (don't increase). |
Training-check and [watchdog.py](../../tools/watchdog.py) operate at different levels:
| Layer | Tool | What it checks | Frequency | |-------|------|----------------|-----------| | Process health | watchdog.py | Session alive? GPU active? | Every 60s (continuous) | | Training quality | training-check | Loss trend? Metrics improving? | Every 10-60 min (periodic) |
Use both together:
After training is confirmed stable:
CronCreate (recurring, every 10 minutes initially):
"Run /aris-training-check for wandb run <entity>/<project>/<run_id>"As the check interval increases, delete the old CronCreate job and create a new one with the longer interval.
A Super AI Lab with massive AI Doctors as Assistants. Best IDE for Research via AI Power.
Repo: OpenLAIR/dr-claw
Dr. Claw skill for OpenClaw project discovery, idea intake, waiting-session triage, structured session control, event-driven notifications, and mobile…
Academic research assistant for literature reviews, paper analysis, and scholarly writing. Use when: reviewing academic papers, conducting literature reviews,…
Autonomous AI agent platform for building and deploying continuous agents. Use when creating visual workflow agents, deploying persistent autonomous agents, or…
Multi-agent orchestration framework for autonomous AI collaboration. Use when building teams of specialized agents working together on complex tasks, when you…
Framework for building LLM-powered applications with agents, chains, and RAG. Supports multiple providers (OpenAI, Anthropic, Google), 500+ integrations, ReAct…
Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices,…