/checkpoint-and-recover
Checkpoint state before risky operations, detect anomalies, and recover
$ npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill checkpoint-and-recover --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/checkpoint-and-recover
Context preview
The summary Claude sees to decide when to auto-load this skill.
Checkpoint state before risky operations, detect anomalies, and recover
SKILL.md
checkpoint-and-recover.SKILL.mdname: checkpoint-and-recover
description: Checkpoint state before risky operations, detect anomalies, and recover
gracefully
version: 1.0.0
category: experiment-execution
type: tactic
orchestrates:
- execution-monitoring
- result-collection
dependencies:
sops:
- execution-monitoring
- result-collection
Tactic: Checkpoint and Recover
Orchestration Pattern
FUNCTION checkpoint_and_recover(task, execute_fn):
// Pre-execution checkpoint
checkpoint = {
timestamp: now(),
task_id: task.id,
state: capture_current_state(),
files_modified: [],
outputs_produced: []
}
save_checkpoint(checkpoint)
TRY:
// Execute with monitoring
monitor = SPAWN execution-monitoring(task)
result = execute_fn(task)
// Post-execution validation
IF monitor.anomalies_detected:
RAISE AnomalyError(monitor.anomalies)
END
// Validate result integrity
validated = SPAWN result-collection(result, task.success_criterion)
IF validated.complete AND validated.consistent:
// Success — archive checkpoint (keep for audit trail)
archive_checkpoint(checkpoint)
RETURN {status: DONE, result: validated}
ELSE:
// Partial success — decide whether to keep or rollback
IF validated.partial_value > threshold:
archive_checkpoint(checkpoint)
RETURN {status: PARTIAL, result: validated, missing: validated.gaps}
ELSE:
restore_state(checkpoint)
RETURN {status: ROLLED_BACK, reason: validated.failure_reason}
END
END
CATCH error:
// Failure — diagnose and recover
diagnosis = diagnose_failure(error, checkpoint, task)
SWITCH diagnosis.severity:
CASE TRANSIENT:
// Retry without rollback (e.g., network timeout)
RETURN {status: RETRY, reason: diagnosis}
CASE CORRUPTING:
// Rollback to checkpoint
restore_state(checkpoint)
RETURN {status: ROLLED_BACK, reason: diagnosis}
CASE FATAL:
// Rollback and escalate
restore_state(checkpoint)
RETURN {status: FATAL, reason: diagnosis, escalate: true}
END
END
ENDDecision Criteria
| Condition | Action | |-----------|--------| | Task modifies existing files | MUST checkpoint before | | Task is read-only/analysis | Checkpoint optional | | Anomaly detected during execution | Pause, diagnose, decide | | Result partially valid | Keep if value > threshold | | Result invalid | Rollback to checkpoint | | Transient error (timeout, rate limit) | Retry without rollback | | Corrupting error (bad state) | Rollback then retry | | Fatal error (impossible task) | Rollback and escalate |
Checkpoint Contents
A checkpoint captures:
- Timestamp and task ID
- File system state (modified files' contents before modification)
- Execution context (variables, intermediate results)
- Dependencies state (which tasks were complete)
Recovery Strategies
1. **Retry**: Same task, same parameters (for transient failures) 2. **Retry with modification**: Same task, adjusted parameters (for NEEDS_CONTEXT) 3. **Rollback and skip**: Restore state, mark task BLOCKED, continue 4. **Rollback and escalate**: Restore state, report to orchestrator for human decision
<!-- BEGIN available-tables (generated) -->
Available SOPs
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | execution-monitoring | Monitor execution progress, detect anomalies, and report status | | result-collection | Collect experiment outputs — metrics, logs, artifacts — into structured result set |
<!-- END available-tables (generated) -->
Read more
name: checkpoint-and-recover description: Checkpoint state before risky operations, detect anomalies, and recover gracefully version: 1.0.0 category: experiment-execution type: tactic orchestrates: - execution-monitoring - result-collection dependencies: sops: - execution-monitoring - result-collection
Tactic: Checkpoint and Recover
Orchestration Pattern
FUNCTION checkpoint_and_recover(task, execute_fn):
// Pre-execution checkpoint
checkpoint = {
timestamp: now(),
task_id: task.id,
state: capture_current_state(),
files_modified: [],
outputs_produced: []
}
save_checkpoint(checkpoint)
TRY:
// Execute with monitoring
monitor = SPAWN execution-monitoring(task)
result = execute_fn(task)
// Post-execution validation
IF monitor.anomalies_detected:
RAISE AnomalyError(monitor.anomalies)
END
// Validate result integrity
validated = SPAWN result-collection(result, task.success_criterion)
IF validated.complete AND validated.consistent:
// Success — archive checkpoint (keep for audit trail)
archive_checkpoint(checkpoint)
RETURN {status: DONE, result: validated}
ELSE:
// Partial success — decide whether to keep or rollback
IF validated.partial_value > threshold:
archive_checkpoint(checkpoint)
RETURN {status: PARTIAL, result: validated, missing: validated.gaps}
ELSE:
restore_state(checkpoint)
RETURN {status: ROLLED_BACK, reason: validated.failure_reason}
END
END
CATCH error:
// Failure — diagnose and recover
diagnosis = diagnose_failure(error, checkpoint, task)
SWITCH diagnosis.severity:
CASE TRANSIENT:
// Retry without rollback (e.g., network timeout)
RETURN {status: RETRY, reason: diagnosis}
CASE CORRUPTING:
// Rollback to checkpoint
restore_state(checkpoint)
RETURN {status: ROLLED_BACK, reason: diagnosis}
CASE FATAL:
// Rollback and escalate
restore_state(checkpoint)
RETURN {status: FATAL, reason: diagnosis, escalate: true}
END
END
ENDDecision Criteria
| Condition | Action | |-----------|--------| | Task modifies existing files | MUST checkpoint before | | Task is read-only/analysis | Checkpoint optional | | Anomaly detected during execution | Pause, diagnose, decide | | Result partially valid | Keep if value > threshold | | Result invalid | Rollback to checkpoint | | Transient error (timeout, rate limit) | Retry without rollback | | Corrupting error (bad state) | Rollback then retry | | Fatal error (impossible task) | Rollback and escalate |
Checkpoint Contents
A checkpoint captures:
- Timestamp and task ID
- File system state (modified files' contents before modification)
- Execution context (variables, intermediate results)
- Dependencies state (which tasks were complete)
Recovery Strategies
1. **Retry**: Same task, same parameters (for transient failures) 2. **Retry with modification**: Same task, adjusted parameters (for NEEDS_CONTEXT) 3. **Rollback and skip**: Restore state, mark task BLOCKED, continue 4. **Rollback and escalate**: Restore state, report to orchestrator for human decision
<!-- BEGIN available-tables (generated) -->
Available SOPs
Optional, no fixed order; the final leaf is always a sop.
| SOP | When to use | | --- | --- | | execution-monitoring | Monitor execution progress, detect anomalies, and report status | | result-collection | Collect experiment outputs — metrics, logs, artifacts — into structured result set |
<!-- END available-tables (generated) -->
The complete research orchestration system for AI-native science. What It Does Design Philosophy Architecture (v3.2.2) Quick Start Configuration Roadmap License DARE is not a tool that helps you do research. It is the researcher.
Repo: yogsoth-ai/de-anthropocentric-research-engine
Other skills on de-anthropocentric-research-engine.
- /formated-results
Closing skill for the research-executor, loaded as the last step of formated-specs. Summarize the design just produced into one research-result JSON fenced block in your reply. Do not execute the research.
Open skill - /formated-specs
Spec-slot skill for the research-executor. Emit the 4-layer DARE orchestration of the assigned topic as one research-graph JSON fenced block in your reply. Replaces the generic spec-writing step.
Open skill - /injection-fidelity
Loss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's per-axis pressure. Judge enactment of the card, never whether the research is good.
Open skill - /ladder-quality-order
Loss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.
Open skill - /optimization-loop
The optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to one weight (attribute-first), and recovers from disk after compaction. Control flow is fully scripted; only the
Open skill - /acu-nugget-recall
Tactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall checks; cannot run without a target summary.
Open skill

