ai-output-validation
Validates, parses, and sanitizes AI-generated outputs before they reach end users or downstream systems. Structured output enforcement, schema validation, and…
Transforms imperative instructions into declarative goals with verifiable success criteria. Enables autonomous looping until verified completion.
$ npx -y skills add DevelopersGlobal/ai-agent-skills --skill goal-driven-execution --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/goal-driven-executionContext preview
The summary Claude sees to decide when to auto-load this skill.
Transforms imperative instructions into declarative goals with verifiable success criteria. Enables autonomous looping until verified completion.
name: goal-driven-execution description: Transforms imperative instructions into declarative goals with verifiable success criteria. Enables autonomous looping until verified completion. category: think applies-to: [claude, gemini, cursor, copilot, any] version: 1.0.0
Andrej Karpathy's key insight: *"LLMs are exceptionally good at looping until they meet specific goals. Don't tell it what to do — give it success criteria and watch it go."*
This skill converts vague imperative instructions ("make the login work") into declarative goals with concrete, testable success criteria. Agents with clear goals self-correct autonomously. Agents with vague goals produce vague results and require constant intervention.
1. Read the full request. 2. Ask: *What is the user trying to achieve, not just what they asked for?* 3. Write the goal as: **"The task is complete when [observable, verifiable outcome]."**
Example transformation:
**Verify:** The goal statement is observable and testable by a third party.
4. List 3–7 specific, binary success criteria:
Success when: - [ ] All existing tests pass - [ ] New behavior X is demonstrated by test Y - [ ] No regressions in file Z - [ ] Manual check: [describe what to look for]
5. Each criterion must be **falsifiable** — you can clearly state when it passes or fails.
**Verify:** Every criterion can be checked without the original author.
6. Break the goal into ordered steps, each with its own verify check:
1. [Step] → verify: [command or check] 2. [Step] → verify: [command or check] 3. [Step] → verify: [command or check]
7. Identify the **first failure mode** — what's most likely to go wrong? Plan for it.
**Verify:** The plan is readable and each step is independently verifiable.
8. Follow the plan step-by-step. 9. At each verify checkpoint — actually run the check. Do not skip. 10. If a check fails: diagnose, fix, re-verify. Do not proceed past a failing check. 11. When all checks pass: report completion with evidence.
| Excuse | Rebuttal | |--------|----------| | "The goal is obvious" | Obvious goals still need explicit success criteria. What's obvious to you is ambiguous to an agent. | | "I'll know when it's done" | That's not a verifiable criterion. Write it down. | | "The tests will tell me" | Which tests? What do they cover? What don't they cover? | | "It's too simple for a plan" | Simple tasks rarely fail. Complex tasks without a plan always do. |
AI agent skills for production grade applications
Validates, parses, and sanitizes AI-generated outputs before they reach end users or downstream systems. Structured output enforcement, schema validation, and…
Design stable, versioned, self-documenting APIs. Easy to use correctly, hard to use incorrectly. Apply Hyrum's Law from day one.
Automated quality gates from commit to production. Every merge to main is potentially shippable. No manual steps in the deployment path.
Get layered, context-aware explanations of unfamiliar code. Understand what it does, why it was written that way, and how to work with it safely.
Structured code review focusing on correctness, security, and maintainability. Correctness before style. Every reviewer comment must be actionable.
Load minimum necessary context into agent context windows. Prevents token bloat, reduces cost, and improves focus. Only load what the current task needs.