00-andruia-consultant
Arquitecto de Soluciones Principal y Consultor Tecnológico de Andru.ia. Diagnostica y traza…
Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.
$ npx -y skills add sickn33/agentic-awesome-skills --skill ai-prompt-regression-testing --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/ai-prompt-regression-testingContext preview
The summary Claude sees to decide when to auto-load this skill.
Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.
name: ai-prompt-regression-testing description: 'Prompt engineering regression test matrix register: baseline outputs, semantic drift thresholds, judge evaluations, and golden dataset.' category: engineering risk: safe source: self source_type: self date_added: "2026-10-01" author: Ranjeet2063 tags: [ai, prompt-engineering, testing, evaluation, llm, benchmarking] tools: [] source_repo: Ranjeet2063/agentic-awesome-skills
**What it is:** Tracks regression baselines, evaluation rubrics, and automated judge verdicts to prevent output degradation across prompt revisions.
Provides a standardized, auditable framework and data model for **AI Prompt Regression Test Matrix** operations across distributed engineering and decentralized application systems.
1. Define the parameters, thresholds, and identity bindings required for the target operational register. 2. Select appropriate boundary enforcement values from validated enum select sets. 3. Export standardized artifacts (CSV table, SQL DDL, JSON Schema) to integrate into validation CI pipelines.
| # | Field Name | Type | SQL Type | JSON Schema Type | Notion Property Type | Example Value | |---|------------|------|----------|------------------|----------------------|---------------| | 1 | Prompt TestCase ID | `id` | `SERIAL PRIMARY KEY` | `integer` | Text | `PTEST-001` | | 2 | Prompt Identifier | `text` | `VARCHAR(64)` | `string` | Text | `soroban_code_refactor_v2` | | 3 | Target LLM Model Family | `select` | `VARCHAR(64)` | `string` | Select | `Claude 3.5 Sonnet` | | 4 | Evaluation Metric | `select` | `VARCHAR(64)` | `string` | Select | `AST Code Correctness` | | 5 | Semantic Drift Threshold | `number` | `NUMERIC(5,2)` | `number` | Number | `0.05` | | 6 | Golden Baseline Match % | `number` | `NUMERIC(5,2)` | `number` | Number | `98.50` | | 7 | Judge Model Evaluator | `text` | `VARCHAR(64)` | `string` | Text | `Gemini 1.5 Pro` | | 8 | Zero-Shot Reasoning Verified | `select` | `VARCHAR(16)` | `string` | Select | `Yes` | | 9 | Latency Bound Seconds | `number` | `NUMERIC(6,2)` | `number` | Number | `3.20` | | 10 | Test Suite Verdict | `select` | `VARCHAR(32)` | `string` | Select | `Passed` | | 11 | Benchmarking Date | `date` | `DATE` | `string, format: date` | Date | `2026-10-01` |
**Target LLM Model Family**
Claude 3.5 Sonnet | GPT-4o | Gemini 1.5 Pro | DeepSeek Coder
**Evaluation Metric**
AST Code Correctness | Semantic Embedding Cosine | Exact Match | Rubric Scoring
**Zero-Shot Reasoning Verified**
Yes | No
**Test Suite Verdict**
Passed | Degraded | Failed Regression
**Prompt**
How do I configure and track AI Prompt Regression Test Matrix for our production environment?
**Recommended Next Step**
> Generate the unified field schema, SQL DDL migration, and JSON validation schema to register into your system catalog. > > Workflow: Define criteria -> Run automated verification -> Record baseline -> Monitor invariants.
**Solution:** Always verify decimals using the explicit field mapping in this reference.
**Solution:** Cross-validate against the Security Audit register before deployment.
I want to establish a verified AI Prompt Regression Test Matrix register for our production protocol. Guide me through the required field parameters and output the corresponding SQL DDL and JSON Schema.
Find reusable instructions for your project, inspect their complete files, and keep an exact skill set you can review and reuse. Agentic Awesome Skills is a library of 2,652+ installable SKILL.md playbooks.
Repo: sickn33/agentic-awesome-skills
Arquitecto de Soluciones Principal y Consultor Tecnológico de Andru.ia. Diagnostica y traza…
Security audit, hardening, threat modeling (STRIDE/PASTA), Red/Blue Team, OWASP checks, code…
Ingeniero de Sistemas de Andru.ia. Diseña, redacta y despliega nuevas habilidades (skills)…
Estratega de Inteligencia de Dominio de Andru.ia. Analiza el nicho específico de un proyecto…
AI-powered presentation generation via the 2slides API — create slides from text, match a…
360 feedback register: reviewer, subject, review cycle, visibility, due date and score, as…