Skip to content
Development
Skill

/canary-deploy-patterns

Traffic splitting, health checks, automated rollback, progressive delivery, and canary analysis for safe deployments.

From plugin
vibecosystem
532200 skills138 agents7 hooks
Install
$ npx -y skills add vibeeval/vibecosystem --skill canary-deploy-patterns --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/canary-deploy-patterns

Context preview

The summary Claude sees to decide when to auto-load this skill.

Traffic splitting, health checks, automated rollback, progressive delivery, and canary analysis for safe deployments.

SKILL.md

canary-deploy-patterns.SKILL.md
name: canary-deploy-patterns
description: Traffic splitting, health checks, automated rollback, progressive delivery, and canary analysis for safe deployments.

Canary Deploy Patterns

Progressive delivery patterns for safe, automated production deployments.

Traffic Splitting Strategy

# Istio VirtualService: gradual traffic shift
apiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
  name: api-canary
spec:
  hosts:
    - api.example.com
  http:
    - route:
        - destination:
            host: api-stable
            port:
              number: 80
          weight: 95          # 95% to stable version
        - destination:
            host: api-canary
            port:
              number: 80
          weight: 5           # 5% to canary version

---
# Progressive rollout schedule
# Step 1:  5% canary, observe 10 minutes
# Step 2: 25% canary, observe 10 minutes
# Step 3: 50% canary, observe 10 minutes
# Step 4: 75% canary, observe 10 minutes
# Step 5: 100% canary → promote to stable

Argo Rollouts Canary

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: api-server
spec:
  replicas: 10
  strategy:
    canary:
      canaryService: api-canary-svc
      stableService: api-stable-svc
      trafficRouting:
        istio:
          virtualService:
            name: api-vsvc
      steps:
        # Step 1: 5% traffic to canary
        - setWeight: 5
        - pause: { duration: 10m }

        # Step 2: Run analysis (automated health check)
        - analysis:
            templates:
              - templateName: canary-success-rate
            args:
              - name: service-name
                value: api-canary-svc

        # Step 3: Increase to 25%
        - setWeight: 25
        - pause: { duration: 10m }

        # Step 4: Another analysis gate
        - analysis:
            templates:
              - templateName: canary-success-rate
              - templateName: canary-latency

        # Step 5: Increase to 50%
        - setWeight: 50
        - pause: { duration: 15m }

        # Step 6: Final analysis before full promotion
        - analysis:
            templates:
              - templateName: canary-success-rate
              - templateName: canary-latency
              - templateName: canary-error-rate

        # Step 7: Full rollout
        - setWeight: 100

      # Auto-rollback on analysis failure
      rollbackWindow:
        revisions: 2

---
# Analysis template: success rate must stay above 99%
apiVersion: argoproj.io/v1alpha1
kind: AnalysisTemplate
metadata:
  name: canary-success-rate
spec:
  metrics:
    - name: success-rate
      interval: 60s
      count: 5
      successCondition: result[0] >= 0.99
      failureLimit: 2
      provider:
        prometheus:
          address: http://prometheus:9090
          query: |
            sum(rate(http_requests_total{
              service="{{args.service-name}}",
              status=~"2.."
            }[2m]))
            /
            sum(rate(http_requests_total{
              service="{{args.service-name}}"
            }[2m]))

Health Check Design

// Multi-level health checks for canary validation
interface HealthCheckResult {
  status: 'healthy' | 'degraded' | 'unhealthy'
  checks: Record<string, {
    status: 'pass' | 'fail'
    latencyMs: number
    message?: string
  }>
  version: string
  uptime: number
}

async function deepHealthCheck(): Promise<HealthCheckResult> {
  const checks: HealthCheckResult['checks'] = {}

  // Database connectivity
  const dbStart = Date.now()
  try {
    await db.$queryRaw`SELECT 1`
    checks.database = { status: 'pass', latencyMs: Date.now() - dbStart }
  } catch (err) {
    checks.database = {
      status: 'fail',
      latencyMs: Date.now() - dbStart,
      message: (err as Error).message
    }
  }

  // Redis connectivity
  const redisStart = Date.now()
  try {
    await redis.ping()
    checks.redis = { status: 'pass', latencyMs: Date.now() - redisStart }
  } catch (err) {
    checks.redis = {
      status: 'fail',
      latencyMs: Date.now() - redisStart,
      message: (err as Error).message
    }
  }

  // Downstream service
  const apiStart = Date.now()
  try {
    const res = await fetch('http://payment-service/health', { signal: AbortSignal.timeout(3000) })
    checks.paymentService = {
      status: res.ok ? 'pass' : 'fail',
      latencyMs: Date.now() - apiStart,
    }
  } catch (err) {
    checks.paymentService = {
      status: 'fail',
      latencyMs: Date.now() - apiStart,
      message: (err as Error).message
    }
  }

  const allPassing = Object.values(checks).every(c => c.status === 'pass')
  const anyFailing = Object.values(checks).some(c => c.status === 'fail')

  return {
    status: allPassing ? 'healthy' : anyFailing ? 'unhealthy' : 'degraded',
    checks,
    version: process.env.APP_VERSION ?? 'unknown',
    uptime: process.uptime(),
  }
}

Automated Rollback

// Canary controller: monitor metrics and auto-rollback
interface CanaryConfig {
  maxErrorRate: number        // e.g., 0.02 (2%)
  maxP95LatencyMs: number     // e.g., 500
  minSuccessRate: number      // e.g., 0.99
  evaluationIntervalMs: number // e.g., 60000 (1 minute)
  warmupPeriodMs: number      // e.g., 120000 (2 minutes, ignore initial spike)
}

class CanaryController {
  private startTime: number = Date.now()

  constructor(
    private config: CanaryConfig,
    private metrics: MetricsClient,
    private deployer: DeployClient,
  ) {}

  async evaluate(): Promise<'continue' | 'promote' | 'rollback'> {
    // Skip evaluation during warmup
    if (Date.now() - this.startTime < this.config.warmupPeriodMs) {
      return 'continue'
    }

    const [errorRate, p95Latency, successRate] = await Promise.all([
      this.metrics.getErrorRate('canary', '5m'),
      this.metrics.getP95Latency('canary', '5m'),
      this.metrics.getSuccessRate('canary', '5m'),
    ])

    // Automatic rollback
Read more
Ships withvibecosystem

Your AI software team. Built on Claude Code. vibecosystem turns Claude Code into a full AI software team — 138 specialized agents that plan, build, review, test, and learn from every mistake. No configuration needed — just install and code.

Get the whole plugin

Other skills on vibecosystem.