/diagnosing-ci-and-merge-bottlenecks
Diagnoses CI and pull-request pipeline health for a GitHub repo using the engineering analytics MCP tools — pull-requests (PR list with CI status), workflow-health (per-workflow CI trends), and pr-lifecycle (a single PR's timeline). Use when asked whether CI is getting faster or
$ npx -y skills add posthog/posthog --skill diagnosing-ci-and-merge-bottlenecks --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/diagnosing-ci-and-merge-bottlenecks
Context preview
The summary Claude sees to decide when to auto-load this skill.
Diagnoses CI and pull-request pipeline health for a GitHub repo using the engineering analytics MCP tools — pull-requests (PR list with CI status), workflow-health (per-workflow CI trends), and pr-lifecycle (a single PR's timeline). Use when asked whether CI is getting faster or
SKILL.md
diagnosing-ci-and-merge-bottlenecks.SKILL.mdname: diagnosing-ci-and-merge-bottlenecks
description: >
Diagnoses CI and pull-request pipeline health for a GitHub repo using the engineering analytics MCP tools —
pull-requests (PR list with CI status), workflow-health (per-workflow CI trends), and pr-lifecycle (a single PR's
timeline). Use when asked whether CI is getting faster or slower, which GitHub Actions workflow is the slow or
flaky long-pole, how long PRs take from open to merge, how an author's merge time compares to the cohort, which
open PRs have failing or pending CI, or where a specific pull request is stuck. Triggers on "engineering
analytics", "is CI getting slower", "slow workflow", "flaky CI", "time to merge", "cycle time", "PR throughput",
"failing checks", "where is PR <n> stuck", "CI long pole", "what's holding up this PR".
Diagnosing CI and merge bottlenecks
Engineering analytics treats a pull request like product analytics treats a user: a PR moves through a pipeline (`opened → CI → review → merged → deployed`) and the job is to find where it slows down. The surface is **named MCP tools** — you call them, you don't write SQL. Dogfooded on `PostHog/posthog`; the same tools serve autonomous agents (e.g. PostHog Desktop) reasoning about their own PRs.
The tools
- **`pull-requests`** — the PR workhorse. Open PRs plus anything merged or closed since `date_from` (default
`-30d`), newest first. Each row carries `author` (nested object: `handle`, `display_name`, `is_bot`), `repo` (nested: `owner`, `name`), `state`, `is_draft`, `labels`, `open_to_merge_seconds`, `ready_to_merge_seconds`, and a `ci` rollup (`runs` / `passing` / `failing` / `pending`) from the head-SHA join. Answers most PR-level questions: which PRs have failing or pending CI, which are stuck open longest, per-author or per-repo triage, and time-to-merge stats (aggregate over the returned merged rows yourself, median and p95, never a mean; prefer `ready_to_merge_seconds` where non-null, it excludes draft time).
- **`workflow-health`** — per-workflow CI health over a window (`date_from` / `date_to`, default last 24 hours):
`run_count`, `success_rate`, `p50_seconds`, `p95_seconds`, `last_failure_at`. Answers "is CI getting faster or slower" and "which workflow is the slow or flaky long pole". There is no built-in trend — call it over two adjacent windows and compare. `success_rate` covers completed runs; `p50_seconds` / `p95_seconds` cover successful runs only (cancelled and failed runs end early and would bias the duration trend). Each is `null` when a window has no qualifying runs — guard for null before comparing two windows (a workflow can have runs in one and none in the other). `run_scope=pull_request` scopes to PR-attributed runs, excluding master/main (same-repo PRs only — fork runs carry no PR attribution).
- **`pr-lifecycle`** — a single PR's timeline: a header plus ordered events — opened, ready-for-review and
converted-to-draft transitions (when the issue-events table is synced), then a CI started/finished pair **per workflow run** (many on a multi-workflow repo, interleaved by time), then merged/closed. Answers "where is PR N stuck". `metric_quality` is `partial` (no review or comment events).
- **`engineering-analytics-flaky-tests`** — the active test-health queue from the per-test CI spans, over a
window (`date_from` default `-7d`, max 30 days). Evidence is counted per CI run, never per span or run attempt. `classification` is `confirmed_flake` only where the evidence proves nondeterminism (`same_commit_recovery_run_count > 0`: one commit both failed and passed the test, via a "Re-run failed jobs" attempt going green or an in-job retry); `quarantined` means a tolerated failure was recorded while masked; `suspected_regression` means only failures were recorded, which is absence of proof, not proof of a real break. A test qualifies on any same-commit recovery, a quarantined failure, any master/main failure, or failures on ≥ `min_failed_prs` distinct PRs (`failed_pr_count`). Answers "what is this failing test costing us" and picks quarantine candidates. **It does not answer "which tests are flaky"**: this queue only sees the main Backend pytest and Frontend Jest suites, and recovery proof only arrives when someone re-runs failed jobs (or a pytest test is hand-marked `@pytest.mark.flaky(reruns=N)`). Counts are absolute signal, never rates: passing runs are mostly not emitted, so there is no honest denominator.
There is no aggregate time-to-merge tool and no "counts" tool — derive those from `pull-requests` (the stuck/failing counts, the merge-time percentiles).
Caveats you must carry into every answer
These are structural limits of today's snapshot data — state them, don't paper over them.
- **`open_to_merge_seconds` is coarse.** It fuses _draft_ time and _ready-for-review_ time into one figure. Report
it as "open to merge", never "cycle time" or "review time". Flag it when long-lived drafts inflate a number.
- **`ready_to_merge_seconds` is the precise companion**: merged_at minus the last observed ready-for-review
transition (only the last draft/ready switch counts), or minus created_at for a merged PR verifiably never drafted. Null means "not observed" (the PR's life isn't fully inside the synced issue-event window, or the table isn't synced), never zero, so aggregate only over non-null rows and say how many were observable.
- **CI status can be stale.** The CI source syncs on a watermark and does not refresh a run that completes after
newer runs land (until the `workflow_run` webhook ships). Treat a `pending` count as unsettled, not as a settled failure; lead with status, not a verdict.
- **CI for a PR is the head-SHA join, nothing else.** The `ci` rollup reflects only the latest commit's runs. There
is no other link between a PR and its checks.
- **No reviews, approvals, per-check/job, or deploys yet.** Don't infer review behaviour or DORA
Read more
name: diagnosing-ci-and-merge-bottlenecks description: > Diagnoses CI and pull-request pipeline health for a GitHub repo using the engineering analytics MCP tools — pull-requests (PR list with CI status), workflow-health (per-workflow CI trends), and pr-lifecycle (a single PR's timeline). Use when asked whether CI is getting faster or slower, which GitHub Actions workflow is the slow or flaky long-pole, how long PRs take from open to merge, how an author's merge time compares to the cohort, which open PRs have failing or pending CI, or where a specific pull request is stuck. Triggers on "engineering analytics", "is CI getting slower", "slow workflow", "flaky CI", "time to merge", "cycle time", "PR throughput", "failing checks", "where is PR <n> stuck", "CI long pole", "what's holding up this PR".
Diagnosing CI and merge bottlenecks
Engineering analytics treats a pull request like product analytics treats a user: a PR moves through a pipeline (`opened → CI → review → merged → deployed`) and the job is to find where it slows down. The surface is **named MCP tools** — you call them, you don't write SQL. Dogfooded on `PostHog/posthog`; the same tools serve autonomous agents (e.g. PostHog Desktop) reasoning about their own PRs.
The tools
- **`pull-requests`** — the PR workhorse. Open PRs plus anything merged or closed since `date_from` (default
`-30d`), newest first. Each row carries `author` (nested object: `handle`, `display_name`, `is_bot`), `repo` (nested: `owner`, `name`), `state`, `is_draft`, `labels`, `open_to_merge_seconds`, `ready_to_merge_seconds`, and a `ci` rollup (`runs` / `passing` / `failing` / `pending`) from the head-SHA join. Answers most PR-level questions: which PRs have failing or pending CI, which are stuck open longest, per-author or per-repo triage, and time-to-merge stats (aggregate over the returned merged rows yourself, median and p95, never a mean; prefer `ready_to_merge_seconds` where non-null, it excludes draft time).
- **`workflow-health`** — per-workflow CI health over a window (`date_from` / `date_to`, default last 24 hours):
`run_count`, `success_rate`, `p50_seconds`, `p95_seconds`, `last_failure_at`. Answers "is CI getting faster or slower" and "which workflow is the slow or flaky long pole". There is no built-in trend — call it over two adjacent windows and compare. `success_rate` covers completed runs; `p50_seconds` / `p95_seconds` cover successful runs only (cancelled and failed runs end early and would bias the duration trend). Each is `null` when a window has no qualifying runs — guard for null before comparing two windows (a workflow can have runs in one and none in the other). `run_scope=pull_request` scopes to PR-attributed runs, excluding master/main (same-repo PRs only — fork runs carry no PR attribution).
- **`pr-lifecycle`** — a single PR's timeline: a header plus ordered events — opened, ready-for-review and
converted-to-draft transitions (when the issue-events table is synced), then a CI started/finished pair **per workflow run** (many on a multi-workflow repo, interleaved by time), then merged/closed. Answers "where is PR N stuck". `metric_quality` is `partial` (no review or comment events).
- **`engineering-analytics-flaky-tests`** — the active test-health queue from the per-test CI spans, over a
window (`date_from` default `-7d`, max 30 days). Evidence is counted per CI run, never per span or run attempt. `classification` is `confirmed_flake` only where the evidence proves nondeterminism (`same_commit_recovery_run_count > 0`: one commit both failed and passed the test, via a "Re-run failed jobs" attempt going green or an in-job retry); `quarantined` means a tolerated failure was recorded while masked; `suspected_regression` means only failures were recorded, which is absence of proof, not proof of a real break. A test qualifies on any same-commit recovery, a quarantined failure, any master/main failure, or failures on ≥ `min_failed_prs` distinct PRs (`failed_pr_count`). Answers "what is this failing test costing us" and picks quarantine candidates. **It does not answer "which tests are flaky"**: this queue only sees the main Backend pytest and Frontend Jest suites, and recovery proof only arrives when someone re-runs failed jobs (or a pytest test is hand-marked `@pytest.mark.flaky(reruns=N)`). Counts are absolute signal, never rates: passing runs are mostly not emitted, so there is no honest denominator.
There is no aggregate time-to-merge tool and no "counts" tool — derive those from `pull-requests` (the stuck/failing counts, the merge-time percentiles).
Caveats you must carry into every answer
These are structural limits of today's snapshot data — state them, don't paper over them.
- **`open_to_merge_seconds` is coarse.** It fuses _draft_ time and _ready-for-review_ time into one figure. Report
it as "open to merge", never "cycle time" or "review time". Flag it when long-lived drafts inflate a number.
- **`ready_to_merge_seconds` is the precise companion**: merged_at minus the last observed ready-for-review
transition (only the last draft/ready switch counts), or minus created_at for a merged PR verifiably never drafted. Null means "not observed" (the PR's life isn't fully inside the synced issue-event window, or the table isn't synced), never zero, so aggregate only over non-null rows and say how many were observable.
- **CI status can be stale.** The CI source syncs on a watermark and does not refresh a run that completes after
newer runs land (until the `workflow_run` webhook ships). Treat a `pending` count as unsettled, not as a settled failure; lead with status, not a verdict.
- **CI for a PR is the head-SHA join, nothing else.** The `ci` rollup reflects only the latest commit's runs. There
is no other link between a PR and its checks.
- **No reviews, approvals, per-check/job, or deploys yet.** Don't infer review behaviour or DORA
:hedgehog: PostHog is the leading platform for building self-driving products. Our developer tools – AI observability, analytics, session replay, flags, experiments, error tracking, logs, and more – capture all the context agents need to diagnose problems, uncover opportunities, and ship fixes. Steer it all from Slack, web, desktop, or the MCP.
Repo: posthog/posthog
Other skills on posthog.
- /analyzing-expensive-users
Analyze the most expensive users in AI observability and explain why they cost so much. Use when the user asks about top spenders, expensive users, per-user LLM cost, user-level cost drivers, or patterns behind high AI observability spend.
Open skill - /creating-online-evaluations
Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified. Use when the user wants evaluations that automatically score new generations or whole traces going forward — "create an eval to catch X", "continuously
Open skill - /exploring-ai-failures
Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces. Use when someone wants to understand what's going wrong with an AI feature, find and categorize failure modes, triage errors, or investigate quality issues
Open skill - /exploring-llm-clusters
Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.
Open skill - /exploring-llm-costs
Investigate LLM spend in PostHog — total cost over time, cost by model, provider, user, trace, or custom dimension, token and cache-hit economics, and cost regressions. Use when the user asks "how much are we spending on LLMs?", "which model / user / feature is most expensive?",
Open skill - /exploring-llm-evaluations
Investigate AI observability evaluations — `hog` (deterministic code-based), `llm_judge` (LLM-prompt-based), and `sentiment` (user-message sentiment). Find existing evaluations, inspect their configuration, run them against specific generations, query individual results, and
Open skill

