Skip to content
Data
Skill

/diagnosing-experiment-health

Runs diagnostics and a health check on one PostHog experiment: its setup, exposures, results and changes during the run. Covers 0 exposures, sample ratio mismatch, lost or uneven exposures, users in multiple variants, identity faults, who the experiment counts, significance

GuideBOOST
From plugin
posthog-posthog
40k149 skills11 agents1 command3 MCP
Install
$ npx -y skills add posthog/posthog --skill diagnosing-experiment-health --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/diagnosing-experiment-health

Context preview

The summary Claude sees to decide when to auto-load this skill.

Runs diagnostics and a health check on one PostHog experiment: its setup, exposures, results and changes during the run. Covers 0 exposures, sample ratio mismatch, lost or uneven exposures, users in multiple variants, identity faults, who the experiment counts, significance

SKILL.md

diagnosing-experiment-health.SKILL.md
name: diagnosing-experiment-health
description: "Runs diagnostics and a health check on one PostHog experiment: its setup, exposures, results and changes during the run. Covers 0 exposures, sample ratio mismatch, lost or uneven exposures, users in multiple variants, identity faults, who the experiment counts, significance traps (peeking, A/A, 'target reached', Bayesian vs Frequentist, sequential, CUPED), mid-run edits, pause and freeze, and survey follow-up.\nTRIGGER when: user asks 'is my experiment healthy / set up right / biased?' or 'why 0 exposures?', asks why it counts different people than their insight, mentions the bias or mismatch warning, an uneven split, significance that flips, numbers that disagree with their SQL, results that changed after an edit, or wants user feedback on an experiment.\nDO NOT TRIGGER when: creating an experiment (use creating-experiments), only configuring rollout (use configuring-experiment-rollout) or metrics (use configuring-experiment-analytics), or only asking lifecycle questions (use managing-experiment-lifecycle)."

Diagnosing experiment health

This skill answers: **My PostHog experiment results look wrong, biased, or empty — what's going on?** It is the diagnostics and health check for one experiment: its setup, its exposures, its results and what changed during the run.

Match the user's complaint in the dispatch table, then read the matching reference file for the diagnostic.

Each diagnostic in the reference files is tagged `[HIGH]`, `[MEDIUM]`, or `[LOW]` based on how strongly it's verified — `[HIGH]` is verified directly in PostHog code, `[MEDIUM]` is partially or team-source verified, `[LOW]` describes SDK/external behavior that wasn't verified here. Treat `[LOW]` items as hypotheses to test, not facts to assert.

Step 1 — Resolve the experiment

If the user refers to an experiment by name or description, load the `finding-experiments` skill first to resolve it to a concrete ID.

Call `experiment-get` and pull these fields. They are inputs for almost every diagnostic:

  • `status` (`draft` / `running` / `paused` / `exposure_frozen` / `stopped`), `start_date`, `end_date`, `is_legacy`
  • `feature_flag.active`, and `feature_flag.filters.multivariate.variants[]` — the variant keys and the split (`rollout_percentage`).

This is the flag as it is now. After a ship or an edit, the split of the run is in the change history (Step 1.5).

  • `feature_flag.filters.aggregation_group_type_index` — when set, the experiment's unit is a group, not a person
  • `feature_flag.filters.groups[]` — for each group read `variant`, `properties`, and

`rollout_percentage` (that group's rollout, % of the matched bucketing units — persons, or groups when the flag or the group is aggregated by a group type — that enter the experiment; each group has its own, so there is no single overall rollout unless the flag has one unconditional group). A non-null `variant` that is one of the flag's variant keys is a forced-variant override on the matched cohort (release-condition assignment, not randomized) — surfaces A7. Watch for the severe shape (A7b): a variant-pinned group with broad/empty `properties` at high rollout, or no group left randomized (`variant: null`) / no release path to one arm — that starves the other variant (one arm gets ~0 analyzable exposures). See `references/bias-and-skew.md`.

  • `feature_flag.bucketing_identifier` and `feature_flag.ensure_experience_continuity` — what the flag hashes (A3, A6, A8, A10)
  • `holdout` and `excluded_variants` — people and arms that are outside the analysis on purpose
  • `exposure_criteria.multiple_variant_handling` — defaults to `"exclude"` if absent
  • `exposure_criteria.exposure_config` — the exposure event and its property filters.

Unset, or naming `$feature_flag_called`, means the default exposure event: read which one from `resolved_exposure_event`. A config that names `$experiment_exposure` counts that event, whatever `resolved_exposure_event` says. Any other event, or an action, is a _custom exposure event_. Property filters in the config apply to either (B12). The table in `references/diagnostic-snapshot.md` says which property carries the variant.

  • `exposure_criteria.activation_config` — an activation event, or an action, on top of the default exposure.

When it is set, a person counts as exposed only from their first activation event at or after their first flag exposure, and the time of that event is the exposure time. It does not combine with a custom exposure event.

  • `exposure_criteria.filterTestAccounts` — defaults to `true`
  • `metrics` and `metrics_secondary`, with shared metrics — for each: the name, the type and the definition.

Describe a metric from its definition, not from its name (D13).

  • `stats_config` — `method` (Bayesian or Frequentist), the confidence level (`bayesian.ci_level`, `frequentist.alpha`), `frequentist.sequential_testing_enabled`, `cuped`, `baseline_variant_key`.

A missing `method` reads as Bayesian and a missing confidence level as 95%. Sequential testing and CUPED fall back to the project default (C13).

  • `running_time_calculation` — the target sample (a total across all variants), the days still missing to reach it, and the minimum detectable effect.

When the estimate comes from numbers the user typed, or was written at creation and never refreshed by the page, the days are the whole estimated length (C12). A duplicated experiment starts with the values of its source until the page refreshes them.

  • `only_count_matured_users` — E11

Step 1.5 — Pull a diagnostic snapshot (verify before asking)

Before asking the user clarifying questions, pull the diagnostic snapshot in [references/diagnostic-snapshot.md](references/diagnostic-snapshot.md). Most diagnostics in this skill can be confirmed or ruled out from that data without an interview.

Read the cheap sources first, in every diagnosis: the stored results, then the change history. When th

Read more
Ships withposthog-posthog

:hedgehog: PostHog is the leading platform for building self-driving products. Our developer tools – AI observability, analytics, session replay, flags, experiments, error tracking, logs, and more – capture all the context agents need to diagnose problems, uncover opportunities, and ship fixes. Steer it all from Slack, web, desktop, or the MCP.

Get the whole plugin

Other skills on posthog-posthog.