/datalineage-summary
Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS. Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report. Don't use for generic
$ npx -y skills add google/skills --skill datalineage-summary --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/datalineage-summary
Context preview
The summary Claude sees to decide when to auto-load this skill.
Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS. Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report. Don't use for generic
SKILL.md
datalineage-summary.SKILL.mdname: datalineage-summary
metadata:
category: BigDataAndAnalytics
description: >-
Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS.
Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report.
Don't use for generic BigQuery queries, editing lineage relationships, or downstream deprecation.
Don't use for downstream blast-radius impact analysis (use datalineage-bigquery-asset-impact-analysis skill instead).
Data Lineage Summary
This skill guides the agent in investigating and summarizing the Data Lineage graph for a specific focal asset (Table-Level Lineage) or specific fields (Column-Level Lineage). It provides an intuitive left-to-right walkthrough of how data enters and leaves the asset, abstracting away complex node and link details into plain English.
Prerequisites
This skill relies on the **Google Cloud Data Lineage (Knowledge Catalog) MCP Server** for graph traversal. Ensure you can run `search_lineage` queries in both upstream and downstream directions. For detailed connection configurations and tool schemas, refer to [MCP Usage](references/mcp-usage.md).
Workflow Logic
1. Get Lineage
Fetch the lineage graph in both directions from the focal point (both upstream and downstream) by making *two separate calls* to the MCP tool: one with `"direction": "UPSTREAM"` and another with `"direction": "DOWNSTREAM"`.
- **Location Strategy**: You **MUST** use the `read_url` tool to fetch the
comprehensive list of locations dynamically from the provided [Knowledge Catalog Locations](https://docs.cloud.google.com/dataplex/docs/locations.md.txt) link. To ensure cross-regional lineage is not missed, always verify the current list of GCP regions using this link before populating the `locations` array. You **MUST** populate the `locations` array with all supported physical regions fetched from this link. You may optionally additionally determine the asset's specific active region (using `bq show` or `gcloud storage ls`).
- **Search Parameters**: Use `maxDepth = 10`, `maxResults = 5000` and
`maxProcessPerLink = 10` as robust defaults when calling `search_lineage`. For example, a DOWNSTREAM call should be formatted like this (expanding the `locations` array as needed):
{
"parent": "projects/project_id/locations/us",
"locations": [
"us",
"us-central1",
"us-east1",
"us-west1",
"europe-west1",
"asia-northeast1"
],
"rootCriteria": {
"entities": {
"entities": [
{
"fullyQualifiedName": "bigquery:project.dataset.table"
}
]
}
},
"direction": "DOWNSTREAM",
"limits": {
"maxDepth": 10,
"maxResults": 5000,
"maxProcessPerLink": 10
}
}Ensure you make a similar call with `"direction": "UPSTREAM"` to fetch the upstream lineage.
- **Column-Level Lineage (CLL)**: The `search_lineage` tool can find all
Column-Level Lineage (CLL) by configuring the `field` array. If Table-Level Lineage (TLL) is requested, configure the call to get CLL links along with the TLL links by exploiting the `"*"` wildcard. For example:
"rootCriteria": {
"entities": {
"entities": [
{
"fullyQualifiedName": "bigquery:project.dataset.table",
"field": [
"*"
]
}
]
}
}If evaluating a specific column, replace `"*"` with the specific column name (e.g., `"efficiency_score"`).
2. Summarize
Generate the summary using the prompt guidelines below.
- **Persona**: Act as an expert Data Lineage Analyst generating a concise,
easy-to-understand left-to-right walkthrough of the data flow.
- **Structure & Flow**: Start immediately with the summary text, structured as
follows:
- **Overall Flow Type**: State the inferred workflow type and data domain
(e.g., "This appears to be a Feature Engineering workflow...").
- **Systems Overview**: List the primary systems involved up front. If the
request is for Column-Level Lineage, you MUST explicitly declare that the scope of the analysis is limited to the specified field up front.
- **Upstream Lineage**: Use the exact bold header `**Upstream Lineage:**`.
Narrative must detail how data arrives at the focal asset, mentioning key source systems, projects, and processing tasks (e.g., Spark on Dataproc).
- **Downstream Lineage**: Use the exact bold header `**Downstream
Lineage:**`. Detail where data goes from the focal asset to final consumer systems.
- **Analysis Metadata**: Display the parameters used for the API call to
provide transparency on the boundaries of the summary. The output must contain:
- **Locations Searched**: `{list_of_locations_queried}`
- **Parent Location**: `{parent_path}`
- **Depth Limit**: `{maxDepth}`
- **Process per Link Limit**: `{maxProcessPerLink}`
- **Tip for User**: A prompt suggesting they can ask to rerun with
expanded locations (if not all were used) or depth.
- **Granularity Constraints**:
- Prioritize flows between Systems, Projects, and Datasets over individual
files/tables.
- You MUST explicitly list specific asset names (e.g., source tables,
intermediate views, consumer tables) if there are fewer than 5. Do not just summarize counts if there are fewer than 5; name them explicitly. Otherwise, if 5 or more, aggregate them by count (e.g., "5 GCS buckets").
- Only mention counts for *ultimate sources*, *final consumers*, and
*total assets*.
Read more
name: datalineage-summary metadata: category: BigDataAndAnalytics description: >- Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS. Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report. Don't use for generic BigQuery queries, editing lineage relationships, or downstream deprecation. Don't use for downstream blast-radius impact analysis (use datalineage-bigquery-asset-impact-analysis skill instead).
Data Lineage Summary
This skill guides the agent in investigating and summarizing the Data Lineage graph for a specific focal asset (Table-Level Lineage) or specific fields (Column-Level Lineage). It provides an intuitive left-to-right walkthrough of how data enters and leaves the asset, abstracting away complex node and link details into plain English.
Prerequisites
This skill relies on the **Google Cloud Data Lineage (Knowledge Catalog) MCP Server** for graph traversal. Ensure you can run `search_lineage` queries in both upstream and downstream directions. For detailed connection configurations and tool schemas, refer to [MCP Usage](references/mcp-usage.md).
Workflow Logic
1. Get Lineage
Fetch the lineage graph in both directions from the focal point (both upstream and downstream) by making *two separate calls* to the MCP tool: one with `"direction": "UPSTREAM"` and another with `"direction": "DOWNSTREAM"`.
- **Location Strategy**: You **MUST** use the `read_url` tool to fetch the
comprehensive list of locations dynamically from the provided [Knowledge Catalog Locations](https://docs.cloud.google.com/dataplex/docs/locations.md.txt) link. To ensure cross-regional lineage is not missed, always verify the current list of GCP regions using this link before populating the `locations` array. You **MUST** populate the `locations` array with all supported physical regions fetched from this link. You may optionally additionally determine the asset's specific active region (using `bq show` or `gcloud storage ls`).
- **Search Parameters**: Use `maxDepth = 10`, `maxResults = 5000` and
`maxProcessPerLink = 10` as robust defaults when calling `search_lineage`. For example, a DOWNSTREAM call should be formatted like this (expanding the `locations` array as needed):
{
"parent": "projects/project_id/locations/us",
"locations": [
"us",
"us-central1",
"us-east1",
"us-west1",
"europe-west1",
"asia-northeast1"
],
"rootCriteria": {
"entities": {
"entities": [
{
"fullyQualifiedName": "bigquery:project.dataset.table"
}
]
}
},
"direction": "DOWNSTREAM",
"limits": {
"maxDepth": 10,
"maxResults": 5000,
"maxProcessPerLink": 10
}
}Ensure you make a similar call with `"direction": "UPSTREAM"` to fetch the upstream lineage.
- **Column-Level Lineage (CLL)**: The `search_lineage` tool can find all
Column-Level Lineage (CLL) by configuring the `field` array. If Table-Level Lineage (TLL) is requested, configure the call to get CLL links along with the TLL links by exploiting the `"*"` wildcard. For example:
"rootCriteria": {
"entities": {
"entities": [
{
"fullyQualifiedName": "bigquery:project.dataset.table",
"field": [
"*"
]
}
]
}
}If evaluating a specific column, replace `"*"` with the specific column name (e.g., `"efficiency_score"`).
2. Summarize
Generate the summary using the prompt guidelines below.
- **Persona**: Act as an expert Data Lineage Analyst generating a concise,
easy-to-understand left-to-right walkthrough of the data flow.
- **Structure & Flow**: Start immediately with the summary text, structured as
follows:
- **Overall Flow Type**: State the inferred workflow type and data domain
(e.g., "This appears to be a Feature Engineering workflow...").
- **Systems Overview**: List the primary systems involved up front. If the
request is for Column-Level Lineage, you MUST explicitly declare that the scope of the analysis is limited to the specified field up front.
- **Upstream Lineage**: Use the exact bold header `**Upstream Lineage:**`.
Narrative must detail how data arrives at the focal asset, mentioning key source systems, projects, and processing tasks (e.g., Spark on Dataproc).
- **Downstream Lineage**: Use the exact bold header `**Downstream
Lineage:**`. Detail where data goes from the focal asset to final consumer systems.
- **Analysis Metadata**: Display the parameters used for the API call to
provide transparency on the boundaries of the summary. The output must contain:
- **Locations Searched**: `{list_of_locations_queried}`
- **Parent Location**: `{parent_path}`
- **Depth Limit**: `{maxDepth}`
- **Process per Link Limit**: `{maxProcessPerLink}`
- **Tip for User**: A prompt suggesting they can ask to rerun with
expanded locations (if not all were used) or depth.
- **Granularity Constraints**:
- Prioritize flows between Systems, Projects, and Datasets over individual
files/tables.
- You MUST explicitly list specific asset names (e.g., source tables,
intermediate views, consumer tables) if there are fewer than 5. Do not just summarize counts if there are fewer than 5; name them explicitly. Otherwise, if 5 or more, aggregate them by count (e.g., "5 GCS buckets").
- Only mention counts for *ultimate sources*, *final consumers*, and
*total assets*.
This repository contains Agent Skills for Google products and technologies, including Google Cloud. This repository is under active development.
Repo: google/skills
Other skills on google-skills.
- /data-manager-api-audience-ingestion
Guides developers through managing (adding, removing, and clearing) audience members for Google products using the Data Manager API and its associated client libraries. Use this skill when the user wants to upload audience members, remove specific users, or clear/replace an
Open skill - /data-manager-api-event-ingestion
Guides developers through implementing event and conversion ingestion to Google products using the Data Manager API /v1/events/ingest endpoint and its associated client libraries. Use this skill when the user wants to upload offline conversions, enhanced conversions for leads,
Open skill - /data-manager-api-setup
Guides developers through client library installation and authentication setup steps for the Data Manager API. Use this skill when a user is getting started with the Data Manager API and needs to setup their local environment, install the client library, or setup access to the
Open skill - /google-ads-api-account-diagnostics
Diagnoses Google Ads account performance issues such as conversion loss (value or volume), low lead flow/volume, and lost impression share (opportunities) due to ad rank, bids, or budgets. Use when troubleshooting sudden performance drops, analyzing campaign impression share
Open skill - /google-ads-api-mcp-setup
Guides developers through downloading, configuring, and installing the official open-source Google Ads MCP Server. Use this skill when a user wants to connect their AI assistant (such as Gemini, Claude Code, or Cursor) to their Google Ads account to query campaigns or retrieve
Open skill - /google-ads-api-quickstart
Guides developers through Google Ads API quickstart: credential setup, choosing from 6 client libraries/REST, configuring environments, and running a "retrieve campaigns" script. Troubleshoots common setup errors: USER_PERMISSION_DENIED, login_customer_id issues, and
Open skill

