finding-google-skills
Locates and loads the right Google product skill on demand from a remote catalog index, instead of preloading every skill. Use at the START of any request…
Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS. Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report. Don't use for generic
$ npx -y skills add google/skills --skill datalineage-summary --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/datalineage-summaryContext preview
The summary Claude sees to decide when to auto-load this skill.
Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS. Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report. Don't use for generic
name: datalineage-summary metadata: category: BigDataAndAnalytics description: >- Summarizes Google Cloud Data Lineage graphs to help users debug data quality issues and understand data provenance for BQ/GCS. Use when summarizing upstream and downstream data flows, and presenting complex lineage data as an intuitive Markdown report. Don't use for generic BigQuery queries, editing lineage relationships, or downstream deprecation. Don't use for downstream blast-radius impact analysis (use datalineage-bigquery-asset-impact-analysis skill instead).
This skill guides the agent in investigating and summarizing the Data Lineage graph for a specific focal asset (Table-Level Lineage) or specific fields (Column-Level Lineage). It provides an intuitive left-to-right walkthrough of how data enters and leaves the asset, abstracting away complex node and link details into plain English.
This skill relies on the **Google Cloud Data Lineage (Knowledge Catalog) MCP Server** for graph traversal. Ensure you can run `search_lineage` queries in both upstream and downstream directions. For detailed connection configurations and tool schemas, refer to [MCP Usage](references/mcp-usage.md).
Fetch the lineage graph in both directions from the focal point (both upstream and downstream) by making *two separate calls* to the MCP tool: one with `"direction": "UPSTREAM"` and another with `"direction": "DOWNSTREAM"`.
comprehensive list of locations dynamically from the provided [Knowledge Catalog Locations](https://docs.cloud.google.com/dataplex/docs/locations.md.txt) link. To ensure cross-regional lineage is not missed, always verify the current list of GCP regions using this link before populating the `locations` array. You **MUST** populate the `locations` array with all supported physical regions fetched from this link. You may optionally additionally determine the asset's specific active region (using `bq show` or `gcloud storage ls`).
`maxProcessPerLink = 10` as robust defaults when calling `search_lineage`. For example, a DOWNSTREAM call should be formatted like this (expanding the `locations` array as needed):
{
"parent": "projects/project_id/locations/us",
"locations": [
"us",
"us-central1",
"us-east1",
"us-west1",
"europe-west1",
"asia-northeast1"
],
"rootCriteria": {
"entities": {
"entities": [
{
"fullyQualifiedName": "bigquery:project.dataset.table"
}
]
}
},
"direction": "DOWNSTREAM",
"limits": {
"maxDepth": 10,
"maxResults": 5000,
"maxProcessPerLink": 10
}
}Ensure you make a similar call with `"direction": "UPSTREAM"` to fetch the upstream lineage.
Column-Level Lineage (CLL) by configuring the `field` array. If Table-Level Lineage (TLL) is requested, configure the call to get CLL links along with the TLL links by exploiting the `"*"` wildcard. For example:
"rootCriteria": {
"entities": {
"entities": [
{
"fullyQualifiedName": "bigquery:project.dataset.table",
"field": [
"*"
]
}
]
}
}If evaluating a specific column, replace `"*"` with the specific column name (e.g., `"efficiency_score"`).
Generate the summary using the prompt guidelines below.
easy-to-understand left-to-right walkthrough of the data flow.
follows:
(e.g., "This appears to be a Feature Engineering workflow...").
request is for Column-Level Lineage, you MUST explicitly declare that the scope of the analysis is limited to the specified field up front.
Narrative must detail how data arrives at the focal asset, mentioning key source systems, projects, and processing tasks (e.g., Spark on Dataproc).
Lineage:**`. Detail where data goes from the focal asset to final consumer systems.
provide transparency on the boundaries of the summary. The output must contain:
expanded locations (if not all were used) or depth.
files/tables.
intermediate views, consumer tables) if there are fewer than 5. Do not just summarize counts if there are fewer than 5; name them explicitly. Otherwise, if 5 or more, aggregate them by count (e.g., "5 GCS buckets").
*total assets*.
This repository contains Agent Skills for Google products and technologies, including Google Cloud.
Repo: google/skills
Locates and loads the right Google product skill on demand from a remote catalog index, instead of preloading every skill. Use at the START of any request…
Provides safety-critical validation, guardrails, and data reduction for gcloud CLI operations across Google Cloud Platform (GCP) services and infrastructure.…
Provides expert guidance on authenticating and authorizing to Google Cloud services and APIs, covering human users, service identities, Application Default…
Guides a developer's first steps on Google Cloud, covering account creation, billing setup, project management, and deploying a first resource. Use when a new…
Searches, retrieves, and synthesizes official Google developer documentation across Google Cloud, AI/Gemini, Android, Chrome, Web, Flutter, Go, Firebase, and…
Guides developers through managing (adding, removing, and clearing) audience members for Google products using the Data Manager API and its associated client…