Skip to content
Marketing
Skill

/github-repo-signals

Extract and score leads from GitHub repositories by analyzing stars, forks, issues, PRs, comments, and contributions. Produces unified multi-repo CSV with deduplicated user profiles. No paid API credits required.

From plugin
goose-skills
1.2k200 skills
Install
$ npx -y skills add gooseworks-ai/goose-skills --skill github-repo-signals --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/github-repo-signals

Context preview

The summary Claude sees to decide when to auto-load this skill.

Extract and score leads from GitHub repositories by analyzing stars, forks, issues, PRs, comments, and contributions. Produces unified multi-repo CSV with deduplicated user profiles. No paid API credits required.

SKILL.md

github-repo-signals.SKILL.md
name: github-repo-signals
description: Extract and score leads from GitHub repositories by analyzing stars, forks, issues, PRs, comments, and contributions. Produces unified multi-repo CSV with deduplicated user profiles. No paid API credits required.
user-invocable: true
allowed-tools: Bash, Read, Write, Edit, Grep, Glob
argument-hint: "[owner/repo1,owner/repo2] [limit]"

GitHub Repository Signals

Extract high-intent leads from one or more GitHub repositories by analyzing every type of user interaction. This skill uses only free GitHub API data — no enrichment credits are spent.

When to Use

  • User wants to find leads from open-source GitHub repositories
  • User wants to identify people who interact with competitor or category repos
  • User wants cross-repo interaction analysis to find high-intent prospects
  • User asks for GitHub-based lead generation without paid enrichment
  • User says their ICP, target audience, or buyers are developers, engineers, or technical people who are active on GitHub
  • User describes prospects who use open-source tools, contribute to open source, or build with specific technologies — and those technologies have public GitHub repos
  • User wants to find leads in a technical space (e.g., "real-time communication", "AI agents", "infrastructure") where the community congregates around GitHub repositories

**Note:** If the user describes their ICP as GitHub-active but hasn't identified specific repositories yet, this skill still applies. In that case, ask the user which repositories their ICP is likely to interact with, or help them identify relevant repos based on the technology/space they describe.

Prerequisites

  • `gh` CLI authenticated (`gh auth status` to verify)
  • Python 3.9+ with `PyYAML` installed
  • Working directory: the project root containing this skill

Inputs to Collect from User

Before running, ask the user for:

1. **Repositories** (required): One or more GitHub repository URLs or `owner/repo` strings 2. **User limit** (required): How many top users to include in the output. Explain that more users = longer runtime due to GitHub profile fetching (~5,000 profiles/hour). Suggest 500 as a good starting point for testing.

Execution Steps

Step 1: Verify Environment

gh auth status

Step 2: Run the Tool

python3 ${CLAUDE_SKILL_DIR}/scripts/gh_repo_signals.py \
    --repos "owner1/repo1,owner2/repo2" \
    --limit <USER_LIMIT> \
    --output ${CLAUDE_SKILL_DIR}/../.tmp/repo_signals.csv

Replace the repos and limit with user-provided values.

The tool will: 1. **Extract** all interaction types per repo (stars, forks, contributors, issues, PRs, comments, watchers, commit emails) 2. **Filter out** bots and org members automatically (fetches org member lists and detects org email domains) 3. **Score** each user by interaction depth using these weights:

  • Issue opener: 5 points
  • PR author: 5 points
  • Contributor: 4 points
  • Issue commenter: 3 points
  • Forker: 3 points
  • Watcher: 2 points
  • Stargazer: 1 point

4. **Rank** users by (repos_interacted desc, total_score desc) — multi-repo users surface first 5. **Fetch** GitHub profiles for the top N users (name, email, company, location, blog, twitter, bio, followers) 6. **Export** two CSV files: `_users.csv` and `_interactions.csv`

Step 3: Review Output

The tool produces two CSV files:

**`repo_signals_users.csv`** — One row per person, deduplicated across all repos | Column | Description | |--------|-------------| | username | GitHub login | | name | Display name | | email | Public GitHub email | | commit_email | Email from git commits (if different from public) | | company | Company from GitHub profile | | location | Location from GitHub profile | | blog | Website/blog URL | | twitter | Twitter/X handle | | bio | GitHub bio | | followers | Follower count | | public_repos | Number of public repos | | total_repos_interacted | Number of input repos this user interacted with | | interaction_score | Weighted score across all repos |

**`repo_signals_interactions.csv`** — One row per user x repo combination | Column | Description | |--------|-------------| | username | GitHub login | | repository | Which repo this row is about | | is_contributor | YES/NO | | is_stargazer | YES/NO | | is_forker | YES/NO | | is_watcher | YES/NO | | is_issue_opener | YES/NO | | is_pr_author | YES/NO | | is_issue_commenter | YES/NO | | contribution_count | Number of commits (0 if not contributor) | | starred_at | Date starred (if applicable) | | forked_at | Date forked (if applicable) | | repo_score | Interaction score for this specific repo |

Phase 3: Analyze & Recommend

Once the CSV files are generated, **do not stop**. Immediately proceed to analyze the data and brief the user.

Step 5: Collect Company Context

Check if you already know the user's company and intent from prior conversation. If not, ask:

> "Before I analyze these results, I need to understand who you're finding leads for: > 1. **What does your company/product do?** (one-liner is fine) > 2. **Who is your ideal customer?** (role, company size, industry, tech stack — whatever is relevant) > 3. **What's the goal for these leads?** (outbound sales, partnership, hiring, community building, etc.)"

Do NOT proceed to analysis until you have this context. It directly shapes the recommendations.

Step 6: Analyze the Data

Read the generated .csv file and compute the following analysis. Present it to the user as a structured briefing.

**6a. Overall Stats**

  • Total users in the sheet
  • Score distribution (how many at 15+, 10-14, below 10)
  • Email coverage: how many have any email (public or commit)
  • Company coverage: how many have a company listed

**6b. Multi-Repo Users (if multiple repos were scanned)**

  • How many users interacted with 2+ repos
  • List the top 10 multi-repo users with their names, companies, and which repos they touched
  • This is the highest-signal segment — call it
Read more
Ships withgoose-skills

Put your AI agent on the growth team. Research customers and competitors, analyze what is working, create the next campaign, and learn from the result.

Get the whole plugin

Other skills on goose-skills.