Skip to content
Data
Skill

/connector-building

Build a new OpenMetadata connector from scratch — scaffold JSON Schema, Python boilerplate, and AI context using schema-first architecture with code generation across Python, Java, TypeScript, and auto-rendered UI forms.

BOOST
From plugin
openmetadata
15k24 skills
Install
$ npx -y skills add open-metadata/OpenMetadata --skill connector-building --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/connector-building

Context preview

The summary Claude sees to decide when to auto-load this skill.

Build a new OpenMetadata connector from scratch — scaffold JSON Schema, Python boilerplate, and AI context using schema-first architecture with code generation across Python, Java, TypeScript, and auto-rendered UI forms.

SKILL.md

connector-building.SKILL.md
name: scaffold-connector
description: Build a new OpenMetadata connector from scratch — scaffold JSON Schema, Python boilerplate, and AI context using schema-first architecture with code generation across Python, Java, TypeScript, and auto-rendered UI forms.
user-invocable: true
argument-hint: "[connector name or description]"
allowed-tools:
  - Bash
  - Read
  - Write
  - Edit
  - Glob
  - Grep
  - Agent
hooks:
  SessionStart: |
    Load the OpenMetadata connector standards before starting:
    Read the standards at ${CLAUDE_SKILL_DIR}/standards/main.md

OpenMetadata Connector Building Skill

When to Activate

When a user asks to build, create, add, or scaffold a new connector, source, or integration for OpenMetadata.

Core Insight

**One JSON Schema definition cascades through 6 layers**: Python Pydantic models, Java models, UI forms (RJSF auto-render), API validation, test fixtures, and documentation. Define the schema once — everything else is generated or guided.

Workflow: 7 Phases

Phase 0: ENVIRONMENT — Set Up Python Dev Environment

Before any `make` or `python` commands, set up the environment from the repo root:

python3.11 -m venv env
source env/bin/activate
make install_dev generate

Always activate before running commands: `source env/bin/activate`

Phase 1: SCAFFOLD — Generate Boilerplate

Run the scaffold CLI to collect inputs and generate files:

source env/bin/activate
metadata scaffold-connector

Interactive mode collects: connector name, service type, connection type, auth types, capabilities, docs URL, SDK package, API endpoints, implementation notes, Docker image, container port.

Non-interactive mode:

metadata scaffold-connector \
  --name my_db \
  --service-type database \
  --connection-type sqlalchemy \
  --scheme "mydb+pymydb" \
  --auth-types basic \
  --capabilities metadata lineage usage profiler \
  --docs-url "https://docs.example.com/api" \
  --sdk-package "mydb-sdk" \
  --docker-image "mydb/mydb:latest" \
  --docker-port 5432

**Output**: JSON Schema + test connection JSON + Python files + `CONNECTOR_CONTEXT.md` as an AI working document. SQLAlchemy database connectors get concrete code templates; all others get skeleton files with pointers to reference connectors.

**CONNECTOR_CONTEXT.md handling**: The scaffold generates `CONNECTOR_CONTEXT.md` in the connector directory as a working document for any AI tool (Claude Code, Cursor, Codex, Copilot, Windsurf). It is **gitignored** — it stays local and is never committed to the repo. No cleanup needed.

Phase 2: CLASSIFY — Understand the Source

The scaffold classifies along 3 dimensions. Verify the choices:

**Dimension 1 — Service Type** (determines directory + base class):

| Service Type | Base Class | Reference | |---|---|---| | `database` | `CommonDbSourceService` | `mysql/` | | `dashboard` | `DashboardServiceSource` | `metabase/` | | `pipeline` | `PipelineServiceSource` | `airflow/` | | `messaging` | `MessagingServiceSource` | `kafka/` | | `mlmodel` | `MlModelServiceSource` | `mlflow/` | | `storage` | `StorageServiceSource` | `s3/` | | `search` | `SearchServiceSource` | `elasticsearch/` | | `api` | `ApiServiceSource` | `rest/` |

**Dimension 2 — Connection Type** (database only):

  • `sqlalchemy` → `BaseConnection[Config, Engine]` + SQLAlchemy dialect
  • `rest_api` → `get_connection()` + custom REST client (ref: `salesforce/`)
  • `sdk_client` → `get_connection()` + vendor SDK wrapper

**Dimension 3 — Capabilities** (determines extra files): `metadata` (always), `lineage`, `usage`, `profiler`, `stored_procedures`, `data_diff`

Read the source-type-specific standard at `${CLAUDE_SKILL_DIR}/standards/source_types/{service_type}.md` for detailed patterns.

Phase 3: RESEARCH — API/SDK Discovery

Read the `CONNECTOR_CONTEXT.md` generated by the scaffold. Then research the source's API/SDK.

**If you can dispatch sub-agents** (Claude Code): Launch a `connector-researcher` agent:

Agent: openmetadata-skills:connector-researcher
Prompt: "Research {source_name} for an OpenMetadata {service_type} connector.
Find: API docs, auth methods, key endpoints, pagination, rate limits, SDK packages."

**If you cannot dispatch sub-agents**: Perform the research yourself using WebSearch and WebFetch.

Phase 4: IMPLEMENT — Fill in the TODO Items

The scaffold generates files with `# TODO` markers. Read the relevant standards before implementing:

  • `${CLAUDE_SKILL_DIR}/standards/connection.md` — Connection patterns
  • `${CLAUDE_SKILL_DIR}/standards/patterns.md` — Error handling, pagination, auth
  • `${CLAUDE_SKILL_DIR}/standards/performance.md` — Pagination, lookup optimization, anti-patterns
  • `${CLAUDE_SKILL_DIR}/standards/memory.md` — Memory management, streaming, OOM prevention
  • `${CLAUDE_SKILL_DIR}/standards/source_types/{service_type}.md` — Service-specific patterns

**SQLAlchemy database**: Templates are mostly complete. Customize `_get_client()` if needed. **Non-SQLAlchemy**: Study the reference connector, then implement each skeleton file.

**Critical for JSON Schema**:

  • Make auth fields (`username`, `password`, `token`) **required** when the service needs authentication by default. If omitting a field means an opaque 401 at runtime, make it required so the UI validates upfront.
  • Include SSL/TLS config (`verifySSL` + `sslConfig` `$ref`) for any connector that communicates over HTTPS — enterprise deployments use internal CAs.
  • **SSL must be wired end-to-end**: schema → `connection.py` (resolve with `get_verify_ssl_fn`) → `client.py` (`session.verify = verify_ssl`). Missing wiring triggers SonarQube Security Review failure.
  • See `${CLAUDE_SKILL_DIR}/standards/schema.md` for the `$ref` patterns and required fields guidance.

**Critical for Pydantic API models (models.py)**:

  • Always set `model_config = ConfigDict(populate_by_name=True)` when using `Field(alias=...)` — without this, constructing instances with Python attribute names raises
Read more
Ships withopenmetadata

The Open Context Layer for Data and AI , OpenMetadata is the open platform for building trusted data context and business semantics for humans, AI assistants, and agents.

Get the whole plugin
Stats
15,365
Stars
2,424
Forks
Active
Maintenance
TypeScript
Language
Apache-2.0
License
3h ago
Last commit
5y ago
Created
3h ago
Added

Repo: open-metadata/OpenMetadata

Other skills on openmetadata.