Skip to content
Cloud & Infrastructure
Skill

/create-pipeline-task

通过 CLI 创建集成管道任务(数据同步/数据搬运/ETL pipeline)。 触发场景:创建数据集成任务 / 数据同步任务 / 管道任务 / pipeline / 数据搬运 / reader-writer 配置 / MySQL→MaxCompute / MySQL→Hive / Doris→PostgreSQL / 离线集成 / create-pipeline / update-pipeline / create-pipeline-node。 覆盖两条路径:两步法(create-pipeline-node 建草稿 → update-pipeline

BOOST
From plugin
alibabacloud-aiops-skills
256200 skills
Install
$ npx -y skills add aliyun/alibabacloud-aiops-skills --skill create-pipeline-task --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/create-pipeline-task

Context preview

The summary Claude sees to decide when to auto-load this skill.

通过 CLI 创建集成管道任务(数据同步/数据搬运/ETL pipeline)。 触发场景:创建数据集成任务 / 数据同步任务 / 管道任务 / pipeline / 数据搬运 / reader-writer 配置 / MySQL→MaxCompute / MySQL→Hive / Doris→PostgreSQL / 离线集成 / create-pipeline / update-pipeline / create-pipeline-node。 覆盖两条路径:两步法(create-pipeline-node 建草稿 → update-pipeline

SKILL.md

create-pipeline-task.SKILL.md
name: create-pipeline-task
description: |-
  通过 CLI 创建集成管道任务(数据同步/数据搬运/ETL pipeline)。 触发场景:创建数据集成任务 / 数据同步任务 / 管道任务 / pipeline / 数据搬运 / reader-writer 配置 / MySQL→MaxCompute / MySQL→Hive / Doris→PostgreSQL / 离线集成 / create-pipeline / update-pipeline / create-pipeline-node。 覆盖两条路径:两步法(create-pipeline-node 建草稿 → update-pipeline 填配置提交) 和 一步法(create-pipeline)。 关键坑:PluginConfig 必须是 JSON 字符串;columnMappings 顺序敏感且必填;Hive 输出 Key 必须是 hadoophiveoutput 不能是 hiveoutput;空 PluginConfig 触发 ClassCastException。 触发词:创建管道任务、数据同步、数据集成、数据搬运、pipeline、create-pipeline、update-pipeline、reader writer、PluginConfig、MySQL→MaxCompute、MySQL→Hive、Doris→PostgreSQL。

新建集成管道任务 skill

调用 CLI / SDK 前,继承[父技能 §7](../../../SKILL.md#7-observability) 初始化的 session-id 与套件 `references/manifest.json` 中的 `version`;直接加载时先完成父层初始化。API 命令统一附带 `--user-agent "AlibabaCloud-Agent-Skills/alibabacloud-dataphin-skills/{session-id} skill-version/{version}"`,使用父技能名称与同一会话、版本。

适用场景

  • 通过 CLI 创建一个**离线集成管道任务**(offline pipeline / 实时 / 工作流同理)
  • 典型链路:reader (MySQL/Oracle/Doris/Hive/...) → writer (MaxCompute/Hive/PostgreSQL/...) 一对一搬运

> 💡 **术语**:ODPS(Open Data Processing Service)是 MaxCompute 的旧名称,在 pipeline PluginConfig、API 参数中仍可能出现 `odps` 字样,均指 MaxCompute。

  • 需要把 Steps(reader/writer 插件配置)、Hops(DAG 边)、调度 + 资源 settings 一次性提交

两条 CLI 路径

| 路径 | 命令组合 | 适用 | |---|---|---| | **A. 两步法**(推荐) | `dev create-pipeline-node`(建空草稿) → `dev update-pipeline`(填配置 + 提交) | 想分阶段:先占名/占目录,再慢慢调试 Steps | | **B. 一步法** | `dev create-pipeline`(直接带完整 config 创建并提交) | 配置已稳定、CI 化场景 |

> 共同点:两条路径最终落库的 `pipelineDTO.steps[].pluginConfig` 结构完全相同;本 skill 的 PluginConfig 参考片段对两者通用。

---

通用顶层参数

--tenant-id <租户ID>     必填(profile 已配置可省);多租户共享 endpoint 时必须显式传项目所属租户,否则报 DPN.Filter.ProjectNotFound
--project-id   <项目ID>     必填(profile 已配置可省)
--env          DEV|PROD     create-pipeline 用(必填);create-pipeline-node 不需要
--context      Env+ProjectId  update-pipeline 专用(必填),格式 --context 'Env=DEV ProjectId=<项目ID>';update-pipeline 无 --env 参数

---

路径 A:两步法

A-1. 创建空草稿

aliyun dataphin-public create-pipeline-node \
  --tenant-id <tenant-id> \
  --project-id <project-id> \
  --pipeline-name <task-name> \
  --pipeline-type OFFLINE_PIPELINE \
  --node-type NORMAL \
  --file-info '{"FileName":"<task-name>","Directory":"/"}'

返回:

{
  "Data": {
    "PipelineId": <int>,   // 记下来,下一步要用
    "SubmitId": null,
    "Version": null,
    "NodeId": null
  },
  "Code": "OK", "Success": true
}

| 参数 | 说明 | |---|---| | `--pipeline-type` | `OFFLINE_PIPELINE` / `REAL_TIME_PIPELINE` | | `--node-type` | `NORMAL` / `MANUAL` / `REAL_TIME` | | `FileInfo.Directory` | 默认 `/`;非 `/` 必须先存在(否则报错) |

A-2. 填充 Steps/Hops 并提交

aliyun dataphin-public update-pipeline \
  --tenant-id <tenant-id> \
  --project-id <project-id> \
  --context 'Env=DEV ProjectId=<project-id>' \
  --node-info '{"NodeName":"<task-name>","PipelineId":<上一步PipelineId>}' \
  --pipeline-config '<见下方 JSON>' \
  --schedule-config '<见下方 JSON>' \
  --settings '<见下方 JSON>' \
  --submit=true

---

路径 B:一步法

aliyun dataphin-public create-pipeline \
  --tenant-id <tenant-id> \
  --project-id <project-id> \
  --env DEV \
  --pipeline-type 0 \
  --mode PIPELINE \
  --node-info '{"NodeName":"<task-name>","Directory":"/"}' \
  --pipeline-config '<见下方 JSON>' \
  --schedule-config '<见下方 JSON>' \
  --settings '<见下方 JSON>' \
  --submit

`--pipeline-type` 取值:`0` = 离线集成(默认) / `1` = 实时 / `14` = 工作流。

---

`--pipeline-config` 完整骨架

pipeline-config 是集成任务最复杂的字段(含 reader/writer/transformer/column 映射),按 reader/writer 类型组合的完整骨架抽离到独立 reference:

> 📖 详见 [references/pipeline-config.md](references/pipeline-config.md)(涵盖 MySQL reader / MaxCompute writer / Hive writer / Doris reader / PostgreSQL writer 及 MySQL→Hive、MySQL→MaxCompute、Doris→PG 等常用组合)

关键规则速查:

  • **CLI 的 autocreate 不生效**:目标表必须手动预建。`prodTableNotExistAction` 仅两个合法值:`"ignore"`(默认,忽略)和 `"autocreate"`(自动建表),**不存在 `"error"` 值**(虽 CLI 可能接受,但非 Java `TableNotExistAction` 枚举成员)
  • **类型映射陷阱**:Doris LARGEINT→PG NUMERIC、Doris TINYINT→PG SMALLINT
  • **column 顺序**:reader.column 与 writer.column 必须一一对应、长度一致

⚠️ Hive 输出组件专用陷阱

**Hive writer 的 `Key` 必须是 `hadoophiveoutput`,不是 `hiveoutput`!**

| 错误写法 | 正确写法 | |---|---| | `Step.Key = "hiveoutput"` | `Step.Key = "hadoophiveoutput"` | | `PluginConfig.pluginAlias = "hiveoutput"` | `PluginConfig.pluginAlias = "hadoophiveoutput"` | | `PluginConfig.webPluginKey = "hiveoutput"` | `PluginConfig.webPluginKey = "hadoophiveoutput"` |

> **根因**:Java 模型 `OAHiveOutputConfig.stepKey()` 固定返回 `"hadoophiveoutput"`,UI 通过 `Key` 匹配组件类型。用 `"hiveoutput"` 虽然 API 能接受(服务端自动回填正确字段),但 UI 无法识别该组件类型,导致输出端"数据源"下拉框不显示、页面渲染异常。

**Hive writer 的 `dsId` / `dsName` 指向计算源**(不是数据源):

  • `dsId`:项目的 Hive 计算源 ID(如 `"7004766411885056"`)
  • `dsName`:计算源名称(如 `"mdc_dev"`)
  • `dsProjectId`:必填,String 类型,项目 ID(如 `"7004768582924800"`),服务端通过它解析计算源

> **Hive reader 同理**:`Key` 必须是 `hadoophiveinput`,不是 `hiveinput`。

详细信息见 [references/pipeline-config.md](references/pipeline-config.md) 的 Hive writer 章节。

`--schedule-config` 完整骨架

{
  "ScheduleType": "NORMAL",
  "CronExpression": "0 0 0 * * ?",          // ⚠ 字段名是 CronExpression,不是 ScheduleCron
  "ScheduleStartTime": "1970-01-01 00:00:00",
  "ScheduleEndTime":   "9999-01-01 00:00:00",
  "ScheduleIntervalType": "DAILY",          // DAILY/HOURLY/WEEKLY/MONTHLY/CRON
  "ReRunMode": "ALL_ALLOWED",               // ALL_DENIED | FAILURE_ALLOWED | ALL_ALLOWED
  "NodeStatus": 1,                          // 1 正常 / 2 暂停 / 3 空跑
  "Priority": 5,                            // 1~9
  "ResourceGroupId": "default",
  "DevResourceGroupId": "default",
  "ExecuteTimeOutConfig":  { "FollowSystem": true },
  "ExecuteRerunConfig":    { "FollowSystem": true },
  "UpStreamList": [
    {
      "NodeType": "PHYSICAL",
      "SourceNodeId": "<上游节点ID>",
      "SourceNodeOutputName": "<上游输出名>",
      "PeriodDiff": 0
    }
    // 缺省上游时使用租户虚拟根节点 virtual_root_node_<DagId 数字>
    // 详见 ../../dev/find-tenant-root-node/SKILL.md
  ],
  "NodeOutputNameLis
Read more
Ships withalibabacloud-aiops-skills

Official Alibaba Cloud Agent Skills collection, providing AI agents with rich Alibaba Cloud product capabilities and general-purpose tooling.

Get the whole plugin

Other skills on alibabacloud-aiops-skills.