/sn-da-non-spreadsheet-analysis
Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。**遇到以下任一情况就主动使用本 skill**:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 / 幻灯片分析 / 发票提取 / 合同分析 / 文档统计 / 错别字 / 语病 / 字号检查 / 简历分析 /
$ npx -y skills add OpenSenseNova/SenseNova-Skills --skill sn-da-non-spreadsheet-analysis --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/sn-da-non-spreadsheet-analysis
Context preview
The summary Claude sees to decide when to auto-load this skill.
Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。**遇到以下任一情况就主动使用本 skill**:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 / 幻灯片分析 / 发票提取 / 合同分析 / 文档统计 / 错别字 / 语病 / 字号检查 / 简历分析 /
SKILL.md
sn-da-non-spreadsheet-analysis.SKILL.mdname: sn-da-non-spreadsheet-analysis
description: "Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。**遇到以下任一情况就主动使用本 skill**:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 / 幻灯片分析 / 发票提取 / 合同分析 / 文档统计 / 错别字 / 语病 / 字号检查 / 简历分析 / 多文档对比;③任务涉及从文档中提取表格、数值、图表、格式(颜色/高亮/字号)、组织架构、时间线等结构化信息。仅不用于:Excel/CSV 数据分析(使用 sn-da-excel-workflow)、纯图片分析(使用 sn-da-image-caption)。"
Document Analysis Skill — Word / PDF / PPT
End-to-end workflow for Word, PDF, and PPT document parsing. Each format has specific parsing pitfalls — follow the format-specific sub-skill exactly.
---
Workflow
Step 0 — Identify file type and input scope
import os
input_path = "/mnt/data/..." # from user
# Detect single file vs directory (multi-file scenario)
if os.path.isdir(input_path):
all_files = [
os.path.join(input_path, f)
for f in os.listdir(input_path)
if f.lower().endswith(('.docx', '.doc', '.pdf', '.pptx', '.ppt'))
]
print(f"Found {len(all_files)} documents: {all_files}")
else:
all_files = [input_path]
# Route by extension
ext = os.path.splitext(all_files[0])[-1].lower()
print(f"File type: {ext}")> **Critical rule**: When `input_path` is a directory OR the user says "这些文件" / "所有文档", > process **every file** and aggregate. Never stop at the first file.
---
Step 1 — Load sub-skill by format
| Extension | Sub-skill to load | |-----------|------------------| | `.docx` / `.doc` | `capability/word-analysis/SKILL.md` | | `.pdf` | `capability/pdf-analysis/SKILL.md` | | `.pptx` / `.ppt` | `capability/ppt-analysis/SKILL.md` |
read_file(path="<skills_root>/sn-da-non-spreadsheet-analysis/capability/<format>-analysis/SKILL.md")
Load **only the sub-skill you need** — do not load all three at once.
---
Step 2 — Parse and extract
Follow the sub-skill's extraction pattern. For all formats:
- **Full scan**: iterate all pages/slides/paragraphs — never stop early
- **Table extraction**: get every table, not just the first one
- **Image/chart detection**: if a page/slide yields no text, treat it as image-based and call `caption.py`
---
Step 3 — Answer with verification
After extracting data, verify before answering:
# For count/statistics questions: spot-check 3-5 items
sample = result_list[:3]
print(f"Sample check: {sample}")
print(f"Total count: {len(result_list)}")
# For numeric calculations: print intermediate values
print(f"Max={max_val}, Min={min_val}, Range={max_val - min_val}")
# For unit-sensitive answers: always include the unit
print(f"Answer: {value} {unit}") # e.g., "475 千港元" not just "475"---
Universal Rules
MUST DO
- **Always iterate all pages/slides/paragraphs** — `for page in doc`, `for slide in prs.slides`, `for para in doc.paragraphs`
- **When input is a directory**: collect and process all matching files, then aggregate results
- **For scanned PDFs**: detect empty text → call `caption.py` for OCR
- **For image-only slides**: text extraction returns empty → render slide as PNG → call `caption.py`
- **For calculations**: show intermediate values; confirm unit matches the question
NEVER DO
- Do NOT use `pytesseract` or `easyocr` as primary OCR — they are not installed; use `caption.py`
- Do NOT use PIL pixel analysis to infer chart values — use vision model caption instead
- Do NOT stop at the first file, first page, or first table
- Do NOT guess content from filenames — always parse the actual file
- Do NOT output percentage when the question asks for absolute value (and vice versa)
---
Caption Script (for image/chart content in any document)
When a page, slide, or embedded image needs vision understanding, load the `sn-da-image-caption` skill first, then use its `scripts/caption.py`:
read_file(path="<skills_root>/sn-da-image-caption/SKILL.md")
import subprocess, json
CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py"
def caption_image(image_path, prompt=None):
cmd = ["python3", CAPTION, image_path, "--json"]
if prompt:
cmd += ["--prompt", prompt]
result = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
if result.returncode != 0:
raise RuntimeError(f"caption failed: {result.stderr[:200]}")
return json.loads(result.stdout)["description"]
# Example prompts by content type:
# Table: "提取表格所有内容,Markdown 表格格式,保持行列结构,数值不四舍五入。"
# Chart: "提取图表标题、坐标轴标签、每个数据点的数值。Markdown 表格输出。"
# Diagram: "描述所有节点和连接关系。"---
Available sub-skills
sn-da-non-spreadsheet-analysis/capability/word-analysis/SKILL.md — .docx/.doc
sn-da-non-spreadsheet-analysis/capability/pdf-analysis/SKILL.md — .pdf
sn-da-non-spreadsheet-analysis/capability/ppt-analysis/SKILL.md — .pptx/.ppt
Read more
name: sn-da-non-spreadsheet-analysis description: "Word / PDF / PPT 文档解析与数据分析引擎。覆盖三类文件格式的全量提取、表格数值化、图表理解与跨文档汇总分析。**遇到以下任一情况就主动使用本 skill**:①用户上传或指定了 .docx / .doc / .pdf / .pptx / .ppt 文件并要求分析、提取或统计其中内容;②用户出现触发词:Word分析 / PDF解析 / PPT提取 / 文档分析 / 报告解析 / 幻灯片分析 / 发票提取 / 合同分析 / 文档统计 / 错别字 / 语病 / 字号检查 / 简历分析 / 多文档对比;③任务涉及从文档中提取表格、数值、图表、格式(颜色/高亮/字号)、组织架构、时间线等结构化信息。仅不用于:Excel/CSV 数据分析(使用 sn-da-excel-workflow)、纯图片分析(使用 sn-da-image-caption)。"
Document Analysis Skill — Word / PDF / PPT
End-to-end workflow for Word, PDF, and PPT document parsing. Each format has specific parsing pitfalls — follow the format-specific sub-skill exactly.
---
Workflow
Step 0 — Identify file type and input scope
import os
input_path = "/mnt/data/..." # from user
# Detect single file vs directory (multi-file scenario)
if os.path.isdir(input_path):
all_files = [
os.path.join(input_path, f)
for f in os.listdir(input_path)
if f.lower().endswith(('.docx', '.doc', '.pdf', '.pptx', '.ppt'))
]
print(f"Found {len(all_files)} documents: {all_files}")
else:
all_files = [input_path]
# Route by extension
ext = os.path.splitext(all_files[0])[-1].lower()
print(f"File type: {ext}")> **Critical rule**: When `input_path` is a directory OR the user says "这些文件" / "所有文档", > process **every file** and aggregate. Never stop at the first file.
---
Step 1 — Load sub-skill by format
| Extension | Sub-skill to load | |-----------|------------------| | `.docx` / `.doc` | `capability/word-analysis/SKILL.md` | | `.pdf` | `capability/pdf-analysis/SKILL.md` | | `.pptx` / `.ppt` | `capability/ppt-analysis/SKILL.md` |
read_file(path="<skills_root>/sn-da-non-spreadsheet-analysis/capability/<format>-analysis/SKILL.md")
Load **only the sub-skill you need** — do not load all three at once.
---
Step 2 — Parse and extract
Follow the sub-skill's extraction pattern. For all formats:
- **Full scan**: iterate all pages/slides/paragraphs — never stop early
- **Table extraction**: get every table, not just the first one
- **Image/chart detection**: if a page/slide yields no text, treat it as image-based and call `caption.py`
---
Step 3 — Answer with verification
After extracting data, verify before answering:
# For count/statistics questions: spot-check 3-5 items
sample = result_list[:3]
print(f"Sample check: {sample}")
print(f"Total count: {len(result_list)}")
# For numeric calculations: print intermediate values
print(f"Max={max_val}, Min={min_val}, Range={max_val - min_val}")
# For unit-sensitive answers: always include the unit
print(f"Answer: {value} {unit}") # e.g., "475 千港元" not just "475"---
Universal Rules
MUST DO
- **Always iterate all pages/slides/paragraphs** — `for page in doc`, `for slide in prs.slides`, `for para in doc.paragraphs`
- **When input is a directory**: collect and process all matching files, then aggregate results
- **For scanned PDFs**: detect empty text → call `caption.py` for OCR
- **For image-only slides**: text extraction returns empty → render slide as PNG → call `caption.py`
- **For calculations**: show intermediate values; confirm unit matches the question
NEVER DO
- Do NOT use `pytesseract` or `easyocr` as primary OCR — they are not installed; use `caption.py`
- Do NOT use PIL pixel analysis to infer chart values — use vision model caption instead
- Do NOT stop at the first file, first page, or first table
- Do NOT guess content from filenames — always parse the actual file
- Do NOT output percentage when the question asks for absolute value (and vice versa)
---
Caption Script (for image/chart content in any document)
When a page, slide, or embedded image needs vision understanding, load the `sn-da-image-caption` skill first, then use its `scripts/caption.py`:
read_file(path="<skills_root>/sn-da-image-caption/SKILL.md")
import subprocess, json
CAPTION = "/path/to/skills/sn-da-image-caption/scripts/caption.py"
def caption_image(image_path, prompt=None):
cmd = ["python3", CAPTION, image_path, "--json"]
if prompt:
cmd += ["--prompt", prompt]
result = subprocess.run(cmd, capture_output=True, text=True, timeout=60)
if result.returncode != 0:
raise RuntimeError(f"caption failed: {result.stderr[:200]}")
return json.loads(result.stdout)["description"]
# Example prompts by content type:
# Table: "提取表格所有内容,Markdown 表格格式,保持行列结构,数值不四舍五入。"
# Chart: "提取图表标题、坐标轴标签、每个数据点的数值。Markdown 表格输出。"
# Diagram: "描述所有节点和连接关系。"---
Available sub-skills
sn-da-non-spreadsheet-analysis/capability/word-analysis/SKILL.md — .docx/.doc sn-da-non-spreadsheet-analysis/capability/pdf-analysis/SKILL.md — .pdf sn-da-non-spreadsheet-analysis/capability/ppt-analysis/SKILL.md — .pptx/.ppt
The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent, with the skills in this repository extending the models with concrete, end-to-end office capabilities.
Repo: OpenSenseNova/SenseNova-Skills
Other skills on sensenova-skills.
- /sn-da-excel-workflow
Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子 skill。**遇到以下任一情况就主动使用本 skill,不要自行写几行 pandas 就回答**:①用户出现触发词:Excel 分析 / 表格分析 / 数据分析 /
Open skill - /category-coloring
当Excel文件总行数超过1万行时,通过转换为Parquet格式提升读取性能,提取目标指标并计算最大值,最后将结果输出为Excel并对特定行进行高亮标注。
Open skill - /duplicate-value-coloring
对比Excel多表中的特定系数并对异常值进行颜色标记。
Open skill - /outlier-coloring
识别 Excel 中的超限数值与错误单元格并进行高亮标注。
Open skill - /threshold-cell-coloring
根据Excel总行数自动切换Parquet加速读取,计算特定维度的时间序列平均值,并使用openpyxl输出带有条件格式(如低于均值标绿)和自定义样式的分析报告。
Open skill - /top-value-coloring
根据数据规模动态选择处理策略,对多表数据进行合并、统计筛选,并利用 openpyxl 实现关键指标的自动化样式高亮与格式化导出。
Open skill

