/duplicate-removal
统计多Sheet Excel总行数并根据规模选择处理策略,提取特定维度信息进行去重统计,并生成摘要与明细报表。
$ npx -y skills add OpenSenseNova/SenseNova-Skills --skill duplicate-removal --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
- Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
- You can call itInvoke it directly when you want it.
- Slash command
/duplicate-removal
Context preview
The summary Claude sees to decide when to auto-load this skill.
统计多Sheet Excel总行数并根据规模选择处理策略,提取特定维度信息进行去重统计,并生成摘要与明细报表。
SKILL.md
duplicate-removal.SKILL.mdname: excel-multi-sheet-threshold-analysis
description: "统计多Sheet Excel总行数并根据规模选择处理策略,提取特定维度信息进行去重统计,并生成摘要与明细报表。"
Excel_Multi_Sheet_Deduplication
> This sub-skill covers one capability of the Excel workflow. For reading/counting/Parquet optimization, see the parent workflow SKILL.md.
Step1 加载目标数据表,并进行初步的数据预览与结构检查。
import pandas as pd
file_path = 'input_file.xlsx'
target_sheet = 'Sheet1' # 根据实际情况指定 sheet 名称
# 读取数据,header=None 用于处理无表头或非标准表头文件
df = pd.read_excel(file_path, sheet_name=target_sheet, header=None)
print(f"数据形状: {df.shape}")
print("前 5 行预览:")
print(df.head())Step2 遍历数据行,基于关键词提取目标信息,并执行数据清洗(去除空格、空值过滤)。
import pandas as pd
# 设定目标列索引及过滤关键词
target_col_idx = 1
keywords = ["关键词A", "关键词B"] # 示例:如"综合楼"、"控制中心"
extracted_data = []
for idx, row in df.iterrows():
cell_val = str(row[target_col_idx]) if pd.notna(row[target_col_idx]) else ""
# 数据清洗:去除首尾空格并匹配关键词
clean_val = cell_val.strip()
if any(k in clean_val for k in keywords):
if clean_val and clean_val.lower() not in ["nan", "null", ""]:
extracted_data.append(clean_val)
print(f"提取到相关记录共 {len(extracted_data)} 条")Step3 对提取的信息进行分类去重,统计各维度的唯一项数量。
# 使用 set 进行高效去重
category_a_items = set()
category_b_items = set()
for item in extracted_data:
if "关键词A" in item:
category_a_items.add(item)
elif "关键词B" in item:
category_b_items.add(item)
# 转换为排序后的列表
list_a = sorted(list(category_a_items))
list_b = sorted(list(category_b_items))
print(f"类别A 唯一项数量: {len(list_a)}")
print(f"类别B 唯一项数量: {len(list_b)}")Step4 将统计摘要与详细清单整理为 DataFrame,并导出为 Excel 文件提供下载。
import pandas as pd
# 1. 生成统计摘要
summary_df = pd.DataFrame({
'分类名称': ['类别A', '类别B'],
'唯一项总数': [len(list_a), len(list_b)]
})
# 2. 生成详细清单
detail_list = []
for val in list_a:
detail_list.append({'分类': '类别A', '详细名称': val})
for val in list_b:
detail_list.append({'分类': '类别B', '详细名称': val})
detail_df = pd.DataFrame(detail_list)
# 导出结果
output_summary_path = 'summary_report.xlsx'
output_detail_path = 'detail_list.xlsx'
summary_df.to_excel(output_summary_path, index=False)
detail_df.to_excel(output_detail_path, index=False)
print(f"统计摘要已保存: {output_summary_path}")
print(f"详细清单已保存: {output_detail_path}")Read more
name: excel-multi-sheet-threshold-analysis description: "统计多Sheet Excel总行数并根据规模选择处理策略,提取特定维度信息进行去重统计,并生成摘要与明细报表。"
Excel_Multi_Sheet_Deduplication
> This sub-skill covers one capability of the Excel workflow. For reading/counting/Parquet optimization, see the parent workflow SKILL.md.
Step1 加载目标数据表,并进行初步的数据预览与结构检查。
import pandas as pd
file_path = 'input_file.xlsx'
target_sheet = 'Sheet1' # 根据实际情况指定 sheet 名称
# 读取数据,header=None 用于处理无表头或非标准表头文件
df = pd.read_excel(file_path, sheet_name=target_sheet, header=None)
print(f"数据形状: {df.shape}")
print("前 5 行预览:")
print(df.head())Step2 遍历数据行,基于关键词提取目标信息,并执行数据清洗(去除空格、空值过滤)。
import pandas as pd
# 设定目标列索引及过滤关键词
target_col_idx = 1
keywords = ["关键词A", "关键词B"] # 示例:如"综合楼"、"控制中心"
extracted_data = []
for idx, row in df.iterrows():
cell_val = str(row[target_col_idx]) if pd.notna(row[target_col_idx]) else ""
# 数据清洗:去除首尾空格并匹配关键词
clean_val = cell_val.strip()
if any(k in clean_val for k in keywords):
if clean_val and clean_val.lower() not in ["nan", "null", ""]:
extracted_data.append(clean_val)
print(f"提取到相关记录共 {len(extracted_data)} 条")Step3 对提取的信息进行分类去重,统计各维度的唯一项数量。
# 使用 set 进行高效去重
category_a_items = set()
category_b_items = set()
for item in extracted_data:
if "关键词A" in item:
category_a_items.add(item)
elif "关键词B" in item:
category_b_items.add(item)
# 转换为排序后的列表
list_a = sorted(list(category_a_items))
list_b = sorted(list(category_b_items))
print(f"类别A 唯一项数量: {len(list_a)}")
print(f"类别B 唯一项数量: {len(list_b)}")Step4 将统计摘要与详细清单整理为 DataFrame,并导出为 Excel 文件提供下载。
import pandas as pd
# 1. 生成统计摘要
summary_df = pd.DataFrame({
'分类名称': ['类别A', '类别B'],
'唯一项总数': [len(list_a), len(list_b)]
})
# 2. 生成详细清单
detail_list = []
for val in list_a:
detail_list.append({'分类': '类别A', '详细名称': val})
for val in list_b:
detail_list.append({'分类': '类别B', '详细名称': val})
detail_df = pd.DataFrame(detail_list)
# 导出结果
output_summary_path = 'summary_report.xlsx'
output_detail_path = 'detail_list.xlsx'
summary_df.to_excel(output_summary_path, index=False)
detail_df.to_excel(output_detail_path, index=False)
print(f"统计摘要已保存: {output_summary_path}")
print(f"详细清单已保存: {output_detail_path}")The SenseNova model family plugs directly into agent runtimes such as OpenClaw and hermes-agent, with the skills in this repository extending the models with concrete, end-to-end office capabilities.
Repo: OpenSenseNova/SenseNova-Skills
Other skills on sensenova-skills.
- /sn-da-excel-workflow
Excel 数据分析多步编排器。覆盖:(1) 读取多 Sheet Excel 文件并统计行数,(2) 大文件检测(≥10k 行自动 Parquet 优化),(3) 数据清洗(缺失值、文本标准化、无效字符),(4) 条件筛选与分类提取,(5) 跨 Sheet 统计聚合,(6) 导出 Excel/CSV 并提供下载链接。覆盖从数据读取到报告生成全流程,按步骤编排 capability 子 skill。**遇到以下任一情况就主动使用本 skill,不要自行写几行 pandas 就回答**:①用户出现触发词:Excel 分析 / 表格分析 / 数据分析 /
Open skill - /category-coloring
当Excel文件总行数超过1万行时,通过转换为Parquet格式提升读取性能,提取目标指标并计算最大值,最后将结果输出为Excel并对特定行进行高亮标注。
Open skill - /duplicate-value-coloring
对比Excel多表中的特定系数并对异常值进行颜色标记。
Open skill - /outlier-coloring
识别 Excel 中的超限数值与错误单元格并进行高亮标注。
Open skill - /threshold-cell-coloring
根据Excel总行数自动切换Parquet加速读取,计算特定维度的时间序列平均值,并使用openpyxl输出带有条件格式(如低于均值标绿)和自定义样式的分析报告。
Open skill - /top-value-coloring
根据数据规模动态选择处理策略,对多表数据进行合并、统计筛选,并利用 openpyxl 实现关键指标的自动化样式高亮与格式化导出。
Open skill

