doc-parse
将 PDF/PPT/Excel/Word 等多格式文档解析为结构化 Markdown,并输出元数据与解析置信度,作为 RAG 与四色卡片的数据底座。
综合办公文员与软件开发工程师当需要批量处理PDF表单、提取表格或进行OCR识别时,使用内置脚本一键完成自动化提取与数据校验,彻底告别繁琐的手动录入,让复杂文档工作流高效、稳健落地。
$ npx -y skills add anbeime/skill --skill pdf-processing-pro --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/pdf-processing-proContext preview
The summary Claude sees to decide when to auto-load this skill.
综合办公文员与软件开发工程师当需要批量处理PDF表单、提取表格或进行OCR识别时,使用内置脚本一键完成自动化提取与数据校验,彻底告别繁琐的手动录入,让复杂文档工作流高效、稳健落地。
name: PDF Processing Pro description: 综合办公文员与软件开发工程师当需要批量处理PDF表单、提取表格或进行OCR识别时,使用内置脚本一键完成自动化提取与数据校验,彻底告别繁琐的手动录入,让复杂文档工作流高效、稳健落地。
Production-ready PDF processing toolkit with pre-built scripts, comprehensive error handling, and support for complex workflows.
import pdfplumber
with pdfplumber.open("document.pdf") as pdf:
text = pdf.pages[0].extract_text()
print(text)python scripts/analyze_form.py input.pdf --output fields.json # Returns: JSON with all form fields, types, and positions
python scripts/fill_form.py input.pdf data.json output.pdf # Validates all fields before filling, includes error reporting
python scripts/extract_tables.py report.pdf --output tables.csv # Extracts all tables with automatic column detection
All scripts include:
For complete form workflows including:
See [FORMS.md](FORMS.md)
For complex table extraction:
See [TABLES.md](TABLES.md)
For scanned PDFs and image-based documents:
See [OCR.md](OCR.md)
**analyze_form.py** - Extract form field information
python scripts/analyze_form.py input.pdf [--output fields.json] [--verbose]
**fill_form.py** - Fill PDF forms with data
python scripts/fill_form.py input.pdf data.json output.pdf [--validate]
**validate_form.py** - Validate form data before filling
python scripts/validate_form.py data.json schema.json
**extract_tables.py** - Extract tables to CSV/Excel
python scripts/extract_tables.py input.pdf [--output tables.csv] [--format csv|excel]
**extract_text.py** - Extract text with formatting preservation
python scripts/extract_text.py input.pdf [--output text.txt] [--preserve-formatting]
**merge_pdfs.py** - Merge multiple PDFs
python scripts/merge_pdfs.py file1.pdf file2.pdf file3.pdf --output merged.pdf
**split_pdf.py** - Split PDF into individual pages
python scripts/split_pdf.py input.pdf --output-dir pages/
**validate_pdf.py** - Validate PDF integrity
python scripts/validate_pdf.py input.pdf
# 1. Analyze form structure python scripts/analyze_form.py template.pdf --output schema.json # 2. Validate submission data python scripts/validate_form.py submission.json schema.json # 3. Fill form python scripts/fill_form.py template.pdf submission.json completed.pdf # 4. Validate output python scripts/validate_pdf.py completed.pdf
# 1. Extract tables python scripts/extract_tables.py monthly_report.pdf --output data.csv # 2. Extract text for analysis python scripts/extract_text.py monthly_report.pdf --output report.txt
import glob
from pathlib import Path
import subprocess
# Process all PDFs in directory
for pdf_file in glob.glob("invoices/*.pdf"):
output_file = Path("processed") / Path(pdf_file).name
result = subprocess.run([
"python", "scripts/extract_text.py",
pdf_file,
"--output", str(output_file)
], capture_output=True)
if result.returncode == 0:
print(f"✓ Processed: {pdf_file}")
else:
print(f"✗ Failed: {pdf_file} - {result.stderr}")All scripts follow consistent error patterns:
# Exit codes
# 0 - Success
# 1 - File not found
# 2 - Invalid input
# 3 - Processing error
# 4 - Validation error
# Example usage in automation
result = subprocess.run(["python", "scripts/fill_form.py", ...])
if result.returncode == 0:
print("Success")
elif result.returncode == 4:
print("Validation failed - check input data")
else:
print(f"Error occurred: {result.returncode}")All scripts require:
pip install pdfplumber pypdf pillow pytesseract pandas
Optional for OCR:
# Install tesseract-ocr system package # macOS: brew install tesseract # Ubuntu: apt-get install tesseract-ocr # Windows: Download from GitHub releases
1. **Always validate inputs** before processing 2. **Use try-except** in custom scripts 3. **Log all operations** for debugging 4. **Test with sample PDFs** before production 5. **Set timeouts** for long-running operations 6. **Check exit codes** in automation 7. **Backup originals** before m
收录最全、更新最快的AI Agent技能库,涵盖文档处理、内容创作、编程开发、机器学习、自动化工作流等多个领域的精选技能包。 🧠 知易智能基座 :20+模型自由切换,知识永远留在你手里 —一个对话框调度所有模型,240+技能即插即用,四色卡片让AI真正记住你。 👉 申请体验
将 PDF/PPT/Excel/Word 等多格式文档解析为结构化 Markdown,并输出元数据与解析置信度,作为 RAG 与四色卡片的数据底座。
产品经理与项目管理专家在应对复杂项目时,使用此技能可动态组建包含“执行、指挥、评审”的专属AI团队。实时查看多智能体辩论与决策全过程,共享统一上下文记忆,一键完成从会议决策到系统构建的全流…
电商运营与内容创作者在需要制作带货短视频时,用此技能一键生成9:16竖屏数字人成片。自动编排AI绘画、语音合成与视频生成,快速产出“小省导购员”专属形象与专业配音,让多模态视频制作省时省力…