Skip to content
AI & Agents
Skill

/pdf-processing-pro

综合办公文员与软件开发工程师当需要批量处理PDF表单、提取表格或进行OCR识别时,使用内置脚本一键完成自动化提取与数据校验,彻底告别繁琐的手动录入,让复杂文档工作流高效、稳健落地。

BOOST
From plugin
anbeime-skill
7.4k78 skills
Install
$ npx -y skills add anbeime/skill --skill pdf-processing-pro --agent claude-code

How it fires

How this skill gets triggered: by you, by Claude, or both.

  • Fires itselfAuto-invocation. Claude auto-loads it when your prompt matches the work.Auto-invocation is when the right skill fires by itself at the right moment, driven by a FLOW.md router and a hook, instead of you invoking it by name. It is the difference between a skill being installed and a skill actually getting used.Read the full definition →
  • You can call itInvoke it directly when you want it.
  • Slash command/pdf-processing-pro

Context preview

The summary Claude sees to decide when to auto-load this skill.

综合办公文员与软件开发工程师当需要批量处理PDF表单、提取表格或进行OCR识别时,使用内置脚本一键完成自动化提取与数据校验,彻底告别繁琐的手动录入,让复杂文档工作流高效、稳健落地。

SKILL.md

pdf-processing-pro.SKILL.md
name: PDF Processing Pro
description: 综合办公文员与软件开发工程师当需要批量处理PDF表单、提取表格或进行OCR识别时,使用内置脚本一键完成自动化提取与数据校验,彻底告别繁琐的手动录入,让复杂文档工作流高效、稳健落地。

PDF Processing Pro

Production-ready PDF processing toolkit with pre-built scripts, comprehensive error handling, and support for complex workflows.

Quick start

Extract text from PDF

import pdfplumber

with pdfplumber.open("document.pdf") as pdf:
    text = pdf.pages[0].extract_text()
    print(text)

Analyze PDF form (using included script)

python scripts/analyze_form.py input.pdf --output fields.json
# Returns: JSON with all form fields, types, and positions

Fill PDF form with validation

python scripts/fill_form.py input.pdf data.json output.pdf
# Validates all fields before filling, includes error reporting

Extract tables from PDF

python scripts/extract_tables.py report.pdf --output tables.csv
# Extracts all tables with automatic column detection

Features

✅ Production-ready scripts

All scripts include:

  • **Error handling**: Graceful failures with detailed error messages
  • **Validation**: Input validation and type checking
  • **Logging**: Configurable logging with timestamps
  • **Type hints**: Full type annotations for IDE support
  • **CLI interface**: `--help` flag for all scripts
  • **Exit codes**: Proper exit codes for automation

✅ Comprehensive workflows

  • **PDF Forms**: Complete form processing pipeline
  • **Table Extraction**: Advanced table detection and extraction
  • **OCR Processing**: Scanned PDF text extraction
  • **Batch Operations**: Process multiple PDFs efficiently
  • **Validation**: Pre and post-processing validation

Advanced topics

PDF Form Processing

For complete form workflows including:

  • Field analysis and detection
  • Dynamic form filling
  • Validation rules
  • Multi-page forms
  • Checkbox and radio button handling

See [FORMS.md](FORMS.md)

Table Extraction

For complex table extraction:

  • Multi-page tables
  • Merged cells
  • Nested tables
  • Custom table detection
  • Export to CSV/Excel

See [TABLES.md](TABLES.md)

OCR Processing

For scanned PDFs and image-based documents:

  • Tesseract integration
  • Language support
  • Image preprocessing
  • Confidence scoring
  • Batch OCR

See [OCR.md](OCR.md)

Included scripts

Form processing

**analyze_form.py** - Extract form field information

python scripts/analyze_form.py input.pdf [--output fields.json] [--verbose]

**fill_form.py** - Fill PDF forms with data

python scripts/fill_form.py input.pdf data.json output.pdf [--validate]

**validate_form.py** - Validate form data before filling

python scripts/validate_form.py data.json schema.json

Table extraction

**extract_tables.py** - Extract tables to CSV/Excel

python scripts/extract_tables.py input.pdf [--output tables.csv] [--format csv|excel]

Text extraction

**extract_text.py** - Extract text with formatting preservation

python scripts/extract_text.py input.pdf [--output text.txt] [--preserve-formatting]

Utilities

**merge_pdfs.py** - Merge multiple PDFs

python scripts/merge_pdfs.py file1.pdf file2.pdf file3.pdf --output merged.pdf

**split_pdf.py** - Split PDF into individual pages

python scripts/split_pdf.py input.pdf --output-dir pages/

**validate_pdf.py** - Validate PDF integrity

python scripts/validate_pdf.py input.pdf

Common workflows

Workflow 1: Process form submissions

# 1. Analyze form structure
python scripts/analyze_form.py template.pdf --output schema.json

# 2. Validate submission data
python scripts/validate_form.py submission.json schema.json

# 3. Fill form
python scripts/fill_form.py template.pdf submission.json completed.pdf

# 4. Validate output
python scripts/validate_pdf.py completed.pdf

Workflow 2: Extract data from reports

# 1. Extract tables
python scripts/extract_tables.py monthly_report.pdf --output data.csv

# 2. Extract text for analysis
python scripts/extract_text.py monthly_report.pdf --output report.txt

Workflow 3: Batch processing

import glob
from pathlib import Path
import subprocess

# Process all PDFs in directory
for pdf_file in glob.glob("invoices/*.pdf"):
    output_file = Path("processed") / Path(pdf_file).name

    result = subprocess.run([
        "python", "scripts/extract_text.py",
        pdf_file,
        "--output", str(output_file)
    ], capture_output=True)

    if result.returncode == 0:
        print(f"✓ Processed: {pdf_file}")
    else:
        print(f"✗ Failed: {pdf_file} - {result.stderr}")

Error handling

All scripts follow consistent error patterns:

# Exit codes
# 0 - Success
# 1 - File not found
# 2 - Invalid input
# 3 - Processing error
# 4 - Validation error

# Example usage in automation
result = subprocess.run(["python", "scripts/fill_form.py", ...])

if result.returncode == 0:
    print("Success")
elif result.returncode == 4:
    print("Validation failed - check input data")
else:
    print(f"Error occurred: {result.returncode}")

Dependencies

All scripts require:

pip install pdfplumber pypdf pillow pytesseract pandas

Optional for OCR:

# Install tesseract-ocr system package
# macOS: brew install tesseract
# Ubuntu: apt-get install tesseract-ocr
# Windows: Download from GitHub releases

Performance tips

  • **Use batch processing** for multiple PDFs
  • **Enable multiprocessing** with `--parallel` flag (where supported)
  • **Cache extracted data** to avoid re-processing
  • **Validate inputs early** to fail fast
  • **Use streaming** for large PDFs (>50MB)

Best practices

1. **Always validate inputs** before processing 2. **Use try-except** in custom scripts 3. **Log all operations** for debugging 4. **Test with sample PDFs** before production 5. **Set timeouts** for long-running operations 6. **Check exit codes** in automation 7. **Backup originals** before m

Read more
Ships withanbeime-skill

收录最全、更新最快的AI Agent技能库,涵盖文档处理、内容创作、编程开发、机器学习、自动化工作流等多个领域的精选技能包。 🧠 知易智能基座 :20+模型自由切换,知识永远留在你手里 —一个对话框调度所有模型,240+技能即插即用,四色卡片让AI真正记住你。 👉 申请体验

Get the whole plugin
Stats
7,502
Stars
683
Forks
Active
Maintenance
Python
Language
6h ago
Last commit
8mo ago
Created
3h ago
Added

Repo: anbeime/skill

Other skills on anbeime-skill.

agent-team
Skill

agent-team

产品经理与项目管理专家在应对复杂项目时,使用此技能可动态组建包含“执行、指挥、评审”的专属AI团队。实时查看多智能体辩论与决策全过程,共享统一上下文记忆,一键完成从会议决策到系统构建的全流…

agentkit-multimedia-shopping
Skill

agentkit-multimedia-sh…

电商运营与内容创作者在需要制作带货短视频时,用此技能一键生成9:16竖屏数字人成片。自动编排AI绘画、语音合成与视频生成,快速产出“小省导购员”专属形象与专业配音,让多模态视频制作省时省力…