$npx skillfedfor your agent

pdf

PDF skill provides Python-based tools for extracting text and tables, merging and splitting documents, creating new PDFs, and handling forms. Use it to process, analyze, or generate PDF files at scale.

PDF skill extracts text and tables from PDF documents programmatically using Python libraries.

AI-generated summary based on this skill's SKILL.md

137 9 MITupdated by appautomaton

Decision gist · record as of 2026-07-01

PDF skill extracts text and tables from PDF documents programmatically using Python libraries. PDF skill provides Python-based tools for extracting text and tables, merging and splitting documents, creating new PDFs, and handling forms. Use it to process, analyze, or generate PDF files at scale.

manual: git clone https://github.com/appautomaton/document-SKILLs → cp -r document-SKILLs/pdf ~/.claude/skills/pdf
pdf/SKILL.md · version b49646aa

Use it when

  • PDF skill provides Python-based tools to extract text from PDF documents efficiently.
  • Yes, PDF skill lets you split PDF documents into individual pages or custom page ranges.

Verify before relying

Read SKILL.md below before installing (13 files). Open directory: indexed for reading, not audited.

Same gist for agents: .md · .json

Install

appautomaton/document-SKILLs/pdf · repository language: Python

Open directory. Skills are indexed for reading, not audited. Review a skill's body before installing it.

Frequently asked questions

AI-generated answers based on this skill's SKILL.md and metadata

How can I merge PDF files using the PDF skill?

PDF skill enables you to merge multiple PDF documents programmatically. You can combine pages from different PDFs into a single document, controlling the order and structure of merged content. This is useful for consolidating reports, combining chapters, or batch-processing document collections.

What's the best way to extract text from PDF with Python?

PDF skill provides Python-based tools to extract text from PDF documents efficiently. You can retrieve text content while preserving layout information, making it easier to parse structured data. The skill handles various PDF formats and supports extracting text alongside tables and metadata.

Can PDF skill split a PDF into separate pages?

Yes, PDF skill lets you split PDF documents into individual pages or custom page ranges. You can extract specific pages, reorganize document structure, or separate multi-page PDFs for further processing. This is essential for batch operations and selective document handling.

How do I read PDF tables and convert them to Excel?

PDF skill can extract tables from PDF documents and convert them into structured formats like Excel. It recognizes table layouts and preserves cell relationships, allowing you to transform PDF data into spreadsheets for analysis and reporting.

Can I create and generate new PDF documents from code?

PDF skill supports creating new PDF documents programmatically from Python code. You can generate PDFs from scratch, add content dynamically, and automate document generation workflows. This enables building custom reports, invoices, and data-driven documents at scale.

Does PDF skill support form filling and automation?

PDF skill provides tools for automating PDF form processing and filling. You can populate form fields programmatically, extract form data, and streamline document workflows. This is ideal for batch form processing, data entry automation, and compliance document handling.

SKILL.md

Rendered from the published skill. Quoted content, verbatim.

PDF Processing Guide

Overview

This guide covers essential PDF processing operations using Python libraries and command-line tools. For advanced features, JavaScript libraries, and detailed examples, see reference.md. If you need to fill out a PDF form, read forms.md and follow its instructions.

Prerequisites

Python dependencies are resolved automatically by uv run — every script declares them in its PEP 723 header. Some workflows also need system tools:

  • poppler (brew install poppler / apt-get install poppler-utils) — provides pdftoppm and pdftotext; required by the form-filling workflow, which converts PDF pages to images via pdf2image
  • tesseract (brew install tesseract tesseract-lang) — only for OCR on scanned documents (see ocr.md)
  • qpdf (brew install qpdf) — only for the command-line recipes in the qpdf section below

Quick

(truncated - see the full file via the links below)

File tree — 13 files
pdf/SKILL.md
pdf/forms.md
pdf/ocr.md
pdf/reference.md
pdf/scripts/check_bounding_boxes.py
pdf/scripts/check_bounding_boxes_test.py
pdf/scripts/check_fillable_fields.py
pdf/scripts/convert_pdf_to_images.py
pdf/scripts/create_validation_image.py
pdf/scripts/extract_form_field_info.py
pdf/scripts/fill_fillable_fields.py
pdf/scripts/fill_pdf_form_with_annotations.py
pdf/tables.md

Let your AI agent find skills like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 56,283 agent skills by what they can do, searchable in plain language.

wish › “Extract text and tables from PDF documents programmatically”

Give your agent the search over MCP, or paste the wish link into any chat. No install? Search from any chat →

Related skills

document-skills/pdf
by aitytech · aitytech/agentkits-marketing

This skill provides Python-based PDF processing for extracting content, manipulating documents, and generating new files. Use it to pull text and tables from existing PDFs, merge or split documents, create PDFs from scratch, rotate pages, add watermarks, and handle form filling. Supports both library-based approaches and command-line tools.

MITupdated Jul 2026
★ 573repo stars
Pdf
by samhvw8 · samhvw8/dot-claude

This skill provides comprehensive PDF processing capabilities through Python libraries like pypdf, pdfplumber, and reportlab, plus command-line tools. Extract text and tables, merge or split documents, create PDFs programmatically, add watermarks, rotate pages, and handle form filling. Supports both simple operations and advanced workflows like OCR for scanned documents.

no license declared → metadata onlyupdated Dec 2025
★ 10repo stars
Pdf
by AgentTeam-TaichuAI · AgentTeam-TaichuAI/ScienceClaw

Handle any PDF task—from text and table extraction to merging, splitting, encryption, and form filling. The skill automatically selects the right extraction method based on your document type, whether it's academic papers, simple text files, or scanned images requiring OCR. Generate professional PDF reports from research and analysis work.

no license declared → metadata onlyupdated May 2026
★ 564repo stars
Anthropics Pdf
by family3253 · family3253/skill

Anthropics Pdf handles the full spectrum of PDF operations—reading and extracting text or tables, merging multiple files, splitting pages, rotating content, adding watermarks, creating new documents, filling forms, encrypting files, pulling images, and running OCR on scanned documents. Built on Python libraries like pypdf and pdfplumber plus command-line utilities, it covers everything from simple text extraction to complex document workflows.

no license declared → metadata onlyupdated Apr 2026
★ 0repo stars
Pdf
by Wide-Moat · Wide-Moat/open-computer-use

This skill provides comprehensive PDF manipulation capabilities including text and table extraction, document merging and splitting, form filling, and PDF creation. It supports both Python libraries and command-line tools for handling PDF operations at scale.

no license declared → metadata onlyupdated Jul 2026
★ 107repo stars
Pdf
by dawiddutoit · dawiddutoit/custom-claude

This skill provides Python-based PDF processing workflows using pypdf for core operations, pdfplumber for text and table extraction, and ReportLab for document generation. Handle merging, splitting, watermarking, encryption, metadata work, and form filling through code examples and command-line tools.

no license declared → metadata onlyupdated Jan 2026
★ 1repo stars

More skills pdf-processor (MIT) · Pdf (unlicensed)

Tags
document-automationdata-extractionform-processingbatch-operationsfile-conversionocr-scanningmetadata-handlingencryption-securitytable-parsingpage-manipulation