{"enrichment":{"faq":[{"a":"Document Processing handles PDF, DOCX, PPTX, and XLSX files. The skill processes these formats locally on your machine using the open-source Docling library, converting unstructured content into search-ready chunks without relying on cloud services.","q":"What file formats does Document Processing support?"},{"a":"Document Processing automatically converts your files into indexed-ready JSONL chunks. Each chunk includes extracted text, headings, source file references, and page numbers, making them ready for direct ingestion into OpenSearch or other search engines.","q":"How do I chunk documents for search indexing with this tool?"},{"a":"Yes. Document Processing runs entirely on your machine without cloud services. It transforms unstructured files into search-ready chunks locally, giving you full control over your document data while preparing it for indexing.","q":"Can Document Processing process unstructured documents locally?"},{"a":"Document Processing outputs JSONL format optimized for OpenSearch ingestion. Each record contains extracted text, headings, source file references, and page numbers, allowing you to index documents with full context and traceability.","q":"What output format does Document Processing generate?"},{"a":"Document Processing includes inspection capabilities to review chunk quality before indexing. You can examine the generated JSONL output, verify text extraction accuracy, and validate that headings and metadata are properly captured for your search engine.","q":"How can I evaluate document chunk quality before indexing?"},{"a":"Document Processing is designed to extract and split large documents into manageable indexed segments. It can process multiple PDFs, DOCX, PPTX, and XLSX files, converting them into search-ready chunks that maintain document structure and reference information.","q":"Does Document Processing handle batch processing of multiple files?"}],"shadow_tags":["local-processing","document-extraction","chunking-strategy","search-indexing","format-conversion","batch-operations","quality-evaluation","jsonl-output","no-cloud-required","text-preparation"],"summary_rewrite":"Document Processing transforms unstructured files into indexed-ready JSONL chunks using the open-source Docling library, running entirely on your machine. Output includes text, headings, source file references, and page numbers for direct ingestion into OpenSearch."},"files":[{"bytes":1535,"path":"skills/opensearch-skills/ingest/document-processing/SKILL.md","sha256":"8d08fabf374a597343f2c9cfbd8085613637745220aade35c43cdf806c1e0532","url":"https://skillfed.io/files/opensearch-project/opensearch-agent-skills/document-processing/428b6f76/SKILL.md"}],"id":"opensearch-project/opensearch-agent-skills/document-processing","links":{"html":"https://skillfed.io/opensearch-project/opensearch-agent-skills/document-processing","md":"https://skillfed.io/opensearch-project/opensearch-agent-skills/document-processing.md","repo":"https://github.com/opensearch-project/opensearch-agent-skills"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":30,"language":"Python","last_updated":"2026-07-22","license":"Apache-2.0","name":"document-processing","publisher":"opensearch-project","stars":37},"relations":{"similar":[{"id":"opensearch-project/opensearch-agent-skills/ingest"},{"id":"opensearch-project/opensearch-agent-skills/opensearch-launchpad"},{"id":"opensearch-project/opensearch-agent-skills/managed-ingestion-service"},{"id":"opensearch-project/opensearch-agent-skills/opensearch-skills"},{"id":"opensearch-project/opensearch-agent-skills/cloud"},{"id":"opensearch-project/opensearch-agent-skills/log-analytics"},{"id":"claude-office-skills/skills/data-extractor"},{"id":"opensearch-project/opensearch-agent-skills/trace-analytics"},{"id":"opensearch-project/opensearch-agent-skills/search"},{"id":"opensearch-project/opensearch-agent-skills/observability"}]},"slug":{"owner":"opensearch-project","repo":"opensearch-agent-skills","skill":"document-processing"},"version":"428b6f76"}
