{"enrichment":{"faq":[{"a":"ingest transforms PDF, DOCX, and PPTX files into JSONL chunks using Docling, a document-processing engine that runs locally on your machine. The skill parses unstructured content, structures it into discrete chunks, and outputs them in JSONL format ready for indexing without cloud dependency.","q":"How does ingest convert PDF to JSONL chunks?"},{"a":"Yes. ingest is designed to process unstructured documents for search by converting them into search-ready JSONL chunks. It evaluates chunk quality and structure to optimize search performance, then prepares the output for OpenSearch indexing or other downstream systems.","q":"Can ingest process unstructured documents for search?"},{"a":"ingest supports PDF, DOCX, and PPTX formats. The skill handles all three through its document-processing pipeline, converting each into consistently structured JSONL chunks suitable for ingestion into search systems.","q":"What file formats does ingest support?"},{"a":"No. ingest operates entirely locally on your machine, eliminating cloud dependency. All document processing, chunking, and quality assessment happen without AWS or external services, giving you full control over your data pipeline.","q":"Does ingest require cloud infrastructure?"},{"a":"ingest evaluates chunk quality and structure locally before ingestion, examining factors like coherence, size, and metadata completeness. This assessment ensures optimal search performance when chunks are indexed into OpenSearch or similar systems.","q":"How does ingest assess chunk quality before indexing?"},{"a":"ingest is licensed under Apache-2.0, allowing free use, modification, and distribution under the terms of the Apache License 2.0.","q":"What license does ingest use?"}],"shadow_tags":["local-processing","document-chunking","pdf-conversion","search-indexing","jsonl-export","unstructured-data","batch-ingestion","content-preparation"],"summary_rewrite":"ingest is a category skill for transforming unstructured files into JSONL chunks on your machine. It handles PDF, DOCX, and PPTX formats through the document-processing skill, which uses Docling to produce search-ready output without requiring cloud infrastructure or AWS."},"files":[{"bytes":1598,"path":"skills/opensearch-skills/ingest/SKILL.md","sha256":"07492aec507772967fac73208a2d380b6161d1b84bf1dac16f1c845c1ffecffd","url":"https://skillfed.io/files/opensearch-project/opensearch-agent-skills/ingest/986fda63/SKILL.md"}],"id":"opensearch-project/opensearch-agent-skills/ingest","links":{"html":"https://skillfed.io/opensearch-project/opensearch-agent-skills/ingest","md":"https://skillfed.io/opensearch-project/opensearch-agent-skills/ingest.md","repo":"https://github.com/opensearch-project/opensearch-agent-skills"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":30,"language":"Python","last_updated":"2026-07-22","license":"Apache-2.0","name":"ingest","publisher":"opensearch-project","stars":37},"relations":{"similar":[{"id":"opensearch-project/opensearch-agent-skills/document-processing"},{"id":"opensearch-project/opensearch-agent-skills/managed-ingestion-service"},{"id":"opensearch-project/opensearch-agent-skills/opensearch-skills"},{"id":"opensearch-project/opensearch-agent-skills/cloud"},{"id":"opensearch-project/opensearch-agent-skills/opensearch-launchpad"},{"id":"opensearch-project/opensearch-agent-skills/search"},{"id":"opensearch-project/opensearch-agent-skills/log-analytics"},{"id":"secondsky/sap-skills/sap-btp-cloud-logging"},{"id":"opensearch-project/opensearch-agent-skills/trace-analytics"},{"id":"opensearch-project/opensearch-agent-skills/observability"}]},"slug":{"owner":"opensearch-project","repo":"opensearch-agent-skills","skill":"ingest"},"version":"986fda63"}
