{"categories":[{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence"}],"enrichment":{"capability":"Docling parses diverse document formats\u2014PDF, DOCX, PPTX, XLSX, HTML, EPUB, email, images, video, audio, and more\u2014into a unified representation, with advanced PDF layout understanding and export to Markdown, HTML, JSON, and other formats.","skillfed_tags":["document-processing","pdf-parsing","ai-integration"],"use_cases":["Convert research papers or technical PDFs to structured Markdown for ingestion into RAG or LLM pipelines.","Extract tables, charts, and text from financial reports (XBRL) or patent documents (USPTO) for automated analysis.","Process scanned or image-based documents using OCR and Visual Language Models to recover structured content.","Build document-aware chatbots or Q&A systems by parsing diverse input formats into a unified representation.","Batch-convert email archives (EML, MSG) and office documents (DOCX, XLSX, PPTX) to Markdown or JSON for downstream workflows."],"what_it_does":"Docling is a document processing SDK that converts PDFs, Word documents, spreadsheets, presentations, email, images, video, audio, and other formats into a unified, structured representation. It specializes in advanced PDF understanding\u2014extracting page layout, reading order, table structure, code blocks, formulas, and charts\u2014and can export to Markdown, HTML, JSON, and domain-specific schemas (DocLang, USPTO patents, JATS articles, XBRL financial reports). The library runs locally, supports OCR for scanned documents, integrates with Visual Language Models and ASR systems, and plugs into AI frameworks like LangChain, LlamaIndex, and Haystack.\n\nIt is designed for developers building document-aware AI applications, knowledge extraction pipelines, and content processing workflows. The package offers both a Python API and a command-line interface, with options to run as a service via an API server or as an MCP server for agent integration.","worth_installing":"Yes. Docling is actively maintained, production-stable, permissively licensed (MIT), and solves a real problem\u2014unified parsing of many document formats with strong PDF understanding. Low install friction, no known vulnerabilities, and broad Python version support make it a solid choice for document-heavy AI and data extraction projects. Install if you need to parse or convert diverse document types at scale."},"id":"docling","links":{"html":"https://skillfed.io/packages/docling","md":"https://skillfed.io/packages/docling.md","pypi":"https://pypi.org/project/docling/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-14","license_spdx":"MIT","license_treatment":"permissive","name":"docling","python_support":"supports_current","summary":"SDK and CLI for parsing PDF, DOCX, HTML, and more, to a unified document representation for powering downstream workflows such as gen AI applications."},"popularity":{"monthly_downloads":17994695,"position":1092,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"2.120.1"}
