{"enrichment":{"faq":[{"a":"ocr_kb uses multimodal AI models to process PDF documents page-by-page, extracting text, LaTeX formulas, and independent research figures. The skill applies global numbering to equations, figures, and tables, then generates incremental DOCX output with quality checkpoints every two pages for reliable document digitization.","q":"How does ocr_kb extract text, formulas, and figures from PDF pages?"},{"a":"Yes. ocr_kb converts scanned papers and image-based PDFs to formatted DOCX output by extracting all content through multimodal processing. It applies IEEE/APA formatting and typesetting automatically, preserving formulas, figures, and tables with proper global numbering in the final Word document.","q":"Can ocr_kb convert scanned or image-based PDFs to formatted DOCX documents?"},{"a":"ocr_kb maintains checkpoint files to enable recovery from interruptions. If processing stops mid-document, you can resume from the last checkpoint without reprocessing completed pages. The skill supports environment cleanup and partial reruns, allowing efficient continuation of long document workflows.","q":"What happens if document processing is interrupted\u2014can ocr_kb resume?"},{"a":"Yes. ocr_kb extracts LaTeX equations directly from paper images and PDF pages using its multimodal model. Extracted formulas are preserved in the output DOCX with proper formatting and integrated into the global equation numbering system for scientific document workflows.","q":"Does ocr_kb extract LaTeX equations from paper images?"},{"a":"ocr_kb identifies and extracts independent research figures and tables from each page, assigning them global numbering across the entire document. All extracted figures, tables, and equations are organized with IEEE/APA formatting in the incremental DOCX output, maintaining cross-reference integrity.","q":"How does ocr_kb handle figures and tables in multi-page documents?"},{"a":"ocr_kb generates quality checkpoints every two pages during batch processing to verify extraction accuracy. Combined with checkpoint recovery, this ensures reliable page-by-page digitization of long documents. Progress is tracked through checkpoint files, enabling safe partial reruns and environment cleanup without data loss.","q":"What quality assurance does ocr_kb provide during batch processing?"}],"shadow_tags":["document-digitization","formula-extraction","batch-processing","checkpoint-recovery","figure-cropping","format-conversion","quality-verification","latex-support","workflow-automation","incremental-generation"],"summary_rewrite":"ocr_kb processes PDF documents page-by-page using multimodal models to extract text, LaTeX formulas, and independent research figures with global numbering. It generates incremental DOCX output with quality checkpoints every two pages, supports environment cleanup and partial reruns, and maintains progress through checkpoint files for reliable recovery."},"files":[{"bytes":15514,"path":"ocr_kb/SKILL.md","sha256":"017f5e850e776d777af2fa379fa6572f640c533f4dfb2bfa70103e868164d7b9","url":"https://skillfed.io/files/TFboy1/academic-paper-writer/ocr_kb/c240be59/SKILL.md"}],"id":"TFboy1/academic-paper-writer/ocr_kb","links":{"html":"https://skillfed.io/TFboy1/academic-paper-writer/ocr_kb","md":"https://skillfed.io/TFboy1/academic-paper-writer/ocr_kb.md","repo":"https://github.com/TFboy1/academic-paper-writer"},"meta":{"agents_supported":[],"first_seen":"2026-07-28","forks":14,"language":"Python","last_updated":"2026-06-19","license":"MIT","name":"ocr_kb","publisher":"TFboy1","stars":247},"relations":{"similar":[{"id":"TFboy1/academic-paper-writer/academic-paper-writer"},{"id":"TFboy1/academic-paper-writer/docx"},{"id":"zhu1090093659/deepseek-pp/officecli-academic-paper"},{"id":"alextangson/feishu_skills/feishu-doc-writer"},{"id":"iOfficeAI/OfficeCLI/officecli-academic-paper"},{"id":"TFboy1/academic-paper-writer/content_generation"},{"id":"thvroyal/kimi-skills/kimi-docx"},{"id":"cat-xierluo/legal-skills/pdf-organizer"},{"id":"zhu1090093659/deepseek-pp/officecli-docx"},{"id":"iOfficeAI/OfficeCLI/officecli-docx"}]},"slug":{"owner":"TFboy1","repo":"academic-paper-writer","skill":"ocr_kb"},"version":"c240be59"}
