tika
Apache Tika Python library
Decision gist · record as of 2026-08-14
Yes, if you need to extract text and metadata from diverse document formats in Python and can meet the Java 11+ requirement. The library is actively maintained, has low install friction, carries a permissive license, and no known vulnerabilities. It is well-suited for document processing pipelines, search indexing, and content analysis. Not suitable if you cannot run Java or need to work entirely offline without pre-staging a Tika server JAR.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Java 11+ must be installed on your system; tika-python starts the Tika REST server as a background process.
- In airgap environments, you must manually download tika-server.jar and set TIKA_SERVER_JAR environment variable.
- Low install friction with just two runtime dependencies (beautifulsoup4, requests).
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 is permissive; you can use this freely in commercial and open-source projects with minimal restrictions, provided you include a copy of the license.
last release 2026-08-01 (13 days) · last repo commit 2026-08-01 · 1,666 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 532,562 downloads/mo, #6,147 on PyPI
Alternatives
Verify before relying
pip install tika
from tika import parser
parsed = parser.from_file('/path/to/file')
print(parsed["metadata"])
print(parsed["content"])- Whether the package handles all document formats that Apache Tika supports (fact sheet does not enumerate supported file types).
- Performance characteristics when processing large files or high-volume document batches.
- Stability and resource usage of the background Tika server process across different operating systems.
What it is and what it does
Tika-python is a Python wrapper around Apache Tika that makes document parsing and metadata extraction available through a REST server interface. It handles text extraction, MIME type detection, language identification, and translation for a wide variety of document formats. The library manages a Tika REST server running in the background, so you interact with it as a Python library rather than managing a separate service.
The package provides multiple interfaces: a parser for extracting text and metadata, a detector for MIME type classification, a language detector, a translator, and a config interface to inspect available parsers and detectors. It supports both file paths and in-memory buffers, optional gzip compression, and can output content as plain text or XHTML. For disconnected environments, you can point to a local Tika server JAR file via environment variables.
Use it for
- Extract text and metadata from PDFs, Word documents, and other formats for indexing or archival systems.
- Automatically detect MIME types of uploaded files to validate content before processing.
- Identify the language of document content to route to appropriate downstream processing pipelines.
- Translate extracted text from one language to another as part of a document processing workflow.
- Inspect Tika server configuration to understand which parsers and detectors are available in your deployment.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to extract text and metadata from diverse document formats in Python and can meet the Java 11+ requirement.
The library is actively maintained, has low install friction, carries a permissive license, and no known vulnerabilities. It is well-suited for document processing pipelines, search indexing, and content analysis. Not suitable if you cannot run Java or need to work entirely offline without pre-staging a Tika server JAR.
Install
tika on PyPI
Before you install
Low install friction with just two runtime dependencies (beautifulsoup4, requests). Active maintenance with a recent release and steady commit history. Requires Java 11+ on the system to run the Tika REST server in the background.
Java 11+ must be installed on your system; tika-python starts the Tika REST server as a background process. In airgap environments, you must manually download tika-server.jar and set TIKA_SERVER_JAR environment variable.
License in practice
Apache-2.0 is permissive; you can use this freely in commercial and open-source projects with minimal restrictions, provided you include a copy of the license.
Quickstart
pip install tika
from tika import parser
parsed = parser.from_file('/path/to/file')
print(parsed["metadata"])
print(parsed["content"])
Verify before relying
- Whether the package handles all document formats that Apache Tika supports (fact sheet does not enumerate supported file types).
- Performance characteristics when processing large files or high-volume document batches.
- Stability and resource usage of the background Tika server process across different operating systems.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagesbeautifulsoup4requests |
| Maintenance | Actively maintained 13 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 532,562 / month, #6,147 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Database :: Front-EndsTopic :: Scientific/EngineeringTopic :: Software Development :: Libraries :: Python Modules |
Evidence: tika-3.3.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “apache tika python”
- tikaTika-python provides Python access to Apache Tika's document parsing,…
- apache-airflow-providers-apache-beamIntegrates Apache Beam data processing pipelines into Apache Airflow…
- kylinpyPython client library and SQLAlchemy dialect for querying Apache…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also pdftotext · python-magic · file-magic · kreuzberg · textract · magika · python-magic-bin · wikitextparser · comment-parser · libmagic