$npx skillfedfor your agent

tika

Apache Tika Python library

With conditionsPyPI Scientific/EngineeringReleased Aug 2026532.6K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — tika-3.3.2-py3-none-any.whl
v3.3.2 · released 2026-08-01 · Python >=3.10 · 2 runtime deps: beautifulsoup4, requests

Yes, if you need to extract text and metadata from diverse document formats in Python and can meet the Java 11+ requirement. The library is actively maintained, has low install friction, carries a permissive license, and no known vulnerabilities. It is well-suited for document processing pipelines, search indexing, and content analysis. Not suitable if you cannot run Java or need to work entirely offline without pre-staging a Tika server JAR.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Java 11+ must be installed on your system; tika-python starts the Tika REST server as a background process.
  • In airgap environments, you must manually download tika-server.jar and set TIKA_SERVER_JAR environment variable.
  • Low install friction with just two runtime dependencies (beautifulsoup4, requests).

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 is permissive; you can use this freely in commercial and open-source projects with minimal restrictions, provided you include a copy of the license.

last release 2026-08-01 (13 days) · last repo commit 2026-08-01 · 1,666 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 532,562 downloads/mo, #6,147 on PyPI

Verify before relying

pip install tika

from tika import parser
parsed = parser.from_file('/path/to/file')
print(parsed["metadata"])
print(parsed["content"])
  • Whether the package handles all document formats that Apache Tika supports (fact sheet does not enumerate supported file types).
  • Performance characteristics when processing large files or high-volume document batches.
  • Stability and resource usage of the background Tika server process across different operating systems.
Same gist for agents: .md · .json

What it is and what it does

Tika-python is a Python wrapper around Apache Tika that makes document parsing and metadata extraction available through a REST server interface. It handles text extraction, MIME type detection, language identification, and translation for a wide variety of document formats. The library manages a Tika REST server running in the background, so you interact with it as a Python library rather than managing a separate service.

The package provides multiple interfaces: a parser for extracting text and metadata, a detector for MIME type classification, a language detector, a translator, and a config interface to inspect available parsers and detectors. It supports both file paths and in-memory buffers, optional gzip compression, and can output content as plain text or XHTML. For disconnected environments, you can point to a local Tika server JAR file via environment variables.

Use it for

  • Extract text and metadata from PDFs, Word documents, and other formats for indexing or archival systems.
  • Automatically detect MIME types of uploaded files to validate content before processing.
  • Identify the language of document content to route to appropriate downstream processing pipelines.
  • Translate extracted text from one language to another as part of a document processing workflow.
  • Inspect Tika server configuration to understand which parsers and detectors are available in your deployment.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need to extract text and metadata from diverse document formats in Python and can meet the Java 11+ requirement.

The library is actively maintained, has low install friction, carries a permissive license, and no known vulnerabilities. It is well-suited for document processing pipelines, search indexing, and content analysis. Not suitable if you cannot run Java or need to work entirely offline without pre-staging a Tika server JAR.

Install

tika on PyPI

Before you install

Low install friction with just two runtime dependencies (beautifulsoup4, requests). Active maintenance with a recent release and steady commit history. Requires Java 11+ on the system to run the Tika REST server in the background.

Java 11+ must be installed on your system; tika-python starts the Tika REST server as a background process. In airgap environments, you must manually download tika-server.jar and set TIKA_SERVER_JAR environment variable.

License in practice

Apache-2.0 is permissive; you can use this freely in commercial and open-source projects with minimal restrictions, provided you include a copy of the license.

Quickstart

pip install tika

from tika import parser
parsed = parser.from_file('/path/to/file')
print(parsed["metadata"])
print(parsed["content"])

Verify before relying

  • Whether the package handles all document formats that Apache Tika supports (fact sheet does not enumerate supported file types).
  • Performance characteristics when processing large files or high-volume document batches.
  • Stability and resource usage of the background Tika server process across different operating systems.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
2 packages
beautifulsoup4requests
MaintenanceActively maintained 13 days since the last release
Last repo commit
First released
Downloads532,562 / month, #6,147 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Database :: Front-EndsTopic :: Scientific/EngineeringTopic :: Software Development :: Libraries :: Python Modules

Evidence: tika-3.3.2-py3-none-any.whl

Tags

Capabilities
document text extraction pythonpdf metadata extractionmime type detectionapache tika pythonlanguage detection documentsdocument parser librarycontent extraction from files
Topics
document-parsingmetadata-extractionlanguage-detection
PyPI keywords
tikadigitalbabel fishapache

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “apache tika python”

  • tikaTika-python provides Python access to Apache Tika's document parsing,…
  • apache-airflow-providers-apache-beamIntegrates Apache Beam data processing pipelines into Apache Airflow…
  • kylinpyPython client library and SQLAlchemy dialect for querying Apache…

Give your agent the search over MCP, or paste the wish link into any chat.

More Scientific/Engineering packages

numpy Worth it
PyPI · Software Development · released Aug 2026

NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.

BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0compiled wheel · 3.12+
1.1Bdownloads / mo
pandas Worth it
PyPI · Scientific/Engineering · released Jul 2026

pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.

BSD-3-Clausecompiled wheel · 3.11+
769.1Mdownloads / mo
scipy Worth it
PyPI · Libraries · released Jun 2026

scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.

BSD-3-Clausecompiled wheel · 3.12+
449.0Mdownloads / mo
scikit-learn Worth it
PyPI · Software Development · released Jun 2026

scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.

Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.

BSD-3-Clausecompiled wheel · 3.11+
235.5Mdownloads / mo
dill Worth it
PyPI · Software Development · released Jan 2026

dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.

BSD-3-Clausepure Python · 3.9+
208.1Mdownloads / mo
multiprocess Worth it
PyPI · Software Development · released Jan 2026

Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.

Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.

BSD-3-Clausepure Python · 3.9+
202.7Mdownloads / mo

See also pdftotext · python-magic · file-magic · kreuzberg · textract · magika · python-magic-bin · wikitextparser · comment-parser · libmagic