pagefind
Python API for Pagefind
Decision gist · record as of 2026-08-14
Yes. Low install friction, no runtime dependencies, active maintenance, MIT license, and zero known vulnerabilities make it a safe choice. Install it if you need to index HTML or custom content for search—particularly valuable for static site generators and documentation platforms where you want search without infrastructure overhead.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.9 or later; the [bin] extra installs the underlying binary.
- Low friction install with no runtime dependencies.
- Active maintenance status with recent releases.
License · maintenance · safety
MIT (permissive) — MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
last release 2026-04-12 (124 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,707,561 downloads/mo, #2,930 on PyPI
Alternatives
Verify before relying
pip install 'pagefind[bin]'
import asyncio
from pagefind.index import PagefindIndex, IndexConfig
async def main():
config = IndexConfig(output_path="./output")
async with PagefindIndex(config=config) as index:
result = await index.add_html_file(
content="<html><body>Example</body></html>",
url="https://example.com"
)
asyncio.run(main())- Whether the binary is automatically downloaded or requires separate installation when using the [bin] extra.
- Performance characteristics when indexing large HTML document collections.
- Search query syntax and filtering capabilities available through the API.
- Whether add_directory, add_html_file, and add_custom_record methods support concurrent execution beyond the example shown.
What it is and what it does
Pagefind is an async Python wrapper that lets you build searchable indexes from HTML files and custom content records. It exposes a clean async API for adding HTML files, custom records, or entire directories, and retrieving indexed file metadata. The package handles communication with an underlying binary process, letting you control output paths, logging, and which DOM elements to index through the IndexConfig object.
You use it when you need full-text search on static content—typically for documentation sites, blogs, or knowledge bases where you want to avoid running a separate search service. The async design means you can index multiple sources concurrently, making it suitable for batch indexing pipelines and build-time content processing.
Use it for
- Index a static documentation site and expose search to readers without a backend search service.
- Build a searchable archive of HTML pages or exported content in a batch indexing pipeline.
- Add full-text search to a static site by pre-indexing content during the build step.
- Index custom metadata records alongside HTML files to create a unified search index.
- Programmatically index and re-index content as part of a content management workflow.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Low install friction, no runtime dependencies, active maintenance, MIT license, and zero known vulnerabilities make it a safe choice. Install it if you need to index HTML or custom content for search—particularly valuable for static site generators and documentation platforms where you want search without infrastructure overhead.
Install
pagefind on PyPI
Before you install
Low friction install with no runtime dependencies. Active maintenance status with recent releases.
Requires Python 3.9 or later; the [bin] extra installs the underlying binary.
License in practice
MIT license permits free use, modification, and distribution with minimal restrictions—suitable for both open-source and commercial projects.
Quickstart
pip install 'pagefind[bin]'
import asyncio
from pagefind.index import PagefindIndex, IndexConfig
async def main():
config = IndexConfig(output_path="./output")
async with PagefindIndex(config=config) as index:
result = await index.add_html_file(
content="<html><body>Example</body></html>",
url="https://example.com"
)
asyncio.run(main())
Verify before relying
- Whether the binary is automatically downloaded or requires separate installation when using the [bin] extra.
- Performance characteristics when indexing large HTML document collections.
- Search query syntax and filtering capabilities available through the API.
- Whether add_directory, add_html_file, and add_custom_record methods support concurrent execution beyond the example shown.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 124 days since the last release |
| First released | |
| Downloads | 2,707,561 / month, #2,930 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: MIT LicenseTopic :: Text Processing :: IndexingTopic :: Text Processing :: Markup :: HTML |
Evidence: pagefind-1.5.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “static site search indexing”
- pagefindProvides an async Python API to index and search HTML content,…
- pagefind-binPagefind-bin provides a precompiled binary distribution of Pagefind,…
- lunrLunr is a lightweight, in-memory full-text search library that builds…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also pagefind-bin · typesense · algoliasearch · Whoosh-Reloaded · meilisearch-python-sdk · Whoosh · turbopuffer · lunr · readable-content · meilisearch