turbohtml
A fast, fully typed HTML toolkit for Python, powered by a C-accelerated core.
Decision gist · record as of 2026-08-14
Yes. turbohtml is actively maintained, carries no security vulnerabilities, uses a permissive MIT license, and installs without compilation on modern Python versions (3.10–3.15). It is well-suited for performance-critical HTML processing, web scraping, and content extraction. The main caveat is that it is not API-compatible with BeautifulSoup or lxml, so adoption requires rewriting existing code—but migration guides are provided for 65 libraries. Install it if you need fast, typed HTML handling and are willing to learn its API.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Wheels are pre-built for CPython 3.10–3.15 across major platforms (macOS, Linux, Windows, including free-threading variants), so installation requires no compilation.
- Active maintenance with a release 3 days old.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute turbohtml freely provided you include the license notice.
last release 2026-08-11 (3 days) · last repo commit 2026-08-14 · 20 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 99,408 downloads/mo, #13,024 on PyPI
Alternatives
Verify before relying
pip install turbohtml
import turbohtml
from turbohtml import Html, Formatter
doc = turbohtml.parse("<p>café & cake</p>")
print(doc.select_one("p").text) # café & cake
print(doc.select_one("p").serialize(Html(formatter=Formatter.NAMED_ENTITIES)))- Whether the C-accelerated core and free-threading support deliver the claimed 2–5× parsing speedup and 9–15× tokenization speedup in typical production workloads.
- Whether the WHATWG DOM model (indexing children via node[i], accessing attributes via node.attrs) is sufficiently intuitive for developers migrating from BeautifulSoup or lxml.
- Whether the 65 migration shims cover the specific libraries and patterns your codebase relies on.
What it is and what it does
turbohtml is a high-performance HTML and XML processing library that combines a C-accelerated parser and query engine with a fully-typed Python interface. It handles tokenization, parsing, querying via CSS selectors and XPath, serialization to HTML or markdown, form extraction, sanitization, minification, and rewriting—all without native dependencies beyond the pre-built wheels. The library models the DOM according to the WHATWG standard, exposing child nodes as indexable items and attributes through a dedicated `.attrs` interface.
The package is designed for web scraping, content extraction, and HTML transformation workflows. Common use cases include extracting article text and converting it to markdown for language models, sanitizing untrusted HTML, minifying markup, detecting character encodings, and building or editing HTML programmatically. It is not a drop-in replacement for BeautifulSoup or lxml; the fact sheet provides migration guides for 65 libraries, but you will need to adapt your code to turbohtml's API.
Use it for
- Extract main article content from a web page and convert it to markdown for processing by a language model.
- Sanitize user-submitted HTML to remove scripts and dangerous attributes while preserving safe markup.
- Parse and minify HTML, CSS, and JavaScript in a single streaming pass without building an intermediate tree.
- Query HTML documents using CSS selectors or XPath and extract structured data (tables, JSON-LD, microdata).
- Detect character encoding of HTML documents and parse them with automatic encoding sniffing.
- Build or edit HTML programmatically using Element constructors and live attribute manipulation.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
turbohtml is actively maintained, carries no security vulnerabilities, uses a permissive MIT license, and installs without compilation on modern Python versions (3.10–3.15). It is well-suited for performance-critical HTML processing, web scraping, and content extraction. The main caveat is that it is not API-compatible with BeautifulSoup or lxml, so adoption requires rewriting existing code—but migration guides are provided for 65 libraries. Install it if you need fast, typed HTML handling and are willing to learn its API.
Install
turbohtml on PyPI
Before you install
Wheels are pre-built for CPython 3.10–3.15 across major platforms (macOS, Linux, Windows, including free-threading variants), so installation requires no compilation. Active maintenance with a release 3 days old.
Requires Python 3.10 or later.
License in practice
MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute turbohtml freely provided you include the license notice.
Quickstart
pip install turbohtml
import turbohtml
from turbohtml import Html, Formatter
doc = turbohtml.parse("<p>café & cake</p>")
print(doc.select_one("p").text) # café & cake
print(doc.select_one("p").serialize(Html(formatter=Formatter.NAMED_ENTITIES)))
Verify before relying
- Whether the C-accelerated core and free-threading support deliver the claimed 2–5× parsing speedup and 9–15× tokenization speedup in typical production workloads.
- Whether the WHATWG DOM model (indexing children via node[i], accessing attributes via node.attrs) is sufficiently intuitive for developers migrating from BeautifulSoup or lxml.
- Whether the 65 migration shims cover the specific libraries and patterns your codebase relies on.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Actively maintained 3 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 99,408 / month, #13,024 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.15Programming Language :: Python :: Free Threading :: 1 - UnstableProgramming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyTopic :: InternetTopic :: Software Development :: LibrariesTopic :: Text Processing :: Markup :: HTMLTopic :: Text Processing :: Markup :: MarkdownTopic :: Text Processing :: Markup :: XMLTyping :: Typed |
Evidence: turbohtml-1.6.0-cp310-cp310-macosx_11_0_arm64.whl; turbohtml-1.6.0-cp310-cp310-manylinux_2_28_aarch64.whl; turbohtml-1.6.0-cp310-cp310-manylinux_2_28_x86_64.whl; turbohtml-1.6.0-cp310-cp310-musllinux_1_2_aarch64.whl; turbohtml-1.6.0-cp310-cp310-musllinux_1_2_x86_64.whl; turbohtml-1.6.0-cp310-cp310-win_amd64.whl; turbohtml-1.6.0-cp311-cp311-macosx_11_0_arm64.whl; turbohtml-1.6.0-cp311-cp311-manylinux_2_28_aarch64.whl; turbohtml-1.6.0-cp311-cp311-manylinux_2_28_x86_64.whl; turbohtml-1.6.0-cp311-cp311-musllinux_1_2_aarch64.whl; turbohtml-1.6.0-cp311-cp311-musllinux_1_2_x86_64.whl; turbohtml-1.6.0-cp311-cp311-win_amd64.whl; turbohtml-1.6.0-cp312-cp312-macosx_11_0_arm64.whl; turbohtml-1.6.0-cp312-cp312-manylinux_2_28_aarch64.whl; turbohtml-1.6.0-cp312-cp312-manylinux_2_28_x86_64.whl; turbohtml-1.6.0-cp312-cp312-musllinux_1_2_aarch64.whl; turbohtml-1.6.0-cp312-cp312-musllinux_1_2_x86_64.whl; turbohtml-1.6.0-cp312-cp312-win_amd64.whl; turbohtml-1.6.0-cp313-cp313-macosx_11_0_arm64.whl; turbohtml-1.6.0-cp313-cp313-manylinux_2_28_aarch64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “html parsing and querying”
- turbohtmlParse, query, edit, and serialize HTML and XML documents with a…
- pyquerypyquery lets you query and manipulate XML and HTML documents using a…
- cssselectcssselect parses CSS3 selectors and translates them to XPath 1.0…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also cssselect · w3lib · parsel · pyquery · selectolax · tinyhtml5 · beautifulsoup4 · saxonche · BeautifulSoup · html5lib