tree-sitter-html
HTML grammar for tree-sitter
Decision gist · record as of 2026-08-14
Yes, if you need to parse HTML programmatically with tree-sitter. The package is permissively licensed (MIT), has no known vulnerabilities, and offers prebuilt wheels for common platforms. Install friction is moderate but manageable. Repository is active though aging (641 days since last release). Verify tree-sitter availability in your environment before installing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.9 or later; tree-sitter core library must be available in the environment.
- Medium install friction due to compiled wheels; prebuilt binaries available for common platforms (x86_64, ARM, macOS, Linux, Windows).
- Last release was 641 days ago; repository is active but aging.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal attribution requirements—suitable for both open-source and commercial projects.
last release 2024-11-11 (641 days) · last repo commit 2025-11-24 · 213 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 2,212,176 downloads/mo, #3,206 on PyPI
Alternatives
Verify before relying
pip install tree-sitter-html
import tree_sitter_html- Whether tree-sitter Python bindings are automatically installed as a dependency or must be installed separately.
- Performance characteristics for parsing large or deeply nested HTML documents.
- Conformance level to HTML5 spec beyond the reference link provided.
- Exact usage patterns and API surface for integrating this grammar with tree-sitter.
What it is and what it does
tree-sitter-html is a language grammar for the tree-sitter incremental parser, specifically targeting HTML. It translates HTML source code into a concrete syntax tree that can be traversed and analyzed programmatically. The package is a compiled extension (available as prebuilt wheels for multiple platforms) that integrates with the tree-sitter ecosystem to enable fast, incremental parsing of HTML documents.
Developers use this package when they need to parse HTML programmatically—extracting structure, analyzing document trees, or building tools that reason about HTML syntax. It's particularly useful for code editors, linters, static analysis tools, and any application that needs to understand HTML structure. The grammar is based on the HTML5 specification.
Use it for
- Extract and analyze HTML document structure for web scraping or content extraction tasks.
- Build linting or validation tools that check HTML syntax and structure.
- Implement syntax highlighting or code navigation in HTML editors.
- Generate abstract syntax trees for HTML transformation or code generation pipelines.
- Analyze HTML templates or markup in static site generators or documentation tools.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to parse HTML programmatically with tree-sitter.
The package is permissively licensed (MIT), has no known vulnerabilities, and offers prebuilt wheels for common platforms. Install friction is moderate but manageable. Repository is active though aging (641 days since last release). Verify tree-sitter availability in your environment before installing.
Install
tree-sitter-html on PyPI
Before you install
Medium install friction due to compiled wheels; prebuilt binaries available for common platforms (x86_64, ARM, macOS, Linux, Windows). Last release was 641 days ago; repository is active but aging.
Requires Python 3.9 or later; tree-sitter core library must be available in the environment.
License in practice
MIT license permits unrestricted use, modification, and distribution with minimal attribution requirements—suitable for both open-source and commercial projects.
Quickstart
pip install tree-sitter-html
import tree_sitter_html
Verify before relying
- Whether tree-sitter Python bindings are automatically installed as a dependency or must be installed separately.
- Performance characteristics for parsing large or deeply nested HTML documents.
- Conformance level to HTML5 spec beyond the reference link provided.
- Exact usage patterns and API surface for integrating this grammar with tree-sitter.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | None |
| Maintenance | Aging 641 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 2,212,176 / month, #3,206 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseTopic :: Software Development :: CompilersTopic :: Text Processing :: LinguisticTyping :: Typed |
Evidence: tree_sitter_html-0.23.2-cp39-abi3-macosx_10_9_x86_64.whl; tree_sitter_html-0.23.2-cp39-abi3-macosx_11_0_arm64.whl; tree_sitter_html-0.23.2-cp39-abi3-manylinux_2_17_aarch64.manylinux2014_aarch64.whl; tree_sitter_html-0.23.2-cp39-abi3-manylinux_2_5_x86_64.manylinux1_x86_64.manylinux_2_17_x86_64.manylinux2014_x86_64.whl; tree_sitter_html-0.23.2-cp39-abi3-musllinux_1_2_x86_64.whl; tree_sitter_html-0.23.2-cp39-abi3-win_amd64.whl; tree_sitter_html-0.23.2-cp39-abi3-win_arm64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “html parsing tree-sitter”
- tree-sitter-htmlProvides an HTML grammar for tree-sitter, enabling incremental…
- tree-sitter-regexProvides a tree-sitter grammar for parsing regular expressions in…
- tree-sitter-xmlProvides a tree-sitter parser grammar for XML and DTD files, enabling…
Give your agent the search over MCP, or paste the wish link into any chat.
More Linguistic packages
Detects and normalizes text encoding from unknown or ambiguous sources, supporting all IANA character sets that Python's core library provides codecs for, with the ability to register custom codecs.
tiktoken is a fast BPE tokenizer that converts text into token sequences compatible with OpenAI models, supporting multiple encoding schemes including o200k_base and model-specific encodings.
Install it if you work with OpenAI APIs or need to understand token boundaries in GPT-family models.
Detects character encoding and language in byte sequences with high accuracy, supporting 99 encodings and returning confidence scores, language tags, and MIME types.
Install it if you need to detect character encoding or language in byte data; the rewrite makes it substantially faster and more accurate than its predecessors.
Converts Unicode text to ASCII by transliterating non-ASCII characters into their closest ASCII equivalents, with no runtime dependencies.
However, if transliteration quality or ongoing maintenance matters, consider unidecode instead despite its GPL-only license.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
Python bindings to the tree-sitter parsing library, enabling incremental parsing and syntax tree analysis for source code.
See also tree-sitter-java · tree-sitter-xml · tree-sitter-toml · tree-sitter-embedded-template · tree-sitter-python · tree-sitter-css · tree-sitter-json · tree-sitter-php · tree-sitter-regex · tree-sitter-scala