$npx skillfedfor your agent

html-to-json

Convert html to json.

Worth itPyPI HTMLReleased May 2026321.9K downloads / moMIT LicensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — html_to_json-3.0.0-py3-none-any.whl
v3.0.0 · released 2026-05-13 · Python >=3.10 · 1 runtime deps: beautifulsoup4

Yes. The package is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a concrete problem—converting HTML to JSON—with a straightforward API. It's suitable for production use in web scraping and data extraction workflows where HTML tables or documents need to become JSON.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later.
  • Low friction installation with a single runtime dependency (beautifulsoup4).
  • The package is actively maintained with a recent release and no known vulnerabilities.

License · maintenance · safety

MIT License (permissive) — MIT License permits commercial and private use with minimal restrictions; you may use, modify, and distribute the package freely as long as you include the license notice.

last release 2026-05-13 (93 days) · last repo commit 2026-08-11 · 54 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 321,894 downloads/mo, #7,617 on PyPI

Verify before relying

pip install html-to-json

import html_to_json

html_string = "<html><head><title>Test</title></head></html>"
output_json = html_to_json.convert(html_string)
print(output_json)
  • Performance characteristics when processing large HTML documents or tables with many rows.
  • Behavior with malformed or non-standard HTML input.
  • Memory usage patterns for deeply nested HTML structures.
Same gist for agents: .md · .json

What it is and what it does

html-to-json is a Python library that transforms HTML markup into JSON structures. It provides two main conversion functions: one for general HTML documents that preserves element hierarchy, attributes, and text content, and another specialized for HTML tables that intelligently extracts rows and columns into JSON objects keyed by header names. The library depends on beautifulsoup4 for HTML parsing and supports modern Python versions from 3.10 onward.

The package handles three common table patterns: headers in the first row, headers in the first column, or no headers at all. For both HTML and table conversion, you can control what gets captured—text values, attributes, nested tags as inner HTML, or nested tags as structured JSON—via keyword arguments. This makes it useful for web scraping workflows, data extraction pipelines, and converting HTML reports into machine-readable formats.

Use it for

  • Extract structured data from HTML tables on web pages and convert to JSON for further processing.
  • Parse HTML documents from web scraping operations and transform into JSON for storage or API responses.
  • Convert HTML-based reports or exports into JSON format for integration with downstream data pipelines.
  • Preserve HTML attributes and nested link structures when extracting table data by using record_html or record_children options.
  • Build data ingestion tools that accept HTML input and output standardized JSON for consumption by other services.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The package is actively maintained, has no known vulnerabilities, installs with minimal friction, and solves a concrete problem—converting HTML to JSON—with a straightforward API. It's suitable for production use in web scraping and data extraction workflows where HTML tables or documents need to become JSON.

Install

html-to-json on PyPI

Before you install

Low friction installation with a single runtime dependency (beautifulsoup4). The package is actively maintained with a recent release and no known vulnerabilities.

Requires Python 3.10 or later.

License in practice

MIT License permits commercial and private use with minimal restrictions; you may use, modify, and distribute the package freely as long as you include the license notice.

Quickstart

pip install html-to-json

import html_to_json

html_string = "<html><head><title>Test</title></head></html>"
output_json = html_to_json.convert(html_string)
print(output_json)

Verify before relying

  • Performance characteristics when processing large HTML documents or tables with many rows.
  • Behavior with malformed or non-standard HTML input.
  • Memory usage patterns for deeply nested HTML structures.

Package facts

LicenseMIT License permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
beautifulsoup4
MaintenanceActively maintained 93 days since the last release
Last repo commit
First released
Downloads321,894 / month, #7,617 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: Information TechnologyLicense :: OSI Approved :: MIT LicenseNatural Language :: EnglishProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14

Evidence: html_to_json-3.0.0-py3-none-any.whl

Tags

Capabilities
html to json conversionparse html as jsonhtml table to jsonextract html structure jsonhtml scraping json outputconvert html markup jsonhtml parsing library
Topics
html-parsingdata-extractionweb-scraping
PyPI keywords
conversionhtmlhtml to jsonjson

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “html to json conversion”

  • html-to-jsonConverts HTML documents and HTML tables into JSON structures,…
  • pylint-json2htmlConverts pylint JSON reports into formatted HTML documents,…
  • soup2dictConverts BeautifulSoup4 parsed HTML or XML objects into Python…

Give your agent the search over MCP, or paste the wish link into any chat.

More HTML packages

MarkupSafe Worth it
PyPI · Dynamic Content · released Sep 2025

MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.

BSD-3-Clausecompiled wheel · 3.9+aging
797.1Mdownloads / mo
Jinja2 Worth it
PyPI · Dynamic Content · released Mar 2025

Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.

BSD-3-Clausepure Python · 3.7+aging
718.6Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
Markdown Worth it
PyPI · Python Modules · released Jul 2026

Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.

Install it if you need to parse Markdown in Python.

BSD-3-Clausepure Python · 3.10+
121.7Mdownloads / mo

See also html-for-docx · html-table-parser-python3 · prettierfier · soup2dict · xlsx2html · xmltojson · Selenium-Screenshot · hyperscript · premailer · w3lib