$npx skillfedfor your agent

extruct

Extract embedded metadata from HTML markup

Worth itPyPI HTMLReleased Nov 2024659.4K downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — extruct-0.18.0-py2.py3-none-any.whl
v0.18.0 · released 2024-11-08 · Python >=3.8 · 8 runtime deps: lxml, lxml-html-clean, rdflib, pyrdfa3, mf2py, w3lib, html-text, jstyleson

Yes. extruct is actively maintained, has no known vulnerabilities, supports current Python versions (3.8–3.12), and installs with low friction. It solves a real problem—unified extraction of multiple metadata formats—that would otherwise require juggling separate parsers. The permissive license and 971 repository stars indicate solid community adoption. Install it if you need to reliably extract structured metadata from HTML in production or research contexts.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Low install friction with a pure-Python wheel.
  • Active maintenance with recent commits and 971 repository stars.
  • Supports Python 3.8 through 3.12.

License · maintenance · safety

permissive license (permissive) — Permissive license allows commercial and private use without restriction.

last release 2024-11-08 (644 days) · last repo commit 2026-04-01 · 971 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 659,409 downloads/mo, #5,463 on PyPI

Verify before relying

pip install extruct

import extruct
from w3lib.html import get_base_url

html = '<html>...</html>'
base_url = 'https://example.com'
data = extruct.extract(html, base_url=base_url)
  • Performance characteristics when processing large HTML documents or batch operations
  • Specific version compatibility matrix for each metadata format parser (mf2py, rdflib, pyrdfa3)
  • Whether experimental RDFa support is production-ready
Same gist for agents: .md · .json

What it is and what it does

extruct is a metadata extraction library that parses HTML documents and pulls out structured data embedded in standard formats. It handles W3C Microdata (including schema.org), JSON-LD, Microformats, Open Graph tags, RDFa, and Dublin Core metadata—all through a single unified interface. You pass it an HTML string and optional base URL, and it returns a dictionary keyed by format type, each containing the extracted data.

The library wraps several specialized parsers (lxml, rdflib, mf2py, pyrdfa3) and HTML utilities (w3lib, html-text, jstyleson) to normalize their output. It's commonly used in web scraping, SEO analysis, and data integration pipelines where you need to reliably extract metadata that content publishers have embedded for search engines, social media, or semantic web applications.

Use it for

  • Extract schema.org Microdata from e-commerce product pages to build product catalogs
  • Parse Open Graph tags from web pages for social media link previews and metadata
  • Harvest JSON-LD structured data from news articles and blog posts for content indexing
  • Extract Microformats from personal websites and social profiles for contact/event data
  • Build semantic web applications by parsing RDFa annotations from HTML documents
  • Aggregate Dublin Core metadata from institutional or library websites

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

extruct is actively maintained, has no known vulnerabilities, supports current Python versions (3.8–3.12), and installs with low friction. It solves a real problem—unified extraction of multiple metadata formats—that would otherwise require juggling separate parsers. The permissive license and 971 repository stars indicate solid community adoption. Install it if you need to reliably extract structured metadata from HTML in production or research contexts.

Install

extruct on PyPI

Before you install

Low install friction with a pure-Python wheel. Active maintenance with recent commits and 971 repository stars. Supports Python 3.8 through 3.12.

License in practice

Permissive license allows commercial and private use without restriction.

Quickstart

pip install extruct

import extruct
from w3lib.html import get_base_url

html = '<html>...</html>'
base_url = 'https://example.com'
data = extruct.extract(html, base_url=base_url)

Verify before relying

  • Performance characteristics when processing large HTML documents or batch operations
  • Specific version compatibility matrix for each metadata format parser (mf2py, rdflib, pyrdfa3)
  • Whether experimental RDFa support is production-ready

Package facts

Licensepermissive license permissive
Python supportSupports the current Python release >=3.8
Install frictionLow. Pure-Python wheel
Runtime dependencies
8 packages
lxmllxml-html-cleanrdflibpyrdfa3mf2pyw3libhtml-textjstyleson
MaintenanceActively maintained 644 days since the last release
Last repo commit
First released
Downloads659,409 / month, #5,463 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: BSD LicenseNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: extruct-0.18.0-py2.py3-none-any.whl

Tags

Capabilities
extract metadata from htmlparse microdata schema.orgjson-ld extractionopen graph parserhtml structured data extractionrdfa parsermicroformats parserdublin core metadata
Topics
metadata-extractionweb-scrapingstructured-data
PyPI keywords
extruct

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract metadata from html”

  • extructExtracts structured metadata from HTML markup in multiple formats:…
  • micawberExtracts rich metadata (title, author, thumbnail, embed HTML) from…
  • pyrdfa3Parses RDFa markup embedded in HTML documents and extracts structured…

Give your agent the search over MCP, or paste the wish link into any chat.

More HTML packages

MarkupSafe Worth it
PyPI · Dynamic Content · released Sep 2025

MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.

BSD-3-Clausecompiled wheel · 3.9+aging
797.1Mdownloads / mo
Jinja2 Worth it
PyPI · Dynamic Content · released Mar 2025

Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.

BSD-3-Clausepure Python · 3.7+aging
718.6Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
Markdown Worth it
PyPI · Python Modules · released Jul 2026

Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.

Install it if you need to parse Markdown in Python.

BSD-3-Clausepure Python · 3.10+
121.7Mdownloads / mo

See also linkpreview · mf2py · PyLD · pyrdfa3 · recipe-scrapers · rdflib-jsonld · prov · rdflib · sphinxext-opengraph · Sickle