mf2py
Microformats parser
Decision gist · record as of 2026-08-14
Yes, if you are working with microformats-marked HTML or building IndieWeb tooling. The package is stable, has no known vulnerabilities, and low install friction. However, dormant maintenance (980 days since last release) means you should not expect bug fixes or spec updates; use it only for stable, well-established microformats2 parsing tasks.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.8 or later.
- The parser uses html5lib internally, so BeautifulSoup documents passed to it may be modified in place.
- Low install friction with a pure-Python wheel and three common dependencies (html5lib, requests, beautifulsoup4).
License · maintenance · safety
MIT (permissive) — MIT license is permissive, allowing commercial and private use with minimal restrictions. You may use, modify, and distribute mf2py freely as long as you retain the license notice.
last release 2023-12-08 (980 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 562,651 downloads/mo, #5,986 on PyPI
Alternatives
Verify before relying
import mf2py
mf2json = mf2py.parse(doc="<div class='h-card'><p class='p-name'>Jane</p></div>")
print(mf2json['items'])- Whether dormant status (980 days since last release) affects real-world compatibility with current microformats2 specs or HTML5 parsing edge cases.
- Performance characteristics when parsing large or deeply nested HTML documents.
What it is and what it does
mf2py is a microformats parser that reads HTML and extracts structured data in JSON format. It recognizes microformats2 classes (like h-card, h-entry, h-event) and converts them into a standardized dictionary structure, making it easy to programmatically access semantic markup embedded in web pages. The library also supports the older microformats1 format and has experimental metaformats extraction.
You typically use mf2py when you need to consume or validate microformats-marked HTML—common in IndieWeb applications, social media integrations, and semantic web tooling. It accepts HTML from files, strings, or URLs, and returns a dict with items, relationships, and metadata. The three runtime dependencies (html5lib, requests, beautifulsoup4) handle the heavy lifting of HTML parsing and network fetching.
Use it for
- Extract author and publication metadata from blog posts marked with h-entry and h-card microformats.
- Parse event listings (h-event) from calendar or conference websites to build aggregators or search indexes.
- Validate that your own HTML markup correctly implements microformats2 before publishing.
- Build IndieWeb applications that need to read contact cards (h-card) or social interactions (webmentions) from remote URLs.
- Convert microformats-annotated HTML into JSON for downstream processing or storage in a database.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are working with microformats-marked HTML or building IndieWeb tooling.
The package is stable, has no known vulnerabilities, and low install friction. However, dormant maintenance (980 days since last release) means you should not expect bug fixes or spec updates; use it only for stable, well-established microformats2 parsing tasks.
Install
mf2py on PyPI
Before you install
Low install friction with a pure-Python wheel and three common dependencies (html5lib, requests, beautifulsoup4). Maintenance is dormant—last release was 980 days ago—so expect no active bug fixes or feature updates, though the package has been stable since its 2014 inception.
Requires Python 3.8 or later. The parser uses html5lib internally, so BeautifulSoup documents passed to it may be modified in place.
License in practice
MIT license is permissive, allowing commercial and private use with minimal restrictions. You may use, modify, and distribute mf2py freely as long as you retain the license notice.
Quickstart
import mf2py
mf2json = mf2py.parse(doc="<div class='h-card'><p class='p-name'>Jane</p></div>")
print(mf2json['items'])
Verify before relying
- Whether dormant status (980 days since last release) affects real-world compatibility with current microformats2 specs or HTML5 parsing edge cases.
- Performance characteristics when parsing large or deeply nested HTML documents.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packageshtml5librequestsbeautifulsoup4 |
| Maintenance | Dormant 980 days since the last release |
| First released | |
| Downloads | 562,651 / month, #5,986 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Intended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Topic :: Text Processing :: Markup :: HTML |
Evidence: mf2py-2.0.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “microformats parser”
- mf2pymf2py parses HTML documents to extract microformats2 data structures,…
- extructExtracts structured metadata from HTML markup in multiple formats:…
- PyLDPyLD is a Python implementation of the JSON-LD specification that…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also extruct · html2docx · jira2markdown · BeautifulSoup · img2table · hyperscript · tinyhtml5 · collate-dbt-artifacts-parser