pyrdfa3
pyRdfa distiller/parser library
Decision gist · record as of 2026-08-14
Yes, with conditions. The package has low install friction and no known vulnerabilities, making it technically safe to use. However, the original maintainer has retired and the package is now community-maintained with minimal activity (1 star, aging status, last commit in January 2026). Install it if you need RDFa 1.1 parsing and are comfortable with a stable but not actively developed library; verify the unclear license terms before use in commercial or restricted-license projects.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.8 or higher; requests 2.32.3 or higher, rdflib 7.0.0 or higher, and html5lib 1.1.
- Low install friction with three straightforward dependencies (requests, rdflib, html5lib).
- However, the original maintainer has stepped back and the package is now maintained by a community contributor; the repository is archived but not abandoned.
License · maintenance · safety
(unclear) — License treatment is unclear—the classifier indicates W3C License approval, but no explicit SPDX or raw license text is recorded in the package metadata. Verify the actual license terms before relying on this package in a commercial or restricted-license context.
last release 2026-01-17 (209 days) · last repo commit 2026-01-18 · 1 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 585,794 downloads/mo, #5,885 on PyPI
Alternatives
Verify before relying
pip install pyRdfa3
import pyRdfa
from rdflib import Graph
graph = pyRdfa.parse('http://example.com/page.html')
# graph is an rdflib Graph object containing extracted RDF triples- Whether the W3C License classifier accurately reflects the actual license terms and any restrictions on use.
- Current maintenance status and responsiveness to issues given the original maintainer's retirement and community takeover.
- Whether the package handles all RDFa 1.1 edge cases or if there are known parsing limitations.
What it is and what it does
pyRdfa3 is a Python library that parses RDFa 1.1 markup embedded in HTML documents and distills it into structured RDF data. It reads HTML files or web pages containing RDFa annotations—semantic metadata embedded as HTML attributes—and extracts the underlying RDF triples, which can then be queried or manipulated using rdflib.
The package is built on three core dependencies: requests for fetching remote HTML, rdflib for representing and working with RDF graphs, and html5lib for robust HTML parsing. It supports Python 3.8 and higher and is designed for developers working with semantic web standards, linked data, or applications that need to extract structured metadata from web pages.
Use it for
- Extract structured data from web pages marked up with RDFa to build knowledge graphs or semantic search indexes.
- Convert RDFa-annotated HTML documents into RDF/Turtle or other RDF serialization formats for downstream processing.
- Validate or test RDFa markup compliance in HTML documents against the RDFa 1.1 specification.
- Integrate semantic web data extraction into data pipelines that consume linked data from third-party websites.
- Build tools that harvest schema.org or other RDFa vocabularies from HTML to populate databases or ontologies.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
The package has low install friction and no known vulnerabilities, making it technically safe to use. However, the original maintainer has retired and the package is now community-maintained with minimal activity (1 star, aging status, last commit in January 2026). Install it if you need RDFa 1.1 parsing and are comfortable with a stable but not actively developed library; verify the unclear license terms before use in commercial or restricted-license projects.
Install
pyrdfa3 on PyPI
Before you install
Low install friction with three straightforward dependencies (requests, rdflib, html5lib). However, the original maintainer has stepped back and the package is now maintained by a community contributor; the repository is archived but not abandoned. Last commit was recent, but the aging status and low star count suggest limited active development.
Requires Python 3.8 or higher; requests 2.32.3 or higher, rdflib 7.0.0 or higher, and html5lib 1.1.
License in practice
License treatment is unclear—the classifier indicates W3C License approval, but no explicit SPDX or raw license text is recorded in the package metadata. Verify the actual license terms before relying on this package in a commercial or restricted-license context.
Quickstart
pip install pyRdfa3
import pyRdfa
from rdflib import Graph
graph = pyRdfa.parse('http://example.com/page.html')
# graph is an rdflib Graph object containing extracted RDF triples
Verify before relying
- Whether the W3C License classifier accurately reflects the actual license terms and any restrictions on use.
- Current maintenance status and responsiveness to issues given the original maintainer's retirement and community takeover.
- Whether the package handles all RDFa 1.1 edge cases or if there are known parsing limitations.
Package facts
| License | Not declared unclear |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 3 packagesrequestsrdflibhtml5lib |
| Maintenance | Aging 209 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 585,794 / month, #5,885 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: W3C LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: pyrdfa3-3.6.5-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “rdfa parser”
- pyrdfa3Parses RDFa markup embedded in HTML documents and extracts structured…
- extructExtracts structured metadata from HTML markup in multiple formats:…
- recipe-scrapersExtracts structured recipe data (ingredients, instructions, cooking…
Give your agent the search over MCP, or paste the wish link into any chat.
More HTML packages
MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.
Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.
Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.
Install it if you need to parse or extract data from markup documents.
lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.
Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.
Install it if you need to parse Markdown in Python.
See also extruct · rdflib · html5rdf · rdflib-jsonld · owlrl · CFGraph · SPARQLWrapper · PyLD · PyShEx · pymantic