$npx skillfedfor your agent

htmldate

Fast and robust extraction of original and updated publication dates from URLs and web pages.

Worth itPyPI WWW/HTTPReleased Jun 202614.9M downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — htmldate-1.10.0-py3-none-any.whl
v1.10.0 · released 2026-06-01 · Python >=3.10 · 5 runtime deps: charset_normalizer, dateparser, lxml, python-dateutil, urllib3

Yes. The package is actively maintained, has no known vulnerabilities, low install friction, and a permissive license. It is production-tested on millions of documents and ranks in the top 5000 PyPI packages by download volume. Install it if you need reliable date extraction from web pages; the fast mode offers good speed and the extensive mode provides high recall when accuracy matters most.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; earlier versions (1.8.1 for Python 3.6–3.7, 1.9.x for 3.8–3.9) are available if needed.
  • Low friction installation with a pure-Python wheel and five runtime dependencies.
  • Actively maintained as of 2026-07-31 with recent release activity.

License · maintenance · safety

Apache-2.0 (permissive) — Distributed under Apache 2.0, a permissive license allowing commercial and private use with minimal restrictions; versions prior to 1.8.0 used GPLv3+.

last release 2026-06-01 (74 days) · last repo commit 2026-07-31 · 154 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 14,949,222 downloads/mo, #1,210 on PyPI

Verify before relying

pip install htmldate

from htmldate import find_date
date = find_date('http://blog.python.org/2016/12/python-360-is-now-available.html')
print(date)  # '2016-12-23'
  • Whether the package handles redirects, authentication, or rate-limiting when fetching URLs.
  • Performance characteristics on pages with malformed or ambiguous date markup.
  • Whether batch processing supports concurrent requests or is sequential only.
Same gist for agents: .md · .json

What it is and what it does

Htmldate is a Python library and command-line tool for finding publication and update dates on web pages. It works by examining HTML markup (meta tags, Open Graph attributes, structural elements like `time` and `abbr`), then falling back to heuristic text analysis when metadata is absent. The package includes both a fast mode for quick extraction and an extensive mode that collects all candidate dates and uses a disambiguation algorithm to select the most likely one.

The library handles flexible input (URLs, HTML files, or parsed trees) and outputs dates in customizable formats, defaulting to ISO 8601. It is multilingual and has been deployed in production on millions of documents. The package depends on lxml for parsing, dateparser for date normalization, charset_normalizer for encoding detection, python-dateutil for date manipulation, and urllib3 for HTTP requests.

Use it for

  • Automated metadata extraction for web corpora and text databases in research or archival projects.
  • Enriching web scraping pipelines with reliable publication dates when server headers are missing or unreliable.
  • Batch processing of archived or crawled web pages to extract and standardize publication timestamps.
  • Building content aggregation systems that need to sort or filter articles by publication date.
  • Detecting content updates by comparing original and updated publication dates on news or blog sites.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The package is actively maintained, has no known vulnerabilities, low install friction, and a permissive license. It is production-tested on millions of documents and ranks in the top 5000 PyPI packages by download volume. Install it if you need reliable date extraction from web pages; the fast mode offers good speed and the extensive mode provides high recall when accuracy matters most.

Install

htmldate on PyPI

Before you install

Low friction installation with a pure-Python wheel and five runtime dependencies. Actively maintained as of 2026-07-31 with recent release activity.

Requires Python 3.10 or later; earlier versions (1.8.1 for Python 3.6–3.7, 1.9.x for 3.8–3.9) are available if needed.

License in practice

Distributed under Apache 2.0, a permissive license allowing commercial and private use with minimal restrictions; versions prior to 1.8.0 used GPLv3+.

Quickstart

pip install htmldate

from htmldate import find_date
date = find_date('http://blog.python.org/2016/12/python-360-is-now-available.html')
print(date)  # '2016-12-23'

Verify before relying

  • Whether the package handles redirects, authentication, or rate-limiting when fetching URLs.
  • Performance characteristics on pages with malformed or ambiguous date markup.
  • Whether batch processing supports concurrent requests or is sequential only.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
5 packages
charset_normalizerdateparserlxmlpython-dateutilurllib3
MaintenanceActively maintained 74 days since the last release
Last repo commit
First released
Downloads14,949,222 / month, #1,210 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Information TechnologyIntended Audience :: Science/ResearchOperating System :: MacOS :: MacOS XOperating System :: Microsoft :: WindowsOperating System :: POSIX :: LinuxProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Topic :: Internet :: WWW/HTTPTopic :: Scientific/Engineering :: Information AnalysisTopic :: Text Processing :: LinguisticTopic :: Text Processing :: Markup :: HTML

Evidence: htmldate-1.10.0-py3-none-any.whl

Tags

Capabilities
extract publication date from htmlweb page date extractionfind article publish datehtml metadata date parsingweb scraping date detectionpublication date finderarticle date extractionwebpage timestamp extraction
Topics
web-scrapingmetadata-extractiondate-parsing
PyPI keywords
datetimedate-parserentity-extractionhtml-extractionhtml-parsingmetadata-extractionwebarchivesweb-scraping

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “extract publication date from html”

  • htmldateExtracts original and updated publication dates from web pages by…
  • date-guesserExtracts publication dates from web pages by analyzing URL patterns…
  • dateparserParses dates from text in multiple languages and formats, handling…

Give your agent the search over MCP, or paste the wish link into any chat.

More WWW/HTTP packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
h11 With conditions
PyPI · WWW/HTTP · released Apr 2025

h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.

MITpure Python · 3.8+aging
894.9Mdownloads / mo
httpx Worth it
PyPI · WWW/HTTP · released Dec 2024

HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.

Install it if you are building new projects or modernizing existing ones that rely on HTTP.

BSD-3-Clausepure Python · 3.8+
797.0Mdownloads / mo
httpcore With conditions
PyPI · WWW/HTTP · released Apr 2025

A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.

BSD-3-Clausepure Python · 3.8+aging
783.6Mdownloads / mo
aiohttp Worth it
PyPI · WWW/HTTP · released Jul 2026

aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.

Install it if you need async HTTP client or server capabilities in asyncio-based applications.

permissive licensecompiled wheel · 3.10+
643.6Mdownloads / mo

See also date-guesser · trafilatura · datefinder · dateparser · readable-content · nr-date · newspaper3k · prefixdate · requests-html · mkdocs-rss-plugin