date-guesser
Extract publication dates from web pages
Decision gist · record as of 2026-08-14
No. While the package has low install friction and a permissive license, it has been abandoned since August 2019 with no maintenance or updates. Any bugs, compatibility issues with modern Python or dependency versions, or gaps in date format support will not be addressed. For active projects requiring publication date extraction, consider maintained alternatives or build a custom solution tailored to your specific content sources.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with four straightforward runtime dependencies.
- However, the package is archived and abandoned as of August 2019, with no maintenance or updates since its last release over five years ago.
License · maintenance · safety
MIT (permissive) — MIT license permits free use, modification, and distribution with minimal restrictions, making it legally safe to adopt for most projects.
last release 2019-08-13 (2558 days) · last repo commit 2019-08-13 · 41 stars · archived
0 known vulnerabilities (OSV.dev, 2026-08-14) · 88,539 downloads/mo, #13,718 on PyPI
Alternatives
Verify before relying
pip install date-guesser
from date_guesser import guess_date, Accuracy
guess = guess_date(url='https://www.example.com/2017/10/13/article.html', html='')
print(guess.date) # datetime object
print(guess.accuracy) # Accuracy enum value
print(guess.method) # string describing extraction method- Whether the package handles modern web page structures and date formats reliably, given its abandonment since 2019
- Compatibility with current versions of beautifulsoup4 and lxml, which may have breaking changes since the package's last release
- Performance and accuracy on non-English language content, acknowledged as a known limitation in the documentation
What it is and what it does
date-guesser is a library that attempts to identify publication dates from web pages by combining heuristics across multiple signals: URL path patterns, HTML metadata tags, and embedded date strings. It returns not just a date but also an accuracy level (ranging from full datetime precision to partial date or no match) and a description of which method successfully extracted the date.
The package was developed for the mediacloud project to handle real-world news and blog content where publication dates are scattered across different page elements with varying reliability. It prioritizes trustworthy sources over more precise but less reliable ones—for instance, preferring a date found in the URL structure over a more specific timestamp buried in user comments. However, the package is no longer maintained; its last release was in August 2019.
Use it for
- Bulk-extract publication dates from news articles and blog posts during web scraping or content archival workflows
- Determine article freshness and temporal relevance when indexing or searching web content
- Validate or fill missing publication metadata in content management systems by cross-checking multiple date sources
- Benchmark date extraction accuracy in media analysis pipelines, as the package includes comparison metrics against other tools
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
No.
While the package has low install friction and a permissive license, it has been abandoned since August 2019 with no maintenance or updates. Any bugs, compatibility issues with modern Python or dependency versions, or gaps in date format support will not be addressed. For active projects requiring publication date extraction, consider maintained alternatives or build a custom solution tailored to your specific content sources.
Install
date-guesser on PyPI
Before you install
Low install friction with four straightforward runtime dependencies. However, the package is archived and abandoned as of August 2019, with no maintenance or updates since its last release over five years ago.
License in practice
MIT license permits free use, modification, and distribution with minimal restrictions, making it legally safe to adopt for most projects.
Quickstart
pip install date-guesser
from date_guesser import guess_date, Accuracy
guess = guess_date(url='https://www.example.com/2017/10/13/article.html', html='')
print(guess.date) # datetime object
print(guess.accuracy) # Accuracy enum value
print(guess.method) # string describing extraction method
Verify before relying
- Whether the package handles modern web page structures and date formats reliably, given its abandonment since 2019
- Compatibility with current versions of beautifulsoup4 and lxml, which may have breaking changes since the package's last release
- Performance and accuracy on non-English language content, acknowledged as a known limitation in the documentation
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packagesarrowbeautifulsoup4lxmlpytz |
| Maintenance | Abandoned 2,558 days since the last release |
| Last repo commit | repository archived |
| First released | |
| Downloads | 88,539 / month, #13,718 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: DevelopersLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.6 |
Evidence: date_guesser-2.1.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “extract publication date from webpage”
- date-guesserExtracts publication dates from web pages by analyzing URL patterns…
- htmldateExtracts original and updated publication dates from web pages by…
- dateparserParses dates from text in multiple languages and formats, handling…
Give your agent the search over MCP, or paste the wish link into any chat.
More WWW/HTTP packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
h11 is a pure-Python HTTP/1.1 protocol implementation that handles parsing and serializing HTTP messages without any built-in I/O, letting you integrate it with any network layer you choose.
HTTPX is a fully featured HTTP client library for Python that provides both sync and async APIs, with support for HTTP/1.1 and HTTP/2, plus an integrated command-line client.
Install it if you are building new projects or modernizing existing ones that rely on HTTP.
A minimal low-level HTTP client library that sends HTTP requests with thread-safe and task-safe connection pooling, supporting HTTP/1.1, HTTP/2, proxies, and both sync and async interfaces.
aiohttp is an async HTTP client and server framework built on asyncio, supporting both WebSockets and middleware-based routing for building concurrent web applications.
Install it if you need async HTTP client or server capabilities in asyncio-based applications.
See also htmldate · readable-content · datefinder · nr-date · newspaper4k · w3lib · newspaper3k · python-dateutil · python-whois · stringparser