$npx skillfedfor your agent

newspaper3k

Simplified python article discovery & extraction.

With conditionsPyPI HTMLReleased Sep 2018845.9K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — newspaper3k-0.2.8-py3-none-any.whl
v0.2.8 · released 2018-09-28 · 13 runtime deps: tinysegmenter, beautifulsoup4, Pillow, PyYAML, cssselect, lxml, nltk, requests

Yes, if you need reliable article extraction from web pages and news sites. The package is actively maintained, has low install friction, carries no security vulnerabilities, and is widely used (top 5000 on PyPI). The MIT license poses no restrictions. Install it if you're building a news aggregator, content pipeline, or article analysis tool; skip it if you only need simple HTML parsing without article-specific extraction logic.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • On Debian/Ubuntu, system libraries are required for lxml and Pillow: libxml2-dev, libxslt-dev, libjpeg-dev, zlib1g-dev, libpng-dev.
  • NLP features require downloading language corpora.
  • Low install friction with a pure-Python wheel distribution.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute the package freely provided you include the original license notice.

last release 2018-09-28 (2877 days) · last repo commit 2026-08-09 · 15,141 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 845,877 downloads/mo, #4,918 on PyPI

Verify before relying

pip3 install newspaper3k

from newspaper import Article

url = 'http://example.com/article'
article = Article(url)
article.download()
article.parse()
print(article.title)
print(article.text)
  • Whether all 13 runtime dependencies are required for basic article extraction or if some are optional for specific features.
  • Performance characteristics and memory usage when processing large numbers of articles or very large HTML documents.
  • Current maintenance status and whether the 2018-09-28 release is the final stable version or if newer development exists.
Same gist for agents: .md · .json

What it is and what it does

Newspaper3k is a Python library for discovering, downloading, and extracting structured content from web articles and news sites. It automates the process of fetching HTML from a URL, parsing the DOM to isolate article text, and extracting metadata like title, author, publish date, top image, and embedded videos. The library also performs natural language processing to identify keywords and generate summaries.

The package is designed around simplicity and speed, relying on lxml for fast HTML parsing and requests for HTTP operations. It supports multi-threaded article downloads and can work with news sources in multiple languages. Core use cases include building news aggregators, content curation pipelines, and automated article analysis workflows. Installation requires several system-level dependencies on Linux, and NLP features require downloading language corpora.

Use it for

  • Build a news aggregator that discovers and extracts articles from multiple news sites automatically.
  • Extract article text and metadata from URLs for content curation or archival systems.
  • Perform bulk text analysis on news articles across multiple languages for research or trend detection.
  • Automate extraction of article images and videos for content republishing or media analysis.
  • Generate article summaries and keyword extraction for search indexing or content recommendation.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need reliable article extraction from web pages and news sites.

The package is actively maintained, has low install friction, carries no security vulnerabilities, and is widely used (top 5000 on PyPI). The MIT license poses no restrictions. Install it if you're building a news aggregator, content pipeline, or article analysis tool; skip it if you only need simple HTML parsing without article-specific extraction logic.

Install

newspaper3k on PyPI

Before you install

Low install friction with a pure-Python wheel distribution. The package is actively maintained with recent commits as of 2026-08-09, though the latest release dates to 2018-09-28. Requires 13 runtime dependencies including lxml, beautifulsoup4, and nltk, which may need system libraries on Linux.

On Debian/Ubuntu, system libraries are required for lxml and Pillow: libxml2-dev, libxslt-dev, libjpeg-dev, zlib1g-dev, libpng-dev. NLP features require downloading language corpora.

License in practice

MIT license permits commercial and private use with minimal restrictions; you may use, modify, and distribute the package freely provided you include the original license notice.

Quickstart

pip3 install newspaper3k

from newspaper import Article

url = 'http://example.com/article'
article = Article(url)
article.download()
article.parse()
print(article.title)
print(article.text)

Verify before relying

  • Whether all 13 runtime dependencies are required for basic article extraction or if some are optional for specific features.
  • Performance characteristics and memory usage when processing large numbers of articles or very large HTML documents.
  • Current maintenance status and whether the 2018-09-28 release is the final stable version or if newer development exists.

Package facts

LicenseMIT permissive
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
13 packages
tinysegmenterbeautifulsoup4PillowPyYAMLcssselectlxmlnltkrequestsfeedparsertldextractfeedfinder2jieba3kpython-dateutil
MaintenanceActively maintained 2,877 days since the last release
Last repo commit
First released
Downloads845,877 / month, #4,918 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Intended Audience :: DevelopersNatural Language :: EnglishProgramming Language :: Python :: 3

Evidence: newspaper3k-0.2.8-py3-none-any.whl

Tags

Capabilities
web scraping article extractionnews content parsinghtml to article textmultilingual article extractionweb page text miningnews site scrapingarticle metadata extraction
Topics
web-scrapingnlpmultilingual

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “web scraping article extraction”

  • newspaper3kNewspaper3k downloads and extracts article text, metadata, images,…
  • newspaper4kExtracts article text, metadata, and images from web pages and news…
  • readable-contentExtracts the main article content from web pages, handling both…

Give your agent the search over MCP, or paste the wish link into any chat.

More HTML packages

MarkupSafe Worth it
PyPI · Dynamic Content · released Sep 2025

MarkupSafe provides a text object that escapes special characters so untrusted strings can be safely embedded in HTML and XML without injection attacks.

BSD-3-Clausecompiled wheel · 3.9+aging
797.1Mdownloads / mo
Jinja2 Worth it
PyPI · Dynamic Content · released Mar 2025

Jinja2 is a templating engine that renders dynamic content by combining templates with Python-like syntax and data, supporting template inheritance, macros, autoescaping, and sandboxed execution.

BSD-3-Clausepure Python · 3.7+aging
718.6Mdownloads / mo
beautifulsoup4 Worth it
PyPI · Python Modules · released Jun 2026

Beautiful Soup parses HTML and XML documents into a navigable tree, providing Pythonic methods to search, iterate, and modify the parsed content.

Install it if you need to parse or extract data from markup documents.

MITpure Python · 3.7.0+
432.1Mdownloads / mo
lxml Worth it
PyPI · Python Modules · released May 2026

lxml provides Python bindings to libxml2 and libxslt, enabling parsing, validation, and transformation of XML and HTML documents through an ElementTree-compatible API with support for XPath, XSLT, and schema validation.

Install it if you need robust XML/HTML parsing, validation, or transformation; avoid it only if you must stay pure-Python and can accept slower performance.

permissive licensecompiled wheel · 3.8+
416.8Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
Markdown Worth it
PyPI · Python Modules · released Jul 2026

Converts Markdown text to HTML using a Python implementation of John Gruber's Markdown specification, with support for extensions.

Install it if you need to parse Markdown in Python.

BSD-3-Clausepure Python · 3.10+
121.7Mdownloads / mo

See also newspaper4k · goose3 · boilerpy3 · gnews · readable-content · breadability · date-guesser · htmldate · readabilipy · readability-lxml

Further reading