$npx skillfedfor your agent

sumy

Module for automatic summarization of text documents and HTML pages.

Worth itPyPI InternetReleased Aug 2026173.5K downloads / moApache License, Version 2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — sumy-0.13.0-py3-none-any.whl
v0.13.0 · released 2026-08-12 · Python >=3.8 · 7 runtime deps: breadability, docopt-ng, lxml-html-clean, nltk, pycountry, requests, setuptools

Yes. Sumy is actively maintained, has no known vulnerabilities, low install friction, and a permissive Apache License, Version 2.0. It offers multiple summarization algorithms and multilingual support, making it suitable for both simple summarization tasks and more complex NLP pipelines. The recent release and strong GitHub presence indicate ongoing development and community use.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.8+; nltk and lxml-html-clean are compiled/system-dependent dependencies that may need additional setup on some platforms.
  • Low friction install with 7 runtime dependencies including nltk, lxml-html-clean, and requests.
  • Package is actively maintained with a recent release 2 days ago and 3701 GitHub stars.

License · maintenance · safety

Apache License, Version 2.0 (permissive) — Licensed under Apache License, Version 2.0 (permissive), allowing commercial and private use with minimal restrictions beyond attribution and liability disclaimers.

last release 2026-08-12 (2 days) · last repo commit 2026-08-14 · 3,701 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 173,532 downloads/mo, #10,305 on PyPI

Verify before relying

pip install sumy

from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer
from sumy.utils import get_stop_words

parser = HtmlParser.from_url("https://example.com", Tokenizer("english"))
summarizer = LsaSummarizer()
summarizer.stop_words = get_stop_words("english")
for sentence in summarizer(parser.document, 10):
    print(sentence)
  • Whether all listed natural languages are fully supported or require additional tokenizer configuration.
  • Performance characteristics and memory usage for large documents or batch processing.
  • Accuracy or ROUGE scores for the implemented summarization algorithms relative to alternatives.
Same gist for agents: .md · .json

What it is and what it does

Sumy is a Python library for automatic text summarization that works with both HTML pages and plain text. It implements multiple summarization algorithms—LSA, LexRank, Luhn, and Edmundson—allowing you to choose the approach that fits your use case. The package includes a command-line tool for quick summarization tasks and an evaluation framework to measure summary quality against reference texts.

The library handles multilingual content and provides tokenizers for language-specific processing. It depends on nltk for natural language processing, lxml-html-clean for HTML parsing, requests for fetching web content, and breadability for content extraction. You can use it as a library in your Python code or invoke it from the command line, making it flexible for both programmatic and ad-hoc summarization tasks.

Use it for

  • Summarize web articles or Wikipedia pages to extract key information without manual reading.
  • Build a bot that automatically generates TL;DR summaries for forum discussions or social media posts.
  • Reduce large document collections to key sentences for quick review or archival purposes.
  • Evaluate the quality of generated summaries against reference texts using built-in evaluation methods.
  • Integrate multilingual summarization into applications serving non-English-speaking users.
  • Extract summaries from video transcripts or meeting notes by parsing plain text output.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Sumy is actively maintained, has no known vulnerabilities, low install friction, and a permissive Apache License, Version 2.0. It offers multiple summarization algorithms and multilingual support, making it suitable for both simple summarization tasks and more complex NLP pipelines. The recent release and strong GitHub presence indicate ongoing development and community use.

Install

sumy on PyPI

Before you install

Low friction install with 7 runtime dependencies including nltk, lxml-html-clean, and requests. Package is actively maintained with a recent release 2 days ago and 3701 GitHub stars. Supports Python 3.8 through 3.14.

Requires Python 3.8+; nltk and lxml-html-clean are compiled/system-dependent dependencies that may need additional setup on some platforms.

License in practice

Licensed under Apache License, Version 2.0 (permissive), allowing commercial and private use with minimal restrictions beyond attribution and liability disclaimers.

Quickstart

pip install sumy

from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer
from sumy.utils import get_stop_words

parser = HtmlParser.from_url("https://example.com", Tokenizer("english"))
summarizer = LsaSummarizer()
summarizer.stop_words = get_stop_words("english")
for sentence in summarizer(parser.document, 10):
    print(sentence)

Verify before relying

  • Whether all listed natural languages are fully supported or require additional tokenizer configuration.
  • Performance characteristics and memory usage for large documents or batch processing.
  • Accuracy or ROUGE scores for the implemented summarization algorithms relative to alternatives.

Package facts

LicenseApache License, Version 2.0 permissive
Python supportSupports the current Python release >=3.8
Install frictionLow. Pure-Python wheel
Runtime dependencies
7 packages
breadabilitydocopt-nglxml-html-cleannltkpycountryrequestssetuptools
MaintenanceActively maintained 2 days since the last release
Last repo commit
First released
Downloads173,532 / month, #10,305 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: EducationLicense :: OSI Approved :: Apache Software LicenseNatural Language :: ArabicNatural Language :: Chinese (Simplified)Natural Language :: CzechNatural Language :: EnglishNatural Language :: FrenchNatural Language :: GermanNatural Language :: GreekNatural Language :: HebrewNatural Language :: ItalianNatural Language :: JapaneseNatural Language :: PolishNatural Language :: PortugueseNatural Language :: SlovakNatural Language :: SpanishNatural Language :: SwedishNatural Language :: ThaiNatural Language :: UkrainianOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyTopic :: EducationTopic :: InternetTopic :: Scientific/Engineering :: Information AnalysisTopic :: Text Processing :: FiltersTopic :: Text Processing :: LinguisticTopic :: Text Processing :: Markup :: HTML

Evidence: sumy-0.13.0-py3-none-any.whl

Tags

Capabilities
automatic text summarizationextract summary from HTMLNLP text summarizerdocument summarization libraryLSA LexRank summarizermultilingual text summarizationcommand-line summarization tool
Topics
nlptext-summarizationmultilingual
PyPI keywords
LSALexRankNLPTextRankautomatic summarizationdata miningdata reductionlatent semantic analysisnatural language processingweb-data extraction

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “automatic text summarization”

  • sumySumy extracts summaries from HTML pages or plain text using multiple…
  • rouge-metricComputes ROUGE metrics (ROUGE-N, ROUGE-L, ROUGE-W, ROUGE-S, ROUGE-SU)…
  • rake-nltkExtracts keywords and key phrases from text using the RAKE algorithm,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Internet packages

botocore Worth it
PyPI · Internet · released Aug 2026

Botocore provides low-level, data-driven access to Amazon Web Services APIs, serving as the foundation for the AWS CLI and boto3 libraries.

Install it if you need programmatic access to AWS services.

permissive licensepure Python · 3.10+
1.5Bdownloads / mo
aiobotocore Worth it
PyPI · Internet · released Aug 2026

Provides an async client for AWS services using botocore and aiohttp, allowing you to call AWS APIs asynchronously within asyncio-based applications.

Install it if you need to call AWS services from async Python code; it is the standard way to do so.

Apache-2.0pure Python · 3.10+
1.2Bdownloads / mo
pydantic Worth it
PyPI · Python Modules · released May 2026

Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.

MITpure Python · 3.9+
1.1Bdownloads / mo
filelock Worth it
PyPI · Libraries · released Aug 2026

Provides a platform-independent file locking mechanism to coordinate access to files across processes and threads.

MITpure Python · 3.10+
717.1Mdownloads / mo
fastapi Worth it
PyPI · Software Development · released Jul 2026

FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.

MITpure Python · 3.10+
568.6Mdownloads / mo
googleapis-common-protos Worth it
PyPI · Internet · released Aug 2026

Provides common Protocol Buffer message definitions used across Google Cloud APIs, enabling Python clients to interact with Google services.

Apache-2.0pure Python · 3.10+
513.7Mdownloads / mo

See also jusText · pytextrank · semantic-text-splitter · rouge-chinese · rouge · readabilipy · textacy · gensim · rouge-metric · parsel