sumy
Module for automatic summarization of text documents and HTML pages.
Decision gist · record as of 2026-08-14
Yes. Sumy is actively maintained, has no known vulnerabilities, low install friction, and a permissive Apache License, Version 2.0. It offers multiple summarization algorithms and multilingual support, making it suitable for both simple summarization tasks and more complex NLP pipelines. The recent release and strong GitHub presence indicate ongoing development and community use.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.8+; nltk and lxml-html-clean are compiled/system-dependent dependencies that may need additional setup on some platforms.
- Low friction install with 7 runtime dependencies including nltk, lxml-html-clean, and requests.
- Package is actively maintained with a recent release 2 days ago and 3701 GitHub stars.
License · maintenance · safety
Apache License, Version 2.0 (permissive) — Licensed under Apache License, Version 2.0 (permissive), allowing commercial and private use with minimal restrictions beyond attribution and liability disclaimers.
last release 2026-08-12 (2 days) · last repo commit 2026-08-14 · 3,701 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 173,532 downloads/mo, #10,305 on PyPI
Alternatives
Verify before relying
pip install sumy
from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer
from sumy.utils import get_stop_words
parser = HtmlParser.from_url("https://example.com", Tokenizer("english"))
summarizer = LsaSummarizer()
summarizer.stop_words = get_stop_words("english")
for sentence in summarizer(parser.document, 10):
print(sentence)- Whether all listed natural languages are fully supported or require additional tokenizer configuration.
- Performance characteristics and memory usage for large documents or batch processing.
- Accuracy or ROUGE scores for the implemented summarization algorithms relative to alternatives.
What it is and what it does
Sumy is a Python library for automatic text summarization that works with both HTML pages and plain text. It implements multiple summarization algorithms—LSA, LexRank, Luhn, and Edmundson—allowing you to choose the approach that fits your use case. The package includes a command-line tool for quick summarization tasks and an evaluation framework to measure summary quality against reference texts.
The library handles multilingual content and provides tokenizers for language-specific processing. It depends on nltk for natural language processing, lxml-html-clean for HTML parsing, requests for fetching web content, and breadability for content extraction. You can use it as a library in your Python code or invoke it from the command line, making it flexible for both programmatic and ad-hoc summarization tasks.
Use it for
- Summarize web articles or Wikipedia pages to extract key information without manual reading.
- Build a bot that automatically generates TL;DR summaries for forum discussions or social media posts.
- Reduce large document collections to key sentences for quick review or archival purposes.
- Evaluate the quality of generated summaries against reference texts using built-in evaluation methods.
- Integrate multilingual summarization into applications serving non-English-speaking users.
- Extract summaries from video transcripts or meeting notes by parsing plain text output.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Sumy is actively maintained, has no known vulnerabilities, low install friction, and a permissive Apache License, Version 2.0. It offers multiple summarization algorithms and multilingual support, making it suitable for both simple summarization tasks and more complex NLP pipelines. The recent release and strong GitHub presence indicate ongoing development and community use.
Install
sumy on PyPI
Before you install
Low friction install with 7 runtime dependencies including nltk, lxml-html-clean, and requests. Package is actively maintained with a recent release 2 days ago and 3701 GitHub stars. Supports Python 3.8 through 3.14.
Requires Python 3.8+; nltk and lxml-html-clean are compiled/system-dependent dependencies that may need additional setup on some platforms.
License in practice
Licensed under Apache License, Version 2.0 (permissive), allowing commercial and private use with minimal restrictions beyond attribution and liability disclaimers.
Quickstart
pip install sumy
from sumy.parsers.html import HtmlParser
from sumy.nlp.tokenizers import Tokenizer
from sumy.summarizers.lsa import LsaSummarizer
from sumy.utils import get_stop_words
parser = HtmlParser.from_url("https://example.com", Tokenizer("english"))
summarizer = LsaSummarizer()
summarizer.stop_words = get_stop_words("english")
for sentence in summarizer(parser.document, 10):
print(sentence)
Verify before relying
- Whether all listed natural languages are fully supported or require additional tokenizer configuration.
- Performance characteristics and memory usage for large documents or batch processing.
- Accuracy or ROUGE scores for the implemented summarization algorithms relative to alternatives.
Package facts
| License | Apache License, Version 2.0 permissive |
| Python support | Supports the current Python release >=3.8 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagesbreadabilitydocopt-nglxml-html-cleannltkpycountryrequestssetuptools |
| Maintenance | Actively maintained 2 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 173,532 / month, #10,305 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersIntended Audience :: EducationLicense :: OSI Approved :: Apache Software LicenseNatural Language :: ArabicNatural Language :: Chinese (Simplified)Natural Language :: CzechNatural Language :: EnglishNatural Language :: FrenchNatural Language :: GermanNatural Language :: GreekNatural Language :: HebrewNatural Language :: ItalianNatural Language :: JapaneseNatural Language :: PolishNatural Language :: PortugueseNatural Language :: SlovakNatural Language :: SpanishNatural Language :: SwedishNatural Language :: ThaiNatural Language :: UkrainianOperating System :: OS IndependentProgramming Language :: PythonProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPyTopic :: EducationTopic :: InternetTopic :: Scientific/Engineering :: Information AnalysisTopic :: Text Processing :: FiltersTopic :: Text Processing :: LinguisticTopic :: Text Processing :: Markup :: HTML |
Evidence: sumy-0.13.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “automatic text summarization”
- sumySumy extracts summaries from HTML pages or plain text using multiple…
- rouge-metricComputes ROUGE metrics (ROUGE-N, ROUGE-L, ROUGE-W, ROUGE-S, ROUGE-SU)…
- rake-nltkExtracts keywords and key phrases from text using the RAKE algorithm,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Internet packages
Botocore provides low-level, data-driven access to Amazon Web Services APIs, serving as the foundation for the AWS CLI and boto3 libraries.
Install it if you need programmatic access to AWS services.
Provides an async client for AWS services using botocore and aiohttp, allowing you to call AWS APIs asynchronously within asyncio-based applications.
Install it if you need to call AWS services from async Python code; it is the standard way to do so.
Pydantic validates Python data structures against type hints, coercing and checking input at runtime to ensure it matches a declared schema.
Provides a platform-independent file locking mechanism to coordinate access to files across processes and threads.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides common Protocol Buffer message definitions used across Google Cloud APIs, enabling Python clients to interact with Google services.
See also jusText · pytextrank · semantic-text-splitter · rouge-chinese · rouge · readabilipy · textacy · gensim · rouge-metric · parsel