yake
Keyword extraction Python package
Decision gist · record as of 2026-08-14
Yes, with conditions. YAKE is worth installing if you need unsupervised keyword extraction without training overhead and accept its aging maintenance status (186 days since last release). The copyleft license requires careful review if you plan proprietary use. No known security vulnerabilities, low install friction, and active repository support make it a reasonable choice for single-document keyword analysis across languages.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later.
- Low install friction with a pure Python wheel distribution.
- Maintenance status is aging—last release was 186 days ago—but the repository remains active with recent commits and a modest user base of 1877 stars.
License · maintenance · safety
LGPLv3 (copyleft) — Licensed under LGPLv3 (copyleft). Derivative works and modifications must be released under the same license; proprietary applications using this package may face licensing constraints.
last release 2026-02-09 (186 days) · last repo commit 2026-02-11 · 1,877 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 453,507 downloads/mo, #6,577 on PyPI
Alternatives
Verify before relying
pip install yake
import yake
text = "Your document text here"
kw_extractor = yake.KeywordExtractor()
keywords = kw_extractor.extract_keywords(text)
for kw, score in keywords:
print(f"{kw} ({score})")- Performance characteristics on very large documents or streaming text.
- Accuracy comparison with other unsupervised keyword extraction methods.
- Memory footprint and scalability with high-dimensional feature sets.
What it is and what it does
YAKE is an unsupervised keyword extraction library that identifies important terms and phrases in text by analyzing statistical features like word frequency, position, and co-occurrence patterns. It works on single documents without requiring training, external corpora, or language-specific resources, making it applicable across languages and domains.
The package provides both command-line and Python APIs for keyword extraction, with configurable parameters for n-gram size, deduplication strategies, context window size, and result count. It includes multilingual support, lemmatization to normalize morphological variants, and text highlighting capabilities. Dependencies are lightweight: click for CLI, jellyfish and networkx for graph-based deduplication, numpy for numerical operations, segtok for sentence segmentation, and tabulate for formatted output.
Use it for
- Extract key topics from research papers or articles for indexing and search.
- Identify main concepts in customer feedback or support tickets for categorization.
- Generate tag suggestions for blog posts or documents without manual annotation.
- Analyze multilingual documents to find domain-relevant terms across languages.
- Build automated content summarization pipelines that preserve semantic focus.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, with conditions.
YAKE is worth installing if you need unsupervised keyword extraction without training overhead and accept its aging maintenance status (186 days since last release). The copyleft license requires careful review if you plan proprietary use. No known security vulnerabilities, low install friction, and active repository support make it a reasonable choice for single-document keyword analysis across languages.
Install
yake on PyPI
Before you install
Low install friction with a pure Python wheel distribution. Maintenance status is aging—last release was 186 days ago—but the repository remains active with recent commits and a modest user base of 1877 stars.
Requires Python 3.10 or later.
License in practice
Licensed under LGPLv3 (copyleft). Derivative works and modifications must be released under the same license; proprietary applications using this package may face licensing constraints.
Quickstart
pip install yake
import yake
text = "Your document text here"
kw_extractor = yake.KeywordExtractor()
keywords = kw_extractor.extract_keywords(text)
for kw, score in keywords:
print(f"{kw} ({score})")
Verify before relying
- Performance characteristics on very large documents or streaming text.
- Accuracy comparison with other unsupervised keyword extraction methods.
- Memory footprint and scalability with high-dimensional feature sets.
Package facts
| License | LGPLv3 copyleft |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 6 packagesclickjellyfishnetworkxnumpysegtoktabulate |
| Maintenance | Aging 186 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 453,507 / month, #6,577 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaLicense :: OSI Approved :: GNU General Public License v3 (GPLv3)Programming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Topic :: Scientific/Engineering :: Information AnalysisTopic :: Software Development :: LibrariesTopic :: Text ProcessingTopic :: Text Processing :: Linguistic |
Evidence: yake-0.7.3-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “unsupervised keyword extraction”
- yakeYAKE extracts keywords from text documents using unsupervised…
- keyphrase-vectorizersExtracts keyphrases from text documents using part-of-speech patterns…
- soynlpUnsupervised Korean natural language processing toolkit that extracts…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also keybert · keyphrase-vectorizers · rake-nltk · flashtext · polyglot · textract · bertopic · simplemma · geotext · goose3