rank-bm25
Various BM25 algorithms for document ranking
Decision gist · record as of 2026-08-14
Yes. Low install friction, no security vulnerabilities, permissive license, active maintenance, and a focused, well-documented implementation of a standard algorithm. Install if you need BM25 ranking and want to control preprocessing yourself; skip if you need a full-featured search engine with built-in text processing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Package expects pre-tokenized input (list of token lists); you must handle text preprocessing (lowercasing, stopword removal, stemming) yourself before passing to BM25.
- Low friction install with a single runtime dependency (numpy).
- Repository is active with recent commits and steady maintenance since 2019.
License · maintenance · safety
Apache2.0 (permissive) — Apache2.0 permissive license allows commercial and private use with minimal restrictions.
last release 2022-02-16 (1640 days) · last repo commit 2026-05-02 · 1,375 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 8,915,190 downloads/mo, #1,578 on PyPI
Alternatives
Verify before relying
pip install rank-bm25
from rank_bm25 import BM25Okapi
corpus = ["Hello there good man!", "It is quite windy in London"]
tokenized_corpus = [doc.split(" ") for doc in corpus]
bm25 = BM25Okapi(tokenized_corpus)
query = "windy London"
tokenized_query = query.split(" ")
scores = bm25.get_scores(tokenized_query)- Whether BM25-Adpt and BM25T algorithms mentioned as unimplemented are planned for future releases.
- Current test coverage and benchmarking against other ranking libraries.
- Performance characteristics on large corpora (memory usage, query latency).
What it is and what it does
Rank-BM25 provides implementations of the BM25 family of ranking algorithms—Okapi BM25, BM25L, and BM25+—for scoring how relevant documents are to a search query. It takes a corpus of pre-tokenized documents, builds an index, and then scores or ranks documents against tokenized queries using probabilistic relevance models. The package is intentionally minimal: it does not handle text preprocessing like lowercasing, stemming, or stopword removal, leaving those decisions to the caller.
The typical workflow is to tokenize your document corpus and query using your chosen preprocessing pipeline, initialize a BM25 class with the tokenized corpus, then call get_scores() to retrieve relevance scores or get_top_n() to retrieve the highest-ranking documents. It's commonly used to build search engines or to rank candidate documents in information retrieval pipelines.
Use it for
- Build a lightweight full-text search engine for a document collection without external infrastructure.
- Rank candidate documents in a retrieval-augmented generation (RAG) pipeline before passing to a language model.
- Score document relevance in a question-answering system to find the most relevant passages.
- Implement search functionality in a web application where you control the preprocessing and indexing.
- Benchmark BM25 variants against each other on your own corpus to evaluate ranking quality.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Low install friction, no security vulnerabilities, permissive license, active maintenance, and a focused, well-documented implementation of a standard algorithm. Install if you need BM25 ranking and want to control preprocessing yourself; skip if you need a full-featured search engine with built-in text processing.
Install
rank-bm25 on PyPI
Before you install
Low friction install with a single runtime dependency (numpy). Repository is active with recent commits and steady maintenance since 2019.
Package expects pre-tokenized input (list of token lists); you must handle text preprocessing (lowercasing, stopword removal, stemming) yourself before passing to BM25.
License in practice
Apache2.0 permissive license allows commercial and private use with minimal restrictions.
Quickstart
pip install rank-bm25
from rank_bm25 import BM25Okapi
corpus = ["Hello there good man!", "It is quite windy in London"]
tokenized_corpus = [doc.split(" ") for doc in corpus]
bm25 = BM25Okapi(tokenized_corpus)
query = "windy London"
tokenized_query = query.split(" ")
scores = bm25.get_scores(tokenized_query)
Verify before relying
- Whether BM25-Adpt and BM25T algorithms mentioned as unimplemented are planned for future releases.
- Current test coverage and benchmarking against other ranking libraries.
- Performance characteristics on large corpora (memory usage, query latency).
Package facts
| License | Apache2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagenumpy |
| Maintenance | Actively maintained 1,640 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 8,915,190 / month, #1,578 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: rank_bm25-0.2.2-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “BM25 document ranking”
- rank-bm25Implements BM25 ranking algorithms (Okapi BM25, BM25L, BM25+) to…
- llama-index-retrievers-bm25Integrates BM25 full-text search retrieval into LlamaIndex…
- bm25sBM25S implements the BM25 ranking algorithm in pure Python with…
Give your agent the search over MCP, or paste the wish link into any chat.
More Text Processing packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.
Install it if you need to parse text or define grammars programmatically.
fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.
Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.
Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.
tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.
Install it if your project requires CSS tokenization or syntax manipulation.
See also bm25s · sqlite-fts4 · llama-index-retrievers-bm25 · voyageai · rouge-score · colbert-ai · lunr · ir-measures · implicit · mrmr-selection