pytrec-eval-terrier
Provides Python bindings for popular Information Retrieval measures implemented within trec_eval.
Decision gist · record as of 2026-08-14
Yes. The package is actively maintained, has no known vulnerabilities, installs cleanly on modern Python versions via pre-built wheels, and solves a real problem—standardized IR evaluation—with a proven, well-cited implementation. Use it if you need to evaluate ranked search results against relevance judgments; the MIT license and permissive treatment pose no barrier.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3 or later; numpy and scipy must be installed.
- Medium install friction due to compiled C++ components; pre-built wheels available for Python 3.10–3.14 on Linux, macOS, and Windows.
- Actively maintained with recent releases.
License · maintenance · safety
permissive license (permissive) — Licensed under the MIT license (permissive), allowing broad reuse. Note that the underlying trec_eval tool is licensed separately; check its terms if you modify or redistribute.
last release 2025-10-20 (298 days) · last repo commit 2026-06-25 · 7 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,240,259 downloads/mo, #4,173 on PyPI
Alternatives
Verify before relying
pip install pytrec_eval_terrier
import pytrec_eval
qrel = {'q1': {'d1': 0, 'd2': 1}}
run = {'q1': {'d1': 1.0, 'd2': 0.0}}
evaluator = pytrec_eval.RelevanceEvaluator(qrel, {'map', 'ndcg'})
print(evaluator.evaluate(run))- Whether the fork maintains full API compatibility with the original pytrec_eval project.
- Performance characteristics when evaluating large-scale ranking runs.
- Availability of documentation beyond the GitHub README and examples.
What it is and what it does
pytrec_eval_terrier is a Python wrapper around TREC's trec_eval evaluation tool, designed to eliminate custom implementations of Information Retrieval metrics in Python. It exposes a simple API for computing standard IR measures—such as MAP (Mean Average Precision) and NDCG (Normalized Discounted Cumulative Gain)—on ranked search results against relevance judgments. The package accepts query-document relevance labels and ranked runs as nested dictionaries, then returns computed metrics per query.
This fork, maintained by the University of Glasgow, provides pre-built wheels for modern Python versions (3.10–3.14) and multiple platforms (Linux, macOS, Windows), reducing installation friction compared to building from source. It depends on numpy and scipy for numerical operations and wraps C++ code from the original trec_eval project, making it both fast and reliable for research and production IR evaluation workflows.
Use it for
- Evaluate search engine or ranking model performance using standard TREC metrics without writing custom metric code.
- Compute statistical significance between two ranked runs to compare IR system improvements.
- Benchmark information retrieval systems in research papers or academic projects.
- Integrate IR evaluation into automated testing or continuous evaluation pipelines.
- Validate ranking quality during development of search applications or recommendation systems.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
The package is actively maintained, has no known vulnerabilities, installs cleanly on modern Python versions via pre-built wheels, and solves a real problem—standardized IR evaluation—with a proven, well-cited implementation. Use it if you need to evaluate ranked search results against relevance judgments; the MIT license and permissive treatment pose no barrier.
Install
pytrec-eval-terrier on PyPI
Before you install
Medium install friction due to compiled C++ components; pre-built wheels available for Python 3.10–3.14 on Linux, macOS, and Windows. Actively maintained with recent releases. Depends on numpy and scipy.
Requires Python 3 or later; numpy and scipy must be installed.
License in practice
Licensed under the MIT license (permissive), allowing broad reuse. Note that the underlying trec_eval tool is licensed separately; check its terms if you modify or redistribute.
Quickstart
pip install pytrec_eval_terrier
import pytrec_eval
qrel = {'q1': {'d1': 0, 'd2': 1}}
run = {'q1': {'d1': 1.0, 'd2': 0.0}}
evaluator = pytrec_eval.RelevanceEvaluator(qrel, {'map', 'ndcg'})
print(evaluator.evaluate(run))
Verify before relying
- Whether the fork maintains full API compatibility with the original pytrec_eval project.
- Performance characteristics when evaluating large-scale ranking runs.
- Availability of documentation beyond the GitHub README and examples.
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3 |
| Install friction | Medium. Platform-specific wheel |
| Runtime dependencies | 2 packagesnumpyscipy |
| Maintenance | Actively maintained 298 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 1,240,259 / month, #4,173 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 3 - AlphaIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseOperating System :: POSIX :: LinuxProgramming Language :: C++Programming Language :: PythonTopic :: Text Processing :: General |
Evidence: pytrec_eval_terrier-0.5.10-cp310-cp310-macosx_10_9_universal2.whl; pytrec_eval_terrier-0.5.10-cp310-cp310-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; pytrec_eval_terrier-0.5.10-cp310-cp310-musllinux_1_2_x86_64.whl; pytrec_eval_terrier-0.5.10-cp310-cp310-win_amd64.whl; pytrec_eval_terrier-0.5.10-cp311-cp311-macosx_10_9_universal2.whl; pytrec_eval_terrier-0.5.10-cp311-cp311-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; pytrec_eval_terrier-0.5.10-cp311-cp311-musllinux_1_2_x86_64.whl; pytrec_eval_terrier-0.5.10-cp311-cp311-win_amd64.whl; pytrec_eval_terrier-0.5.10-cp312-cp312-macosx_10_13_universal2.whl; pytrec_eval_terrier-0.5.10-cp312-cp312-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; pytrec_eval_terrier-0.5.10-cp312-cp312-musllinux_1_2_x86_64.whl; pytrec_eval_terrier-0.5.10-cp312-cp312-win_amd64.whl; pytrec_eval_terrier-0.5.10-cp313-cp313-macosx_10_13_universal2.whl; pytrec_eval_terrier-0.5.10-cp313-cp313-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; pytrec_eval_terrier-0.5.10-cp313-cp313-musllinux_1_2_x86_64.whl; pytrec_eval_terrier-0.5.10-cp313-cp313-win_amd64.whl; pytrec_eval_terrier-0.5.10-cp314-cp314-macosx_10_15_universal2.whl; pytrec_eval_terrier-0.5.10-cp314-cp314-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl; pytrec_eval_terrier-0.5.10-cp314-cp314-musllinux_1_2_x86_64.whl; pytrec_eval_terrier-0.5.10-cp314-cp314t-manylinux_2_24_x86_64.manylinux_2_28_x86_64.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “search result evaluation”
- pytrec-eval-terrierProvides Python bindings to TREC's trec_eval tool for computing…
- ir-measuresProvides a unified Python interface to compute standard information…
- pytrec-evalProvides Python bindings to compute standard Information Retrieval…
Give your agent the search over MCP, or paste the wish link into any chat.
More General packages
A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.
Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.
Sphinx generates professional documentation from reStructuredText source files, producing HTML, PDF, EPUB, and other formats with automatic cross-references, code highlighting, and hierarchical navigation.
Lark is a parsing library that builds abstract syntax trees from context-free grammars, supporting multiple parsing algorithms (Earley, LALR(1), CYK) with automatic line and column tracking.
NLTK is a Python library for natural language processing tasks including tokenization, parsing, tagging, and linguistic analysis, with built-in datasets and educational resources.
Install it if you need foundational NLP tools, linguistic datasets, or are learning the field; consider specialized libraries (spaCy, transformers) if you need…
Converts numbers, dates, times, and file sizes into human-readable text formats, with support for fuzzy durations like "3 minutes ago" and localization to multiple languages.
Install it if you need to display human-readable numbers, durations, or sizes to end users.
See also pytrec-eval · trec-car-tools · ir-measures · ranx · mir-eval · ir-datasets · colbert-ai · py-expression-eval · deepeval · nemo-evaluator