evalplus
"EvalPlus for rigourous evaluation of LLM-synthesized code"
Decision gist · record as of 2026-08-14
Yes, if you are evaluating or benchmarking LLM code-generation models. The extended test suites and multi-backend support make it the standard tool for rigorous code-generation assessment. The aging maintenance status (663 days since release) is a minor concern but not a blocker—the repository remains active and the framework is stable. Install friction is low and there are no known vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >= 3.9.
- Code execution evaluation may require Docker for safety; EvalPerf (efficiency evaluation) is Unix-only and requires perf event access.
- Low install friction with a pure-Python wheel.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 is permissive; you can use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.
last release 2024-10-20 (663 days) · last repo commit 2025-10-02 · 1,798 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 159,707 downloads/mo, #10,690 on PyPI
Alternatives
Verify before relying
pip install evalplus
# Evaluate a model on HumanEval+
evalplus.evaluate --model "ise-uiuc/Magicoder-S-DS-6.7B" \
--dataset humaneval \
--backend hf \
--greedy- Whether all 18 runtime dependencies are required for basic evaluation or only for specific backends (vllm, openai, anthropic, etc.)
- Current status of the leaderboard and whether benchmark data is actively maintained post-v0.3.1
- Compatibility of tree-sitter-python with recent Python versions and whether it adds significant install complexity
What it is and what it does
EvalPlus is a benchmarking and evaluation framework designed to rigorously test code generated by large language models. It extends two popular code-generation benchmarks—HumanEval and MBPP—with significantly more test cases (80x and 35x respectively) to catch edge cases and fragile implementations that simpler test suites miss. Beyond correctness, it includes EvalPerf, a dataset for measuring the efficiency and performance characteristics of generated code.
The framework supports multiple LLM backends (HuggingFace transformers, vLLM, OpenAI, Anthropic, Google Gemini) and can run end-to-end evaluation pipelines: code generation, post-processing, and test execution. It includes Docker-based execution for safe sandboxing of untrusted generated code and provides detailed ranking and scoring to help identify which models produce more robust, production-ready implementations.
Use it for
- Benchmark and rank LLM code-generation models on standardized test suites with rigorous correctness criteria.
- Evaluate the efficiency of LLM-generated code via performance-exercising tasks and measure runtime/memory characteristics.
- Compare model robustness before and after using EvalPlus tests to identify fragile implementations that pass simple tests.
- Run safe, isolated evaluation of untrusted model-generated code using Docker sandboxing.
- Integrate code-generation evaluation into CI/CD pipelines for continuous model quality monitoring.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are evaluating or benchmarking LLM code-generation models.
The extended test suites and multi-backend support make it the standard tool for rigorous code-generation assessment. The aging maintenance status (663 days since release) is a minor concern but not a blocker—the repository remains active and the framework is stable. Install friction is low and there are no known vulnerabilities.
Install
evalplus on PyPI
Before you install
Low install friction with a pure-Python wheel. Maintenance status is aging (663 days since last release), though the repository remains active with recent commits and moderate community engagement (1798 stars).
Requires Python >= 3.9. Code execution evaluation may require Docker for safety; EvalPerf (efficiency evaluation) is Unix-only and requires perf event access.
License in practice
Apache-2.0 is permissive; you can use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.
Quickstart
pip install evalplus
# Evaluate a model on HumanEval+
evalplus.evaluate --model "ise-uiuc/Magicoder-S-DS-6.7B" \
--dataset humaneval \
--backend hf \
--greedy
Verify before relying
- Whether all 18 runtime dependencies are required for basic evaluation or only for specific backends (vllm, openai, anthropic, etc.)
- Current status of the leaderboard and whether benchmark data is actively maintained post-v0.3.1
- Compatibility of tree-sitter-python with recent Python versions and whether it adds significant install complexity
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 18 packageswgettempdirmultipledispatchappdirsnumpytqdmtermcolorfireopenaitree-sittertree-sitter-pythonrichtransformersstop-sequenceranthropicgoogle-generativeaidatasetspsutil |
| Maintenance | Aging 663 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 159,707 / month, #10,690 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: evalplus-0.3.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “llm code generation evaluation”
- evalplusEvalPlus provides a rigorous evaluation framework for code generated…
- openevalsOpenEvals provides a framework for writing and running evaluators to…
- swebenchSWE-bench is a benchmark framework for evaluating language models on…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also lm-eval · deepeval · inspect-evals · arize-phoenix-evals · bfcl-eval · autoevals · vllm · vllm-tpu · openevals · unitxt