$npx skillfedfor your agent

evalplus

"EvalPlus for rigourous evaluation of LLM-synthesized code"

With conditionsPyPI TestingReleased Oct 2024159.7K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — evalplus-0.3.1-py3-none-any.whl
v0.3.1 · released 2024-10-20 · Python >=3.9 · 18 runtime deps: wget, tempdir, multipledispatch, appdirs, numpy, tqdm, termcolor, fire

Yes, if you are evaluating or benchmarking LLM code-generation models. The extended test suites and multi-backend support make it the standard tool for rigorous code-generation assessment. The aging maintenance status (663 days since release) is a minor concern but not a blocker—the repository remains active and the framework is stable. Install friction is low and there are no known vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >= 3.9.
  • Code execution evaluation may require Docker for safety; EvalPerf (efficiency evaluation) is Unix-only and requires perf event access.
  • Low install friction with a pure-Python wheel.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 is permissive; you can use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.

last release 2024-10-20 (663 days) · last repo commit 2025-10-02 · 1,798 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 159,707 downloads/mo, #10,690 on PyPI

Verify before relying

pip install evalplus

# Evaluate a model on HumanEval+
evalplus.evaluate --model "ise-uiuc/Magicoder-S-DS-6.7B" \
                  --dataset humaneval \
                  --backend hf \
                  --greedy
  • Whether all 18 runtime dependencies are required for basic evaluation or only for specific backends (vllm, openai, anthropic, etc.)
  • Current status of the leaderboard and whether benchmark data is actively maintained post-v0.3.1
  • Compatibility of tree-sitter-python with recent Python versions and whether it adds significant install complexity
Same gist for agents: .md · .json

What it is and what it does

EvalPlus is a benchmarking and evaluation framework designed to rigorously test code generated by large language models. It extends two popular code-generation benchmarks—HumanEval and MBPP—with significantly more test cases (80x and 35x respectively) to catch edge cases and fragile implementations that simpler test suites miss. Beyond correctness, it includes EvalPerf, a dataset for measuring the efficiency and performance characteristics of generated code.

The framework supports multiple LLM backends (HuggingFace transformers, vLLM, OpenAI, Anthropic, Google Gemini) and can run end-to-end evaluation pipelines: code generation, post-processing, and test execution. It includes Docker-based execution for safe sandboxing of untrusted generated code and provides detailed ranking and scoring to help identify which models produce more robust, production-ready implementations.

Use it for

  • Benchmark and rank LLM code-generation models on standardized test suites with rigorous correctness criteria.
  • Evaluate the efficiency of LLM-generated code via performance-exercising tasks and measure runtime/memory characteristics.
  • Compare model robustness before and after using EvalPlus tests to identify fragile implementations that pass simple tests.
  • Run safe, isolated evaluation of untrusted model-generated code using Docker sandboxing.
  • Integrate code-generation evaluation into CI/CD pipelines for continuous model quality monitoring.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are evaluating or benchmarking LLM code-generation models.

The extended test suites and multi-backend support make it the standard tool for rigorous code-generation assessment. The aging maintenance status (663 days since release) is a minor concern but not a blocker—the repository remains active and the framework is stable. Install friction is low and there are no known vulnerabilities.

Install

evalplus on PyPI

Before you install

Low install friction with a pure-Python wheel. Maintenance status is aging (663 days since last release), though the repository remains active with recent commits and moderate community engagement (1798 stars).

Requires Python >= 3.9. Code execution evaluation may require Docker for safety; EvalPerf (efficiency evaluation) is Unix-only and requires perf event access.

License in practice

Apache-2.0 is permissive; you can use, modify, and distribute this package freely in commercial and private projects with minimal restrictions.

Quickstart

pip install evalplus

# Evaluate a model on HumanEval+
evalplus.evaluate --model "ise-uiuc/Magicoder-S-DS-6.7B" \
                  --dataset humaneval \
                  --backend hf \
                  --greedy

Verify before relying

  • Whether all 18 runtime dependencies are required for basic evaluation or only for specific backends (vllm, openai, anthropic, etc.)
  • Current status of the leaderboard and whether benchmark data is actively maintained post-v0.3.1
  • Compatibility of tree-sitter-python with recent Python versions and whether it adds significant install complexity

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
18 packages
wgettempdirmultipledispatchappdirsnumpytqdmtermcolorfireopenaitree-sittertree-sitter-pythonrichtransformersstop-sequenceranthropicgoogle-generativeaidatasetspsutil
MaintenanceAging 663 days since the last release
Last repo commit
First released
Downloads159,707 / month, #10,690 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3

Evidence: evalplus-0.3.1-py3-none-any.whl

Tags

Capabilities
llm code generation evaluationhumaneval+ mbpp+ benchmarkcode correctness testing frameworkllm4code evaluation harnesscode efficiency performance testing
Topics
llm-evaluationcode-benchmarkingml-ops

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “llm code generation evaluation”

  • evalplusEvalPlus provides a rigorous evaluation framework for code generated…
  • openevalsOpenEvals provides a framework for writing and running evaluators to…
  • swebenchSWE-bench is a benchmark framework for evaluating language models on…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also lm-eval · deepeval · inspect-evals · arize-phoenix-evals · bfcl-eval · autoevals · vllm · vllm-tpu · openevals · unitxt

Further reading