openevals
Open-source evaluators for LLM applications
Decision gist · record as of 2026-08-14
Yes. OpenEvals is actively maintained, has low install friction, carries a permissive MIT license, and addresses a real need in LLM application development. It provides both prebuilt evaluation patterns and extensibility for custom evals. The four runtime dependencies are all standard LLM/LangChain ecosystem packages. No security vulnerabilities are known.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires OPENAI_API_KEY environment variable set and Python 3.10 or later.
- Low install friction with a pure Python wheel and four runtime dependencies (langchain, langchain-openai, langsmith, rich).
- Package is actively maintained with recent releases.
License · maintenance · safety
MIT (permissive) — MIT license is permissive, allowing use in commercial and proprietary projects with minimal restrictions.
last release 2026-04-07 (129 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,251,240 downloads/mo, #4,160 on PyPI
Alternatives
Verify before relying
pip install openevals
from openevals.llm import create_llm_as_judge
from openevals.prompts import CONCISENESS_PROMPT
evaluator = create_llm_as_judge(
prompt=CONCISENESS_PROMPT,
model="openai:gpt-5.4",
)
result = evaluator(inputs="How is the weather?", outputs="It is sunny.")- Whether prebuilt prompts cover all common evaluation use cases or if custom prompt writing is frequently required.
- Performance characteristics when evaluating large batches of outputs or with different LLM providers.
- Maturity and stability of the multimodal and sandboxed code evaluation features.
What it is and what it does
OpenEvals is a framework for evaluating LLM application outputs using a variety of evaluation strategies. It centers on LLM-as-judge evaluators, which use another LLM to score outputs against custom or prebuilt prompts, but also includes deterministic evaluators for code quality, exact matching, embedding similarity, and agent trajectory validation. The package integrates with LangChain for model access and LangSmith for logging and tracking evaluation results.
The package is designed as a starting point for building custom evaluations specific to your application. It provides prebuilt prompts for common scenarios (correctness, safety, security, RAG quality, code evaluation) and allows flexible customization of scoring, output schemas, and evaluation criteria. It supports async evaluation, multimodal inputs, and multiturn simulation for testing conversational systems.
Use it for
- Score LLM outputs for quality dimensions like conciseness, correctness, or safety using an LLM-as-judge with prebuilt or custom prompts.
- Evaluate RAG system components (retrieval relevance, groundedness, helpfulness) to measure retrieval and generation quality.
- Validate structured outputs and tool calls from LLM applications using exact-match or LLM-as-judge evaluation.
- Test code generation outputs by extracting and type-checking generated code with Pyright or Mypy.
- Assess agent behavior by matching execution trajectories against expected tool call sequences.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
OpenEvals is actively maintained, has low install friction, carries a permissive MIT license, and addresses a real need in LLM application development. It provides both prebuilt evaluation patterns and extensibility for custom evals. The four runtime dependencies are all standard LLM/LangChain ecosystem packages. No security vulnerabilities are known.
Install
openevals on PyPI
Before you install
Low install friction with a pure Python wheel and four runtime dependencies (langchain, langchain-openai, langsmith, rich). Package is actively maintained with recent releases.
Requires OPENAI_API_KEY environment variable set and Python 3.10 or later.
License in practice
MIT license is permissive, allowing use in commercial and proprietary projects with minimal restrictions.
Quickstart
pip install openevals
from openevals.llm import create_llm_as_judge
from openevals.prompts import CONCISENESS_PROMPT
evaluator = create_llm_as_judge(
prompt=CONCISENESS_PROMPT,
model="openai:gpt-5.4",
)
result = evaluator(inputs="How is the weather?", outputs="It is sunny.")
Verify before relying
- Whether prebuilt prompts cover all common evaluation use cases or if custom prompt writing is frequently required.
- Performance characteristics when evaluating large batches of outputs or with different LLM providers.
- Maturity and stability of the multimodal and sandboxed code evaluation features.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 4 packageslangchainlangchain-openailangsmithrich |
| Maintenance | Actively maintained 129 days since the last release |
| First released | |
| Downloads | 1,251,240 / month, #4,160 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: openevals-0.2.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “LLM evaluation framework”
- openevalsOpenEvals provides a framework for writing and running evaluators to…
- ragasRagas provides objective metrics, test data generation, and…
- rubricRubric is a Python library for evaluating LLM outputs against…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also agentevals · autoevals · deepeval · rubric · pydantic-evals · arize-phoenix-evals · strands-agents-evals · ragas · nvidia-lm-eval · sybil-extras