$npx skillfedfor your agent

openevals

Open-source evaluators for LLM applications

Worth itPyPI TestingReleased Apr 20261.3M downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — openevals-0.2.0-py3-none-any.whl
v0.2.0 · released 2026-04-07 · Python >=3.10 · 4 runtime deps: langchain, langchain-openai, langsmith, rich

Yes. OpenEvals is actively maintained, has low install friction, carries a permissive MIT license, and addresses a real need in LLM application development. It provides both prebuilt evaluation patterns and extensibility for custom evals. The four runtime dependencies are all standard LLM/LangChain ecosystem packages. No security vulnerabilities are known.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires OPENAI_API_KEY environment variable set and Python 3.10 or later.
  • Low install friction with a pure Python wheel and four runtime dependencies (langchain, langchain-openai, langsmith, rich).
  • Package is actively maintained with recent releases.

License · maintenance · safety

MIT (permissive) — MIT license is permissive, allowing use in commercial and proprietary projects with minimal restrictions.

last release 2026-04-07 (129 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,251,240 downloads/mo, #4,160 on PyPI

Verify before relying

pip install openevals

from openevals.llm import create_llm_as_judge
from openevals.prompts import CONCISENESS_PROMPT

evaluator = create_llm_as_judge(
    prompt=CONCISENESS_PROMPT,
    model="openai:gpt-5.4",
)
result = evaluator(inputs="How is the weather?", outputs="It is sunny.")
  • Whether prebuilt prompts cover all common evaluation use cases or if custom prompt writing is frequently required.
  • Performance characteristics when evaluating large batches of outputs or with different LLM providers.
  • Maturity and stability of the multimodal and sandboxed code evaluation features.
Same gist for agents: .md · .json

What it is and what it does

OpenEvals is a framework for evaluating LLM application outputs using a variety of evaluation strategies. It centers on LLM-as-judge evaluators, which use another LLM to score outputs against custom or prebuilt prompts, but also includes deterministic evaluators for code quality, exact matching, embedding similarity, and agent trajectory validation. The package integrates with LangChain for model access and LangSmith for logging and tracking evaluation results.

The package is designed as a starting point for building custom evaluations specific to your application. It provides prebuilt prompts for common scenarios (correctness, safety, security, RAG quality, code evaluation) and allows flexible customization of scoring, output schemas, and evaluation criteria. It supports async evaluation, multimodal inputs, and multiturn simulation for testing conversational systems.

Use it for

  • Score LLM outputs for quality dimensions like conciseness, correctness, or safety using an LLM-as-judge with prebuilt or custom prompts.
  • Evaluate RAG system components (retrieval relevance, groundedness, helpfulness) to measure retrieval and generation quality.
  • Validate structured outputs and tool calls from LLM applications using exact-match or LLM-as-judge evaluation.
  • Test code generation outputs by extracting and type-checking generated code with Pyright or Mypy.
  • Assess agent behavior by matching execution trajectories against expected tool call sequences.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

OpenEvals is actively maintained, has low install friction, carries a permissive MIT license, and addresses a real need in LLM application development. It provides both prebuilt evaluation patterns and extensibility for custom evals. The four runtime dependencies are all standard LLM/LangChain ecosystem packages. No security vulnerabilities are known.

Install

openevals on PyPI

Before you install

Low install friction with a pure Python wheel and four runtime dependencies (langchain, langchain-openai, langsmith, rich). Package is actively maintained with recent releases.

Requires OPENAI_API_KEY environment variable set and Python 3.10 or later.

License in practice

MIT license is permissive, allowing use in commercial and proprietary projects with minimal restrictions.

Quickstart

pip install openevals

from openevals.llm import create_llm_as_judge
from openevals.prompts import CONCISENESS_PROMPT

evaluator = create_llm_as_judge(
    prompt=CONCISENESS_PROMPT,
    model="openai:gpt-5.4",
)
result = evaluator(inputs="How is the weather?", outputs="It is sunny.")

Verify before relying

  • Whether prebuilt prompts cover all common evaluation use cases or if custom prompt writing is frequently required.
  • Performance characteristics when evaluating large batches of outputs or with different LLM providers.
  • Maturity and stability of the multimodal and sandboxed code evaluation features.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
langchainlangchain-openailangsmithrich
MaintenanceActively maintained 129 days since the last release
First released
Downloads1,251,240 / month, #4,160 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: openevals-0.2.0-py3-none-any.whl

Tags

Capabilities
LLM evaluation frameworkLLM-as-judge evaluatorLLM application testingeval framework for language modelsLLM output assessmentevaluation prompts for LLMLLM quality assurance
Topics
llm-evaluationquality-assurancelangchain-integration

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “LLM evaluation framework”

  • openevalsOpenEvals provides a framework for writing and running evaluators to…
  • ragasRagas provides objective metrics, test data generation, and…
  • rubricRubric is a Python library for evaluating LLM outputs against…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also agentevals · autoevals · deepeval · rubric · pydantic-evals · arize-phoenix-evals · strands-agents-evals · ragas · nvidia-lm-eval · sybil-extras

Further reading