agentevals
Open-source evaluators for LLM agents
Decision gist · record as of 2026-08-14
Yes, if you are actively building and testing agentic LLM applications and need to validate agent behavior at the trajectory level. The low install friction and permissive license make adoption easy. However, the aging maintenance status (386 days since last release) means you should verify that it remains compatible with your LLM client versions and check whether the repository is still actively maintained before committing to it for critical production workflows.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires OPENAI_API_KEY environment variable set and an LLM client (langchain_openai comes by default, or install openai directly).
- Low friction installation with a single runtime dependency (openevals).
- Maintenance status is aging—last release was 386 days ago—so expect infrequent updates and consider checking the repository for active development before relying on it for production evals.
License · maintenance · safety
MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects without licensing concerns.
last release 2025-07-24 (386 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 374,429 downloads/mo, #7,145 on PyPI
Alternatives
Verify before relying
pip install agentevals
from agentevals.trajectory.llm import create_trajectory_llm_as_judge, TRAJECTORY_ACCURACY_PROMPT
trajectory_evaluator = create_trajectory_llm_as_judge(
prompt=TRAJECTORY_ACCURACY_PROMPT,
model="openai:o3-mini",
)
result = trajectory_evaluator(outputs=[
{"role": "user", "content": "What is the weather in SF?"},
{"role": "assistant", "content": "", "tool_calls": [{"function": {"name": "get_weather", "arguments": "{\"city\": \"SF\"}"}}]},
{"role": "tool", "content": "It's 80 degrees and sunny in SF."},
{"role": "assistant", "content": "The weather in SF is 80 degrees and sunny."}
])- Whether the package actively maintains compatibility with current LangChain and OpenAI API versions
- Performance characteristics when evaluating large agent trajectories or high-volume eval runs
- Availability of community support or whether issues are actively triaged
What it is and what it does
AgentEvals is a toolkit for evaluating agentic LLM applications by inspecting the sequence of steps (trajectory) an agent takes to solve a problem. It addresses the challenge of understanding how changes to an agent's configuration or tools affect its behavior downstream, which is difficult because LLMs operate as black boxes. The package provides evaluators that compare actual agent trajectories against expected ones using strict, unordered, or subset/superset matching modes, as well as LLM-as-judge evaluators that use an LLM to reason about trajectory quality.
The package depends on openevals for general evaluation utilities and integrates with LangChain chat models and LangSmith for test orchestration. It represents trajectories as lists of OpenAI-format message dicts or LangChain BaseMessage objects, and supports both synchronous and asynchronous evaluation. The quickstart example shows using an LLM-as-judge to evaluate whether an agent's sequence of tool calls and responses logically solves a user query.
Use it for
- Validate that an agent calls tools in the correct order and with expected arguments before deploying to production
- Compare agent behavior across model versions or configuration changes to detect regressions in reasoning
- Use an LLM to judge whether an agent's intermediate steps are reasonable and efficient for a given task
- Integrate agent evals into a pytest or LangSmith workflow to catch trajectory issues in CI/CD
- Debug why an agent takes unexpected paths by examining and comparing actual vs. reference trajectories
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are actively building and testing agentic LLM applications and need to validate agent behavior at the trajectory level.
The low install friction and permissive license make adoption easy. However, the aging maintenance status (386 days since last release) means you should verify that it remains compatible with your LLM client versions and check whether the repository is still actively maintained before committing to it for critical production workflows.
Install
agentevals on PyPI
Before you install
Low friction installation with a single runtime dependency (openevals). Maintenance status is aging—last release was 386 days ago—so expect infrequent updates and consider checking the repository for active development before relying on it for production evals.
Requires OPENAI_API_KEY environment variable set and an LLM client (langchain_openai comes by default, or install openai directly).
License in practice
MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects without licensing concerns.
Quickstart
pip install agentevals
from agentevals.trajectory.llm import create_trajectory_llm_as_judge, TRAJECTORY_ACCURACY_PROMPT
trajectory_evaluator = create_trajectory_llm_as_judge(
prompt=TRAJECTORY_ACCURACY_PROMPT,
model="openai:o3-mini",
)
result = trajectory_evaluator(outputs=[
{"role": "user", "content": "What is the weather in SF?"},
{"role": "assistant", "content": "", "tool_calls": [{"function": {"name": "get_weather", "arguments": "{\"city\": \"SF\"}"}}]},
{"role": "tool", "content": "It's 80 degrees and sunny in SF."},
{"role": "assistant", "content": "The weather in SF is 80 degrees and sunny."}
])
Verify before relying
- Whether the package actively maintains compatibility with current LangChain and OpenAI API versions
- Performance characteristics when evaluating large agent trajectories or high-volume eval runs
- Availability of community support or whether issues are actively triaged
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packageopenevals |
| Maintenance | Aging 386 days since the last release |
| First released | |
| Downloads | 374,429 / month, #7,145 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: agentevals-0.0.9-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “agent trajectory evaluation”
- agentevalsProvides evaluators and utilities to assess agent trajectories—the…
- strands-agents-evalsStrands Evals SDK provides a comprehensive evaluation framework for…
- evoevo evaluates and compares trajectory output from odometry and SLAM…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also agent-lifecycle-toolkit · langwatch-scenario · openevals · strands-agents-evals · arize-phoenix-evals · azure-ai-evaluation · autoevals · pydantic-evals · deepeval · evo