$npx skillfedfor your agent

agentevals

Open-source evaluators for LLM agents

With conditionsPyPI TestingReleased Jul 2025374.4K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — agentevals-0.0.9-py3-none-any.whl
v0.0.9 · released 2025-07-24 · Python >=3.9 · 1 runtime deps: openevals

Yes, if you are actively building and testing agentic LLM applications and need to validate agent behavior at the trajectory level. The low install friction and permissive license make adoption easy. However, the aging maintenance status (386 days since last release) means you should verify that it remains compatible with your LLM client versions and check whether the repository is still actively maintained before committing to it for critical production workflows.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires OPENAI_API_KEY environment variable set and an LLM client (langchain_openai comes by default, or install openai directly).
  • Low friction installation with a single runtime dependency (openevals).
  • Maintenance status is aging—last release was 386 days ago—so expect infrequent updates and consider checking the repository for active development before relying on it for production evals.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects without licensing concerns.

last release 2025-07-24 (386 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 374,429 downloads/mo, #7,145 on PyPI

Verify before relying

pip install agentevals

from agentevals.trajectory.llm import create_trajectory_llm_as_judge, TRAJECTORY_ACCURACY_PROMPT

trajectory_evaluator = create_trajectory_llm_as_judge(
    prompt=TRAJECTORY_ACCURACY_PROMPT,
    model="openai:o3-mini",
)

result = trajectory_evaluator(outputs=[
    {"role": "user", "content": "What is the weather in SF?"},
    {"role": "assistant", "content": "", "tool_calls": [{"function": {"name": "get_weather", "arguments": "{\"city\": \"SF\"}"}}]},
    {"role": "tool", "content": "It's 80 degrees and sunny in SF."},
    {"role": "assistant", "content": "The weather in SF is 80 degrees and sunny."}
])
  • Whether the package actively maintains compatibility with current LangChain and OpenAI API versions
  • Performance characteristics when evaluating large agent trajectories or high-volume eval runs
  • Availability of community support or whether issues are actively triaged
Same gist for agents: .md · .json

What it is and what it does

AgentEvals is a toolkit for evaluating agentic LLM applications by inspecting the sequence of steps (trajectory) an agent takes to solve a problem. It addresses the challenge of understanding how changes to an agent's configuration or tools affect its behavior downstream, which is difficult because LLMs operate as black boxes. The package provides evaluators that compare actual agent trajectories against expected ones using strict, unordered, or subset/superset matching modes, as well as LLM-as-judge evaluators that use an LLM to reason about trajectory quality.

The package depends on openevals for general evaluation utilities and integrates with LangChain chat models and LangSmith for test orchestration. It represents trajectories as lists of OpenAI-format message dicts or LangChain BaseMessage objects, and supports both synchronous and asynchronous evaluation. The quickstart example shows using an LLM-as-judge to evaluate whether an agent's sequence of tool calls and responses logically solves a user query.

Use it for

  • Validate that an agent calls tools in the correct order and with expected arguments before deploying to production
  • Compare agent behavior across model versions or configuration changes to detect regressions in reasoning
  • Use an LLM to judge whether an agent's intermediate steps are reasonable and efficient for a given task
  • Integrate agent evals into a pytest or LangSmith workflow to catch trajectory issues in CI/CD
  • Debug why an agent takes unexpected paths by examining and comparing actual vs. reference trajectories

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are actively building and testing agentic LLM applications and need to validate agent behavior at the trajectory level.

The low install friction and permissive license make adoption easy. However, the aging maintenance status (386 days since last release) means you should verify that it remains compatible with your LLM client versions and check whether the repository is still actively maintained before committing to it for critical production workflows.

Install

agentevals on PyPI

Before you install

Low friction installation with a single runtime dependency (openevals). Maintenance status is aging—last release was 386 days ago—so expect infrequent updates and consider checking the repository for active development before relying on it for production evals.

Requires OPENAI_API_KEY environment variable set and an LLM client (langchain_openai comes by default, or install openai directly).

License in practice

MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects without licensing concerns.

Quickstart

pip install agentevals

from agentevals.trajectory.llm import create_trajectory_llm_as_judge, TRAJECTORY_ACCURACY_PROMPT

trajectory_evaluator = create_trajectory_llm_as_judge(
    prompt=TRAJECTORY_ACCURACY_PROMPT,
    model="openai:o3-mini",
)

result = trajectory_evaluator(outputs=[
    {"role": "user", "content": "What is the weather in SF?"},
    {"role": "assistant", "content": "", "tool_calls": [{"function": {"name": "get_weather", "arguments": "{\"city\": \"SF\"}"}}]},
    {"role": "tool", "content": "It's 80 degrees and sunny in SF."},
    {"role": "assistant", "content": "The weather in SF is 80 degrees and sunny."}
])

Verify before relying

  • Whether the package actively maintains compatibility with current LangChain and OpenAI API versions
  • Performance characteristics when evaluating large agent trajectories or high-volume eval runs
  • Availability of community support or whether issues are actively triaged

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
1 package
openevals
MaintenanceAging 386 days since the last release
First released
Downloads374,429 / month, #7,145 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: agentevals-0.0.9-py3-none-any.whl

Tags

Capabilities
agent trajectory evaluationLLM agent testingagentic application evalsagent step validationtool call sequence matchingagent behavior assessmentagentic workflow evaluation
Topics
agent-evaluationllm-testingtrajectory-analysis

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “agent trajectory evaluation”

  • agentevalsProvides evaluators and utilities to assess agent trajectories—the…
  • strands-agents-evalsStrands Evals SDK provides a comprehensive evaluation framework for…
  • evoevo evaluates and compares trajectory output from odometry and SLAM…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also agent-lifecycle-toolkit · langwatch-scenario · openevals · strands-agents-evals · arize-phoenix-evals · azure-ai-evaluation · autoevals · pydantic-evals · deepeval · evo

Further reading