$npx skillfedfor your agent

strands-agents-evals

Evaluation framework for Strands

With conditionsPyPI TestingReleased Aug 2026155.1K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — strands_agents_evals-1.1.1-py3-none-any.whl
v1.1.1 · released 2026-08-12 · Python >=3.10 · 10 runtime deps: boto3, opentelemetry-api, opentelemetry-instrumentation-threading, opentelemetry-sdk, pydantic, rich, strands-agents-tools, strands-agents

Yes, if you are actively developing or evaluating AI agents and LLM applications. The framework is actively maintained, has low install friction, carries no security vulnerabilities, and offers a broad evaluation toolkit covering output, trajectory, trace, and adversarial testing. Install with caution if strands-agents and strands-agents-tools are not yet available in your environment.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; depends on strands-agents and strands-agents-tools packages being available.
  • Low install friction with a pure Python wheel distribution.
  • Active maintenance with a release 2 days old.

License · maintenance · safety

Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions—suitable for commercial and open-source projects.

last release 2026-08-12 (2 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 155,077 downloads/mo, #10,832 on PyPI

Verify before relying

pip install strands-agents-evals

from strands_agents_evals import Case, Experiment
from strands_agents_evals.evaluators import OutputEvaluator

test_cases = [Case(name="test-1", input="What is 2+2?", expected_output="4")]
evaluators = [OutputEvaluator(rubric="Score 1.0 if correct.")]
experiment = Experiment(cases=test_cases, evaluators=evaluators)
report = experiment.run_evaluations(lambda case: "4")
  • Whether strands-agents and strands-agents-tools are publicly available or proprietary internal packages.
  • Specific LLM models supported by OutputEvaluator and other judge-based evaluators.
  • Performance characteristics and scalability limits for large experiment suites.
  • Whether the CLI subcommands are fully documented and stable.
Same gist for agents: .md · .json

What it is and what it does

Strands Evals SDK is a Python framework for systematically evaluating AI agents and language model applications. It provides multiple evaluation modes—output validation with custom rubrics, trajectory analysis of tool usage, trace-based assessment via OpenTelemetry, and automated test generation—allowing developers to measure agent correctness, safety, and behavior across complex interactions. The framework includes built-in LLM-as-a-judge evaluators, multimodal evaluation support, dynamic conversation simulators, failure detection with root-cause analysis, and chaos testing via fault injection. Experiments can be serialized to JSON, versioned, and executed via Python API or CLI.

The package depends on boto3, the OpenTelemetry stack (opentelemetry-api, opentelemetry-sdk, opentelemetry-instrumentation-threading), pydantic, rich, tenacity, typing-extensions, and two internal Strands packages (strands-agents and strands-agents-tools). It targets Python 3.10+ and is actively maintained, with a permissive Apache-2.0 license suitable for both commercial and open-source use.

Use it for

  • Validate LLM output quality and factual accuracy using custom rubrics and LLM judges.
  • Analyze agent tool usage sequences and verify correct action ordering in multi-step tasks.
  • Detect and diagnose failures in agent sessions with automated root-cause analysis.
  • Generate comprehensive test suites from high-level tool or task descriptions.
  • Simulate multi-turn conversations with realistic user behavior to stress-test agent resilience.
  • Perform adversarial safety testing with built-in attack strategies.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you are actively developing or evaluating AI agents and LLM applications.

The framework is actively maintained, has low install friction, carries no security vulnerabilities, and offers a broad evaluation toolkit covering output, trajectory, trace, and adversarial testing. Install with caution if strands-agents and strands-agents-tools are not yet available in your environment.

Install

strands-agents-evals on PyPI

Before you install

Low install friction with a pure Python wheel distribution. Active maintenance with a release 2 days old. Depends on well-established libraries (boto3, pydantic, rich, opentelemetry stack, tenacity) and two internal Strands packages.

Requires Python 3.10 or later; depends on strands-agents and strands-agents-tools packages being available.

License in practice

Licensed under Apache-2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions—suitable for commercial and open-source projects.

Quickstart

pip install strands-agents-evals

from strands_agents_evals import Case, Experiment
from strands_agents_evals.evaluators import OutputEvaluator

test_cases = [Case(name="test-1", input="What is 2+2?", expected_output="4")]
evaluators = [OutputEvaluator(rubric="Score 1.0 if correct.")]
experiment = Experiment(cases=test_cases, evaluators=evaluators)
report = experiment.run_evaluations(lambda case: "4")

Verify before relying

  • Whether strands-agents and strands-agents-tools are publicly available or proprietary internal packages.
  • Specific LLM models supported by OutputEvaluator and other judge-based evaluators.
  • Performance characteristics and scalability limits for large experiment suites.
  • Whether the CLI subcommands are fully documented and stable.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
10 packages
boto3opentelemetry-apiopentelemetry-instrumentation-threadingopentelemetry-sdkpydanticrichstrands-agents-toolsstrands-agentstenacitytyping-extensions
MaintenanceActively maintained 2 days since the last release
First released
Downloads155,077 / month, #10,832 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14

Evidence: strands_agents_evals-1.1.1-py3-none-any.whl

Tags

Capabilities
LLM evaluation frameworkAI agent testingLLM-as-a-judge evaluationtrajectory analysis for agentsexperiment generation for AIagent behavior assessmentOpenTelemetry trace evaluation
Topics
llm-evaluationagent-testingai-quality-assurance

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “trajectory analysis for agents”

  • strands-agents-evalsStrands Evals SDK provides a comprehensive evaluation framework for…
  • MDAnalysisMDAnalysis reads and analyzes molecular dynamics simulation…
  • mdtrajMDTraj reads, writes, and analyzes molecular dynamics trajectories in…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also agentevals · azure-ai-evaluation · dreadnode · openevals · arize-phoenix-evals · autoevals · strands-agents · strands-agents-builder · strands-agents-tools · nemo-evaluator

Further reading