strands-agents-evals
Evaluation framework for Strands
Decision gist · record as of 2026-08-14
Yes, if you are actively developing or evaluating AI agents and LLM applications. The framework is actively maintained, has low install friction, carries no security vulnerabilities, and offers a broad evaluation toolkit covering output, trajectory, trace, and adversarial testing. Install with caution if strands-agents and strands-agents-tools are not yet available in your environment.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.10 or later; depends on strands-agents and strands-agents-tools packages being available.
- Low install friction with a pure Python wheel distribution.
- Active maintenance with a release 2 days old.
License · maintenance · safety
Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions—suitable for commercial and open-source projects.
last release 2026-08-12 (2 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 155,077 downloads/mo, #10,832 on PyPI
Alternatives
Verify before relying
pip install strands-agents-evals
from strands_agents_evals import Case, Experiment
from strands_agents_evals.evaluators import OutputEvaluator
test_cases = [Case(name="test-1", input="What is 2+2?", expected_output="4")]
evaluators = [OutputEvaluator(rubric="Score 1.0 if correct.")]
experiment = Experiment(cases=test_cases, evaluators=evaluators)
report = experiment.run_evaluations(lambda case: "4")- Whether strands-agents and strands-agents-tools are publicly available or proprietary internal packages.
- Specific LLM models supported by OutputEvaluator and other judge-based evaluators.
- Performance characteristics and scalability limits for large experiment suites.
- Whether the CLI subcommands are fully documented and stable.
What it is and what it does
Strands Evals SDK is a Python framework for systematically evaluating AI agents and language model applications. It provides multiple evaluation modes—output validation with custom rubrics, trajectory analysis of tool usage, trace-based assessment via OpenTelemetry, and automated test generation—allowing developers to measure agent correctness, safety, and behavior across complex interactions. The framework includes built-in LLM-as-a-judge evaluators, multimodal evaluation support, dynamic conversation simulators, failure detection with root-cause analysis, and chaos testing via fault injection. Experiments can be serialized to JSON, versioned, and executed via Python API or CLI.
The package depends on boto3, the OpenTelemetry stack (opentelemetry-api, opentelemetry-sdk, opentelemetry-instrumentation-threading), pydantic, rich, tenacity, typing-extensions, and two internal Strands packages (strands-agents and strands-agents-tools). It targets Python 3.10+ and is actively maintained, with a permissive Apache-2.0 license suitable for both commercial and open-source use.
Use it for
- Validate LLM output quality and factual accuracy using custom rubrics and LLM judges.
- Analyze agent tool usage sequences and verify correct action ordering in multi-step tasks.
- Detect and diagnose failures in agent sessions with automated root-cause analysis.
- Generate comprehensive test suites from high-level tool or task descriptions.
- Simulate multi-turn conversations with realistic user behavior to stress-test agent resilience.
- Perform adversarial safety testing with built-in attack strategies.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are actively developing or evaluating AI agents and LLM applications.
The framework is actively maintained, has low install friction, carries no security vulnerabilities, and offers a broad evaluation toolkit covering output, trajectory, trace, and adversarial testing. Install with caution if strands-agents and strands-agents-tools are not yet available in your environment.
Install
strands-agents-evals on PyPI
Before you install
Low install friction with a pure Python wheel distribution. Active maintenance with a release 2 days old. Depends on well-established libraries (boto3, pydantic, rich, opentelemetry stack, tenacity) and two internal Strands packages.
Requires Python 3.10 or later; depends on strands-agents and strands-agents-tools packages being available.
License in practice
Licensed under Apache-2.0 (permissive), allowing free use, modification, and distribution with minimal restrictions—suitable for commercial and open-source projects.
Quickstart
pip install strands-agents-evals
from strands_agents_evals import Case, Experiment
from strands_agents_evals.evaluators import OutputEvaluator
test_cases = [Case(name="test-1", input="What is 2+2?", expected_output="4")]
evaluators = [OutputEvaluator(rubric="Score 1.0 if correct.")]
experiment = Experiment(cases=test_cases, evaluators=evaluators)
report = experiment.run_evaluations(lambda case: "4")
Verify before relying
- Whether strands-agents and strands-agents-tools are publicly available or proprietary internal packages.
- Specific LLM models supported by OutputEvaluator and other judge-based evaluators.
- Performance characteristics and scalability limits for large experiment suites.
- Whether the CLI subcommands are fully documented and stable.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 10 packagesboto3opentelemetry-apiopentelemetry-instrumentation-threadingopentelemetry-sdkpydanticrichstrands-agents-toolsstrands-agentstenacitytyping-extensions |
| Maintenance | Actively maintained 2 days since the last release |
| First released | |
| Downloads | 155,077 / month, #10,832 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14 |
Evidence: strands_agents_evals-1.1.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “trajectory analysis for agents”
- strands-agents-evalsStrands Evals SDK provides a comprehensive evaluation framework for…
- MDAnalysisMDAnalysis reads and analyzes molecular dynamics simulation…
- mdtrajMDTraj reads, writes, and analyzes molecular dynamics trajectories in…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also agentevals · azure-ai-evaluation · dreadnode · openevals · arize-phoenix-evals · autoevals · strands-agents · strands-agents-builder · strands-agents-tools · nemo-evaluator