{"categories":[{"label":"Testing","url":"https://skillfed.io/packages/category/software-development-testing/5"}],"enrichment":{"capability":"Strands Evals SDK provides a comprehensive evaluation framework for testing and measuring AI agents and LLM applications, supporting output validation, trajectory analysis, trace-based evaluation, and automated experiment generation.","skillfed_tags":["llm-evaluation","agent-testing","ai-quality-assurance"],"use_cases":["Validate LLM output quality and factual accuracy using custom rubrics and LLM judges.","Analyze agent tool usage sequences and verify correct action ordering in multi-step tasks.","Detect and diagnose failures in agent sessions with automated root-cause analysis.","Generate comprehensive test suites from high-level tool or task descriptions.","Simulate multi-turn conversations with realistic user behavior to stress-test agent resilience.","Perform adversarial safety testing with built-in attack strategies."],"what_it_does":"Strands Evals SDK is a Python framework for systematically evaluating AI agents and language model applications. It provides multiple evaluation modes\u2014output validation with custom rubrics, trajectory analysis of tool usage, trace-based assessment via OpenTelemetry, and automated test generation\u2014allowing developers to measure agent correctness, safety, and behavior across complex interactions. The framework includes built-in LLM-as-a-judge evaluators, multimodal evaluation support, dynamic conversation simulators, failure detection with root-cause analysis, and chaos testing via fault injection. Experiments can be serialized to JSON, versioned, and executed via Python API or CLI.\n\nThe package depends on boto3, the OpenTelemetry stack (opentelemetry-api, opentelemetry-sdk, opentelemetry-instrumentation-threading), pydantic, rich, tenacity, typing-extensions, and two internal Strands packages (strands-agents and strands-agents-tools). It targets Python 3.10+ and is actively maintained, with a permissive Apache-2.0 license suitable for both commercial and open-source use.","worth_installing":"Yes, if you are actively developing or evaluating AI agents and LLM applications. The framework is actively maintained, has low install friction, carries no security vulnerabilities, and offers a broad evaluation toolkit covering output, trajectory, trace, and adversarial testing. Install with caution if strands-agents and strands-agents-tools are not yet available in your environment."},"id":"strands-agents-evals","links":{"html":"https://skillfed.io/packages/strands-agents-evals","md":"https://skillfed.io/packages/strands-agents-evals.md","pypi":"https://pypi.org/project/strands-agents-evals/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-12","license_spdx":null,"license_treatment":"permissive","name":"strands-agents-evals","python_support":"supports_current","summary":"Evaluation framework for Strands"},"popularity":{"monthly_downloads":155077,"position":10832,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"1.1.1"}
