$npx skillfedfor your agent

langwatch-scenario

The end-to-end agent testing library

Worth itPyPI TestingReleased Aug 202697.3K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — langwatch_scenario-1.1.1-py3-none-any.whl
v1.1.1 · released 2026-08-13 · Python >=3.10 · 28 runtime deps: pytest, pytest-rerunfailures, litellm, openai, python-dotenv, termcolor, pydantic, joblib

Yes. The package is actively maintained, has no known vulnerabilities, uses a permissive Apache-2.0 license, and solves a real problem—testing agent behavior in realistic multi-turn scenarios. Low install friction and a large dependency set (28 runtime packages) are typical for LLM testing frameworks. Install if you need to validate agent behavior beyond unit tests or if you're building agents that must handle complex, multi-turn interactions reliably.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later.
  • Depends on 28 runtime packages including litellm, openai, fastapi, and opentelemetry-sdk; ensure your environment can resolve all transitive dependencies.
  • Low install friction with a pure-Python wheel distribution.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes.

last release 2026-08-13 (1 days) · last repo commit 2026-08-13 · 951 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 97,271 downloads/mo, #13,159 on PyPI

Verify before relying

pip install langwatch-scenario pytest

import scenario
import pytest

@pytest.mark.asyncio
async def test_agent():
    class MyAgent(scenario.AgentAdapter):
        async def call(self, input: scenario.AgentInput):
            return {"role": "assistant", "content": "response"}
    
    result = await scenario.run(
        name="test",
        description="Test scenario",
        agents=[MyAgent(), scenario.UserSimulatorAgent()]
    )
    assert result.success
  • Whether the framework supports custom evaluation metrics beyond the built-in judge agent criteria.
  • Performance characteristics when running many concurrent simulations or long multi-turn conversations.
  • Compatibility with non-OpenAI LLM providers beyond what litellm abstracts.
Same gist for agents: .md · .json

What it is and what it does

Langwatch-scenario is an agent testing framework designed to validate LLM-based agent behavior through automated simulation. It runs multi-turn conversations between your agent implementation and simulated users (powered by LLMs), optionally with judge agents that evaluate outcomes against predefined criteria. You integrate your agent by implementing a single `call()` method, then define scenarios as pytest tests or standalone scripts that describe the context and expected behavior.

The framework handles the conversation loop, message passing, and evaluation. It supports both scripted control (where you define exact turn sequences) and autopilot mode (where the user simulator drives the conversation until success or max turns). Runtime dependencies include pytest, litellm, openai, pydantic, fastapi, and opentelemetry-sdk, reflecting its design for LLM integration, async execution, and observability. The package is actively maintained and available in Python, TypeScript, and Go.

Use it for

  • Test that a customer support agent asks follow-up questions and provides accurate information before resolving tickets.
  • Validate a recipe recommendation agent generates vegetarian options and includes ingredient lists and cooking instructions.
  • Verify a weather agent calls the correct tool and handles edge cases like missing location data.
  • Evaluate a coding assistant's ability to explain code changes and respond to clarification requests.
  • Benchmark agent behavior across different LLM models to compare response quality and tool usage.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

The package is actively maintained, has no known vulnerabilities, uses a permissive Apache-2.0 license, and solves a real problem—testing agent behavior in realistic multi-turn scenarios. Low install friction and a large dependency set (28 runtime packages) are typical for LLM testing frameworks. Install if you need to validate agent behavior beyond unit tests or if you're building agents that must handle complex, multi-turn interactions reliably.

Install

langwatch-scenario on PyPI

Before you install

Low install friction with a pure-Python wheel distribution. Active maintenance with a recent release (1 day old) and 951 GitHub stars. Requires Python 3.10 or later.

Requires Python 3.10 or later. Depends on 28 runtime packages including litellm, openai, fastapi, and opentelemetry-sdk; ensure your environment can resolve all transitive dependencies.

License in practice

Apache-2.0 permissive license allows commercial and private use with minimal restrictions; you must include a copy of the license and state significant changes.

Quickstart

pip install langwatch-scenario pytest

import scenario
import pytest

@pytest.mark.asyncio
async def test_agent():
    class MyAgent(scenario.AgentAdapter):
        async def call(self, input: scenario.AgentInput):
            return {"role": "assistant", "content": "response"}
    
    result = await scenario.run(
        name="test",
        description="Test scenario",
        agents=[MyAgent(), scenario.UserSimulatorAgent()]
    )
    assert result.success

Verify before relying

  • Whether the framework supports custom evaluation metrics beyond the built-in judge agent criteria.
  • Performance characteristics when running many concurrent simulations or long multi-turn conversations.
  • Compatibility with non-OpenAI LLM providers beyond what litellm abstracts.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
28 packages
pytestpytest-rerunfailureslitellmopenaipython-dotenvtermcolorpydanticjoblibwraptpytest-asynciorichpksuidhttpxrxpython-dateutilpydantic-settingslangwatchopentelemetry-sdkimageio-ffmpegnumpywebrtcvad-wheelswebsocketstwiliofastapiuvicornaudioop-ltsgoogle-genaielevenlabs
MaintenanceActively maintained 1 days since the last release
Last repo commit
First released
Downloads97,271 / month, #13,159 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 5 - Production/StableIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11

Evidence: langwatch_scenario-1.1.1-py3-none-any.whl

Tags

Capabilities
agent testing frameworkLLM agent simulationmulti-turn conversation testingagent evaluation frameworkAI agent behavior testingscenario-based agent testingLLM agent benchmarking
Topics
agent-testingllm-evaluationmulti-turn-simulation

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “agent testing framework”

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also agentevals · dreadnode · deepeval · llama-index-agent-openai · agentops · langchain · judgeval · openevals · openai-agents · policyengine-core

Further reading