$npx skillfedfor your agent

terminal-bench

Terminal-bench is a collection of tasks and evaluation harness for evaluating AI agents' ability to complete complex tasks in terminal environments.

With conditionsPyPI TestingReleased Sep 202590.9K downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — terminal_bench-0.2.18-py3-none-any.whl
v0.2.18 · released 2025-09-26 · Python >=3.12 · 20 runtime deps: anthropic, asciinema, boto3, docker, openai, pandas, psycopg2-binary, sqlalchemy

Yes, with conditions. Install if you are actively developing or evaluating AI agents for terminal tasks and can meet the Python 3.12+ and Docker requirements. The package has low install friction and no known vulnerabilities, but verify the license terms before use in commercial contexts. The aging maintenance status and unclear license are minor concerns; the framework is still actively used for agent benchmarking and leaderboard submissions.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.12+, Docker, and the uv package manager for full functionality.
  • The execution harness also needs a sandboxed terminal environment to run tasks.
  • Low install friction with a pure-Python wheel distribution.

License · maintenance · safety

(unclear) — License treatment is unclear—no SPDX identifier or raw license text is available in the metadata. Before relying on this package in a commercial or copyleft-sensitive context, verify the actual license from the project repository or documentation.

last release 2025-09-26 (322 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 90,914 downloads/mo, #13,556 on PyPI

Verify before relying

pip install terminal-bench

from terminal_bench import run

# Run benchmark with an agent and model
run(agent='terminus', model='anthropic/claude-3-7-latest', dataset_name='terminal-bench-core')
  • Whether the package is suitable for production use or remains experimental despite beta status
  • Exact license terms and any restrictions on commercial or derivative use
  • Performance characteristics and typical runtime for the ~100 tasks in the current dataset
Same gist for agents: .md · .json

What it is and what it does

Terminal-Bench is a benchmarking framework designed to evaluate how well AI agents can autonomously complete real-world terminal tasks. It consists of a dataset of approximately 100 tasks (each with an English instruction, test script, and reference solution) and an execution harness that connects language models to a sandboxed terminal environment. The harness supports multiple LLM providers via integrations with anthropic, openai, and litellm, and can run tasks concurrently across Docker containers.

The package is aimed at researchers and developers building LLM agents, benchmarking frameworks, or stress-testing system-level reasoning capabilities. It provides a reproducible, practical evaluation suite for terminal-based workflows—tasks range from compiling code to training models to setting up servers. The CLI tool `tb` lets you run evaluations, submit results to a leaderboard, and contribute new tasks or adapters to the community.

Use it for

  • Evaluate an LLM agent's ability to autonomously complete real-world terminal tasks and compare performance across models
  • Benchmark a custom AI agent framework against a standardized task suite with reproducible results
  • Stress-test system-level reasoning in language models by running end-to-end workflows like server setup or model training
  • Contribute new terminal-based tasks to the Terminal-Bench dataset to expand the evaluation suite
  • Submit agent evaluation results to the Terminal-Bench leaderboard for public comparison

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

Install if you are actively developing or evaluating AI agents for terminal tasks and can meet the Python 3.12+ and Docker requirements. The package has low install friction and no known vulnerabilities, but verify the license terms before use in commercial contexts. The aging maintenance status and unclear license are minor concerns; the framework is still actively used for agent benchmarking and leaderboard submissions.

Install

terminal-bench on PyPI

Before you install

Low install friction with a pure-Python wheel distribution. The package is aging (322 days since release) but remains actively maintained; however, it requires Python 3.12+ and depends on 20 runtime packages including Docker, which adds system-level prerequisites beyond the Python ecosystem.

Requires Python 3.12+, Docker, and the uv package manager for full functionality. The execution harness also needs a sandboxed terminal environment to run tasks.

License in practice

License treatment is unclear—no SPDX identifier or raw license text is available in the metadata. Before relying on this package in a commercial or copyleft-sensitive context, verify the actual license from the project repository or documentation.

Quickstart

pip install terminal-bench

from terminal_bench import run

# Run benchmark with an agent and model
run(agent='terminus', model='anthropic/claude-3-7-latest', dataset_name='terminal-bench-core')

Verify before relying

  • Whether the package is suitable for production use or remains experimental despite beta status
  • Exact license terms and any restrictions on commercial or derivative use
  • Performance characteristics and typical runtime for the ~100 tasks in the current dataset

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release >=3.12
Install frictionLow. Pure-Python wheel
Runtime dependencies
20 packages
anthropicasciinemaboto3dockeropenaipandaspsycopg2-binarysqlalchemystreamlittenacitytqdmruamel-yamltabulatemcplitellmpydanticsupabaseinquirertyperjinja2
MaintenanceAging 322 days since the last release
First released
Downloads90,914 / month, #13,556 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: terminal_bench-0.2.18-py3-none-any.whl

Tags

Capabilities
ai agent benchmarking terminalllm evaluation frameworkautonomous task testingagent performance evaluationterminal environment sandboxai agent stress testingend-to-end task benchmark
Topics
ai-agent-evalbenchmark-suiteterminal-automation

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “ai agent benchmarking terminal”

  • terminal-benchTerminal-Bench provides a benchmark suite and execution harness for…
  • openhandsOpenHands is a CLI tool that runs an AI agent in your terminal, IDE,…
  • inspect-sweInspect SWE provides a suite of software engineering agents built on…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also harbor · nemo-gym · lm-eval · TextArena · unitxt · asv-runner · deepagents · swebench · harbor-rewardkit · bfcl-eval

Further reading