{"categories":[{"label":"Testing","url":"https://skillfed.io/packages/category/software-development-testing/6"}],"enrichment":{"capability":"Terminal-Bench provides a benchmark suite and execution harness for evaluating AI agents' ability to complete real-world terminal tasks autonomously, from code compilation to server setup.","skillfed_tags":["ai-agent-eval","benchmark-suite","terminal-automation"],"use_cases":["Evaluate an LLM agent's ability to autonomously complete real-world terminal tasks and compare performance across models","Benchmark a custom AI agent framework against a standardized task suite with reproducible results","Stress-test system-level reasoning in language models by running end-to-end workflows like server setup or model training","Contribute new terminal-based tasks to the Terminal-Bench dataset to expand the evaluation suite","Submit agent evaluation results to the Terminal-Bench leaderboard for public comparison"],"what_it_does":"Terminal-Bench is a benchmarking framework designed to evaluate how well AI agents can autonomously complete real-world terminal tasks. It consists of a dataset of approximately 100 tasks (each with an English instruction, test script, and reference solution) and an execution harness that connects language models to a sandboxed terminal environment. The harness supports multiple LLM providers via integrations with anthropic, openai, and litellm, and can run tasks concurrently across Docker containers.\n\nThe package is aimed at researchers and developers building LLM agents, benchmarking frameworks, or stress-testing system-level reasoning capabilities. It provides a reproducible, practical evaluation suite for terminal-based workflows\u2014tasks range from compiling code to training models to setting up servers. The CLI tool `tb` lets you run evaluations, submit results to a leaderboard, and contribute new tasks or adapters to the community.","worth_installing":"Yes, with conditions. Install if you are actively developing or evaluating AI agents for terminal tasks and can meet the Python 3.12+ and Docker requirements. The package has low install friction and no known vulnerabilities, but verify the license terms before use in commercial contexts. The aging maintenance status and unclear license are minor concerns; the framework is still actively used for agent benchmarking and leaderboard submissions."},"id":"terminal-bench","links":{"html":"https://skillfed.io/packages/terminal-bench","md":"https://skillfed.io/packages/terminal-bench.md","pypi":"https://pypi.org/project/terminal-bench/"},"maintenance":{"status":"aging"},"meta":{"latest_release":"2025-09-26","license_spdx":null,"license_treatment":"unclear","name":"terminal-bench","python_support":"supports_current","summary":"Terminal-bench is a collection of tasks and evaluation harness for evaluating AI agents' ability to complete complex tasks in terminal environments."},"popularity":{"monthly_downloads":90914,"position":13556,"tier":"top_15000"},"security":{"n_vulnerabilities":0},"version":"0.2.18"}
