$npx skillfedfor your agent

inspect-evals

Collection of large language model evaluations

Worth itPyPI Artificial IntelligenceReleased Aug 2026856.9K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — inspect_evals-0.17.0-py3-none-any.whl
v0.17.0 · released 2026-08-14 · Python >=3.11 · 13 runtime deps: backoff, datasets, huggingface_hub, hf_xet, inspect_ai, jinja2, numpy, pillow

Yes. Inspect Evals is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and offers low-friction installation. It's worth installing if you need standardized LLM benchmarking without building evaluation infrastructure yourself. The main constraint is disk space (35–100 GB depending on eval scope) and Python version requirements (3.11–3.12 preferred). Suitable for research, safety assessment, and model selection workflows.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.11 or 3.12; some evaluations need extra dependencies or disk space (35 GB minimum recommended, up to 100 GB for Docker-based evals).
  • API keys for model providers required to run evaluations.
  • Low friction install with a pure Python wheel and 13 runtime dependencies.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making the package suitable for both research and production evaluation pipelines.

last release 2026-08-14 (0 days) · last repo commit 2026-08-14 · 625 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 856,912 downloads/mo, #4,886 on PyPI

Verify before relying

pip install inspect-evals

from inspect_ai import eval
from inspect_evals.arc import arc_easy
eval(arc_easy)
  • Whether all 13 runtime dependencies are strictly required or if some are optional for specific evals only
  • Performance characteristics when running multiple evaluations concurrently with eval-set
  • Compatibility status with Python 3.13 for the full eval suite beyond the sciknoweval exception
Same gist for agents: .md · .json

What it is and what it does

Inspect Evals is a curated repository of LLM evaluation tasks maintained by the UK AISI, Arcadia Impact, and the Vector Institute. It provides ready-to-run benchmarks for testing language models across multiple providers (OpenAI, Anthropic, Google, Mistral, AWS Bedrock, and others) using the Inspect AI framework. The package bundles evaluation implementations for tasks like ARC and other standardized benchmarks, handling dataset loading, model interaction, and result logging.

Developers use Inspect Evals to systematically assess model capabilities without building evaluation infrastructure from scratch. You can run evals via command line or import them as Python objects for programmatic use. The package manages caching of datasets and evaluation artifacts, supports parallel evaluation runs, and integrates with Inspect AI's log viewer for result analysis. Community contributions are accepted through a GitHub-based submission process that validates and registers new evaluations.

Use it for

  • Run standardized benchmarks like ARC against your LLM to compare performance across model providers
  • Build a continuous evaluation pipeline by importing eval tasks as Python objects and logging results to track model improvements
  • Contribute new domain-specific evaluations to the community register by submitting your eval implementation with documentation
  • Assess AI safety and capability properties using evaluations designed by UK AISI and partner institutions
  • Compare multiple language models in parallel using eval-set to identify which provider best suits your use case

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Inspect Evals is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and offers low-friction installation. It's worth installing if you need standardized LLM benchmarking without building evaluation infrastructure yourself. The main constraint is disk space (35–100 GB depending on eval scope) and Python version requirements (3.11–3.12 preferred). Suitable for research, safety assessment, and model selection workflows.

Install

inspect-evals on PyPI

Before you install

Low friction install with a pure Python wheel and 13 runtime dependencies. Active maintenance with a release on 2026-08-14 and 625 repository stars. Requires Python 3.11 or 3.12 for full compatibility; Python 3.13 works for most evals except sciknoweval.

Requires Python 3.11 or 3.12; some evaluations need extra dependencies or disk space (35 GB minimum recommended, up to 100 GB for Docker-based evals). API keys for model providers required to run evaluations.

License in practice

MIT license permits commercial and private use with minimal restrictions, making the package suitable for both research and production evaluation pipelines.

Quickstart

pip install inspect-evals

from inspect_ai import eval
from inspect_evals.arc import arc_easy
eval(arc_easy)

Verify before relying

  • Whether all 13 runtime dependencies are strictly required or if some are optional for specific evals only
  • Performance characteristics when running multiple evaluations concurrently with eval-set
  • Compatibility status with Python 3.13 for the full eval suite beyond the sciknoweval exception

Package facts

LicenseMIT permissive
Python supportSupports the current Python release >=3.11
Install frictionLow. Pure-Python wheel
Runtime dependencies
13 packages
backoffdatasetshuggingface_hubhf_xetinspect_aijinja2numpypillowpydanticpyyamlrequeststiktokentoml
MaintenanceActively maintained 0 days since the last release
Last repo commit
First released
Downloads856,912 / month, #4,886 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: Python :: 3Topic :: Scientific/Engineering :: Artificial IntelligenceTyping :: Typed

Evidence: inspect_evals-0.17.0-py3-none-any.whl

Tags

Capabilities
llm evaluation frameworkai model benchmarkinginspect ai evaluationslanguage model testingbenchmark suite for llmsai safety evaluationmodel capability assessment
Topics
llm-benchmarkingai-evaluationmodel-testing

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “inspect ai evaluations”

  • inspect-evalsInspect Evals provides a repository of community-contributed LLM…
  • inspect-sweInspect SWE provides a suite of software engineering agents built on…
  • inspect-aiInspect is a framework for evaluating large language models,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also inspect-ai · nemo-evaluator · autoevals · inspect-swe · evalplus · evidently · deepeval · arize-phoenix-evals · ragas · harbor

Further reading