{"categories":[{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence/4"}],"enrichment":{"capability":"Inspect Evals provides a repository of community-contributed LLM evaluations built on the Inspect AI framework, allowing developers to run standardized benchmarks against language models from multiple providers.","skillfed_tags":["llm-benchmarking","ai-evaluation","model-testing"],"use_cases":["Run standardized benchmarks like ARC against your LLM to compare performance across model providers","Build a continuous evaluation pipeline by importing eval tasks as Python objects and logging results to track model improvements","Contribute new domain-specific evaluations to the community register by submitting your eval implementation with documentation","Assess AI safety and capability properties using evaluations designed by UK AISI and partner institutions","Compare multiple language models in parallel using eval-set to identify which provider best suits your use case"],"what_it_does":"Inspect Evals is a curated repository of LLM evaluation tasks maintained by the UK AISI, Arcadia Impact, and the Vector Institute. It provides ready-to-run benchmarks for testing language models across multiple providers (OpenAI, Anthropic, Google, Mistral, AWS Bedrock, and others) using the Inspect AI framework. The package bundles evaluation implementations for tasks like ARC and other standardized benchmarks, handling dataset loading, model interaction, and result logging.\n\nDevelopers use Inspect Evals to systematically assess model capabilities without building evaluation infrastructure from scratch. You can run evals via command line or import them as Python objects for programmatic use. The package manages caching of datasets and evaluation artifacts, supports parallel evaluation runs, and integrates with Inspect AI's log viewer for result analysis. Community contributions are accepted through a GitHub-based submission process that validates and registers new evaluations.","worth_installing":"Yes. Inspect Evals is actively maintained, has no known vulnerabilities, uses a permissive MIT license, and offers low-friction installation. It's worth installing if you need standardized LLM benchmarking without building evaluation infrastructure yourself. The main constraint is disk space (35\u2013100 GB depending on eval scope) and Python version requirements (3.11\u20133.12 preferred). Suitable for research, safety assessment, and model selection workflows."},"id":"inspect-evals","links":{"html":"https://skillfed.io/packages/inspect-evals","md":"https://skillfed.io/packages/inspect-evals.md","pypi":"https://pypi.org/project/inspect-evals/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-14","license_spdx":"MIT","license_treatment":"permissive","name":"inspect-evals","python_support":"supports_current","summary":"Collection of large language model evaluations"},"popularity":{"monthly_downloads":856912,"position":4886,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"0.17.0"}
