$npx skillfedfor your agent

deepeval

The LLM Evaluation Framework

Worth itPyPI TestingReleased Aug 20266.5M downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — deepeval-4.1.8-py3-none-any.whl
v4.1.8 · released 2026-08-12 · Python <4.0,>=3.9 · 28 runtime deps: aiohttp, click, grpcio, jinja2, nest_asyncio, openai, opentelemetry-api, opentelemetry-sdk

Yes. DeepEval is actively maintained, has no known vulnerabilities, installs with low friction, and is permissively licensed. It's well-suited for teams building LLM applications who need systematic evaluation beyond manual testing. The large dependency footprint and reliance on external LLM APIs for some metrics are trade-offs for comprehensive evaluation coverage; verify that the specific metrics you need match your local-vs.-API execution preferences before committing.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires an LLM provider (OpenAI, Claude, etc.) or local NLP models to run metrics; some metrics depend on external API calls unless configured for local execution.
  • Low friction install with a pure Python wheel.
  • Active maintenance with recent releases (2 days old) and strong community signal (17597 stars).

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial use, modification, and distribution with minimal restrictions—suitable for most production and proprietary projects.

last release 2026-08-12 (2 days) · last repo commit 2026-08-13 · 17,597 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 6,466,639 downloads/mo, #1,906 on PyPI

Verify before relying

pip install deepeval

from deepeval import evaluate
from deepeval.metrics import AnswerRelevancy

metric = AnswerRelevancy()
result = metric.measure(prediction="answer", input="question")
print(result.score)
  • Whether all 28 runtime dependencies are required for basic use or if many are optional (testing, CLI, integrations).
  • Performance characteristics and latency for running metrics on large evaluation batches.
  • Which metrics actually run locally versus requiring external LLM API calls by default.
Same gist for agents: .md · .json

What it is and what it does

DeepEval is a pytest-like testing framework designed specifically for evaluating large language model applications. It provides a collection of pre-built metrics—including G-Eval, answer relevancy, faithfulness, hallucination detection, and agentic metrics—that measure the quality of LLM outputs against criteria like factual accuracy, relevance, and task completion. The framework supports end-to-end evaluation of black-box LLM systems, component-level testing of individual steps (retrieval, tool use, agent handoffs), and multi-turn conversation assessment.

The package integrates with any LLM framework (OpenAI, LangChain, Claude) and CI/CD pipeline. Many of its metrics run locally on your machine using NLP models or statistical methods, though some can also delegate to external LLMs for judgment. It includes tools for synthetic dataset generation, prompt optimization, and benchmarking against standard LLM benchmarks. With 28 runtime dependencies covering async HTTP, CLI tooling, observability (OpenTelemetry, PostHog), and testing utilities, it's designed as a comprehensive evaluation platform rather than a minimal library.

Use it for

  • Evaluate RAG pipeline output quality by measuring answer relevancy, faithfulness, and retrieval context precision.
  • Test AI agent behavior across decision trees by checking task completion, tool correctness, and plan adherence.
  • Monitor chatbot consistency and factual grounding across multi-turn conversations using turn-level metrics.
  • Detect hallucinations and bias in LLM outputs before deploying to production.
  • Compare model performance (OpenAI vs. Claude) or prompt variations using standardized metrics.
  • Benchmark your LLM against public benchmarks like MMLU or HumanEval in minimal code.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

DeepEval is actively maintained, has no known vulnerabilities, installs with low friction, and is permissively licensed. It's well-suited for teams building LLM applications who need systematic evaluation beyond manual testing. The large dependency footprint and reliance on external LLM APIs for some metrics are trade-offs for comprehensive evaluation coverage; verify that the specific metrics you need match your local-vs.-API execution preferences before committing.

Install

deepeval on PyPI

Before you install

Low friction install with a pure Python wheel. Active maintenance with recent releases (2 days old) and strong community signal (17597 stars). Supports Python 3.9 through 3.14.

Requires an LLM provider (OpenAI, Claude, etc.) or local NLP models to run metrics; some metrics depend on external API calls unless configured for local execution.

License in practice

Apache-2.0 permissive license allows commercial use, modification, and distribution with minimal restrictions—suitable for most production and proprietary projects.

Quickstart

pip install deepeval

from deepeval import evaluate
from deepeval.metrics import AnswerRelevancy

metric = AnswerRelevancy()
result = metric.measure(prediction="answer", input="question")
print(result.score)

Verify before relying

  • Whether all 28 runtime dependencies are required for basic use or if many are optional (testing, CLI, integrations).
  • Performance characteristics and latency for running metrics on large evaluation batches.
  • Which metrics actually run locally versus requiring external LLM API calls by default.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release <4.0,>=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
28 packages
aiohttpclickgrpciojinja2nest_asyncioopenaiopentelemetry-apiopentelemetry-sdkportalockerposthogpydanticpydantic-settingspyfigletpytestpytest-asynciopytest-repeatpytest-rerunfailurespytest-xdistpython-dotenvquestionaryrequestsrichsetuptoolstabulatetenacitytqdmtyperwheel
MaintenanceActively maintained 2 days since the last release
Last repo commit
First released
Downloads6,466,639 / month, #1,906 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9

Evidence: deepeval-4.1.8-py3-none-any.whl

Tags

Capabilities
LLM evaluation frameworktest language model outputsRAG pipeline evaluationLLM-as-a-judge metricsAI agent testingprompt quality assessmenthallucination detection
Topics
llm-evaluationai-testingrag-metrics

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “test language model outputs”

  • deepevalDeepEval is an LLM evaluation framework that runs unit tests on…
  • ragasRagas provides objective metrics, test data generation, and…
  • openevalsOpenEvals provides a framework for writing and running evaluators to…

Give your agent the search over MCP, or paste the wish link into any chat.

More Testing packages

pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo
virtualenv Worth it
PyPI · Libraries · released Aug 2026

virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.

MITpure Python · 3.9+
532.9Mdownloads / mo
coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
pytest-asyncio Worth it
PyPI · Testing · released May 2026

pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.

Install it if you write tests for any asyncio-based code.

Apache-2.0pure Python · 3.10+
275.9Mdownloads / mo
pytest-json-ctrf Worth it
PyPI · Testing · released Jul 2026

A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.

Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.

MITpure Python · 3.8+
273.0Mdownloads / mo

See also openevals · autoevals · evalplus · deepteam · cleanlab-tlm · ragas · arize-phoenix-evals · judgeval · unitxt · rubric

Further reading