deepeval
The LLM Evaluation Framework
Decision gist · record as of 2026-08-14
Yes. DeepEval is actively maintained, has no known vulnerabilities, installs with low friction, and is permissively licensed. It's well-suited for teams building LLM applications who need systematic evaluation beyond manual testing. The large dependency footprint and reliance on external LLM APIs for some metrics are trade-offs for comprehensive evaluation coverage; verify that the specific metrics you need match your local-vs.-API execution preferences before committing.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires an LLM provider (OpenAI, Claude, etc.) or local NLP models to run metrics; some metrics depend on external API calls unless configured for local execution.
- Low friction install with a pure Python wheel.
- Active maintenance with recent releases (2 days old) and strong community signal (17597 stars).
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial use, modification, and distribution with minimal restrictions—suitable for most production and proprietary projects.
last release 2026-08-12 (2 days) · last repo commit 2026-08-13 · 17,597 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 6,466,639 downloads/mo, #1,906 on PyPI
Alternatives
Verify before relying
pip install deepeval
from deepeval import evaluate
from deepeval.metrics import AnswerRelevancy
metric = AnswerRelevancy()
result = metric.measure(prediction="answer", input="question")
print(result.score)- Whether all 28 runtime dependencies are required for basic use or if many are optional (testing, CLI, integrations).
- Performance characteristics and latency for running metrics on large evaluation batches.
- Which metrics actually run locally versus requiring external LLM API calls by default.
What it is and what it does
DeepEval is a pytest-like testing framework designed specifically for evaluating large language model applications. It provides a collection of pre-built metrics—including G-Eval, answer relevancy, faithfulness, hallucination detection, and agentic metrics—that measure the quality of LLM outputs against criteria like factual accuracy, relevance, and task completion. The framework supports end-to-end evaluation of black-box LLM systems, component-level testing of individual steps (retrieval, tool use, agent handoffs), and multi-turn conversation assessment.
The package integrates with any LLM framework (OpenAI, LangChain, Claude) and CI/CD pipeline. Many of its metrics run locally on your machine using NLP models or statistical methods, though some can also delegate to external LLMs for judgment. It includes tools for synthetic dataset generation, prompt optimization, and benchmarking against standard LLM benchmarks. With 28 runtime dependencies covering async HTTP, CLI tooling, observability (OpenTelemetry, PostHog), and testing utilities, it's designed as a comprehensive evaluation platform rather than a minimal library.
Use it for
- Evaluate RAG pipeline output quality by measuring answer relevancy, faithfulness, and retrieval context precision.
- Test AI agent behavior across decision trees by checking task completion, tool correctness, and plan adherence.
- Monitor chatbot consistency and factual grounding across multi-turn conversations using turn-level metrics.
- Detect hallucinations and bias in LLM outputs before deploying to production.
- Compare model performance (OpenAI vs. Claude) or prompt variations using standardized metrics.
- Benchmark your LLM against public benchmarks like MMLU or HumanEval in minimal code.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
DeepEval is actively maintained, has no known vulnerabilities, installs with low friction, and is permissively licensed. It's well-suited for teams building LLM applications who need systematic evaluation beyond manual testing. The large dependency footprint and reliance on external LLM APIs for some metrics are trade-offs for comprehensive evaluation coverage; verify that the specific metrics you need match your local-vs.-API execution preferences before committing.
Install
deepeval on PyPI
Before you install
Low friction install with a pure Python wheel. Active maintenance with recent releases (2 days old) and strong community signal (17597 stars). Supports Python 3.9 through 3.14.
Requires an LLM provider (OpenAI, Claude, etc.) or local NLP models to run metrics; some metrics depend on external API calls unless configured for local execution.
License in practice
Apache-2.0 permissive license allows commercial use, modification, and distribution with minimal restrictions—suitable for most production and proprietary projects.
Quickstart
pip install deepeval
from deepeval import evaluate
from deepeval.metrics import AnswerRelevancy
metric = AnswerRelevancy()
result = metric.measure(prediction="answer", input="question")
print(result.score)
Verify before relying
- Whether all 28 runtime dependencies are required for basic use or if many are optional (testing, CLI, integrations).
- Performance characteristics and latency for running metrics on large evaluation batches.
- Which metrics actually run locally versus requiring external LLM API calls by default.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release <4.0,>=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 28 packagesaiohttpclickgrpciojinja2nest_asyncioopenaiopentelemetry-apiopentelemetry-sdkportalockerposthogpydanticpydantic-settingspyfigletpytestpytest-asynciopytest-repeatpytest-rerunfailurespytest-xdistpython-dotenvquestionaryrequestsrichsetuptoolstabulatetenacitytqdmtyperwheel |
| Maintenance | Actively maintained 2 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 6,466,639 / month, #1,906 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9 |
Evidence: deepeval-4.1.8-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “test language model outputs”
- deepevalDeepEval is an LLM evaluation framework that runs unit tests on…
- ragasRagas provides objective metrics, test data generation, and…
- openevalsOpenEvals provides a framework for writing and running evaluators to…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also openevals · autoevals · evalplus · deepteam · cleanlab-tlm · ragas · arize-phoenix-evals · judgeval · unitxt · rubric