$npx skillfedfor your agent

evidently

Open-source tools to analyze, monitor, and debug machine learning model in production.

With conditionsPyPI Quality AssuranceReleased Mar 20261.3M downloads / moApache License 2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — evidently-0.7.21-py3-none-any.whl
v0.7.21 · released 2026-03-10 · Python >=3.10 · 26 runtime deps: certifi, cryptography, deprecation, dynaconf, fsspec, iterative-telemetry, litestar, nltk

Yes, if you need systematic ML/LLM evaluation and monitoring. The framework is actively maintained, permissively licensed, and offers both lightweight one-off evals and production-grade monitoring. The large dependency footprint and requirement for Python 3.10+ are manageable trade-offs for the breadth of built-in metrics and modular architecture. No known vulnerabilities. Best suited for teams building or operating ML systems that require rigorous testing and observability.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later; 26 runtime dependencies including pandas, scikit-learn, and numpy mean a substantial environment footprint.
  • Low friction install with a pure Python wheel.
  • Active maintenance with recent releases; 7808 repository stars and current support for Python 3.10–3.13 indicate a well-maintained project.

License · maintenance · safety

Apache License 2.0 (permissive) — Apache License 2.0 (permissive) allows free use, modification, and distribution with minimal restrictions, making it suitable for both open-source and commercial projects.

last release 2026-03-10 (157 days) · last repo commit 2026-08-05 · 7,808 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,279,811 downloads/mo, #4,118 on PyPI

Verify before relying

pip install evidently

import pandas as pd
from evidently import Report
from evidently.presets import DataDriftPreset

report = Report([DataDriftPreset(method="psi")])
my_eval = report.run(reference_data, current_data)
my_eval.save_html("report.html")
  • Whether the 100+ built-in metrics cover your specific evaluation domain or if custom metric development is required.
  • Performance characteristics and scalability limits for large-scale production monitoring workloads.
  • Integration complexity with existing ML pipelines and whether the open-source version meets your monitoring SLA needs.
Same gist for agents: .md · .json

What it is and what it does

Evidently is a Python framework for evaluating and monitoring machine learning and LLM systems across the full lifecycle—from offline experiments to live production. It provides Reports for summarizing evaluations (with presets for common tasks like data drift detection and text analysis), Test Suites for adding pass/fail conditions to those reports, and an optional Monitoring Dashboard for tracking metrics over time. The framework works with tabular data, text, and generative outputs, offering both built-in metrics (covering classification, regression, ranking, data quality, and LLM-specific evals) and a Python interface for custom metrics.

The package is modular: you can run one-off evaluations in a notebook or deploy a full monitoring service. It exports results as JSON, HTML, or Python dictionaries and integrates with existing tools through an open architecture. With 26 runtime dependencies—including pandas, scikit-learn, numpy, and plotly—it brings substantial analytical and visualization capabilities but also a large dependency footprint. Active maintenance, permissive licensing, and support for modern Python versions make it production-ready.

Use it for

  • Run data drift detection on production ML models to identify when input distributions shift from training data.
  • Evaluate LLM outputs (e.g., RAG systems, summarization) using semantic similarity, retrieval relevance, and custom LLM-as-judge evals.
  • Build CI/CD test suites that auto-generate pass/fail conditions from reference datasets to catch model regressions.
  • Monitor classification or regression model performance metrics over time via the self-hosted or cloud dashboard.
  • Validate data quality in pipelines by checking for missing values, duplicates, new categories, and correlation changes.
  • Compare ranking and recommendation system quality using metrics like NDCG, MAP, and diversity scores.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need systematic ML/LLM evaluation and monitoring.

The framework is actively maintained, permissively licensed, and offers both lightweight one-off evals and production-grade monitoring. The large dependency footprint and requirement for Python 3.10+ are manageable trade-offs for the breadth of built-in metrics and modular architecture. No known vulnerabilities. Best suited for teams building or operating ML systems that require rigorous testing and observability.

Install

evidently on PyPI

Before you install

Low friction install with a pure Python wheel. Active maintenance with recent releases; 7808 repository stars and current support for Python 3.10–3.13 indicate a well-maintained project.

Requires Python 3.10 or later; 26 runtime dependencies including pandas, scikit-learn, and numpy mean a substantial environment footprint.

License in practice

Apache License 2.0 (permissive) allows free use, modification, and distribution with minimal restrictions, making it suitable for both open-source and commercial projects.

Quickstart

pip install evidently

import pandas as pd
from evidently import Report
from evidently.presets import DataDriftPreset

report = Report([DataDriftPreset(method="psi")])
my_eval = report.run(reference_data, current_data)
my_eval.save_html("report.html")

Verify before relying

  • Whether the 100+ built-in metrics cover your specific evaluation domain or if custom metric development is required.
  • Performance characteristics and scalability limits for large-scale production monitoring workloads.
  • Integration complexity with existing ML pipelines and whether the open-source version meets your monitoring SLA needs.

Package facts

LicenseApache License 2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
26 packages
certificryptographydeprecationdynaconffsspeciterative-telemetrylitestarnltknumpyopentelemetry-protopandasplotlypydanticpyyamlrequestsrichscikit-learnscipystatsmodelstypertyping-inspectujsonurllib3uuid6uvicornwatchdog
MaintenanceActively maintained 157 days since the last release
Last repo commit
First released
Downloads1,279,811 / month, #4,118 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13

Evidence: evidently-0.7.21-py3-none-any.whl

Tags

Capabilities
ml model monitoring and evaluationllm output evaluation frameworkdata drift detectionml test suite and regression testingmodel performance metrics dashboardproduction ml system monitoringdata quality and validation checks
Topics
ml-monitoringmodel-evaluationllm-evals

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “ml model monitoring and evaluation”

  • evidentlyEvidently evaluates, tests, and monitors ML and LLM systems through…
  • mlflowMLflow is an open-source platform for managing the complete lifecycle…
  • arizeArize is a Python client library for interacting with the Arize AI…

Give your agent the search over MCP, or paste the wish link into any chat.

More Quality Assurance packages

coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
ruff Worth it
PyPI · Python Modules · released Aug 2026

Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.

MITcompiled wheel · 3.7+
316.1Mdownloads / mo
pexpect With conditions
PyPI · Software Development · released Nov 2023

Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.

ISCpure Pythonaging
200.8Mdownloads / mo
black Worth it
PyPI · Python Modules · released May 2026

Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.

MITpure Python · 3.10+
179.9Mdownloads / mo
pytest-xdist Worth it
PyPI · Utilities · released Jul 2025

pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.

Install it if your test suite takes long enough that parallelization would save meaningful time.

MITpure Python · 3.9+
177.1Mdownloads / mo
cfn-lint Worth it
PyPI · Quality Assurance · released Aug 2026

Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.

Install it if you work with CloudFormation templates.

MIT-0pure Python
114.9Mdownloads / mo

See also ragas · whylogs · opik · inspect-evals · trulens · mlflow · azure-ai-evaluation · deepeval · arize-phoenix-evals · arize

Further reading