{"categories":[{"label":"Quality Assurance","url":"https://skillfed.io/packages/category/software-development-quality-assurance"}],"enrichment":{"capability":"Evidently evaluates, tests, and monitors ML and LLM systems through reports, test suites, and a monitoring dashboard, supporting tabular data, text, and generative tasks with 100+ built-in metrics.","skillfed_tags":["ml-monitoring","model-evaluation","llm-evals"],"use_cases":["Run data drift detection on production ML models to identify when input distributions shift from training data.","Evaluate LLM outputs (e.g., RAG systems, summarization) using semantic similarity, retrieval relevance, and custom LLM-as-judge evals.","Build CI/CD test suites that auto-generate pass/fail conditions from reference datasets to catch model regressions.","Monitor classification or regression model performance metrics over time via the self-hosted or cloud dashboard.","Validate data quality in pipelines by checking for missing values, duplicates, new categories, and correlation changes.","Compare ranking and recommendation system quality using metrics like NDCG, MAP, and diversity scores."],"what_it_does":"Evidently is a Python framework for evaluating and monitoring machine learning and LLM systems across the full lifecycle\u2014from offline experiments to live production. It provides Reports for summarizing evaluations (with presets for common tasks like data drift detection and text analysis), Test Suites for adding pass/fail conditions to those reports, and an optional Monitoring Dashboard for tracking metrics over time. The framework works with tabular data, text, and generative outputs, offering both built-in metrics (covering classification, regression, ranking, data quality, and LLM-specific evals) and a Python interface for custom metrics.\n\nThe package is modular: you can run one-off evaluations in a notebook or deploy a full monitoring service. It exports results as JSON, HTML, or Python dictionaries and integrates with existing tools through an open architecture. With 26 runtime dependencies\u2014including pandas, scikit-learn, numpy, and plotly\u2014it brings substantial analytical and visualization capabilities but also a large dependency footprint. Active maintenance, permissive licensing, and support for modern Python versions make it production-ready.","worth_installing":"Yes, if you need systematic ML/LLM evaluation and monitoring. The framework is actively maintained, permissively licensed, and offers both lightweight one-off evals and production-grade monitoring. The large dependency footprint and requirement for Python 3.10+ are manageable trade-offs for the breadth of built-in metrics and modular architecture. No known vulnerabilities. Best suited for teams building or operating ML systems that require rigorous testing and observability."},"id":"evidently","links":{"html":"https://skillfed.io/packages/evidently","md":"https://skillfed.io/packages/evidently.md","pypi":"https://pypi.org/project/evidently/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-03-10","license_spdx":null,"license_treatment":"permissive","name":"evidently","python_support":"supports_current","summary":"Open-source tools to analyze, monitor, and debug machine learning model in production."},"popularity":{"monthly_downloads":1279811,"position":4118,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"0.7.21"}
