{"categories":[{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence/4"}],"enrichment":{"capability":"Phoenix Evals provides composable building blocks for evaluating LLM applications, including pre-built evaluators for hallucination detection, relevance, toxicity, and other common assessment tasks.","skillfed_tags":["llm-evaluation","observability","quality-assurance"],"use_cases":["Detect hallucinations in LLM outputs by checking whether responses are grounded in provided context.","Score retrieved documents for relevance to user queries in retrieval-augmented generation systems.","Evaluate whether an LLM selected and invoked the correct tool with appropriate arguments.","Run batch evaluations on large datasets of LLM interactions stored in pandas DataFrames.","Build custom evaluators with templated prompts and multi-choice scoring for domain-specific assessment tasks."],"what_it_does":"Phoenix Evals is a framework for building and running evaluations on language model applications. It provides both pre-built evaluators (for tasks like hallucination detection, relevance scoring, and tool invocation checking) and tools to compose custom evaluators using your choice of LLM provider. The package handles input mapping for complex nested data structures and integrates with OpenTelemetry for tracing and observability.\n\nThe framework is designed to work with pandas DataFrames for batch evaluation and supports both synchronous and asynchronous evaluation modes. It has nine runtime dependencies including jsonpath-ng, openinference-instrumentation, openinference-semantic-conventions, opentelemetry-api, pandas, pydantic, pystache, tqdm, and typing-extensions. The package is actively maintained, supports Python 3.10 through 3.14, and has no known security vulnerabilities.","worth_installing":"Yes, with conditions. The package is actively maintained, has low installation friction, and provides a practical framework for LLM evaluation with both pre-built and custom evaluators. However, the Elastic-2.0 license treatment is flagged as unclear\u2014verify the license terms for your use case before production deployment. If you need LLM evaluation capabilities and can clarify the license, this is a solid choice."},"id":"arize-phoenix-evals","links":{"html":"https://skillfed.io/packages/arize-phoenix-evals","md":"https://skillfed.io/packages/arize-phoenix-evals.md","pypi":"https://pypi.org/project/arize-phoenix-evals/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-08","license_spdx":null,"license_treatment":"unclear","name":"arize-phoenix-evals","python_support":"supports_current","summary":"LLM Evaluations"},"popularity":{"monthly_downloads":860405,"position":4875,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"3.4.0"}
