judgeval
The open source post-building layer for Agent Behavior Monitoring.
Decision gist · record as of 2026-08-14
Yes. Judgeval is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and solves a real problem—observing and improving LLM agents in production. It requires Python >=3.10 and API credentials (JUDGMENT_API_KEY, JUDGMENT_ORG_ID), which implies a hosted backend service. No known vulnerabilities. Install it if you need production tracing and behavior-based evaluation for agents.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >=3.10; JUDGMENT_API_KEY and JUDGMENT_ORG_ID environment variables must be set to send traces to the Judgment backend.
- Low friction: pure Python wheel with eight runtime dependencies.
- Active maintenance—last commit 2026-08-10, 1057 GitHub stars, released 2026-08-03.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 (permissive): you may use, modify, and distribute this package freely in commercial and private projects, provided you include the license notice and state significant changes.
last release 2026-08-03 (11 days) · last repo commit 2026-08-10 · 1,057 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 269,024 downloads/mo, #8,265 on PyPI
Alternatives
Verify before relying
pip install judgeval
from judgeval import Tracer
Tracer.init(project_name="my-project")
@Tracer.observe(span_type="agent")
def run_agent(question: str) -> str:
return "response"- Specific latency overhead of server-side scoring on live production traffic.
- Whether Slack alert configuration is built-in or requires external setup.
- Retention and query performance characteristics of the trace storage backend.
- Cost model for the Judgment backend service (API key required suggests hosted service).
- Auto-instrumentation behavior and token usage capture mechanics with specific LLM providers.
What it is and what it does
Judgeval is a Python SDK that instruments LLM-powered agents with OpenTelemetry-based tracing and evaluation. It captures function inputs, outputs, and token usage automatically via decorators like @Tracer.observe(), then runs prompt-based judges—custom scorers you define—to evaluate agent behaviors at scale. These judges produce structured, scored outputs that accumulate into a searchable record of how your agent behaved over time.
The package integrates with multiple LLM providers and frameworks out of the box. You can run judges against live production traffic server-side, or replay them on historical traces to validate fixes before shipping. It includes a CLI for managing traces, judges, and behaviors from the terminal, and an MCP server for querying traces and invoking judges from AI assistants or IDEs.
Use it for
- Detect and triage agent failures in production by running judges on live traffic and surfacing regressions.
- Validate agent fixes by replaying judges against historical traces before shipping to production.
- Build a searchable record of agent behavior over time to understand patterns and recurring issues.
- Instrument custom tools and functions alongside LLM calls to trace end-to-end agent execution.
- Query and inspect agent traces from the CLI or MCP server without writing code.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Judgeval is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and solves a real problem—observing and improving LLM agents in production. It requires Python >=3.10 and API credentials (JUDGMENT_API_KEY, JUDGMENT_ORG_ID), which implies a hosted backend service. No known vulnerabilities. Install it if you need production tracing and behavior-based evaluation for agents.
Install
judgeval on PyPI
Before you install
Low friction: pure Python wheel with eight runtime dependencies. Active maintenance—last commit 2026-08-10, 1057 GitHub stars, released 2026-08-03.
Requires Python >=3.10; JUDGMENT_API_KEY and JUDGMENT_ORG_ID environment variables must be set to send traces to the Judgment backend.
License in practice
Apache-2.0 (permissive): you may use, modify, and distribute this package freely in commercial and private projects, provided you include the license notice and state significant changes.
Quickstart
pip install judgeval
from judgeval import Tracer
Tracer.init(project_name="my-project")
@Tracer.observe(span_type="agent")
def run_agent(question: str) -> str:
return "response"
Verify before relying
- Specific latency overhead of server-side scoring on live production traffic.
- Whether Slack alert configuration is built-in or requires external setup.
- Retention and query performance characteristics of the trace storage backend.
- Cost model for the Judgment backend service (API key required suggests hosted service).
- Auto-instrumentation behavior and token usage capture mechanics with specific LLM providers.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 8 packagesdotenvhttpxopentelemetry-exporter-otlpopentelemetry-sdkorjsonpackagingpathspectyper |
| Maintenance | Actively maintained 11 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 269,024 / month, #8,265 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Operating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: judgeval-1.3.1-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “agent behavior scoring and detection”
- judgevalJudgeval is an SDK for tracing and evaluating LLM-powered agents in…
- cleanlab-tlmCleanlab TLM scores the trustworthiness of LLM responses in…
- detect_agentDetects whether code is running in an AI agent or automated…
Give your agent the search over MCP, or paste the wish link into any chat.
More Monitoring packages
Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.
Provides generated Python code for OpenTelemetry semantic conventions, enabling standardized attribute naming and constant definitions for instrumentation and telemetry collection.
Install it if you are using OpenTelemetry and want to follow semantic conventions correctly.
Provides the reference implementation of the OpenTelemetry API for collecting and exporting traces, metrics, and logs from Python applications.
Provides the abstract API and interfaces for OpenTelemetry instrumentation in Python, defining how to emit traces, metrics, and logs without tying code to a specific SDK implementation.
Exports OpenTelemetry observability data to an OpenTelemetry Collector using Protobuf-encoded messages over HTTP.
Install it if you are using OpenTelemetry in Python and need to send data to a Collector over HTTP.
Provides automatic instrumentation commands and programmatic APIs to inject distributed tracing into Python applications without code changes, detecting and instrumenting packages used by your program.
Install it if you need distributed tracing without code changes and have compatible instrumented packages in your environment.
See also trulens · trulens-core · deepeval · harbor-rewardkit · agentops · opentelemetry-instrumentation-openai-agents · opik · openlit · mlflow · dreadnode