$npx skillfedfor your agent

judgeval

The open source post-building layer for Agent Behavior Monitoring.

Worth itPyPI MonitoringReleased Aug 2026269.0K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — judgeval-1.3.1-py3-none-any.whl
v1.3.1 · released 2026-08-03 · Python >=3.10 · 8 runtime deps: dotenv, httpx, opentelemetry-exporter-otlp, opentelemetry-sdk, orjson, packaging, pathspec, typer

Yes. Judgeval is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and solves a real problem—observing and improving LLM agents in production. It requires Python >=3.10 and API credentials (JUDGMENT_API_KEY, JUDGMENT_ORG_ID), which implies a hosted backend service. No known vulnerabilities. Install it if you need production tracing and behavior-based evaluation for agents.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >=3.10; JUDGMENT_API_KEY and JUDGMENT_ORG_ID environment variables must be set to send traces to the Judgment backend.
  • Low friction: pure Python wheel with eight runtime dependencies.
  • Active maintenance—last commit 2026-08-10, 1057 GitHub stars, released 2026-08-03.

License · maintenance · safety

Apache-2.0 (permissive) — Apache-2.0 (permissive): you may use, modify, and distribute this package freely in commercial and private projects, provided you include the license notice and state significant changes.

last release 2026-08-03 (11 days) · last repo commit 2026-08-10 · 1,057 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 269,024 downloads/mo, #8,265 on PyPI

Verify before relying

pip install judgeval

from judgeval import Tracer

Tracer.init(project_name="my-project")

@Tracer.observe(span_type="agent")
def run_agent(question: str) -> str:
    return "response"
  • Specific latency overhead of server-side scoring on live production traffic.
  • Whether Slack alert configuration is built-in or requires external setup.
  • Retention and query performance characteristics of the trace storage backend.
  • Cost model for the Judgment backend service (API key required suggests hosted service).
  • Auto-instrumentation behavior and token usage capture mechanics with specific LLM providers.
Same gist for agents: .md · .json

What it is and what it does

Judgeval is a Python SDK that instruments LLM-powered agents with OpenTelemetry-based tracing and evaluation. It captures function inputs, outputs, and token usage automatically via decorators like @Tracer.observe(), then runs prompt-based judges—custom scorers you define—to evaluate agent behaviors at scale. These judges produce structured, scored outputs that accumulate into a searchable record of how your agent behaved over time.

The package integrates with multiple LLM providers and frameworks out of the box. You can run judges against live production traffic server-side, or replay them on historical traces to validate fixes before shipping. It includes a CLI for managing traces, judges, and behaviors from the terminal, and an MCP server for querying traces and invoking judges from AI assistants or IDEs.

Use it for

  • Detect and triage agent failures in production by running judges on live traffic and surfacing regressions.
  • Validate agent fixes by replaying judges against historical traces before shipping to production.
  • Build a searchable record of agent behavior over time to understand patterns and recurring issues.
  • Instrument custom tools and functions alongside LLM calls to trace end-to-end agent execution.
  • Query and inspect agent traces from the CLI or MCP server without writing code.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Judgeval is actively maintained, has low install friction, carries a permissive Apache-2.0 license, and solves a real problem—observing and improving LLM agents in production. It requires Python >=3.10 and API credentials (JUDGMENT_API_KEY, JUDGMENT_ORG_ID), which implies a hosted backend service. No known vulnerabilities. Install it if you need production tracing and behavior-based evaluation for agents.

Install

judgeval on PyPI

Before you install

Low friction: pure Python wheel with eight runtime dependencies. Active maintenance—last commit 2026-08-10, 1057 GitHub stars, released 2026-08-03.

Requires Python >=3.10; JUDGMENT_API_KEY and JUDGMENT_ORG_ID environment variables must be set to send traces to the Judgment backend.

License in practice

Apache-2.0 (permissive): you may use, modify, and distribute this package freely in commercial and private projects, provided you include the license notice and state significant changes.

Quickstart

pip install judgeval

from judgeval import Tracer

Tracer.init(project_name="my-project")

@Tracer.observe(span_type="agent")
def run_agent(question: str) -> str:
    return "response"

Verify before relying

  • Specific latency overhead of server-side scoring on live production traffic.
  • Whether Slack alert configuration is built-in or requires external setup.
  • Retention and query performance characteristics of the trace storage backend.
  • Cost model for the Judgment backend service (API key required suggests hosted service).
  • Auto-instrumentation behavior and token usage capture mechanics with specific LLM providers.

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
8 packages
dotenvhttpxopentelemetry-exporter-otlpopentelemetry-sdkorjsonpackagingpathspectyper
MaintenanceActively maintained 11 days since the last release
Last repo commit
First released
Downloads269,024 / month, #8,265 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Operating System :: OS IndependentProgramming Language :: Python :: 3

Evidence: judgeval-1.3.1-py3-none-any.whl

Tags

Capabilities
LLM agent monitoring and evaluationOpenTelemetry tracing for agentsagent behavior scoring and detectionproduction LLM observabilityagent failure detection and triageprompt-based agent judgesagent trace replay and validation
Topics
llm-observabilityagent-evaluationopentelemetry

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “agent behavior scoring and detection”

  • judgevalJudgeval is an SDK for tracing and evaluating LLM-powered agents in…
  • cleanlab-tlmCleanlab TLM scores the trustworthiness of LLM responses in…
  • detect_agentDetects whether code is running in an AI agent or automated…

Give your agent the search over MCP, or paste the wish link into any chat.

More Monitoring packages

tqdm Worth it
PyPI · Libraries · released Jul 2026

Wraps any iterable to display a real-time progress bar in the terminal or Jupyter notebook, showing iteration count, elapsed time, and estimated time remaining.

copyleftpure Python · 3.8+
648.6Mdownloads / mo
opentelemetry-semantic-conventions Worth it
PyPI · Monitoring · released Jul 2026

Provides generated Python code for OpenTelemetry semantic conventions, enabling standardized attribute naming and constant definitions for instrumentation and telemetry collection.

Install it if you are using OpenTelemetry and want to follow semantic conventions correctly.

Apache-2.0pure Python · 3.10+
542.9Mdownloads / mo
opentelemetry-sdk Worth it
PyPI · Monitoring · released Jul 2026

Provides the reference implementation of the OpenTelemetry API for collecting and exporting traces, metrics, and logs from Python applications.

Apache-2.0pure Python · 3.10+
521.8Mdownloads / mo
opentelemetry-api With conditions
PyPI · Monitoring · released Jul 2026

Provides the abstract API and interfaces for OpenTelemetry instrumentation in Python, defining how to emit traces, metrics, and logs without tying code to a specific SDK implementation.

Apache-2.0pure Python · 3.10+
463.8Mdownloads / mo
opentelemetry-exporter-otlp-proto-http Worth it
PyPI · Monitoring · released Jul 2026

Exports OpenTelemetry observability data to an OpenTelemetry Collector using Protobuf-encoded messages over HTTP.

Install it if you are using OpenTelemetry in Python and need to send data to a Collector over HTTP.

Apache-2.0pure Python · 3.10+
409.9Mdownloads / mo
opentelemetry-instrumentation Worth it
PyPI · Monitoring · released Jul 2026

Provides automatic instrumentation commands and programmatic APIs to inject distributed tracing into Python applications without code changes, detecting and instrumenting packages used by your program.

Install it if you need distributed tracing without code changes and have compatible instrumented packages in your environment.

Apache-2.0pure Python · 3.10+
393.5Mdownloads / mo

See also trulens · trulens-core · deepeval · harbor-rewardkit · agentops · opentelemetry-instrumentation-openai-agents · opik · openlit · mlflow · dreadnode

Further reading