$npx skillfedfor your agent

trulens-core

Library to systematically track and evaluate LLM based applications.

Worth itPyPI Quality AssuranceReleased Aug 2026126.1K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — trulens_core-2.13.0-py3-none-any.whl
v2.13.0 · released 2026-08-14 · Python <4.0,>=3.9 · 19 runtime deps: alembic, dill, importlib-resources, munch, nest-asyncio, numpy, opentelemetry-api, opentelemetry-proto

Yes. TruLens is actively maintained, has no known vulnerabilities, and solves a real problem for LLM application teams: making agent behavior traceable and measurable. The OpenTelemetry foundation ensures portability, and the built-in evaluators are benchmarked against human annotations. Install it if you need to instrument and evaluate LLM apps; skip it if you only need basic logging.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later; LLM provider packages (e.g., trulens-providers-openai) must be installed separately for feedback evaluation.
  • Low friction: pure Python wheel with 19 runtime dependencies including well-established libraries like SQLAlchemy, Pydantic, and OpenTelemetry.
  • Active maintenance with release on 2026-08-14 and 3508 GitHub stars.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions; suitable for most projects.

last release 2026-08-14 (0 days) · last repo commit 2026-08-14 · 3,508 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 126,116 downloads/mo, #11,789 on PyPI

Verify before relying

pip install trulens-core

from trulens.core.otel.instrument import instrument
from trulens.otel.semconv.trace import SpanAttributes

class MyRAG:
    @instrument(
        span_type=SpanAttributes.SpanType.RETRIEVAL,
        attributes={
            SpanAttributes.RETRIEVAL.QUERY_TEXT: "query",
            SpanAttributes.RETRIEVAL.RETRIEVED_CONTEXTS: "return",
        },
    )
    def retrieve(self, query: str) -> list:
        pass
  • Whether the package works with Python 3.14 in practice, given it is marked as Alpha status.
  • Performance characteristics when evaluating large-scale batch runs with the Run API.
  • Compatibility with OTLP backends beyond those mentioned (Jaeger, Grafana Tempo, Datadog).
Same gist for agents: .md · .json

What it is and what it does

TruLens is an instrumentation and evaluation framework for LLM-based applications built on OpenTelemetry. It decorates functions to capture structured traces of every step—retrieval, LLM calls, tool invocations—recording latency, tokens, and cost per operation. Once traces are collected, the package evaluates them using LLM judges that score and explain their reasoning, enabling you to identify where applications fail, compare versions, and measure quality tradeoffs against cost.

The package ships with seven agentic evaluators (LogicalConsistency, ExecutionEfficiency, PlanAdherence, ToolSelection, and others) and supports both inline evaluation as your app runs and batch evaluation over pre-collected datasets. Because tracing is OpenTelemetry-native, traces are portable to any OTLP-compatible backend—Jaeger, Grafana Tempo, Datadog—making it interoperable with existing observability infrastructure. Runtime dependencies include SQLAlchemy for storage, Pydantic for validation, and OpenTelemetry libraries for instrumentation.

Use it for

  • Instrument a RAG pipeline to trace retrieval, ranking, and generation steps, then score groundedness and context relevance per query.
  • Compare two versions of an agent by running both against a dataset and measuring latency, token cost, and evaluation scores side-by-side.
  • Debug agent failures by examining structured traces that show which tool was called, what arguments were passed, and where reasoning broke down.
  • Evaluate agentic systems for plan quality and tool selection by running purpose-built judges that flag redundant steps or incorrect tool choices.
  • Export traces to an existing observability backend (Datadog, Grafana) while running evaluations in TruLens without duplicating instrumentation.
  • Batch-evaluate a historical dataset of app runs using the Run API to compute metrics offline and identify systematic quality issues.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

TruLens is actively maintained, has no known vulnerabilities, and solves a real problem for LLM application teams: making agent behavior traceable and measurable. The OpenTelemetry foundation ensures portability, and the built-in evaluators are benchmarked against human annotations. Install it if you need to instrument and evaluate LLM apps; skip it if you only need basic logging.

Install

trulens-core on PyPI

Before you install

Low friction: pure Python wheel with 19 runtime dependencies including well-established libraries like SQLAlchemy, Pydantic, and OpenTelemetry. Active maintenance with release on 2026-08-14 and 3508 GitHub stars.

Requires Python 3.9 or later; LLM provider packages (e.g., trulens-providers-openai) must be installed separately for feedback evaluation.

License in practice

MIT license permits commercial and private use with minimal restrictions; suitable for most projects.

Quickstart

pip install trulens-core

from trulens.core.otel.instrument import instrument
from trulens.otel.semconv.trace import SpanAttributes

class MyRAG:
    @instrument(
        span_type=SpanAttributes.SpanType.RETRIEVAL,
        attributes={
            SpanAttributes.RETRIEVAL.QUERY_TEXT: "query",
            SpanAttributes.RETRIEVAL.RETRIEVED_CONTEXTS: "return",
        },
    )
    def retrieve(self, query: str) -> list:
        pass

Verify before relying

  • Whether the package works with Python 3.14 in practice, given it is marked as Alpha status.
  • Performance characteristics when evaluating large-scale batch runs with the Run API.
  • Compatibility with OTLP backends beyond those mentioned (Jaeger, Grafana Tempo, Datadog).

Package facts

LicenseMIT permissive
Python supportSupports the current Python release <4.0,>=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
19 packages
alembicdillimportlib-resourcesmunchnest-asyncionumpyopentelemetry-apiopentelemetry-protoopentelemetry-sdkpackagingpandaspydanticpython-dotenvrequestsrichsqlalchemytrulens-otel-semconvtyping_extensionswrapt
MaintenanceActively maintained 0 days since the last release
Last repo commit
First released
Downloads126,116 / month, #11,789 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 3 - AlphaLicense :: OSI Approved :: MIT LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Programming Language :: Python :: 3.14Programming Language :: Python :: 3.9

Evidence: trulens_core-2.13.0-py3-none-any.whl

Tags

Capabilities
llm application tracing and evaluationopentelemetry instrumentation for agentsrag evaluation and monitoringllm judge feedback functionsagent performance debuggingtrace-based cost and quality analysisagentic system evaluation
Topics
llm-evaluationobservabilityagent-debugging

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “llm judge feedback functions”

  • trulens-coreInstruments LLM applications with OpenTelemetry-native tracing and…
  • trulensTruLens instruments LLM applications to trace execution step-by-step,…
  • harbor-rewardkitHarbor Rewardkit defines and runs verifiers for task evaluation,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Quality Assurance packages

coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
ruff Worth it
PyPI · Python Modules · released Aug 2026

Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.

MITcompiled wheel · 3.7+
316.1Mdownloads / mo
pexpect With conditions
PyPI · Software Development · released Nov 2023

Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.

ISCpure Pythonaging
200.8Mdownloads / mo
black Worth it
PyPI · Python Modules · released May 2026

Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.

MITpure Python · 3.10+
179.9Mdownloads / mo
pytest-xdist Worth it
PyPI · Utilities · released Jul 2025

pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.

Install it if your test suite takes long enough that parallelization would save meaningful time.

MITpure Python · 3.9+
177.1Mdownloads / mo
cfn-lint Worth it
PyPI · Quality Assurance · released Aug 2026

Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.

Install it if you work with CloudFormation templates.

MIT-0pure Python
114.9Mdownloads / mo

See also trulens · judgeval · cleanlab-tlm · ragas · opentelemetry-instrumentation-openai-agents · opentelemetry-instrumentation-remoulade · azure-core-tracing-opentelemetry · opentelemetry-instrumentation-sklearn · opentelemetry-instrumentation-celery · opentelemetry-instrumentation-llamaindex