cleanlab-tlm
Python client library for Cleanlab Trustworthy Language Model
Decision gist · record as of 2026-08-14
Yes, if you need real-time hallucination detection for production LLM systems and can accept a remote API dependency. The package is stable, has low install friction, and covers a genuine gap in local hallucination detection. However, the aging maintenance signal (only 24 stars, no recent activity beyond the last commit date) and reliance on an external API key and service mean you should verify the API's stability and cost structure for your use case before committing to it at scale.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires a free API key from https://tlm.cleanlab.ai/ set as the CLEANLAB_TLM_API_KEY environment variable.
- Low friction install with standard dependencies.
- The package is marked Production/Stable and supports Python 3.9 through 3.12.
License · maintenance · safety
MIT (permissive) — MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for commercial and open-source projects alike.
last release 2025-11-21 (266 days) · last repo commit 2025-12-09 · 24 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 163,721 downloads/mo, #10,564 on PyPI
Alternatives
Verify before relying
pip install cleanlab-tlm
from cleanlab_tlm import TLM
tlm = TLM(options={"log": ["explanation"]})
result = tlm.get_trustworthiness_score(
prompt="What's the third month of the year alphabetically?",
response="August"
)
print(result["trustworthiness_score"])- Actual latency and throughput characteristics for production-scale deployments.
- Whether the API key tier affects rate limits or feature availability.
- Specific LLM models and frameworks the package has been tested with.
- Cost structure for API usage beyond the free tier.
What it is and what it does
Cleanlab TLM is a Python client for a remote trustworthiness-scoring service that evaluates LLM outputs in real-time. It wraps HTTP calls to Cleanlab's API to assign a confidence score (0–1) to any LLM response, flagging likely hallucinations or incorrect answers by comparing the response against the prompt. The package can either score responses you've already generated from another LLM, or generate and score responses itself by calling a model (GPT, Claude, etc.) on your behalf.
The service is designed for production use in RAG systems, agents, and data-extraction pipelines where incorrect LLM outputs carry real cost. It depends on aiohttp, requests, pandas, tqdm, and a few utility libraries—all standard, low-friction dependencies. You must obtain and configure a free API key before use, and all scoring happens remotely on Cleanlab's servers, not locally.
Use it for
- Score responses from your own LLM pipeline before returning them to users, filtering out low-confidence outputs.
- Build a RAG system that flags uncertain retrieval-augmented answers and triggers fallback logic or human review.
- Evaluate data-extraction or tagging tasks where hallucinated field values would corrupt downstream processes.
- Monitor agent outputs in real-time to detect when an agent has generated an unreliable response.
- Benchmark the quality of different LLM models or prompts by comparing their trustworthiness score distributions.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need real-time hallucination detection for production LLM systems and can accept a remote API dependency.
The package is stable, has low install friction, and covers a genuine gap in local hallucination detection. However, the aging maintenance signal (only 24 stars, no recent activity beyond the last commit date) and reliance on an external API key and service mean you should verify the API's stability and cost structure for your use case before committing to it at scale.
Install
cleanlab-tlm on PyPI
Before you install
Low friction install with standard dependencies. The package is marked Production/Stable and supports Python 3.9 through 3.12. Last commit was recent (2025-12-09), but the project shows aging maintenance signals with only 24 repository stars.
Requires a free API key from https://tlm.cleanlab.ai/ set as the CLEANLAB_TLM_API_KEY environment variable.
License in practice
MIT license permits unrestricted use, modification, and distribution with minimal restrictions—suitable for commercial and open-source projects alike.
Quickstart
pip install cleanlab-tlm
from cleanlab_tlm import TLM
tlm = TLM(options={"log": ["explanation"]})
result = tlm.get_trustworthiness_score(
prompt="What's the third month of the year alphabetically?",
response="August"
)
print(result["trustworthiness_score"])
Verify before relying
- Actual latency and throughput characteristics for production-scale deployments.
- Whether the API key tier affects rate limits or feature availability.
- Specific LLM models and frameworks the package has been tested with.
- Cost structure for API usage beyond the free tier.
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagesaiohttpnest-asynciopandasrequestssemvertqdmtyping-extensions |
| Maintenance | Aging 266 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 163,721 / month, #10,564 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableProgramming Language :: PythonProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9Programming Language :: Python :: Implementation :: CPythonProgramming Language :: Python :: Implementation :: PyPy |
Evidence: cleanlab_tlm-1.1.39-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “trustworthiness scoring for language models”
- cleanlab-tlmCleanlab TLM scores the trustworthiness of LLM responses in…
- nemo-evaluatorNeMo Evaluator runs standardized benchmarks against language models…
- detoxifyDetoxify classifies text comments for toxic content using pre-trained…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also cleanlab · deepeval · trulens · arize-phoenix-evals · ragas · trulens-core · opik · unbabel-comet · llama-index · strands-agents-evals