$npx skillfedfor your agent

dtlpymetrics

Scoring and metrics app

With conditionsPyPI Quality AssuranceReleased Feb 2026506.0K downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — dtlpymetrics-1.2.32-py3-none-any.whl
v1.2.32 · released 2026-02-26 · 4 runtime deps: dtlpy, shapely, matplotlib, seaborn

Yes, if you use Dataloop for annotation workflows and need automated quality scoring. The package is actively maintained, has low install friction, and fills a specific role in Dataloop's quality assurance ecosystem. However, verify the license before use in proprietary contexts, and confirm that dtlpy is available and configured in your environment.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires dtlpy, which is Dataloop's Python SDK; you must have a Dataloop account and environment to use this package meaningfully.
  • Low install friction with a pure-Python wheel.
  • Actively maintained as of February 2026.

License · maintenance · safety

(unclear) — License status is unclear; no SPDX identifier or raw license text is available in the package metadata. Verify the actual license before using in proprietary or copyleft-sensitive contexts.

last release 2026-02-26 (169 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 506,019 downloads/mo, #6,290 on PyPI

Verify before relying

pip install dtlpymetrics

import dtlpymetrics
# Use functions to calculate scores for quality tasks and model predictions
# See docs/dtlpymetrics_fxns.md for available functions
  • Whether the package requires specific versions of dtlpy, shapely, matplotlib, or seaborn beyond what pip resolves.
  • Whether confusion scores are actually calculated for video items, despite the note that they are not due to multi-frame nature.
  • Performance characteristics when scoring large datasets or high-resolution videos.
Same gist for agents: .md · .json

What it is and what it does

dtlpymetrics is a scoring and metrics application for the Dataloop platform that automates quality assessment of annotations. It calculates multiple layers of scores—raw annotation scores (geometry overlap via IOU or distance, label matching, and attribute matching), per-annotation overall scores, per-user confusion scores, per-item label confusion counts, and per-item overall scores—across image and video datasets. It supports classification, bounding box, polygon, segmentation, and point annotations.

The package integrates with Dataloop's quality workflows (qualification tasks, honeypot tasks, and consensus tasks) and can be added as custom nodes to pipelines to compute scores automatically when quality items are completed. For videos, scores are calculated frame-by-frame and then aggregated per annotation; confusion scores are omitted for videos due to their multi-frame nature. Any calculated scores replace previous scores for all items in a task.

Use it for

  • Measure inter-annotator agreement in consensus tasks by comparing each annotator's annotations against all others and generating confusion matrices.
  • Validate annotator quality in honeypot and qualification tasks by scoring their annotations against ground truth and tracking user confusion scores over time.
  • Automate quality scoring in image annotation pipelines by adding scoring nodes that trigger when quality task assignments are completed.
  • Assess label confusion for specific classes to identify which labels annotators confuse most often, guiding retraining or clarification.
  • Evaluate geometry accuracy (bounding box overlap, polygon IOU, point distance) alongside label correctness to identify systematic annotation errors.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you use Dataloop for annotation workflows and need automated quality scoring.

The package is actively maintained, has low install friction, and fills a specific role in Dataloop's quality assurance ecosystem. However, verify the license before use in proprietary contexts, and confirm that dtlpy is available and configured in your environment.

Install

dtlpymetrics on PyPI

Before you install

Low install friction with a pure-Python wheel. Actively maintained as of February 2026. Depends on dtlpy, shapely, matplotlib, and seaborn—all stable, widely-used libraries.

Requires dtlpy, which is Dataloop's Python SDK; you must have a Dataloop account and environment to use this package meaningfully.

License in practice

License status is unclear; no SPDX identifier or raw license text is available in the package metadata. Verify the actual license before using in proprietary or copyleft-sensitive contexts.

Quickstart

pip install dtlpymetrics

import dtlpymetrics
# Use functions to calculate scores for quality tasks and model predictions
# See docs/dtlpymetrics_fxns.md for available functions

Verify before relying

  • Whether the package requires specific versions of dtlpy, shapely, matplotlib, or seaborn beyond what pip resolves.
  • Whether confusion scores are actually calculated for video items, despite the note that they are not due to multi-frame nature.
  • Performance characteristics when scoring large datasets or high-resolution videos.

Package facts

LicenseNot declared unclear
Python supportNot specified
Install frictionLow. Pure-Python wheel
Runtime dependencies
4 packages
dtlpyshapelymatplotlibseaborn
MaintenanceActively maintained 169 days since the last release
First released
Downloads506,019 / month, #6,290 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3.10Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: dtlpymetrics-1.2.32-py3-none-any.whl

Tags

Capabilities
annotation quality scoringannotator agreement metricsconsensus task evaluationhoneypot qualification scoringannotation confusion matrixlabel agreement measurementgeometry overlap scoring
Topics
annotation-qualitydataloop-integrationmetrics

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “annotation quality scoring”

  • dtlpymetricsCalculates and generates quality scores for annotations in image and…
  • udtoolsValidates CoNLL-U format files against Universal Dependencies…
  • unbabel-cometEvaluates machine translation quality using neural metrics, scoring…

Give your agent the search over MCP, or paste the wish link into any chat.

More Quality Assurance packages

coverage Worth it
PyPI · Testing · released Aug 2026

Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.

Install it if you want to measure test completeness or enforce coverage thresholds in your project.

permissive licensepure Python · 3.10+
335.8Mdownloads / mo
ruff Worth it
PyPI · Python Modules · released Aug 2026

Ruff is a Python linter and code formatter written in Rust that combines linting, formatting, and code fixing into a single tool, replacing Flake8, Black, isort, and related utilities.

MITcompiled wheel · 3.7+
316.1Mdownloads / mo
pexpect With conditions
PyPI · Software Development · released Nov 2023

Pexpect spawns and controls interactive console applications by sending input and matching output patterns, automating tasks that would otherwise require manual interaction.

ISCpure Pythonaging
200.8Mdownloads / mo
black Worth it
PyPI · Python Modules · released May 2026

Black reformats Python source code to a consistent style by parsing entire files and rewriting them according to an opinionated, deterministic set of rules, eliminating manual formatting decisions.

MITpure Python · 3.10+
179.9Mdownloads / mo
pytest-xdist Worth it
PyPI · Utilities · released Jul 2025

pytest-xdist distributes pytest tests across multiple CPU cores or machines to speed up test execution, with the simplest usage being `pytest -n auto` to spawn workers equal to available CPUs.

Install it if your test suite takes long enough that parallelization would save meaningful time.

MITpure Python · 3.9+
177.1Mdownloads / mo
cfn-lint Worth it
PyPI · Quality Assurance · released Aug 2026

Validates AWS CloudFormation templates in YAML or JSON format against resource provider schemas and best practices, checking property values and configuration correctness.

Install it if you work with CloudFormation templates.

MIT-0pure Python
114.9Mdownloads / mo

See also unbabel-comet · rouge · dbt-score · label-studio · cleanlab · bert-score · rouge-chinese · sentence-transformers · starlette-exporter · Sift

Further reading