harbor-rewardkit
Lightweight grading toolkit for environment-based tasks.
Decision gist · record as of 2026-08-14
Yes, if you need lightweight task evaluation with LLM or agent judges. The package is actively maintained, has low install friction, and offers a clean API for both simple checks and complex agent-based grading. Best suited for Harbor framework users or teams building evaluation pipelines for AI-generated outputs.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.12 or later.
- Low friction: pure Python wheel with a single runtime dependency (litellm).
- Active maintenance with recent releases; last commit 2026-08-13.
License · maintenance · safety
Apache-2.0 (permissive) — Apache-2.0 permissive license allows commercial and private use with attribution; no copyleft restrictions.
last release 2026-06-27 (48 days) · last repo commit 2026-08-13 · 4,249 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 7,739,787 downloads/mo, #1,700 on PyPI
Alternatives
Verify before relying
# Install
uv tool install harbor-rewardkit
# Use programmatic criteria
from rewardkit import criteria
criteria.file_exists("output.txt")
criteria.file_contains("output.txt", "hello")- Whether litellm's dependencies add meaningful install friction or platform-specific requirements beyond what the wheel distribution handles.
- Whether the package works standalone or requires Harbor task format integration for full functionality.
What it is and what it does
Harbor Rewardkit is a lightweight grading toolkit for evaluating environment-based tasks. It lets you define verifiers using three approaches: simple programmatic criteria (file checks), LLM judges that score outputs using Claude or other models, and agent judges that can call MCP server tools during evaluation. The package integrates with the Harbor task framework but can be used independently.
You define criteria in TOML files or Python code, then run the verifier against task outputs. It handles LLM model selection, tool access control via allowed_tools lists, and binary or scored evaluation results. The single runtime dependency is litellm, which abstracts LLM provider APIs.
Use it for
- Score code quality or correctness by having Claude evaluate generated source files against criteria.
- Verify task outputs with simple file-based checks (existence, content matching) without LLM overhead.
- Build agent judges that navigate web pages or interact with systems via MCP tools to validate rendered output.
- Grade benchmark submissions across multiple criteria in a CI/CD pipeline.
- Evaluate multi-step task completion by combining programmatic checks with LLM judgment.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need lightweight task evaluation with LLM or agent judges.
The package is actively maintained, has low install friction, and offers a clean API for both simple checks and complex agent-based grading. Best suited for Harbor framework users or teams building evaluation pipelines for AI-generated outputs.
Install
harbor-rewardkit on PyPI
Before you install
Low friction: pure Python wheel with a single runtime dependency (litellm). Active maintenance with recent releases; last commit 2026-08-13. Supports Python 3.12 and 3.13.
Requires Python 3.12 or later.
License in practice
Apache-2.0 permissive license allows commercial and private use with attribution; no copyleft restrictions.
Quickstart
# Install
uv tool install harbor-rewardkit
# Use programmatic criteria
from rewardkit import criteria
criteria.file_exists("output.txt")
criteria.file_contains("output.txt", "hello")
Verify before relying
- Whether litellm's dependencies add meaningful install friction or platform-specific requirements beyond what the wheel distribution handles.
- Whether the package works standalone or requires Harbor task format integration for full functionality.
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.12 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 1 packagelitellm |
| Maintenance | Actively maintained 48 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 7,739,787 / month, #1,700 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: DevelopersLicense :: OSI Approved :: Apache Software LicenseProgramming Language :: Python :: 3Programming Language :: Python :: 3.12Programming Language :: Python :: 3.13Topic :: Software Development :: Testing |
Evidence: harbor_rewardkit-0.1.7-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “task evaluation framework”
- harbor-rewardkitHarbor Rewardkit defines and runs verifiers for task evaluation,…
- terminal-benchTerminal-Bench provides a benchmark suite and execution harness for…
- nemo-gymNeMo Gym provides infrastructure for building, running, and scaling…
Give your agent the search over MCP, or paste the wish link into any chat.
More Testing packages
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
virtualenv creates isolated Python environments where packages can be installed independently without affecting the system Python or other projects.
Coverage.py measures which lines of Python code are executed during test runs, reporting coverage percentages and identifying untested code paths.
Install it if you want to measure test completeness or enforce coverage thresholds in your project.
pytest-asyncio is a pytest plugin that enables writing and running async test functions using the asyncio library, allowing developers to await code directly within test cases.
Install it if you write tests for any asyncio-based code.
A pytest plugin that generates test reports in Common Test Report Format (CTRF) as JSON, compatible with pytest-xdist and pytest-playwright for distributed and browser-based testing.
Install it if you need CTRF-formatted test output for CI/CD integration or cross-tool reporting.
See also harbor · judgeval · harborapi · terminal-bench · nemo-evaluator · livekit-plugins-xai · nemo-gym · crewai-tools · fast-agent-mcp · livekit-plugins-groq