$npx skillfedfor your agent

nemo-gym

NeMo Gym is a library for building reinforcement learning environments

With conditionsPyPI LibrariesReleased Aug 2026857.1K downloads / mopermissive licensePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — nemo_gym-0.5.0-py3-none-any.whl
v0.5.0 · released 2026-08-07 · Python >=3.13.14 · 31 runtime deps: openai, anthropic, tqdm, pydantic, pydantic_core, devtools, fastapi, mcp

Yes, if you need to evaluate or train agents in stateful environments at scale with reproducible verifiers and shared benchmarks. Active maintenance, permissive license, low install friction, and integration with major providers make it solid for agentic workflows. However, the strict Python 3.13.14+ requirement and heavy dependency footprint may conflict with existing projects. Early development status and evolving APIs warrant caution in production.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.13.14 or higher; earlier Python versions are not supported.
  • Installation uses uv package manager, not pip.
  • Requires API keys for model providers (openai, anthropic, or self-hosted via vLLM).

License · maintenance · safety

permissive license (permissive) — Apache License 2.0 is permissive—you can use, modify, and distribute NeMo Gym freely in commercial and private projects, provided you include the license and notice of changes. No viral copyleft obligations.

last release 2026-08-07 (7 days) · last repo commit 2026-08-14 · 1,114 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 857,100 downloads/mo, #4,884 on PyPI

Verify before relying

# Install NeMo Gym (requires Python 3.13.14+)
git clone git@github.com:NVIDIA-NeMo/Gym.git
cd Gym
uv venv --python 3.13.14 && source .venv/bin/activate
uv sync

# Configure model (e.g., OpenAI)
echo 'policy_base_url: https://api.openai.com/v1' > env.yaml
echo 'policy_api_key: <your-key>' >> env.yaml
echo 'policy_model_name: gpt-4.1-2025-04-14' >> env.yaml

# Start servers
gym env start --resources-server mcqa --model-type openai_model

# Run evaluation
gym eval run --no-serve --agent mcqa_simple_agent --input data/example.jsonl --output results.jsonl --limit 5 --num-repeats 1
  • Whether the 31 runtime dependencies are all required for basic usage or only for advanced features.
  • Performance characteristics when scaling to thousands of concurrent environments—actual throughput and latency benchmarks.
  • Stability guarantees given the 'early development' status and 'evolving APIs' warning.
  • Whether Windows WSL2 support is fully tested or experimental.
Same gist for agents: .md · .json

What it is and what it does

NeMo Gym is a framework for building and running evaluation environments where language models and agents interact with tasks to solve problems. It abstracts the evaluation loop—task datasets, agent harnesses, verifiers, and execution state—into modular, extensible components. You define or reuse an environment, configure your model (openai, anthropic, local via vLLM, or hosted providers), and run agents against tasks at scale, collecting trajectories and metrics.

The library is designed for teams needing reproducible, stateful evaluation across shared environments, or transitioning from evaluation into agent optimization and training. It includes a hub of popular benchmarks and agent harnesses, integrates with training frameworks, and provides CLI tools for discovery, validation, and diagnostics. With 31 runtime dependencies (ray, mlflow, wandb, fastapi, pydantic, aiohttp, and others), it trades breadth of built-in capability for a heavier dependency footprint.

Use it for

  • Evaluate agents on standardized benchmarks with reproducible verifiers and shared task datasets across teams.
  • Collect training data by running agents on tasks at scale, generating trajectories for supervised fine-tuning or reinforcement learning.
  • Benchmark agent skill impact by running the same tasks with different skill sets to isolate which capabilities drive performance.
  • Run tool-using agents in isolated sandboxes to safely test code execution and external tool interactions.
  • Transition from evaluation to training by using the same environment and verifier definitions with training frameworks.
  • Diagnose evaluation failures to identify which tasks failed, why, and the highest-impact fixes.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need to evaluate or train agents in stateful environments at scale with reproducible verifiers and shared benchmarks.

Active maintenance, permissive license, low install friction, and integration with major providers make it solid for agentic workflows. However, the strict Python 3.13.14+ requirement and heavy dependency footprint may conflict with existing projects. Early development status and evolving APIs warrant caution in production.

Install

nemo-gym on PyPI

Before you install

Low install friction with a pure Python wheel. Active maintenance with a release 7 days old and 1114 repository stars. However, requires Python 3.13.14 or higher, which is a strict version floor that may conflict with existing projects on older Python versions.

Requires Python 3.13.14 or higher; earlier Python versions are not supported. Installation uses uv package manager, not pip. Requires API keys for model providers (openai, anthropic, or self-hosted via vLLM).

License in practice

Apache License 2.0 is permissive—you can use, modify, and distribute NeMo Gym freely in commercial and private projects, provided you include the license and notice of changes. No viral copyleft obligations.

Quickstart

# Install NeMo Gym (requires Python 3.13.14+)
git clone git@github.com:NVIDIA-NeMo/Gym.git
cd Gym
uv venv --python 3.13.14 && source .venv/bin/activate
uv sync

# Configure model (e.g., OpenAI)
echo 'policy_base_url: https://api.openai.com/v1' > env.yaml
echo 'policy_api_key: <your-key>' >> env.yaml
echo 'policy_model_name: gpt-4.1-2025-04-14' >> env.yaml

# Start servers
gym env start --resources-server mcqa --model-type openai_model

# Run evaluation
gym eval run --no-serve --agent mcqa_simple_agent --input data/example.jsonl --output results.jsonl --limit 5 --num-repeats 1

Verify before relying

  • Whether the 31 runtime dependencies are all required for basic usage or only for advanced features.
  • Performance characteristics when scaling to thousands of concurrent environments—actual throughput and latency benchmarks.
  • Stability guarantees given the 'early development' status and 'evolving APIs' warning.
  • Whether Windows WSL2 support is fully tested or experimental.

Package facts

Licensepermissive license permissive
Python supportCapped below the current Python release >=3.13.14
Install frictionLow. Pure-Python wheel
Runtime dependencies
31 packages
openaianthropictqdmpydanticpydantic_coredevtoolsfastapimcpitsdangerousuvicornuvloophydra-coreomegaconfrichmlflow-skinnymlflowaiohttpyappiraypsutildatasetsorjsonurllib3fonttoolspython-multipartpyarrowwandbGitPythonpyasn1gprof2dot
MaintenanceActively maintained 7 days since the last release
Last repo commit
First released
Downloads857,100 / month, #4,884 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.13Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: MathematicsTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules

Evidence: nemo_gym-0.5.0-py3-none-any.whl

Tags

Capabilities
agent evaluation frameworkreinforcement learning environmentsLLM agent benchmarkingscalable task verificationtool-calling agent trainingstateful environment testingagentic AI evaluation
Topics
agent-evaluationrl-trainingbenchmark-hub
PyPI keywords
reinforcement-learningRLagentagenticLLMlarge-language-modelstraining-datarollouttrajectoryverificationverifierrewardtool-callingfunction-callingdata-collectionNeMoNVIDIA

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “scalable task verification”

  • nemo-gymNeMo Gym provides infrastructure for building, running, and scaling…
  • tbbProvides Python bindings to Intel's oneAPI Threading Building Blocks…
  • tbb-develProvides C++ parallelism primitives and task scheduling for…

Give your agent the search over MCP, or paste the wish link into any chat.

More Libraries packages

urllib3 Worth it
PyPI · Libraries · released May 2026

urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.

MITpure Python · 3.10+
1.8Bdownloads / mo
requests Worth it
PyPI · Libraries · released May 2026

Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.

Apache-2.0pure Python · 3.10+
1.8Bdownloads / mo
pluggy Worth it
PyPI · Libraries · released May 2025

Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.

Install it if you're building an extensible application or framework.

MITpure Python · 3.9+aging
1.3Bdownloads / mo
python-dateutil Worth it
PyPI · Libraries · released Mar 2024

Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.

Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.

Apache-2.0pure Python
1.2Bdownloads / mo
six With conditions
PyPI · Libraries · released Dec 2024

Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.

MITpure Python
1.2Bdownloads / mo
pytest Worth it
PyPI · Libraries · released Jun 2026

pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.

MITpure Python · 3.10+
1.1Bdownloads / mo

See also openenv-core · reasoning-gym · fhaviary · gem-llm · verifiers · nvidia-nat-core · nemo-evaluator · kaggle-environments · browsergym-core · gymnasium

Further reading