nemo-gym
NeMo Gym is a library for building reinforcement learning environments
Decision gist · record as of 2026-08-14
Yes, if you need to evaluate or train agents in stateful environments at scale with reproducible verifiers and shared benchmarks. Active maintenance, permissive license, low install friction, and integration with major providers make it solid for agentic workflows. However, the strict Python 3.13.14+ requirement and heavy dependency footprint may conflict with existing projects. Early development status and evolving APIs warrant caution in production.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python 3.13.14 or higher; earlier Python versions are not supported.
- Installation uses uv package manager, not pip.
- Requires API keys for model providers (openai, anthropic, or self-hosted via vLLM).
License · maintenance · safety
permissive license (permissive) — Apache License 2.0 is permissive—you can use, modify, and distribute NeMo Gym freely in commercial and private projects, provided you include the license and notice of changes. No viral copyleft obligations.
last release 2026-08-07 (7 days) · last repo commit 2026-08-14 · 1,114 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 857,100 downloads/mo, #4,884 on PyPI
Alternatives
Verify before relying
# Install NeMo Gym (requires Python 3.13.14+)
git clone git@github.com:NVIDIA-NeMo/Gym.git
cd Gym
uv venv --python 3.13.14 && source .venv/bin/activate
uv sync
# Configure model (e.g., OpenAI)
echo 'policy_base_url: https://api.openai.com/v1' > env.yaml
echo 'policy_api_key: <your-key>' >> env.yaml
echo 'policy_model_name: gpt-4.1-2025-04-14' >> env.yaml
# Start servers
gym env start --resources-server mcqa --model-type openai_model
# Run evaluation
gym eval run --no-serve --agent mcqa_simple_agent --input data/example.jsonl --output results.jsonl --limit 5 --num-repeats 1- Whether the 31 runtime dependencies are all required for basic usage or only for advanced features.
- Performance characteristics when scaling to thousands of concurrent environments—actual throughput and latency benchmarks.
- Stability guarantees given the 'early development' status and 'evolving APIs' warning.
- Whether Windows WSL2 support is fully tested or experimental.
What it is and what it does
NeMo Gym is a framework for building and running evaluation environments where language models and agents interact with tasks to solve problems. It abstracts the evaluation loop—task datasets, agent harnesses, verifiers, and execution state—into modular, extensible components. You define or reuse an environment, configure your model (openai, anthropic, local via vLLM, or hosted providers), and run agents against tasks at scale, collecting trajectories and metrics.
The library is designed for teams needing reproducible, stateful evaluation across shared environments, or transitioning from evaluation into agent optimization and training. It includes a hub of popular benchmarks and agent harnesses, integrates with training frameworks, and provides CLI tools for discovery, validation, and diagnostics. With 31 runtime dependencies (ray, mlflow, wandb, fastapi, pydantic, aiohttp, and others), it trades breadth of built-in capability for a heavier dependency footprint.
Use it for
- Evaluate agents on standardized benchmarks with reproducible verifiers and shared task datasets across teams.
- Collect training data by running agents on tasks at scale, generating trajectories for supervised fine-tuning or reinforcement learning.
- Benchmark agent skill impact by running the same tasks with different skill sets to isolate which capabilities drive performance.
- Run tool-using agents in isolated sandboxes to safely test code execution and external tool interactions.
- Transition from evaluation to training by using the same environment and verifier definitions with training frameworks.
- Diagnose evaluation failures to identify which tasks failed, why, and the highest-impact fixes.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need to evaluate or train agents in stateful environments at scale with reproducible verifiers and shared benchmarks.
Active maintenance, permissive license, low install friction, and integration with major providers make it solid for agentic workflows. However, the strict Python 3.13.14+ requirement and heavy dependency footprint may conflict with existing projects. Early development status and evolving APIs warrant caution in production.
Install
nemo-gym on PyPI
Before you install
Low install friction with a pure Python wheel. Active maintenance with a release 7 days old and 1114 repository stars. However, requires Python 3.13.14 or higher, which is a strict version floor that may conflict with existing projects on older Python versions.
Requires Python 3.13.14 or higher; earlier Python versions are not supported. Installation uses uv package manager, not pip. Requires API keys for model providers (openai, anthropic, or self-hosted via vLLM).
License in practice
Apache License 2.0 is permissive—you can use, modify, and distribute NeMo Gym freely in commercial and private projects, provided you include the license and notice of changes. No viral copyleft obligations.
Quickstart
# Install NeMo Gym (requires Python 3.13.14+)
git clone git@github.com:NVIDIA-NeMo/Gym.git
cd Gym
uv venv --python 3.13.14 && source .venv/bin/activate
uv sync
# Configure model (e.g., OpenAI)
echo 'policy_base_url: https://api.openai.com/v1' > env.yaml
echo 'policy_api_key: <your-key>' >> env.yaml
echo 'policy_model_name: gpt-4.1-2025-04-14' >> env.yaml
# Start servers
gym env start --resources-server mcqa --model-type openai_model
# Run evaluation
gym eval run --no-serve --agent mcqa_simple_agent --input data/example.jsonl --output results.jsonl --limit 5 --num-repeats 1
Verify before relying
- Whether the 31 runtime dependencies are all required for basic usage or only for advanced features.
- Performance characteristics when scaling to thousands of concurrent environments—actual throughput and latency benchmarks.
- Stability guarantees given the 'early development' status and 'evolving APIs' warning.
- Whether Windows WSL2 support is fully tested or experimental.
Package facts
| License | permissive license permissive |
| Python support | Capped below the current Python release >=3.13.14 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 31 packagesopenaianthropictqdmpydanticpydantic_coredevtoolsfastapimcpitsdangerousuvicornuvloophydra-coreomegaconfrichmlflow-skinnymlflowaiohttpyappiraypsutildatasetsorjsonurllib3fonttoolspython-multipartpyarrowwandbGitPythonpyasn1gprof2dot |
| Maintenance | Actively maintained 7 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 857,100 / month, #4,884 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaEnvironment :: ConsoleIntended Audience :: DevelopersIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseNatural Language :: EnglishOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.13Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Scientific/Engineering :: MathematicsTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules |
Evidence: nemo_gym-0.5.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “scalable task verification”
- nemo-gymNeMo Gym provides infrastructure for building, running, and scaling…
- tbbProvides Python bindings to Intel's oneAPI Threading Building Blocks…
- tbb-develProvides C++ parallelism primitives and task scheduling for…
Give your agent the search over MCP, or paste the wish link into any chat.
More Libraries packages
urllib3 is an HTTP client library that provides thread-safe connection pooling, SSL/TLS verification, multipart file uploads, request retries, compression support, and proxy handling for Python applications.
Requests is a Python HTTP library that simplifies sending HTTP/1.1 requests with automatic handling of headers, authentication, cookies, and response parsing.
Pluggy provides a plugin system that lets you define hook specifications and register implementations to be called in sequence, enabling extensible Python applications without tight coupling.
Install it if you're building an extensible application or framework.
Provides parsing, arithmetic, and recurrence rule computation for dates and times, with timezone support and iCalendar RFC compliance.
Install it if you need to parse flexible date strings, compute relative dates, handle timezones, or work with recurrence rules—it's the de facto choice for these tasks.
Six provides utility functions to write Python code that runs on both Python 2.7 and Python 3.3+, smoothing over language differences between the two versions.
pytest is a testing framework that lets you write test functions using plain assert statements and automatically discovers and runs them, with detailed failure reporting.
See also openenv-core · reasoning-gym · fhaviary · gem-llm · verifiers · nvidia-nat-core · nemo-evaluator · kaggle-environments · browsergym-core · gymnasium