{"categories":[{"label":"Libraries","url":"https://skillfed.io/packages/category/software-development-libraries/4"},{"label":"Scientific/Engineering","url":"https://skillfed.io/packages/category/scientific-engineering/3"},{"label":"Python Modules","url":"https://skillfed.io/packages/category/software-development-libraries-python-modules/8"},{"label":"Artificial Intelligence","url":"https://skillfed.io/packages/category/scientific-engineering-artificial-intelligence/4"},{"label":"Mathematics","url":"https://skillfed.io/packages/category/scientific-engineering-mathematics/2"}],"enrichment":{"capability":"NeMo Gym provides infrastructure for building, running, and scaling evaluation and training environments where agents interact with tasks, datasets, verifiers, and execution state to solve problems.","skillfed_tags":["agent-evaluation","rl-training","benchmark-hub"],"use_cases":["Evaluate agents on standardized benchmarks with reproducible verifiers and shared task datasets across teams.","Collect training data by running agents on tasks at scale, generating trajectories for supervised fine-tuning or reinforcement learning.","Benchmark agent skill impact by running the same tasks with different skill sets to isolate which capabilities drive performance.","Run tool-using agents in isolated sandboxes to safely test code execution and external tool interactions.","Transition from evaluation to training by using the same environment and verifier definitions with training frameworks.","Diagnose evaluation failures to identify which tasks failed, why, and the highest-impact fixes."],"what_it_does":"NeMo Gym is a framework for building and running evaluation environments where language models and agents interact with tasks to solve problems. It abstracts the evaluation loop\u2014task datasets, agent harnesses, verifiers, and execution state\u2014into modular, extensible components. You define or reuse an environment, configure your model (openai, anthropic, local via vLLM, or hosted providers), and run agents against tasks at scale, collecting trajectories and metrics.\n\nThe library is designed for teams needing reproducible, stateful evaluation across shared environments, or transitioning from evaluation into agent optimization and training. It includes a hub of popular benchmarks and agent harnesses, integrates with training frameworks, and provides CLI tools for discovery, validation, and diagnostics. With 31 runtime dependencies (ray, mlflow, wandb, fastapi, pydantic, aiohttp, and others), it trades breadth of built-in capability for a heavier dependency footprint.","worth_installing":"Yes, if you need to evaluate or train agents in stateful environments at scale with reproducible verifiers and shared benchmarks. Active maintenance, permissive license, low install friction, and integration with major providers make it solid for agentic workflows. However, the strict Python 3.13.14+ requirement and heavy dependency footprint may conflict with existing projects. Early development status and evolving APIs warrant caution in production."},"id":"nemo-gym","links":{"html":"https://skillfed.io/packages/nemo-gym","md":"https://skillfed.io/packages/nemo-gym.md","pypi":"https://pypi.org/project/nemo-gym/"},"maintenance":{"status":"active"},"meta":{"latest_release":"2026-08-07","license_spdx":null,"license_treatment":"permissive","name":"nemo-gym","python_support":"capped_below_current","summary":"NeMo Gym is a library for building reinforcement learning environments"},"popularity":{"monthly_downloads":857100,"position":4884,"tier":"top_5000"},"security":{"n_vulnerabilities":0},"version":"0.5.0"}
