reasoning-gym
A library of procedural dataset generators for training reasoning models
Decision gist · record as of 2026-08-14
Yes. Reasoning Gym is actively maintained, has no known vulnerabilities, carries a permissive Apache-2.0 license, and requires only Python >= 3.10 with low install friction. It is purpose-built for RL reasoning model training and already adopted by major research labs. Install if you are training reasoning models or need procedurally verifiable reasoning datasets; skip if you need pre-generated static benchmarks or do not work with RL.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python >= 3.10
- Low install friction with a pure-Python wheel and 11 runtime dependencies.
- Active maintenance with a recent release (2026-03-28) and 1485 repository stars.
License · maintenance · safety
Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
last release 2026-03-28 (139 days) · last repo commit 2026-04-17 · 1,485 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 345,015 downloads/mo, #7,369 on PyPI
Alternatives
Verify before relying
pip install reasoning-gym
import reasoning_gym
data = reasoning_gym.create_dataset('leg_counting', size=10, seed=42)
for entry in data:
score = data.score_answer(answer=entry['answer'], entry=entry)
print(f"Question: {entry['question']}, Score: {score}")- Whether all 100+ tasks are equally well-maintained or if some are experimental/unstable
- Performance characteristics and scalability limits for large dataset sizes
- Compatibility with specific RL frameworks beyond the mentioned verifiers library
What it is and what it does
Reasoning Gym is a Python library that generates procedural datasets and verifiable reasoning environments for training reinforcement learning models. It provides more than 100 tasks spanning algebra, arithmetic, computation, cognition, geometry, graph theory, logic, and games—each with adjustable complexity and built-in algorithmic verification of solutions. Some tasks have single correct answers; others like Rubik's Cube or Countdown have multiple valid solutions. The library generates virtually infinite training data on demand via a standard procedural interface, making it suitable for large-scale RL training without pre-generated dataset bottlenecks.
The package is designed for researchers and practitioners building reasoning models. It integrates with RL training frameworks (particularly the verifiers library) and supports both single-task and composite multi-task dataset creation with configurable weightings. Each dataset entry includes a question, answer, and metadata; scoring functions enable reward computation during training. The library is actively maintained, recently released, and already adopted by multiple research organizations including NVIDIA, Meta, and others for production reasoning model training.
Use it for
- Generate infinite training data for reasoning model RL fine-tuning with procedurally controlled difficulty
- Benchmark and evaluate reasoning model performance across diverse task domains with algorithmic verification
- Create composite datasets combining multiple reasoning tasks with custom weightings for curriculum learning
- Build verifiable reward signals for RL training by calling task-specific scoring functions on model outputs
- Prototype new reasoning environments by extending the library's task generators for custom domains
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
Reasoning Gym is actively maintained, has no known vulnerabilities, carries a permissive Apache-2.0 license, and requires only Python >= 3.10 with low install friction. It is purpose-built for RL reasoning model training and already adopted by major research labs. Install if you are training reasoning models or need procedurally verifiable reasoning datasets; skip if you need pre-generated static benchmarks or do not work with RL.
Install
reasoning-gym on PyPI
Before you install
Low install friction with a pure-Python wheel and 11 runtime dependencies. Active maintenance with a recent release (2026-03-28) and 1485 repository stars. Requires Python >= 3.10.
Requires Python >= 3.10
License in practice
Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.
Quickstart
pip install reasoning-gym
import reasoning_gym
data = reasoning_gym.create_dataset('leg_counting', size=10, seed=42)
for entry in data:
score = data.score_answer(answer=entry['answer'], entry=entry)
print(f"Question: {entry['question']}, Score: {score}")
Verify before relying
- Whether all 100+ tasks are equally well-maintained or if some are experimental/unstable
- Performance characteristics and scalability limits for large dataset sizes
- Compatibility with specific RL frameworks beyond the mentioned verifiers library
Package facts
| License | Apache-2.0 permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 11 packagesarckitbficellpylibmagiccubepycosatpyfigletpytzpyyamlsympytabulatezss |
| Maintenance | Actively maintained 139 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 345,015 / month, #7,369 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: reasoning_gym-0.1.25-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “procedural dataset generation”
- reasoning-gymReasoning Gym generates procedurally verifiable reasoning datasets…
- opensimplexGenerates OpenSimplex noise across 2D, 3D, and 4D coordinate spaces,…
- haikunatorGenerates random Heroku-style names composed of adjectives, nouns,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also arckit · nemo-gym · gem-llm · gymnasium · verifiers · rsl-rl-lib · stable-baselines3 · TextArena · tianshou · skrl