$npx skillfedfor your agent

reasoning-gym

A library of procedural dataset generators for training reasoning models

Worth itPyPI Artificial IntelligenceReleased Mar 2026345.0K downloads / moApache-2.0Pure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — reasoning_gym-0.1.25-py3-none-any.whl
v0.1.25 · released 2026-03-28 · Python >=3.10 · 11 runtime deps: arckit, bfi, cellpylib, magiccube, pycosat, pyfiglet, pytz, pyyaml

Yes. Reasoning Gym is actively maintained, has no known vulnerabilities, carries a permissive Apache-2.0 license, and requires only Python >= 3.10 with low install friction. It is purpose-built for RL reasoning model training and already adopted by major research labs. Install if you are training reasoning models or need procedurally verifiable reasoning datasets; skip if you need pre-generated static benchmarks or do not work with RL.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python >= 3.10
  • Low install friction with a pure-Python wheel and 11 runtime dependencies.
  • Active maintenance with a recent release (2026-03-28) and 1485 repository stars.

License · maintenance · safety

Apache-2.0 (permissive) — Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.

last release 2026-03-28 (139 days) · last repo commit 2026-04-17 · 1,485 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 345,015 downloads/mo, #7,369 on PyPI

Verify before relying

pip install reasoning-gym

import reasoning_gym
data = reasoning_gym.create_dataset('leg_counting', size=10, seed=42)
for entry in data:
    score = data.score_answer(answer=entry['answer'], entry=entry)
    print(f"Question: {entry['question']}, Score: {score}")
  • Whether all 100+ tasks are equally well-maintained or if some are experimental/unstable
  • Performance characteristics and scalability limits for large dataset sizes
  • Compatibility with specific RL frameworks beyond the mentioned verifiers library
Same gist for agents: .md · .json

What it is and what it does

Reasoning Gym is a Python library that generates procedural datasets and verifiable reasoning environments for training reinforcement learning models. It provides more than 100 tasks spanning algebra, arithmetic, computation, cognition, geometry, graph theory, logic, and games—each with adjustable complexity and built-in algorithmic verification of solutions. Some tasks have single correct answers; others like Rubik's Cube or Countdown have multiple valid solutions. The library generates virtually infinite training data on demand via a standard procedural interface, making it suitable for large-scale RL training without pre-generated dataset bottlenecks.

The package is designed for researchers and practitioners building reasoning models. It integrates with RL training frameworks (particularly the verifiers library) and supports both single-task and composite multi-task dataset creation with configurable weightings. Each dataset entry includes a question, answer, and metadata; scoring functions enable reward computation during training. The library is actively maintained, recently released, and already adopted by multiple research organizations including NVIDIA, Meta, and others for production reasoning model training.

Use it for

  • Generate infinite training data for reasoning model RL fine-tuning with procedurally controlled difficulty
  • Benchmark and evaluate reasoning model performance across diverse task domains with algorithmic verification
  • Create composite datasets combining multiple reasoning tasks with custom weightings for curriculum learning
  • Build verifiable reward signals for RL training by calling task-specific scoring functions on model outputs
  • Prototype new reasoning environments by extending the library's task generators for custom domains

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Reasoning Gym is actively maintained, has no known vulnerabilities, carries a permissive Apache-2.0 license, and requires only Python >= 3.10 with low install friction. It is purpose-built for RL reasoning model training and already adopted by major research labs. Install if you are training reasoning models or need procedurally verifiable reasoning datasets; skip if you need pre-generated static benchmarks or do not work with RL.

Install

reasoning-gym on PyPI

Before you install

Low install friction with a pure-Python wheel and 11 runtime dependencies. Active maintenance with a recent release (2026-03-28) and 1485 repository stars. Requires Python >= 3.10.

Requires Python >= 3.10

License in practice

Licensed under Apache-2.0 (permissive), allowing commercial use, modification, and distribution with minimal restrictions.

Quickstart

pip install reasoning-gym

import reasoning_gym
data = reasoning_gym.create_dataset('leg_counting', size=10, seed=42)
for entry in data:
    score = data.score_answer(answer=entry['answer'], entry=entry)
    print(f"Question: {entry['question']}, Score: {score}")

Verify before relying

  • Whether all 100+ tasks are equally well-maintained or if some are experimental/unstable
  • Performance characteristics and scalability limits for large dataset sizes
  • Compatibility with specific RL frameworks beyond the mentioned verifiers library

Package facts

LicenseApache-2.0 permissive
Python supportSupports the current Python release >=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
11 packages
arckitbficellpylibmagiccubepycosatpyfigletpytzpyyamlsympytabulatezss
MaintenanceActively maintained 139 days since the last release
Last repo commit
First released
Downloads345,015 / month, #7,369 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3

Evidence: reasoning_gym-0.1.25-py3-none-any.whl

Tags

Capabilities
procedural dataset generationreasoning model training datareinforcement learning environmentsverifiable reasoning tasksalgorithmic problem generatorsRL training dataset libraryreasoning benchmark tasks
Topics
reinforcement-learningdataset-generationreasoning-benchmark

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “procedural dataset generation”

  • reasoning-gymReasoning Gym generates procedurally verifiable reasoning datasets…
  • opensimplexGenerates OpenSimplex noise across 2D, 3D, and 4D coordinate spaces,…
  • haikunatorGenerates random Heroku-style names composed of adjectives, nouns,…

Give your agent the search over MCP, or paste the wish link into any chat.

More Artificial Intelligence packages

litellm With conditions
PyPI · Artificial Intelligence · released Aug 2026

LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.

Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.

MITcompiled wheel
682.8Mdownloads / mo
huggingface-hub Worth it
PyPI · Artificial Intelligence · released Aug 2026

Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.

Install it if you work with Hugging Face Hub models or datasets.

Apache-2.0pure Python · 3.10.0+
442.4Mdownloads / mo
langchain Worth it
PyPI · Python Modules · released Aug 2026

LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.

MITpure Python
315.4Mdownloads / mo
hf-xet With conditions
PyPI · Artificial Intelligence · released Aug 2026

hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.

Apache-2.0compiled wheel · 3.8+
258.4Mdownloads / mo
tokenizers Worth it
PyPI · Artificial Intelligence · released Apr 2026

Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.

Apache-2.0compiled wheel · 3.10+
222.9Mdownloads / mo
transformers Worth it
PyPI · Artificial Intelligence · released Aug 2026

Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.

Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.

permissive licensepure Python · 3.10.0+
186.6Mdownloads / mo

See also arckit · nemo-gym · gem-llm · gymnasium · verifiers · rsl-rl-lib · stable-baselines3 · TextArena · tianshou · skrl

Further reading