pinder
PINDER: The Protein INteraction Dataset and Evaluation Resource
Decision gist · record as of 2026-08-14
Yes, if you are actively developing or benchmarking protein docking algorithms and have approximately 700 GB of disk space available. The dataset is substantially larger than prior resources and uniquely includes paired predicted and apo structures. However, maintenance is dormant, so verify compatibility with your current PyTorch and torch-geometric versions before committing to a production workflow. For exploratory work or small-scale evaluation, the download overhead may not justify the install.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires Python ≥3.10.
- The dataset itself is approximately 700 GB and must be downloaded separately via pinder_download command-line tool or manually from Google Cloud Storage.
- fastpdb may require Rust toolchain on unsupported platforms.
License · maintenance · safety
permissive license (permissive) — Licensed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions.
last release 2024-11-15 (637 days) · last repo commit 2024-11-15 · 158 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 105,620 downloads/mo, #12,688 on PyPI
Alternatives
Verify before relying
pip install pinder
from pinder.core import get_pinder_location
get_pinder_location()- Whether dormant maintenance status affects stability or compatibility with current PyTorch/torch-geometric versions
- Performance characteristics and typical memory footprint when loading subsets of the dataset
- Whether optional dependencies (pytorch-cluster, PRODIGY-cryst) are commonly needed for typical workflows
What it is and what it does
Pinder is a dataset and resource for protein-protein docking research, providing access to a large collection of protein structures and interaction data hosted on Google Cloud Storage. It includes monomer structures, ground-truth dimer complexes, predicted structures, and apo conformations—the first dataset to pair predicted and apo structures for training flexible docking methods. The dataset is approximately 500 times larger than previous state-of-the-art datasets.
The Python API handles automatic downloading and caching of dataset files to a local directory (defaulting to ~/.local/share/pinder), with command-line tools for managing downloads and updates. It depends on 19 runtime packages including torch, torch-geometric, biotite, and pandas, making it suitable for machine learning workflows. The dataset requires approximately 700 GB of disk space when fully unpacked.
Use it for
- Train protein-protein docking models using the large paired dataset of holo and apo structures
- Benchmark docking algorithms against gold-standard test sets included in the dataset
- Access preprocessed protein structure data and metadata for structural biology research
- Develop flexible docking methods using predicted and experimental structure pairs
- Evaluate protein interaction prediction models on standardized benchmarks
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you are actively developing or benchmarking protein docking algorithms and have approximately 700 GB of disk space available.
The dataset is substantially larger than prior resources and uniquely includes paired predicted and apo structures. However, maintenance is dormant, so verify compatibility with your current PyTorch and torch-geometric versions before committing to a production workflow. For exploratory work or small-scale evaluation, the download overhead may not justify the install.
Install
pinder on PyPI
Before you install
Low install friction with a pure Python wheel available. The package depends on fastpdb, which has pre-built wheels for Linux (glibc≥2.34), macOS Sierra or newer, and Windows; other platforms require building from source with the Rust toolchain. Maintenance is dormant—last release was 2024-11-15.
Requires Python ≥3.10. The dataset itself is approximately 700 GB and must be downloaded separately via pinder_download command-line tool or manually from Google Cloud Storage. fastpdb may require Rust toolchain on unsupported platforms.
License in practice
Licensed under Apache 2.0 (permissive), allowing commercial and private use with minimal restrictions.
Quickstart
pip install pinder
from pinder.core import get_pinder_location
get_pinder_location()
Verify before relying
- Whether dormant maintenance status affects stability or compatibility with current PyTorch/torch-geometric versions
- Performance characteristics and typical memory footprint when loading subsets of the dataset
- Whether optional dependencies (pytorch-cluster, PRODIGY-cryst) are commonly needed for typical workflows
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.10 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 19 packagesbiotitefastpdbnumpypandaspyarrowtorchtorchtypingtypeguardtyping-extensionspydantictqdmplotlynbformatgoogle-cloud-storagegcsfstorch-geometrictabulatePyYAMLscikit-learn |
| Maintenance | Dormant 637 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 105,620 / month, #12,688 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3 |
Evidence: pinder-0.5.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “protein docking dataset”
- pinderPinder provides access to a large protein-protein interaction dataset…
- prolifProLIF generates interaction fingerprints from molecular complexes in…
- pdb2pqrPDB2PQR prepares protein structures from PDB files for biomolecular…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also prolif · pdb2pqr · tmtools · mmcif-pdbx · fair-esm · bionty · mmtf-python · biotite · mdtraj · propka