chembl-structure-pipeline
ChEMBL Structure Pipeline
Decision gist · record as of 2026-08-14
Yes, if you work with chemical structures and need ChEMBL-compatible curation. Install friction is low and the license is permissive. The main caveat is aging maintenance (263 days since last release), so if you encounter issues or need active support, you may need to fork or patch it yourself. No known vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- rdkit must be installed and functional; it may require compilation or conda on some platforms.
- Low friction: pure Python wheel with just setuptools and rdkit as runtime dependencies.
- Maintenance status is aging—last release 263 days ago—so expect slower response to issues.
License · maintenance · safety
MIT (permissive) — MIT license is permissive; you can use, modify, and distribute this package with minimal legal friction.
last release 2025-11-24 (263 days)
0 known vulnerabilities (OSV.dev, 2026-08-14) · 194,405 downloads/mo, #9,836 on PyPI
Alternatives
Verify before relying
pip install chembl_structure_pipeline
from chembl_structure_pipeline import standardizer
molblock = """..."""
std_molblock = standardizer.standardize_molblock(molblock)- Whether the package supports modern Python versions (requires_python is unspecified in the fact sheet)
- Whether rdkit is available as a pre-built wheel on all target platforms or requires compilation
What it is and what it does
ChEMBL Structure Pipeline is a molecular curation toolkit that implements the standardization and quality-checking protocols used by the ChEMBL database. It wraps RDKit to perform three main tasks: standardize molecular structures (normalize charges, aromaticity, and stereochemistry), extract parent compounds by removing salts and counterions, and assess structure quality by identifying problematic features and assigning penalty scores. The package is used to prepare chemical data for database ingestion or downstream analysis.
It's a specialized tool for chemoinformatics workflows—useful when you're building a chemical database, cleaning scraped or legacy molecular data, or need to apply ChEMBL's curation rules to your own structures. The API is straightforward: pass a molblock (V2000 format) to one of three main functions and get back a standardized molblock, parent structure, or quality report.
Use it for
- Standardize raw molecular structures before loading them into a chemical database or registry.
- Extract parent compounds from salt forms or multi-component mixtures for deduplication.
- Assess and flag problematic structures in a dataset with penalty scores to prioritize manual review.
- Prepare chemical data for machine learning pipelines that require consistent molecular representation.
- Validate chemical structures submitted by users in a web application or data portal.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work with chemical structures and need ChEMBL-compatible curation.
Install friction is low and the license is permissive. The main caveat is aging maintenance (263 days since last release), so if you encounter issues or need active support, you may need to fork or patch it yourself. No known vulnerabilities.
Install
chembl-structure-pipeline on PyPI
Before you install
Low friction: pure Python wheel with just setuptools and rdkit as runtime dependencies. Maintenance status is aging—last release 263 days ago—so expect slower response to issues.
rdkit must be installed and functional; it may require compilation or conda on some platforms.
License in practice
MIT license is permissive; you can use, modify, and distribute this package with minimal legal friction.
Quickstart
pip install chembl_structure_pipeline
from chembl_structure_pipeline import standardizer
molblock = """..."""
std_molblock = standardizer.standardize_molblock(molblock)
Verify before relying
- Whether the package supports modern Python versions (requires_python is unspecified in the fact sheet)
- Whether rdkit is available as a pre-built wheel on all target platforms or requires compilation
Package facts
| License | MIT permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 2 packagessetuptoolsrdkit |
| Maintenance | Aging 263 days since the last release |
| First released | |
| Downloads | 194,405 / month, #9,836 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: chembl_structure_pipeline-1.2.4-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “molecule standardization pipeline”
- chembl-structure-pipelineStandardizes and salt-strips molecular structures using ChEMBL…
- datamolDatamol provides a pythonic layer on top of RDKit for molecular…
- molecule-vagrantMolecule Vagrant Plugin enables Vagrant-based provisioning and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Scientific/Engineering packages
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
pandas provides fast, flexible data structures (Series and DataFrame) for loading, cleaning, transforming, and analyzing labeled or relational data in Python.
scipy provides numerical algorithms for mathematics, science, and engineering—including optimization, integration, linear algebra, Fourier transforms, signal and image processing, and ODE solvers—built on numpy arrays.
scikit-learn provides a comprehensive Python library for supervised and unsupervised machine learning, including classification, regression, clustering, dimensionality reduction, and model evaluation tools built on NumPy and SciPy.
Install it if you need to train, evaluate, or deploy supervised or unsupervised learning models.
dill extends Python's pickle module to serialize and deserialize a much wider range of Python objects, including functions, lambdas, classes, and interpreter sessions, to byte streams for storage or network transmission.
Multiprocess is an enhanced fork of Python's standard multiprocessing library that uses dill for better serialization, allowing you to spawn processes with a threading-like API and share complex objects between them.
Install it if you use multiprocessing and encounter pickle serialization limits with lambdas or complex objects.
See also rdkit · pdbeccdutils · datamol · aimsim-core · chemprop · PubChemPy · janaf · py2opsin · chemicals