selfies
SELFIES (SELF-referencIng Embedded Strings) is a general-purpose, sequence-based, robust representation of semantically constrained graphs.
Decision gist · record as of 2026-08-14
Yes, if you work with generative chemistry models or need guaranteed-valid molecular representations. The package is well-established (first released 2019, now at 2.2.0) with no known vulnerabilities and permissive licensing. The aging maintenance status (576 days since last release) is not a blocker for stable use, but verify that it meets your specific ML framework and performance requirements before committing to a production pipeline.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low install friction with no runtime dependencies.
- Maintenance status is aging—last release was 576 days ago, though the repository remains active with recent commits and 862 stars.
License · maintenance · safety
permissive license (permissive) — Licensed under Apache License (permissive), allowing commercial and private use with minimal restrictions.
last release 2025-01-15 (576 days) · last repo commit 2025-05-17 · 862 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 181,097 downloads/mo, #10,134 on PyPI
Alternatives
Verify before relying
import selfies as sf
benzene_smiles = "c1ccccc1"
benzene_selfies = sf.encoder(benzene_smiles)
print(benzene_selfies) # [C][=C][C][=C][C][=C][Ring1][=Branch1]
recovered_smiles = sf.decoder(benzene_selfies)
print(recovered_smiles) # C1=CC=CC=C1- Whether the package is actively maintained or in maintenance-only mode given the aging status
- Performance characteristics for large-scale molecular datasets or real-time encoding/decoding
- Compatibility with specific chemistry frameworks or ML pipelines beyond the examples shown
What it is and what it does
SELFIES is a molecular string representation designed to guarantee that every string encodes a valid, semantically meaningful molecule. Unlike SMILES, which can produce invalid molecules through random mutations, SELFIES enforces chemical constraints at the string level, making it particularly useful for generative machine learning models that need to explore molecular space without producing chemically impossible structures.
The package provides bidirectional translation between SELFIES and SMILES formats, tokenization, encoding/decoding for neural networks, and customizable semantic constraints (including hypervalent chemistry). It has no external runtime dependencies and supports Python 3.7 and later, making it straightforward to integrate into existing chemistry or ML pipelines.
Use it for
- Training generative models (VAEs, GANs) on molecular data where every sampled string must decode to a valid molecule
- Converting existing SMILES datasets to SELFIES for more robust chemical exploration and mutation
- Encoding molecules as fixed-length vectors or integer sequences for neural network input
- Creating random valid molecules for high-throughput virtual screening or lead generation
- Analyzing molecular structure by tokenizing and attributing SELFIES symbols to SMILES output tokens
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you work with generative chemistry models or need guaranteed-valid molecular representations.
The package is well-established (first released 2019, now at 2.2.0) with no known vulnerabilities and permissive licensing. The aging maintenance status (576 days since last release) is not a blocker for stable use, but verify that it meets your specific ML framework and performance requirements before committing to a production pipeline.
Install
selfies on PyPI
Before you install
Low install friction with no runtime dependencies. Maintenance status is aging—last release was 576 days ago, though the repository remains active with recent commits and 862 stars.
License in practice
Licensed under Apache License (permissive), allowing commercial and private use with minimal restrictions.
Quickstart
import selfies as sf
benzene_smiles = "c1ccccc1"
benzene_selfies = sf.encoder(benzene_smiles)
print(benzene_selfies) # [C][=C][C][=C][C][=C][Ring1][=Branch1]
recovered_smiles = sf.decoder(benzene_selfies)
print(recovered_smiles) # C1=CC=CC=C1
Verify before relying
- Whether the package is actively maintained or in maintenance-only mode given the aging status
- Performance characteristics for large-scale molecular datasets or real-time encoding/decoding
- Compatibility with specific chemistry frameworks or ML pipelines beyond the examples shown
Package facts
| License | permissive license permissive |
| Python support | Supports the current Python release >=3.7 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Aging 576 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 181,097 / month, #10,134 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | License :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3 :: OnlyProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.7Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9 |
Evidence: selfies-2.2.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “molecular string representation”
- selfiesConverts between SELFIES (Self-Referencing Embedded Strings) and…
- mda-xdrlibProvides the xdrlib module for encoding and decoding XDR (External…
- pytdcPyTDC provides unified access to multimodal biomedical datasets,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also aimsim-core · mordredcommunity · py2opsin · mhfp · chemprop · padelpy · mace-torch · prolif · autogluon.multimodal · PubChemPy