open-clip-torch
Open reproduction of consastive language-image pretraining (CLIP) and related.
Decision gist · record as of 2026-08-14
Yes. OpenCLIP is actively maintained, has no known vulnerabilities, installs with low friction, and offers a mature, well-documented interface to production-grade pretrained models. It is the standard open-source CLIP implementation and worth installing if you need multimodal embeddings or zero-shot classification.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires torch and torchvision installed; GPU recommended for inference speed.
- Models using timm image encoders require the latest timm version to avoid 'Unknown model' errors.
- Low install friction with a pure Python wheel.
License · maintenance · safety
MIT (permissive) — MIT license is permissive, allowing commercial and private use with minimal restrictions—suitable for most projects that can include a license notice.
last release 2026-02-27 (168 days) · last repo commit 2026-08-10 · 14,064 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 3,209,634 downloads/mo, #2,700 on PyPI
Alternatives
Verify before relying
pip install open_clip_torch
import torch
import open_clip
model, _, preprocess = open_clip.create_model_and_transforms('ViT-B-32', pretrained='laion2b_s34b_b79k')
tokenizer = open_clip.get_tokenizer('ViT-B-32')
text = tokenizer(["a dog", "a cat"])
with torch.no_grad():
text_features = model.encode_text(text)- Whether transformers library is required for all models or only those using transformers tokenizers
- Specific memory requirements for different model sizes during inference
- Image preprocessing and loading workflow details beyond tokenization
What it is and what it does
OpenCLIP is an open-source reimplementation of CLIP that learns joint embeddings of images and text through contrastive learning. It provides a collection of pretrained models trained on datasets like LAION-2B and DataComp-1B, ranging from small ConvNext and ViT variants to large models, all loadable through a unified interface.
The package lets you encode images and text into a shared embedding space, enabling zero-shot classification, semantic search, and image-text matching without task-specific fine-tuning. You load a pretrained model, tokenize text with the appropriate tokenizer, and compute embeddings; similarity scores between image and text embeddings reveal semantic alignment. It integrates with torch, torchvision, huggingface-hub, and timm for model architectures and checkpoint management.
Use it for
- Zero-shot image classification by encoding class names as text and comparing to image embeddings without retraining
- Image-text retrieval systems that find images matching natural language queries or vice versa
- Building semantic search engines over image collections using text queries
- Multimodal feature extraction for downstream tasks like clustering or anomaly detection
- Evaluating model robustness across datasets by computing zero-shot accuracy on ImageNet and other benchmarks
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
OpenCLIP is actively maintained, has no known vulnerabilities, installs with low friction, and offers a mature, well-documented interface to production-grade pretrained models. It is the standard open-source CLIP implementation and worth installing if you need multimodal embeddings or zero-shot classification.
Install
open-clip-torch on PyPI
Before you install
Low install friction with a pure Python wheel. The package depends on torch, torchvision, regex, ftfy, tqdm, huggingface-hub, safetensors, and timm. Maintenance is active with recent releases and a well-maintained repository.
Requires torch and torchvision installed; GPU recommended for inference speed. Models using timm image encoders require the latest timm version to avoid 'Unknown model' errors.
License in practice
MIT license is permissive, allowing commercial and private use with minimal restrictions—suitable for most projects that can include a license notice.
Quickstart
pip install open_clip_torch
import torch
import open_clip
model, _, preprocess = open_clip.create_model_and_transforms('ViT-B-32', pretrained='laion2b_s34b_b79k')
tokenizer = open_clip.get_tokenizer('ViT-B-32')
text = tokenizer(["a dog", "a cat"])
with torch.no_grad():
text_features = model.encode_text(text)
Verify before relying
- Whether transformers library is required for all models or only those using transformers tokenizers
- Specific memory requirements for different model sizes during inference
- Image preprocessing and loading workflow details beyond tokenization
Package facts
| License | MIT permissive |
| Python support | Supports the current Python release >=3.9 |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 8 packagestorchtorchvisionregexftfytqdmhuggingface-hubsafetensorstimm |
| Maintenance | Actively maintained 168 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 3,209,634 / month, #2,700 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 4 - BetaIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: MIT LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9Topic :: Scientific/EngineeringTopic :: Scientific/Engineering :: Artificial IntelligenceTopic :: Software DevelopmentTopic :: Software Development :: LibrariesTopic :: Software Development :: Libraries :: Python Modules |
Evidence: open_clip_torch-3.3.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “clip image text embedding”
- open-clip-torchOpenCLIP provides open-source implementations of CLIP (Contrastive…
- clip-anytorchLoads and runs OpenAI's CLIP model to encode images and text into a…
- clip-interrogatorGenerates natural-language prompts from images by combining CLIP and…
Give your agent the search over MCP, or paste the wish link into any chat.
More Software Development packages
Provides backported and experimental type hints for Python 3.9+, allowing use of newer typing features on older Python versions and enabling early experimentation with type system PEPs before they enter the standard library.
NumPy provides an N-dimensional array object and a comprehensive suite of mathematical, linear algebra, Fourier transform, and random number functions for scientific computing in Python.
FastAPI is a Python web framework for building REST APIs using type hints, with automatic request validation, serialization, and interactive API documentation.
Provides a way to document function parameters, class attributes, return types, and variables inline using Python's `Annotated` type hint syntax instead of traditional docstrings.
Typer builds command-line applications from Python functions using type hints, automatically generating help text, argument parsing, and shell completion.
Install it if you are building CLIs in Python.
Distlib provides low-level packaging utilities for building, distributing, and managing Python software—including metadata handling, version specifiers, wheel support, script installation, and dependency resolution.
See also clip-anytorch · clip-benchmark · clip-interrogator · laion-clap · timm · segmentation-models-pytorch · pytorch-pretrained-bert · torchtext · pytorchcv · lightly