clip-anytorch
# CLIP
Decision gist · record as of 2026-08-14
Yes, if you need zero-shot image classification or image-text matching and can accept dormant maintenance. The package is functional, has low install friction, and no known vulnerabilities. However, verify the unclear license before commercial use, and be aware that the last release was 2024-01-13—expect no active bug fixes or feature updates.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Requires torch and torchvision installed; if using torch versions other than 1.7.1, must pass jit=False to clip.load() to avoid JIT compilation errors.
- Low install friction with a pure-Python wheel.
- Maintenance is dormant—last release was 2024-01-13 and last commit 2024-07-08—but the package remains functional.
License · maintenance · safety
(unclear) — License treatment is unclear; no SPDX identifier or raw license text is available in the metadata. Verify the actual license before using in a commercial or restricted context.
last release 2024-01-13 (944 days) · last repo commit 2024-07-08 · 38 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 131,936 downloads/mo, #11,570 on PyPI
Alternatives
Verify before relying
pip install clip-anytorch
import torch
import clip
device = "cuda" if torch.cuda.is_available() else "cpu"
model, preprocess = clip.load("ViT-B/32", device=device, jit=False)
text = clip.tokenize(["a dog", "a cat"]).to(device)
with torch.no_grad():
text_features = model.encode_text(text)- Whether the unclear license permits commercial or proprietary use without restriction.
- Current compatibility with recent PyTorch and torchvision versions beyond what the description explicitly covers.
- Whether PIL or other image-loading libraries are required as implicit dependencies for typical usage.
What it is and what it does
clip-anytorch is a PyPI package wrapping OpenAI's CLIP (Contrastive Language-Image Pre-Training) model. CLIP is a neural network trained on image-text pairs that learns a shared embedding space, allowing it to match images to natural-language descriptions without being explicitly trained on any specific classification task. The package provides methods to load a pretrained model, encode images and text into feature vectors, and compute similarity scores between them.
The main use case is zero-shot image classification: given an image and a list of text labels, CLIP ranks the labels by how well they match the image without needing any labeled training data. It also enables image-text retrieval and similarity search. This fork removes the strict torch version dependency from the original repo and adds a truncate_text option for longer sequences, making it faster to install in environments like Google Colab.
Use it for
- Zero-shot image classification: rank candidate labels for an image without task-specific training data.
- Image-text retrieval: find images matching a natural-language query or vice versa.
- Feature extraction for downstream tasks: encode images or text into fixed-size vectors for use in other models.
- Content moderation or tagging: classify or describe image content using natural-language prompts.
- Cross-modal search: build search systems that match images to text descriptions.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need zero-shot image classification or image-text matching and can accept dormant maintenance.
The package is functional, has low install friction, and no known vulnerabilities. However, verify the unclear license before commercial use, and be aware that the last release was 2024-01-13—expect no active bug fixes or feature updates.
Install
clip-anytorch on PyPI
Before you install
Low install friction with a pure-Python wheel. Maintenance is dormant—last release was 2024-01-13 and last commit 2024-07-08—but the package remains functional. It relaxes the strict torch version constraint of the original repo, making it easier to install on modern environments.
Requires torch and torchvision installed; if using torch versions other than 1.7.1, must pass jit=False to clip.load() to avoid JIT compilation errors.
License in practice
License treatment is unclear; no SPDX identifier or raw license text is available in the metadata. Verify the actual license before using in a commercial or restricted context.
Quickstart
pip install clip-anytorch
import torch
import clip
device = "cuda" if torch.cuda.is_available() else "cpu"
model, preprocess = clip.load("ViT-B/32", device=device, jit=False)
text = clip.tokenize(["a dog", "a cat"]).to(device)
with torch.no_grad():
text_features = model.encode_text(text)
Verify before relying
- Whether the unclear license permits commercial or proprietary use without restriction.
- Current compatibility with recent PyTorch and torchvision versions beyond what the description explicitly covers.
- Whether PIL or other image-loading libraries are required as implicit dependencies for typical usage.
Package facts
| License | Not declared unclear |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 5 packagesftfyregextqdmtorchtorchvision |
| Maintenance | Dormant 944 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 131,936 / month, #11,570 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: clip_anytorch-2.6.0-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “image text matching”
- clip-anytorchLoads and runs OpenAI's CLIP model to encode images and text into a…
- open-clip-torchOpenCLIP provides open-source implementations of CLIP (Contrastive…
- ddddocrRecognizes text and detects objects in captcha images using offline…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also open-clip-torch · clip-benchmark · clip-interrogator · pytorch-pretrained-bert · torchtext · pytorchcv · mosaicml-streaming · chatterbox-tts · setfit · fastai