InstructorEmbedding
Text embedding tool
Decision gist · record as of 2026-08-14
Yes, if you need task-specific embeddings without fine-tuning and can work with a dormant package. The zero-dependency install and permissive license are advantages. However, the lack of maintenance since May 2023 and unspecified Python version support mean you should test compatibility in your environment and be prepared to maintain a fork if critical issues arise.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Model weights are downloaded from Hugging Face on first use; requires internet access and sufficient disk space for the chosen checkpoint.
- Installation is straightforward with no runtime dependencies and a pure-Python wheel distribution.
- The package is dormant (last release May 2023, last commit January 2025), so expect no active maintenance or bug fixes.
License · maintenance · safety
Apache License 2.0 (permissive) — Apache License 2.0 is permissive, allowing commercial use, modification, and distribution with minimal restrictions—suitable for most projects.
last release 2023-05-26 (1176 days) · last repo commit 2025-01-15 · 2,023 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 636,038 downloads/mo, #5,631 on PyPI
Alternatives
Verify before relying
pip install InstructorEmbedding
from InstructorEmbedding import INSTRUCTOR
model = INSTRUCTOR('hkunlp/instructor-large')
embeddings = model.encode([['Represent the Science title:', 'Example text']])- Whether the package works with modern Python versions (requires_python is unspecified in metadata).
- Performance characteristics and memory requirements for different model sizes (base, large, xl).
- Compatibility with recent versions of underlying NLP libraries after 18+ months of dormancy.
What it is and what it does
InstructorEmbedding is a Python wrapper around instruction-finetuned embedding models that generate text representations tailored to specific tasks and domains. Instead of using a one-size-fits-all embedding model, you provide a natural-language instruction alongside your text—for example, 'Represent the Science title:' or 'Represent the Financial statement for retrieval:'—and the model produces embeddings optimized for that context. The package handles model loading and encoding, supporting multiple checkpoint sizes (base, large, xl) hosted on Hugging Face.
The typical workflow is to instantiate a model, prepare text-instruction pairs, call encode(), and receive numpy arrays of embeddings suitable for downstream tasks like similarity computation, clustering, or information retrieval. No fine-tuning or training is required; the instruction acts as a prompt to steer the pre-trained model's output.
Use it for
- Build domain-specific semantic search systems by instructing the model to optimize embeddings for retrieval in science, finance, or medicine.
- Compute similarity scores between text pairs using task-aware embeddings (e.g., 'for duplicate detection' vs. 'for paraphrase matching').
- Cluster documents or sentences with embeddings tailored to your classification or grouping objective.
- Implement information retrieval pipelines where queries and documents are encoded with matching instructions for better ranking.
- Generate embeddings for text evaluation tasks by specifying the evaluation criterion in the instruction.
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes, if you need task-specific embeddings without fine-tuning and can work with a dormant package.
The zero-dependency install and permissive license are advantages. However, the lack of maintenance since May 2023 and unspecified Python version support mean you should test compatibility in your environment and be prepared to maintain a fork if critical issues arise.
Install
instructorembedding on PyPI
Before you install
Installation is straightforward with no runtime dependencies and a pure-Python wheel distribution. The package is dormant (last release May 2023, last commit January 2025), so expect no active maintenance or bug fixes.
Model weights are downloaded from Hugging Face on first use; requires internet access and sufficient disk space for the chosen checkpoint.
License in practice
Apache License 2.0 is permissive, allowing commercial use, modification, and distribution with minimal restrictions—suitable for most projects.
Quickstart
pip install InstructorEmbedding
from InstructorEmbedding import INSTRUCTOR
model = INSTRUCTOR('hkunlp/instructor-large')
embeddings = model.encode([['Represent the Science title:', 'Example text']])
Verify before relying
- Whether the package works with modern Python versions (requires_python is unspecified in metadata).
- Performance characteristics and memory requirements for different model sizes (base, large, xl).
- Compatibility with recent versions of underlying NLP libraries after 18+ months of dormancy.
Package facts
| License | Apache License 2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | None |
| Maintenance | Dormant 1,176 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 636,038 / month, #5,631 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
Evidence: InstructorEmbedding-1.0.1-py2.py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “instruction-based text embeddings”
- InstructorEmbeddingInstructorEmbedding generates task-specific text embeddings by…
- qwen-ttsQwen-TTS generates speech from text using Qwen3-TTS models,…
- llama-index-embeddings-bedrockProvides Amazon Bedrock embedding models integration for LlamaIndex,…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also sentence-transformers · model2vec · mteb · FlagEmbedding · fastembed · setfit · voyageai · llama-index-embeddings-huggingface · swesmith · axial-positional-embedding