setfit
Efficient few-shot learning with Sentence Transformers
Decision gist · record as of 2026-08-14
Yes. SetFit is production-stable (Development Status 5), actively maintained, has no known vulnerabilities, and solves a real problem—achieving strong text classification with minimal labeled data. The low install friction, permissive license, and integration with Hugging Face Hub make it a practical choice for few-shot classification tasks. Install it if you need to classify text with limited labeled examples.AI-flagged interpretation of the facts on this page — verify before relying
Before you install
- Low friction: pure Python wheel with seven runtime dependencies (datasets, sentence-transformers, transformers, evaluate, huggingface_hub, scikit-learn, packaging).
- Repository is active with recent commits and no archived status.
License · maintenance · safety
Apache 2.0 (permissive) — Apache 2.0 permissive license allows commercial and private use with minimal restrictions—suitable for most production and research contexts.
last release 2025-08-05 (374 days) · last repo commit 2026-05-26 · 2,779 stars
0 known vulnerabilities (OSV.dev, 2026-08-14) · 261,172 downloads/mo, #8,386 on PyPI
Alternatives
Verify before relying
pip install setfit
from setfit import SetFitModel, Trainer, TrainingArguments
from datasets import load_dataset
dataset = load_dataset("sst2")
model = SetFitModel.from_pretrained("sentence-transformers/paraphrase-mpnet-base-v2", labels=["negative", "positive"])
trainer = Trainer(model=model, args=TrainingArguments(batch_size=16, num_epochs=4), train_dataset=dataset["train"])
trainer.train()
preds = model.predict(["i loved the spiderman movie!"])- Minimum Python version requirement (classifiers list 3.9–3.12 but requires_python field is unspecified)
- Whether multilingual support requires specific Sentence Transformer checkpoints or works automatically
- Typical memory and compute requirements for training on different dataset sizes
What it is and what it does
SetFit is a framework for few-shot text classification that combines a pretrained Sentence Transformer body with a lightweight classification head (either scikit-learn's LogisticRegression or a differentiable PyTorch-based head). It achieves competitive accuracy on classification tasks using only a small number of labeled examples per class—the documentation cites achieving results comparable to fine-tuning RoBERTa Large on a full 3k-example training set using only 8 labeled examples per class on sentiment data.
The package integrates with Hugging Face Hub for model discovery, training, and sharing. It wraps the fine-tuning process in a Trainer class that handles dataset sampling, evaluation, and model persistence. Runtime dependencies include datasets for data loading, transformers and sentence-transformers for embeddings, evaluate for metrics, huggingface_hub for Hub integration, scikit-learn for the default classification head, and packaging for version handling. No prompts or verbalizers are required—the framework generates embeddings directly from raw text.
Use it for
- Train a sentiment classifier on a customer review dataset with only 8 labeled examples per sentiment class
- Build a multilingual text classifier by fine-tuning a multilingual Sentence Transformer checkpoint on a small labeled corpus
- Rapidly prototype a document categorization system when labeled training data is scarce or expensive to obtain
- Evaluate few-shot learning performance on benchmark datasets and compare against other methods using the provided training scripts
Worth the install?
AI-flagged interpretation of the facts on this page. Verify before relying on it.
Yes.
SetFit is production-stable (Development Status 5), actively maintained, has no known vulnerabilities, and solves a real problem—achieving strong text classification with minimal labeled data. The low install friction, permissive license, and integration with Hugging Face Hub make it a practical choice for few-shot classification tasks. Install it if you need to classify text with limited labeled examples.
Install
setfit on PyPI
Before you install
Low friction: pure Python wheel with seven runtime dependencies (datasets, sentence-transformers, transformers, evaluate, huggingface_hub, scikit-learn, packaging). Repository is active with recent commits and no archived status.
License in practice
Apache 2.0 permissive license allows commercial and private use with minimal restrictions—suitable for most production and research contexts.
Quickstart
pip install setfit
from setfit import SetFitModel, Trainer, TrainingArguments
from datasets import load_dataset
dataset = load_dataset("sst2")
model = SetFitModel.from_pretrained("sentence-transformers/paraphrase-mpnet-base-v2", labels=["negative", "positive"])
trainer = Trainer(model=model, args=TrainingArguments(batch_size=16, num_epochs=4), train_dataset=dataset["train"])
trainer.train()
preds = model.predict(["i loved the spiderman movie!"])
Verify before relying
- Minimum Python version requirement (classifiers list 3.9–3.12 but requires_python field is unspecified)
- Whether multilingual support requires specific Sentence Transformer checkpoints or works automatically
- Typical memory and compute requirements for training on different dataset sizes
Package facts
| License | Apache 2.0 permissive |
| Python support | Not specified |
| Install friction | Low. Pure-Python wheel |
| Runtime dependencies | 7 packagesdatasetssentence-transformerstransformersevaluatehuggingface_hubscikit-learnpackaging |
| Maintenance | Actively maintained 374 days since the last release |
| Last repo commit | |
| First released | |
| Downloads | 261,172 / month, #8,386 on PyPI 30-day window, as of 2026-08-14 |
| Known vulnerabilities | None known OSV.dev, checked 2026-08-14 |
| Classifiers | Development Status :: 5 - Production/StableIntended Audience :: DevelopersIntended Audience :: EducationIntended Audience :: Science/ResearchLicense :: OSI Approved :: Apache Software LicenseOperating System :: OS IndependentProgramming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.12Programming Language :: Python :: 3.9Topic :: Scientific/Engineering :: Artificial Intelligence |
Evidence: setfit-1.1.3-py3-none-any.whl
Tags
Let your AI agent find packages like this
Example. Real query, live index.
You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.
wish › “few-shot text classification”
- setfitSetFit fine-tunes Sentence Transformers for text classification using…
- torchxrayvisionTorchXRayVision provides pre-trained deep learning models and unified…
- rf100vlProvides programmatic access to RF100-VL, a multi-domain object…
Give your agent the search over MCP, or paste the wish link into any chat.
More Artificial Intelligence packages
LiteLLM provides a unified Python interface to call 100+ LLM providers (OpenAI, Anthropic, Gemini, Bedrock, Azure, and others) using OpenAI-compatible API format, available as both a Python SDK and a self-hosted AI Gateway proxy server.
Install it if you need to work with multiple LLM providers or want to centralize LLM routing in your organization.
Client library and CLI tool for downloading, uploading, and managing models, datasets, and repositories on the Hugging Face Hub platform.
Install it if you work with Hugging Face Hub models or datasets.
LangChain provides a framework for building agents and LLM-powered applications by composing language models, tools, and memory through a unified API that abstracts over multiple model providers.
hf-xet provides chunk-based deduplication and efficient file transfer for the Hugging Face Hub, enabling faster uploads and downloads of large files with local disk caching.
Tokenizers converts raw text into token sequences for NLP models, with support for training custom vocabularies and using pre-built tokenizers (BPE, WordPiece) optimized for speed via Rust.
Transformers provides a unified framework for loading, fine-tuning, and running state-of-the-art pretrained models across text, vision, audio, video, and multimodal tasks using PyTorch, JAX, or TensorFlow.
Install it if you need to run or train any transformer-based model for NLP, vision, audio, or multimodal tasks.
See also trl · transformer-smaller-training-vocab · model2vec · sentence-transformers · peft · flair · detoxify · gliner · InstructorEmbedding