$npx skillfedfor your agent

pinecone-text

Text utilities library by Pinecone.io

With conditionsPyPI Text ProcessingReleased Aug 2025514.0K downloads / moPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — pinecone_text-0.11.0-py3-none-any.whl
v0.11.0 · released 2025-08-11 · Python <4.0,>=3.9 · 6 runtime deps: mmh3, nltk, numpy, requests, types-requests, python-dotenv

Yes, with conditions. The package is useful for developers building hybrid search on Pinecone and want to avoid writing encoding logic, but maintenance is aging (368 days since last release) and license terms are unclear. Install if you are committed to Pinecone's platform and can work within Python 3.9–3.11 (avoid 3.12 for SPLADE and Sentence Transformers). Verify the license before use in proprietary contexts. No known security vulnerabilities.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • SPLADE and SentenceTransformerEncoder are incompatible with Python 3.12 due to PyTorch compatibility issues; optional extras (splade, dense, openai) require separate installation and their own dependencies.
  • Low install friction with six runtime dependencies.
  • Maintenance status is aging—last release was 368 days ago—but the package remains functional for current Python versions (3.9–3.11); note that SPLADE and Sentence Transformers encoders have known incompatibilities with Python 3.12 due to PyTorch constraints.

License · maintenance · safety

(unclear) — License treatment is unclear; no SPDX or raw license metadata is available in the package metadata. Verify the actual license terms before use in proprietary or restricted contexts.

last release 2025-08-11 (368 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 514,004 downloads/mo, #6,242 on PyPI

Verify before relying

pip install pinecone-text

from pinecone_text.sparse import BM25Encoder

corpus = ["The quick brown fox", "The lazy dog"]
bm25 = BM25Encoder()
bm25.fit(corpus)
vector = bm25.encode_documents("brown fox")
  • Whether the package is still actively maintained or in maintenance-only mode given the 368-day gap since last release.
  • Specific performance characteristics or throughput limits for encoding large document batches.
  • Whether BM25's static document frequency model is suitable for your use case or if dynamic retraining is needed.
Same gist for agents: .md · .json

What it is and what it does

Pinecone Text is a utility library that bridges text data and Pinecone's vector search engine by providing encoders that convert documents and queries into sparse or dense vectors. It wraps multiple encoding strategies—BM25 for traditional sparse vectors, SPLADE for learned sparse representations, and integrations with Sentence Transformers and OpenAI's embedding models for dense vectors—allowing developers to prepare text for hybrid search without writing encoding logic themselves.

The package is designed for use with Pinecone's hybrid search, which combines sparse and dense vectors for improved retrieval. It handles tokenization, model loading, and vector formatting, but requires explicit installation of optional dependencies for SPLADE, Sentence Transformers, or OpenAI support. BM25 is available by default and can be initialized with precomputed parameters or fitted to a custom corpus; SPLADE uses a fixed HuggingFace model; dense encoders delegate to external services or local models.

Use it for

  • Prepare a corpus of documents for BM25-based sparse vector indexing in Pinecone without implementing tokenization and IDF calculation yourself.
  • Encode queries and documents using SPLADE for learned sparse retrieval when BM25 alone is insufficient.
  • Generate dense embeddings via OpenAI's API and store them in Pinecone for semantic search without managing API calls directly.
  • Combine BM25 sparse vectors with Sentence Transformer dense vectors for hybrid search in a single pipeline.
  • Load and reuse precomputed BM25 parameters (fitted on MS MARCO) to encode new documents without retraining.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, with conditions.

The package is useful for developers building hybrid search on Pinecone and want to avoid writing encoding logic, but maintenance is aging (368 days since last release) and license terms are unclear. Install if you are committed to Pinecone's platform and can work within Python 3.9–3.11 (avoid 3.12 for SPLADE and Sentence Transformers). Verify the license before use in proprietary contexts. No known security vulnerabilities.

Install

pinecone-text on PyPI

Before you install

Low install friction with six runtime dependencies. Maintenance status is aging—last release was 368 days ago—but the package remains functional for current Python versions (3.9–3.11); note that SPLADE and Sentence Transformers encoders have known incompatibilities with Python 3.12 due to PyTorch constraints.

SPLADE and SentenceTransformerEncoder are incompatible with Python 3.12 due to PyTorch compatibility issues; optional extras (splade, dense, openai) require separate installation and their own dependencies.

License in practice

License treatment is unclear; no SPDX or raw license metadata is available in the package metadata. Verify the actual license terms before use in proprietary or restricted contexts.

Quickstart

pip install pinecone-text

from pinecone_text.sparse import BM25Encoder

corpus = ["The quick brown fox", "The lazy dog"]
bm25 = BM25Encoder()
bm25.fit(corpus)
vector = bm25.encode_documents("brown fox")

Verify before relying

  • Whether the package is still actively maintained or in maintenance-only mode given the 368-day gap since last release.
  • Specific performance characteristics or throughput limits for encoding large document batches.
  • Whether BM25's static document frequency model is suitable for your use case or if dynamic retraining is needed.

Package facts

LicenseNot declared unclear
Python supportSupports the current Python release <4.0,>=3.9
Install frictionLow. Pure-Python wheel
Runtime dependencies
6 packages
mmh3nltknumpyrequeststypes-requestspython-dotenv
MaintenanceAging 368 days since the last release
First released
Downloads514,004 / month, #6,242 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Programming Language :: Python :: 3Programming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.9

Evidence: pinecone_text-0.11.0-py3-none-any.whl

Tags

Capabilities
text to sparse vector encodingBM25 encoder for searchhybrid semantic search vectorsdense embeddings for pineconeSPLADE sparse encodingsentence transformer embeddingsopenai embedding wrapper
Topics
vector-embeddingssemantic-searchpinecone-integration

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “text to sparse vector encoding”

Give your agent the search over MCP, or paste the wish link into any chat.

More Text Processing packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyparsing Worth it
PyPI · Text Processing · released Jan 2026

pyparsing provides a library for building text parsers directly in Python code using composable grammar classes, handling quoted strings, whitespace variation, and embedded comments without regex or lex/yacc.

Install it if you need to parse text or define grammars programmatically.

MITpure Python · 3.9+
412.7Mdownloads / mo
fonttools Worth it
PyPI · Text Processing · released May 2026

fonttools manipulates font files in multiple formats (TrueType, OpenType, AFM, Type 1, Mac-specific) and includes TTX, a tool to convert fonts to and from XML text format.

Install it if you need to read, write, or manipulate fonts programmatically or via the TTX command-line tool.

permissive licensepure Python · 3.10+
235.9Mdownloads / mo
docutils With conditions
PyPI · Software Development · released May 2026

Docutils converts plaintext documentation in reStructuredText format into multiple output formats including HTML, XML, and LaTeX using a modular processing system.

BSD-3-Clausepure Python · 3.9+
225.6Mdownloads / mo
RapidFuzz Worth it
PyPI · Text Processing · released Apr 2026

RapidFuzz provides fast fuzzy string matching using Levenshtein Distance and related metrics, implemented mostly in C++ with Python bindings for rapid similarity scoring and approximate string matching.

Install it if you need fuzzy string matching; it's a solid replacement for FuzzyWuzzy with better licensing and performance.

MITcompiled wheel · 3.10+
184.2Mdownloads / mo
tinycss2 Worth it
PyPI · Text Processing · released Nov 2025

tinycss2 parses CSS strings into token and block objects, and generates CSS strings from those objects, following the CSS Syntax Level 3 specification without enforcing specific properties or values.

Install it if your project requires CSS tokenization or syntax manipulation.

BSD-3-Clausepure Python · 3.10+
113.2Mdownloads / mo

See also langchain-pinecone · sentence-transformers · model2vec · pinecone-plugin-interface · llama-index-vector-stores-pinecone · pinecone-plugin-inference · bm25s · floret · voyageai · meilisearch

Further reading