$npx skillfedfor your agent

string-grouper

String grouper contains functions to do string matching using TF-IDF and the cossine similarity.

Worth itPyPI Information AnalysisReleased Jul 2026116.0K downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — string_grouper-0.8.0-py3-none-any.whl
v0.8.0 · released 2026-07-26 · Python <4.0,>=3.10 · 7 runtime deps: loguru, numpy, pandas, scikit-learn, scipy, sp-matmul-rs, sparse-dot-topn

Yes. String Grouper is actively maintained, has no known vulnerabilities, installs with low friction, and solves a concrete problem in data cleaning and deduplication. It is well-suited for anyone working with messy text data in pandas workflows. The MIT license removes licensing friction. Install it if you need fuzzy string matching or deduplication at scale.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.10 or later (and less than 4.0).
  • Low install friction with a pure Python wheel.
  • Active maintenance as of 19 days ago.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects without licensing concerns.

last release 2026-07-26 (19 days)

0 known vulnerabilities (OSV.dev, 2026-08-14) · 115,963 downloads/mo, #12,228 on PyPI

Verify before relying

pip install string-grouper

from string_grouper import match_strings, group_similar_strings
import pandas as pd

names = pd.Series(['Apple Inc', 'Apple Inc.', 'Microsoft Corp'])
matches = match_strings(names)
groups = group_similar_strings(names)
  • Whether the Rust-based sp_matmul_rs backend is pre-compiled for all common platforms or requires build tools.
  • Performance characteristics on datasets smaller than the 663,000-name example cited in the description.
Same gist for agents: .md · .json

What it is and what it does

String Grouper is a library for finding groups of similar strings within a single list or across multiple lists. It uses TF-IDF vectorization and cosine similarity to identify matches, then groups them into clusters with a centroid representative. The library is built for speed: it leverages sp_matmul_rs, a Rust-based sparse matrix multiplication library, to compute similarities efficiently even on large datasets.

The package is typically used for data cleaning and deduplication tasks—matching company names with typos or formatting variations, finding duplicate entries in databases, or resolving indirect associations between strings through graph-based grouping. It exposes two main functions: match_strings to find pairwise matches above a similarity threshold, and group_similar_strings to cluster strings and identify canonical representatives. The core dependencies are numpy, pandas, scikit-learn, and scipy, making it a natural fit for data-science workflows.

Use it for

  • Deduplicate company or product names in datasets where exact matches fail due to formatting or spelling variations.
  • Find all variations of a customer name across multiple database records for entity resolution.
  • Identify similar addresses or locations in bulk data to consolidate records.
  • Cluster misspelled or abbreviated terms in text datasets to group related concepts.
  • Resolve indirect associations between strings in large datasets where direct pairwise comparison is infeasible.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

String Grouper is actively maintained, has no known vulnerabilities, installs with low friction, and solves a concrete problem in data cleaning and deduplication. It is well-suited for anyone working with messy text data in pandas workflows. The MIT license removes licensing friction. Install it if you need fuzzy string matching or deduplication at scale.

Install

string-grouper on PyPI

Before you install

Low install friction with a pure Python wheel. Active maintenance as of 19 days ago. Depends on numpy, pandas, scikit-learn, scipy, and two specialized sparse-matrix libraries (sp_matmul_rs and sparse_dot_topn), all of which are standard data-science packages.

Requires Python 3.10 or later (and less than 4.0).

License in practice

MIT license permits commercial and private use with minimal restrictions, making it suitable for most projects without licensing concerns.

Quickstart

pip install string-grouper

from string_grouper import match_strings, group_similar_strings
import pandas as pd

names = pd.Series(['Apple Inc', 'Apple Inc.', 'Microsoft Corp'])
matches = match_strings(names)
groups = group_similar_strings(names)

Verify before relying

  • Whether the Rust-based sp_matmul_rs backend is pre-compiled for all common platforms or requires build tools.
  • Performance characteristics on datasets smaller than the 663,000-name example cited in the description.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release <4.0,>=3.10
Install frictionLow. Pure-Python wheel
Runtime dependencies
7 packages
logurunumpypandasscikit-learnscipysp-matmul-rssparse-dot-topn
MaintenanceActively maintained 19 days since the last release
First released
Downloads115,963 / month, #12,228 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: string_grouper-0.8.0-py3-none-any.whl

Tags

Capabilities
fuzzy string matchingfind similar stringsstring deduplicationtf-idf similaritygroup similar textcosine similarity matchingfast string clustering
Topics
text-processingdata-deduplicationsimilarity-search

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “find similar strings”

  • string-grouperString Grouper finds groups of similar strings within or across lists…
  • ImageHashImageHash computes perceptual fingerprints of images using multiple…
  • polylevenPolyleven computes Levenshtein distance between two strings using a…

Give your agent the search over MCP, or paste the wish link into any chat.

More Information Analysis packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyarrow Worth it
PyPI · Information Analysis · released Aug 2026

pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.

Apache-2.0compiled wheel · 3.10+
432.9Mdownloads / mo
networkx Worth it
PyPI · Python Modules · released Dec 2025

NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.

BSD-3-Clausepure Python
290.9Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
contourpy Worth it
PyPI · Information Analysis · released Jul 2025

ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.

BSD-3-Clausecompiled wheel · 3.11+
191.2Mdownloads / mo
snowflake-snowpark-python Worth it
PyPI · Software Development · released Jul 2026

Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.

Install it if you use Snowflake and want to process data without moving it to your application layer.

Apache-2.0pure Python
100.7Mdownloads / mo

See also tfidf-matcher · strsimpy · fuzzyset2 · pysimstring · textdistance · ngram · Levenshtein · sparse-dot-topn · py-tlsh · pfzy