$npx skillfedfor your agent

recordlinkage

A record linkage toolkit for linking and deduplication

With conditionsPyPI Information AnalysisReleased Jul 20233.1M downloads / moBSD-3-ClausePure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — recordlinkage-0.16-py3-none-any.whl
v0.16 · released 2023-07-20 · Python >=3.8 · 6 runtime deps: jellyfish, numpy, pandas, scipy, scikit-learn, joblib

Yes, if you need record linkage or deduplication on small-to-medium datasets and can tolerate dormant maintenance. The package is stable, has low install friction, and integrates cleanly with pandas. However, avoid it for production systems requiring active support or if you need compatibility guarantees with very recent versions of numpy, pandas, or scikit-learn.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.8 or higher.
  • Low install friction with a pure-Python wheel and six common dependencies (numpy, pandas, scipy, scikit-learn, joblib, jellyfish).
  • Maintenance is dormant—last release was 2023-07-20 and last commit 2024-02-21, over 1121 days ago—so expect no active bug fixes or feature updates.

License · maintenance · safety

BSD-3-Clause (permissive) — BSD-3-Clause is permissive and imposes minimal restrictions; you may use, modify, and distribute the package freely in commercial and private projects provided you include the license notice.

last release 2023-07-20 (1121 days) · last repo commit 2024-02-21 · 1,059 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 3,071,659 downloads/mo, #2,766 on PyPI

Verify before relying

import recordlinkage
import pandas

df_a = pandas.DataFrame(data_a)
df_b = pandas.DataFrame(data_b)

indexer = recordlinkage.Index()
indexer.block('surname')
candidate_links = indexer.index(df_a, df_b)

c = recordlinkage.Compare()
c.string('name_a', 'name_b', method='jarowinkler', threshold=0.85)
feature_vectors = c.compute(candidate_links, df_a, df_b)
  • Whether dormant maintenance status affects compatibility with pandas/numpy/scikit-learn versions released after 2024-02-21.
  • Performance characteristics on large datasets (the description mentions 'small or medium sized files').
Same gist for agents: .md · .json

What it is and what it does

RecordLinkage is a Python toolkit for matching records within or across datasets, commonly used for deduplication and entity resolution. It wraps pandas and numpy to provide indexing methods (like blocking and sorted neighbourhood indexing), comparison functions for strings, numbers, and dates, and both supervised and unsupervised classifiers to determine which record pairs are matches. The workflow is: create candidate pairs using an indexer, compute similarity features using a comparator, then classify pairs as matches or non-matches with a classifier like Logistic Regression or ECM.

The package is designed for research and small-to-medium datasets. It integrates directly with pandas DataFrames, making it natural to use in existing data pipelines. Dependencies include scipy and scikit-learn for statistical and machine-learning operations, and jellyfish for string similarity metrics. No known security vulnerabilities are recorded.

Use it for

  • Deduplicate customer records in a CRM by blocking on surname and comparing name, address, and phone fields.
  • Link census records across two time periods to track population changes and identify the same individuals.
  • Match product catalogs from different suppliers to identify equivalent items for price comparison.
  • Resolve duplicate entries in a research database by comparing author names, publication dates, and keywords.
  • Merge employee records from multiple HR systems using blocking on department and comparing names and IDs.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

With conditions

Yes, if you need record linkage or deduplication on small-to-medium datasets and can tolerate dormant maintenance.

The package is stable, has low install friction, and integrates cleanly with pandas. However, avoid it for production systems requiring active support or if you need compatibility guarantees with very recent versions of numpy, pandas, or scikit-learn.

Install

recordlinkage on PyPI

Before you install

Low install friction with a pure-Python wheel and six common dependencies (numpy, pandas, scipy, scikit-learn, joblib, jellyfish). Maintenance is dormant—last release was 2023-07-20 and last commit 2024-02-21, over 1121 days ago—so expect no active bug fixes or feature updates.

Requires Python 3.8 or higher.

License in practice

BSD-3-Clause is permissive and imposes minimal restrictions; you may use, modify, and distribute the package freely in commercial and private projects provided you include the license notice.

Quickstart

import recordlinkage
import pandas

df_a = pandas.DataFrame(data_a)
df_b = pandas.DataFrame(data_b)

indexer = recordlinkage.Index()
indexer.block('surname')
candidate_links = indexer.index(df_a, df_b)

c = recordlinkage.Compare()
c.string('name_a', 'name_b', method='jarowinkler', threshold=0.85)
feature_vectors = c.compute(candidate_links, df_a, df_b)

Verify before relying

  • Whether dormant maintenance status affects compatibility with pandas/numpy/scikit-learn versions released after 2024-02-21.
  • Performance characteristics on large datasets (the description mentions 'small or medium sized files').

Package facts

LicenseBSD-3-Clause permissive
Python supportSupports the current Python release >=3.8
Install frictionLow. Pure-Python wheel
Runtime dependencies
6 packages
jellyfishnumpypandasscipyscikit-learnjoblib
MaintenanceDormant 1,121 days since the last release
Last repo commit
First released
Downloads3,071,659 / month, #2,766 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14
Classifiers
Development Status :: 4 - BetaLicense :: OSI Approved :: BSD LicenseProgramming Language :: Python :: 3.10Programming Language :: Python :: 3.11Programming Language :: Python :: 3.8Programming Language :: Python :: 3.9

Evidence: recordlinkage-0.16-py3-none-any.whl

Tags

Capabilities
record linkage and deduplicationduplicate detection in datasetsentity matching and resolutiondata record matchingfuzzy record linkingpandas-based record comparisonblocking and indexing for records
Topics
data-deduplicationentity-resolutionpandas-native

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “pandas-based record comparison”

  • recordlinkageRecordLinkage identifies and matches records within or across…
  • sweetvizSweetviz generates interactive HTML visualizations for exploratory…
  • python-LevenshteinComputes Levenshtein edit distance, string similarity, and…

Give your agent the search over MCP, or paste the wish link into any chat.

More Information Analysis packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyarrow Worth it
PyPI · Information Analysis · released Aug 2026

pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.

Apache-2.0compiled wheel · 3.10+
432.9Mdownloads / mo
networkx Worth it
PyPI · Python Modules · released Dec 2025

NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.

BSD-3-Clausepure Python
290.9Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
contourpy Worth it
PyPI · Information Analysis · released Jul 2025

ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.

BSD-3-Clausecompiled wheel · 3.11+
191.2Mdownloads / mo
snowflake-snowpark-python Worth it
PyPI · Software Development · released Jul 2026

Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.

Install it if you use Snowflake and want to process data without moving it to your application layer.

Apache-2.0pure Python
100.7Mdownloads / mo

See also splink · semhash · pandas · recursive-diff · records · invenio-records-permissions · simhash · datacompy · imagededup · fastcluster