$npx skillfedfor your agent

splink

Fast probabilistic data linkage at scale

Worth itPyPI Information AnalysisReleased Mar 20261.2M downloads / moMITPure Python

Decision gist · record as of 2026-08-14

pure-Python wheel — splink-4.0.16-py3-none-any.whl
v4.0.16 · released 2026-03-11 · Python <4.0.0,>=3.9.0 · 7 runtime deps: altair, duckdb, igraph, jinja2, numpy, pandas, sqlglot

Yes. Splink is actively maintained, has no known vulnerabilities, installs with low friction, and solves a specific and difficult problem (record linkage without unique identifiers) that has few mature alternatives in Python. The MIT license and strong maintenance signal (recent release, active repository) make it suitable for production use in government, academic, and commercial contexts. Install if you need to deduplicate or link records across datasets.AI-flagged interpretation of the facts on this page — verify before relying

Before you install

  • Requires Python 3.9 or later; DuckDB is a runtime dependency for local linkage execution.
  • Low friction install with a pure Python wheel.
  • Actively maintained with a recent release (156 days ago) and 2338 repository stars.

License · maintenance · safety

MIT (permissive) — MIT license permits commercial and private use with minimal restrictions, making it suitable for government, academic, and private sector deployments.

last release 2026-03-11 (156 days) · last repo commit 2026-08-13 · 2,338 stars

0 known vulnerabilities (OSV.dev, 2026-08-14) · 1,170,073 downloads/mo, #4,271 on PyPI

Verify before relying

pip install splink

import splink.comparison_library as cl
from splink import DuckDBAPI, Linker, SettingsCreator, block_on, splink_datasets

db_api = DuckDBAPI()
df = splink_datasets.fake_1000
settings = SettingsCreator(link_type="dedupe_only", comparisons=[cl.ExactMatch("name")])
linker = Linker(df, settings, db_api)
pairwise_predictions = linker.inference.predict()
  • Whether optional backend installations (Spark, Athena, PostgreSQL) are commonly used or if DuckDB covers most use cases.
  • Performance characteristics on datasets larger than the stated 'million records on a laptop in around a minute' benchmark.
  • Whether the Fellegi-Sunter model's accuracy claims have been independently validated outside government case studies.
Same gist for agents: .md · .json

What it is and what it does

Splink is a Python package for probabilistic record linkage that solves the problem of matching and deduplicating records when no unique identifier exists. It uses the Fellegi-Sunter statistical model to compute match probabilities between record pairs, supporting fuzzy matching, term frequency adjustments, and user-defined comparison logic. The package works by comparing multiple non-correlated columns (such as name, date of birth, and location for persons), estimating model parameters through unsupervised learning, and clustering pairwise predictions to generate estimated entity IDs.

The package is designed for datasets with multiple descriptive columns and runs on a local laptop via DuckDB or scales to 100+ million records on big-data backends like AWS Athena or Spark. It includes interactive visualizations to help diagnose model performance and is widely used in government, academia, and the private sector. Runtime dependencies include altair, duckdb, igraph, jinja2, numpy, pandas, and sqlglot.

Use it for

  • Deduplicate customer or patient records in databases lacking a master identifier.
  • Link census or survey data across years or sources to track population changes.
  • Match company records across datasets with different naming conventions or incomplete information.
  • Resolve entity identity in fraud detection or compliance workflows where records may be partially obscured.
  • Consolidate data from multiple administrative systems without a shared key.

Worth the install?

AI-flagged interpretation of the facts on this page. Verify before relying on it.

Worth it

Yes.

Splink is actively maintained, has no known vulnerabilities, installs with low friction, and solves a specific and difficult problem (record linkage without unique identifiers) that has few mature alternatives in Python. The MIT license and strong maintenance signal (recent release, active repository) make it suitable for production use in government, academic, and commercial contexts. Install if you need to deduplicate or link records across datasets.

Install

splink on PyPI

Before you install

Low friction install with a pure Python wheel. Actively maintained with a recent release (156 days ago) and 2338 repository stars. Supports current Python versions (3.9+).

Requires Python 3.9 or later; DuckDB is a runtime dependency for local linkage execution.

License in practice

MIT license permits commercial and private use with minimal restrictions, making it suitable for government, academic, and private sector deployments.

Quickstart

pip install splink

import splink.comparison_library as cl
from splink import DuckDBAPI, Linker, SettingsCreator, block_on, splink_datasets

db_api = DuckDBAPI()
df = splink_datasets.fake_1000
settings = SettingsCreator(link_type="dedupe_only", comparisons=[cl.ExactMatch("name")])
linker = Linker(df, settings, db_api)
pairwise_predictions = linker.inference.predict()

Verify before relying

  • Whether optional backend installations (Spark, Athena, PostgreSQL) are commonly used or if DuckDB covers most use cases.
  • Performance characteristics on datasets larger than the stated 'million records on a laptop in around a minute' benchmark.
  • Whether the Fellegi-Sunter model's accuracy claims have been independently validated outside government case studies.

Package facts

LicenseMIT permissive
Python supportSupports the current Python release <4.0.0,>=3.9.0
Install frictionLow. Pure-Python wheel
Runtime dependencies
7 packages
altairduckdbigraphjinja2numpypandassqlglot
MaintenanceActively maintained 156 days since the last release
Last repo commit
First released
Downloads1,170,073 / month, #4,271 on PyPI 30-day window, as of 2026-08-14
Known vulnerabilitiesNone known OSV.dev, checked 2026-08-14

Evidence: splink-4.0.16-py3-none-any.whl

Tags

Capabilities
record linkage entity resolutiondeduplication matchingprobabilistic data linkingfuzzy record matchingduplicate detectionentity resolution pythonfellegi sunter linkage
Topics
entity-resolutiondata-qualityunsupervised-learning

Let your AI agent find packages like this

Example. Real query, live index.

You found this page by searching. An agent finds it by wishing: SkillFed indexes 14,416 PyPI packages by what they can do, searchable in plain language.

wish › “record linkage entity resolution”

  • splinkSplink performs probabilistic record linkage and deduplication,…
  • recordlinkageRecordLinkage identifies and matches records within or across…
  • fuzzyset2Performs fuzzy string matching and approximate searching against a…

Give your agent the search over MCP, or paste the wish link into any chat.

More Information Analysis packages

regex Worth it
PyPI · Python Modules · released Jul 2026

A drop-in replacement for Python's standard `re` module that adds advanced regex features like nested sets, fuzzy matching, lookaround in conditionals, and full Unicode case-folding while maintaining backward compatibility.

Apache-2.0 AND CNRI-Pythoncompiled wheel · 3.10+
437.7Mdownloads / mo
pyarrow Worth it
PyPI · Information Analysis · released Aug 2026

pyarrow provides Python bindings to Apache Arrow's C++ libraries for efficient columnar data processing, serialization, and interoperability with pandas, NumPy, and other Python ecosystem tools.

Apache-2.0compiled wheel · 3.10+
432.9Mdownloads / mo
networkx Worth it
PyPI · Python Modules · released Dec 2025

NetworkX provides data structures and algorithms for creating, analyzing, and manipulating graphs and networks, supporting everything from simple undirected graphs to complex directed and weighted networks.

BSD-3-Clausepure Python
290.9Mdownloads / mo
snowflake-connector-python Worth it
PyPI · Software Development · released Aug 2026

Connects Python applications to Snowflake data warehouses using the DB API 2.0 specification, enabling SQL queries, data transfers, and warehouse operations.

Apache-2.0compiled wheel · 3.10+
193.6Mdownloads / mo
contourpy Worth it
PyPI · Information Analysis · released Jul 2025

ContourPy calculates contours of 2D quadrilateral grids using C++11 algorithms wrapped in Python, offering serial and multithreaded implementations without requiring Matplotlib as a dependency.

BSD-3-Clausecompiled wheel · 3.11+
191.2Mdownloads / mo
snowflake-snowpark-python Worth it
PyPI · Software Development · released Jul 2026

Snowpark Python provides APIs to query and process data directly in Snowflake without moving data to your local system, with support for both native Snowpark and pandas-compatible interfaces.

Install it if you use Snowflake and want to process data without moving it to your application layer.

Apache-2.0pure Python
100.7Mdownloads / mo

See also recordlinkage · semhash · probablepeople · hdbscan · diff-match-patch · fingerprints · simhash · datasketch · fastcluster