Splink

B
B tier on Identity Resolution SoftwareScore 7.0 · #2 of 25
Android app
Not listed
Free plan
Yes
Runs on
Linux, Mac, self-hosted, Windows
moj-analytical-services.github.io
The Splink homepage

Summary

Splink is a free Python package for probabilistic record linkage, used to deduplicate and connect records when unique identifiers are unavailable. Its core method follows the Fellegi-Sunter model and can be trained without labeled examples through an unsupervised approach. Matching can use term-frequency adjustments and custom fuzzy logic. Splink creates SQL for a selected backend, with documented options including DuckDB, Spark, SQLite and PostgreSQL. The project recommends DuckDB for most situations and Spark for very large linkages or when a Spark cluster is easier to access; the maker says it can handle linkages of 100+ million records on DuckDB or big-data backends such as Spark. The package works best with standardized data across multiple, not highly correlated columns, and is not intended for a single bag-of-words column. Interactive visualizations help inspect and diagnose models, predictions and clusters. Install with pip or conda, with additional backend-specific setup documented for Spark and PostgreSQL. The maker points users with remaining questions to its GitHub discussion forum.

Who it is for

Splink suits analysts and data teams linking or deduplicating records without unique identifiers, provided their data has multiple standardized columns rather than only bag-of-words text.

What is good

  • Can train without labeled examples.
  • Supports fuzzy logic and term-frequency adjustments.
  • Offers four documented SQL backends.
  • Includes interactive model and cluster diagnostics.
  • Free Python package installable with pip or conda.

What to know first

  • Not designed for a single bag-of-words column.
  • Databricks-specific support may be difficult to provide.

Verdict

Splink is aimed at record linkage work where identifiers are missing and data has useful structured columns. Backend choice and data preparation matter, and users working with Databricks-specific issues may have limited support.

Splink plans and pricing

All plans
Splink Free Open-source Python package · install via pip or conda github.com · 29 Sept 2026

Compared on identity resolution software

Free plan
Yesmoj-analytical-services.github.io
Matching approach
hybridmoj-analytical-services.github.io
Real-time API
Yesmoj-analytical-services.github.io
Batch file import
Yesmoj-analytical-services.github.io
Organization matching
Yesmoj-analytical-services.github.io

Facts

Purpose
Splink is a Python package for probabilistic record linkage that deduplicates and links records without unique identifiers.moj-analytical-services.github.io · 29 Sept 2026
Method
Its core linkage algorithm is based on the Fellegi-Sunter model and can be trained without labeled data using an unsupervised approach.moj-analytical-services.github.io · 29 Sept 2026
Matching
It supports term frequency adjustments and user-defined fuzzy matching logic.moj-analytical-services.github.io · 29 Sept 2026
Scale
The maker says Splink can run on DuckDB or big-data backends such as Spark for linkages of 100+ million records.moj-analytical-services.github.io · 29 Sept 2026
Backends
The documented SQL backends include DuckDB, Spark, SQLite, and PostgreSQL; the library generates SQL for a user-chosen backend.moj-analytical-services.github.io · 29 Sept 2026
Backend guidance
DuckDB is recommended for most users except the largest linkages, while Spark is recommended for very large linkages or where a Spark cluster is easier to access.moj-analytical-services.github.io · 29 Sept 2026
Data requirements
Splink works best with standardized data containing multiple columns that are not highly correlated, and is not designed for a single bag-of-words column.moj-analytical-services.github.io · 29 Sept 2026
Diagnostics
Interactive visualisations help users understand and diagnose linkage models, including dashboards for examining predictions and clusters.moj-analytical-services.github.io · 29 Sept 2026
Install
Splink can be installed using pip or conda, with optional backend-specific installs documented for Spark and PostgreSQL.moj-analytical-services.github.io · 29 Sept 2026
Support
The maker directs users with questions remaining after reading the documentation to its GitHub discussion forum.moj-analytical-services.github.io · 29 Sept 2026
Databricks support
The development team says it lacks access to a Databricks environment and may struggle to help with Databricks-specific issues.moj-analytical-services.github.io · 29 Sept 2026
Use cases
The maker lists users across government, academia, and other sectors, including the Office for National Statistics, NHS England, and the Australian Bureau of Statistics.moj-analytical-services.github.io · 29 Sept 2026

Best Splink alternatives

See all 12

Where it ranks on Everything Xiaomi

Is Splink yours?

Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.

Sources