Splink
- Android app
- Not listed
- Free plan
- Yes
- Runs on
- Linux, Mac, self-hosted, Windows

Summary
Splink is a free Python package for probabilistic record linkage, used to deduplicate and connect records when unique identifiers are unavailable. Its core method follows the Fellegi-Sunter model and can be trained without labeled examples through an unsupervised approach. Matching can use term-frequency adjustments and custom fuzzy logic. Splink creates SQL for a selected backend, with documented options including DuckDB, Spark, SQLite and PostgreSQL. The project recommends DuckDB for most situations and Spark for very large linkages or when a Spark cluster is easier to access; the maker says it can handle linkages of 100+ million records on DuckDB or big-data backends such as Spark. The package works best with standardized data across multiple, not highly correlated columns, and is not intended for a single bag-of-words column. Interactive visualizations help inspect and diagnose models, predictions and clusters. Install with pip or conda, with additional backend-specific setup documented for Spark and PostgreSQL. The maker points users with remaining questions to its GitHub discussion forum.
Who it is for
Splink suits analysts and data teams linking or deduplicating records without unique identifiers, provided their data has multiple standardized columns rather than only bag-of-words text.
What is good
- Can train without labeled examples.
- Supports fuzzy logic and term-frequency adjustments.
- Offers four documented SQL backends.
- Includes interactive model and cluster diagnostics.
- Free Python package installable with pip or conda.
What to know first
- Not designed for a single bag-of-words column.
- Databricks-specific support may be difficult to provide.
Verdict
Splink is aimed at record linkage work where identifiers are missing and data has useful structured columns. Backend choice and data preparation matter, and users working with Databricks-specific issues may have limited support.
Splink plans and pricing
All plansCompared on identity resolution software
- Free plan
- Yesmoj-analytical-services.github.io
- Matching approach
- hybridmoj-analytical-services.github.io
- Real-time API
- Yesmoj-analytical-services.github.io
- Batch file import
- Yesmoj-analytical-services.github.io
- Organization matching
- Yesmoj-analytical-services.github.io
Facts
- Purpose
- Splink is a Python package for probabilistic record linkage that deduplicates and links records without unique identifiers.moj-analytical-services.github.io · 29 Sept 2026
- Method
- Its core linkage algorithm is based on the Fellegi-Sunter model and can be trained without labeled data using an unsupervised approach.moj-analytical-services.github.io · 29 Sept 2026
- Matching
- It supports term frequency adjustments and user-defined fuzzy matching logic.moj-analytical-services.github.io · 29 Sept 2026
- Scale
- The maker says Splink can run on DuckDB or big-data backends such as Spark for linkages of 100+ million records.moj-analytical-services.github.io · 29 Sept 2026
- Backends
- The documented SQL backends include DuckDB, Spark, SQLite, and PostgreSQL; the library generates SQL for a user-chosen backend.moj-analytical-services.github.io · 29 Sept 2026
- Backend guidance
- DuckDB is recommended for most users except the largest linkages, while Spark is recommended for very large linkages or where a Spark cluster is easier to access.moj-analytical-services.github.io · 29 Sept 2026
- Data requirements
- Splink works best with standardized data containing multiple columns that are not highly correlated, and is not designed for a single bag-of-words column.moj-analytical-services.github.io · 29 Sept 2026
- Diagnostics
- Interactive visualisations help users understand and diagnose linkage models, including dashboards for examining predictions and clusters.moj-analytical-services.github.io · 29 Sept 2026
- Install
- Splink can be installed using pip or conda, with optional backend-specific installs documented for Spark and PostgreSQL.moj-analytical-services.github.io · 29 Sept 2026
- Support
- The maker directs users with questions remaining after reading the documentation to its GitHub discussion forum.moj-analytical-services.github.io · 29 Sept 2026
- Databricks support
- The development team says it lacks access to a Databricks environment and may struggle to help with Databricks-specific issues.moj-analytical-services.github.io · 29 Sept 2026
- Use cases
- The maker lists users across government, academia, and other sectors, including the Office for National Statistics, NHS England, and the Australian Bureau of Statistics.moj-analytical-services.github.io · 29 Sept 2026
Best Splink alternatives
See all 12Where it ranks on Everything Xiaomi
Is Splink yours?
Claim it for free: prove the domain, then correct facts, plans and screenshots. An editor reviews every change.
Sources
- moj-analytical-services.github.io/splink/· checked 29 Sept 2026
- moj-analytical-services.github.io/splink/topic_guides/splink_fundamentals· checked 29 Sept 2026
- moj-analytical-services.github.io/splink/api_docs/visualisations.html· checked 29 Sept 2026
- moj-analytical-services.github.io/splink/getting_started.html· checked 29 Sept 2026
- github.com/moj-analytical-services/splink· checked 29 Sept 2026



