UCL Discovery
UCL home » Library Services » Electronic resources » UCL Discovery

Generating synthetic identifiers to support development and evaluation of data linkage methods

Lam, Joseph; Boyd, Andy; Linacre, Robin; Blackburn, Ruth; Harron, Katie; (2024) Generating synthetic identifiers to support development and evaluation of data linkage methods. International Journal of Population Data Science , 9 (1) 10.23889/ijpds.v9i1.2389. Green open access

[thumbnail of ijpds-09-2389.pdf]
Preview
PDF
ijpds-09-2389.pdf - Published Version

Download (1MB) | Preview

Abstract

IntroductionCareful development and evaluation of data linkage methods is limited by researcher access to personal identifiers. One solution is to generate synthetic identifiers, which do not pose equivalent privacy concerns, but can form a 'gold-standard' linkage algorithm training dataset. Such data could help inform choices about appropriate linkage strategies in different settings. ObjectivesWe aimed to develop and demonstrate a framework for generating synthetic identifier datasets to support development and evaluation of data linkage methods. We evaluated whether replicating associations between attributes and identifiers improved the utility of the synthetic data for assessing linkage error. MethodsWe determined the steps required to generate synthetic identifiers that replicate the properties of real-world data collection. We then generated synthetic versions of a large UK cohort study (the Avon Longitudinal Study of Parents and Children; ALSPAC), according to the quality and completeness of identifiers recorded over several waves of the cohort. We evaluated the utility of the synthetic identifier data in terms of assessing linkage quality (false matches and missed matches). ResultsComparing data from two collection points in ALSPAC, we found within-person disagreement in identifiers (differences in recording due to both natural change and non-valid entries) in 18% of surnames and 12% of forenames. Rates of disagreement varied by maternal age and ethnic group. Synthetic data provided accurate estimates of linkage quality metrics compared with the original data (within 0.13-0.55% for missed matches and 0.00-0.04% for false matches). Incorporating associations between identifier errors and maternal age/ethnicity improved synthetic data utility. ConclusionsWe show that replicating dependencies between attribute values (e.g. ethnicity), values of identifiers (e.g. name), identifier disagreements (e.g. missing values, errors or changes over time), and their patterns and distribution structure enables generation of realistic synthetic data that can be used for robust evaluation of linkage methods.

Type: Article
Title: Generating synthetic identifiers to support development and evaluation of data linkage methods
Open access status: An open access version is available from UCL Discovery
DOI: 10.23889/ijpds.v9i1.2389
Publisher version: http://dx.doi.org/10.23889/ijpds.v9i1.2389
Language: English
Additional information: © The Authors. Open Access under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/deed.en)
Keywords: record linkage, data linkage, synthetic data, synthetic identifiers, linkage evaluation, ALSPAC
UCL classification: UCL
UCL > Provost and Vice Provost Offices > School of Life and Medical Sciences
UCL > Provost and Vice Provost Offices > School of Life and Medical Sciences > Faculty of Population Health Sciences > UCL GOS Institute of Child Health
UCL > Provost and Vice Provost Offices > School of Life and Medical Sciences > Faculty of Population Health Sciences > UCL GOS Institute of Child Health > Population, Policy and Practice Dept
URI: https://discovery.ucl.ac.uk/id/eprint/10194216
Downloads since deposit
39Downloads
Download activity - last month
Download activity - last 12 months
Downloads by country - last 12 months

Archive Staff Only

View Item View Item