Re-Identification Risk
Re-identification risk is the likelihood that data which has been anonymized or de-identified can be linked back to the specific individuals it describes. This typically happens when the de-identified dataset is combined with other available information, allowing a third party to figure out who the data subjects are. Even data that appears anonymous can carry some level of this risk.
Re-identification risk is the likelihood that a third party can re-identify data subjects within a de-identified or anonymized dataset, commonly by linking residual data elements to external information sources. As commonly defined, assessment approaches examine the data properties and record attributes that overlap with externally available data to estimate the probability that individual subjects can be re-identified. Note that the term describes a residual privacy risk associated with de-identification rather than a guarantee of anonymity; its treatment and acceptable thresholds vary by sector (for example, healthcare and research contexts) and by applicable jurisdictional requirements, which are out of scope for this core definition.
Why it matters
Re-identification risk matters because de-identification and anonymization are frequently treated as endpoints that make data safe to share, reuse, or release, when in practice they leave a residual privacy risk. As commonly defined, this risk arises when de-identified records retain data elements that overlap with externally available information, allowing a third party to link records back to specific individuals. Organizations that assume anonymization is absolute may under-protect data that can still be traced to the people it describes, exposing data subjects to privacy harm and the organization to regulatory and reputational consequences.
The stakes are particularly pronounced in sectors such as healthcare and research, where sensitive information about individuals is routinely de-identified for secondary use. In these contexts, acceptable re-identification risk thresholds and the treatment of residual risk vary by applicable jurisdictional requirements and sector-specific expectations. Because these thresholds are not universal, teams cannot rely on a single standard of anonymity across all uses; what is considered adequately de-identified in one setting or jurisdiction may not satisfy another.
Understanding re-identification risk as a measurable, residual property—rather than a binary state of anonymous versus identifiable—helps organizations make defensible decisions about data sharing, publication, and reuse. It reframes de-identification as a risk-reduction measure that lowers, but does not necessarily eliminate, the possibility of subjects being re-identified.
Who it's relevant to
Inside Re-Identification Risk
Common questions
Answers to the questions practitioners most commonly ask about Re-Identification Risk.