---
title: "molting(): Hash-Based De-identification"
subtitle: "Stripping identifiers while preserving relinkability"
author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service"
date: "`r Sys.Date()`"
output:
  html_document:
    toc: true
    toc_depth: 3
    toc_float: true
    theme: flatly
  pdf_document:
    toc: true
    toc_depth: 3
    number_sections: true
    latex_engine: xelatex
vignette: >
  %\VignetteIndexEntry{molting(): Hash-Based De-identification}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE)
library(mudnester)
```

## Why de-identify?

Sharing a linelist externally — with collaborators, for CRAN examples, in a research supplement — requires that directly identifying information is removed. `molting()` automates this using a cryptographic hash: each row gets a unique fingerprint derived from its identifiers, the identifiers are stripped, and the fingerprint is stored in a lookup table that only authorised personnel hold.

The name comes from the natural process of moulting: a bird sheds its distinctive, identifiable plumage and temporarily becomes more uniform. The old plumage is not destroyed — it re-grows from the same follicles. The lookup table is those follicles.

---

## What gets removed

`molting()` uses regular-expression pattern matching to detect PII columns. The default patterns cover:

- Names: `name`, `surname`, `firstname`, `lastname`
- Dates: `dob`, `birth`
- Identifiers: `mrn`, `urn`, `medicare`, `patient_id`, `subject_id`, `_id$`
- Contact: `address`, `street`, `phone`, `email`

Age **category** variables (`age2cat`, `age5cat`, `age10cat`, etc.) are automatically preserved — they are not directly identifying.

```{r detect-demo}
patient_data <- data.frame(
  patient_name = c("John Doe", "Jane Smith"),
  dob          = as.Date(c("1980-01-01", "1975-05-15")),
  mrn          = c("12345", "67890"),
  age5cat      = factor(c("18-64", "18-64")),   # preserved automatically
  diagnosis    = c("Condition A", "Condition B"),
  lab_value    = c(120, 95)
)

result <- suppressMessages(molting(patient_data))
names(result$deidentified)   # hash + retained columns
names(result$lookup)         # hash + removed columns
```

---

## The output list

`molting()` returns a named list with two elements when `return_lookup = TRUE` (the default):

- `$deidentified` — the de-identified data frame, with `row_hash` as the first column
- `$lookup` — the lookup table, with `row_hash` plus all removed identifier columns

```{r output-structure}
str(result, max.level = 1)
head(result$lookup)
```

Store `$lookup` securely and separately from `$deidentified`. Consider encrypting the lookup file before archiving. In a Queensland Health context, the lookup table should remain within the Health Service network.

---

## Hash algorithm selection

The default is SHA-256, which provides strong collision resistance for typical surveillance dataset sizes (tens of thousands of rows). For very large datasets where speed matters more than collision resistance, SHA-1 or MD5 are faster but should not be used where linkage integrity is critical.

```{r hash-algos, eval=FALSE}
# SHA-256 (default, recommended)
result_256 <- molting(patient_data, hash_method = "sha256")

# MD5 (shorter hash, faster, lower collision resistance)
result_md5 <- molting(patient_data, hash_method = "md5")

# Blake3 (fast and cryptographically strong — good for large datasets)
result_b3 <- molting(patient_data, hash_method = "blake3")
```

---

## Controlling which columns are hashed

By default, `molting()` hashes all detected PII columns. Supply `id_cols` to override this — useful when you want a shorter, more stable hash based only on a true unique identifier.

```{r id-cols}
# Hash only on MRN and DOB — more stable if name variations exist
result_ids <- suppressMessages(
  molting(patient_data, id_cols = c("mrn", "dob"))
)
result_ids$lookup
```

---

## Adding columns to the removal list

Use `additional_pii_cols` for dataset-specific identifiers that don't match the default patterns.

```{r additional}
patient_data2 <- patient_data
patient_data2$study_code <- c("SC-001","SC-002")

result2 <- suppressMessages(
  molting(patient_data2, additional_pii_cols = "study_code")
)
names(result2$deidentified)
```

---

## Irreversible de-identification

If you genuinely do not need to relink (e.g. producing a public-use file), set `return_lookup = FALSE`. This is irreversible — there is no way to recover the original identifiers.

```{r no-lookup}
deidentified_only <- suppressMessages(
  molting(patient_data, return_lookup = FALSE)
)
class(deidentified_only)   # a data frame, not a list
```

---

## Hash collisions

If two rows produce the same hash (extremely rare with SHA-256 for realistic dataset sizes but possible with very short hashes like CRC32), `molting()` warns you. If a collision is detected, switch to a stronger algorithm or add more columns to `id_cols`.

---

## What comes next

Use [homing()] to relink the de-identified data when authorised (see `vignette("homing")`).

For aggregated outputs — monthly counts from `roost()` — de-identification may not be necessary at all if the counts are not small enough to be re-identifying. The ABS cell-suppression threshold of 5 is a useful rule of thumb.
