Summary
Before a data catalog is published, the Anonymization tab protects it in three layers: a threshold masks the cells counting too few patients, secondary suppression keeps them from being recovered by subtraction, and an optional perturbation makes every count approximate. A disclosure audit then checks, from the published files alone, that no masked cell can be worked out.
Why a threshold is not enough
A cell reading “2 patients, neonatal unit, March 2024” is not a statistic, it is an identification: someone who knows the unit will find those two children. The threshold removes it.
But a catalog also publishes totals — all units together, all periods together — and tables that overlap. If the March 2024 total is published, and every unit but the neonatal one is too, the masked cell comes back by one subtraction. The more crossings a catalog has, the more of these overlaps there are. The sections below describe how Linkr closes them.
The threshold
Threshold sets the minimum number of patients per cell — 10 is a common value. Any cell below it is masked: it carries no number in any published file.
The Mode handles the concept list:
- Show ”< threshold” — concepts used by fewer patients than the threshold stay listed, with ”< 10” in place of their count. The reader knows the concept exists in the warehouse.
- Leave the concept out — they disappear from the list.
In the crossings, cells below the threshold are always masked, whatever the mode.
Secondary suppression
So that a masked cell cannot be recovered by subtraction, Linkr masks other cells of the same group — this is called secondary suppression. On the published page, every masked cell is hatched: those below the threshold show ”< 10”, those masked to protect them show “masked” — they have enough patients, but are hidden to protect a small neighbouring cell.
| Service | 2021 | 2022 | 2023 | 2024 |
|---|---|---|---|---|
| Intensive care | 412 | 438 | 451 | 397 |
| Paediatrics | 96 | 88 | masked | 102 |
| Neonatology | 22 | masked | < 10 | 19 |
| Cardiology | 205 | 231 | 219 | 244 |
Two rules add to the classic “never a single masked cell in a row”:
- The sum counts too. Two small cells masked together do not protect each other if their total can be computed: 3 patients hidden in two cells are still 3 patients. So the masked cells of a group must weigh, together, at least the threshold — otherwise one more cell is masked.
- A masked total is not a hidden total. A total masked in one table can be recovered from another (the grand total, for instance). So Linkr also protects the groups whose total is masked.
Perturbation
Even well masked, overlapping tables sometimes let a cell be guessed to within a few units: the attacker does not know the count, but knows it lies between 1 and 3. Perturbation (±) closes that gap: every published count moves by at most that value. The setting is 0 by default, which publishes exact counts. ±3 is a good value.
The noise is not drawn at random on each publication. It is set by the cell’s patients — the cell key method, the one Eurostat recommends for censuses:
- The same cell always gets the same noise. Republishing the catalog, or publishing another one on the same database, cannot average the noise out.
- A subtraction only gives an order of magnitude. Each term carries its own noise; the difference of two noisy counts is only known to within a few units.
- The threshold applies to the true counts. A cell is masked before any noise, and a published cell never shows less than the threshold.
The published page explains it in its Info tab, and a key figure that adds up several perturbed cells is preceded there by ”≈”.
Approximate counts are enough for a catalog
A catalog tells a team whether your warehouse is worth approaching: about 1,200 patients on mechanical ventilation in 2023 is the answer they are after, not 1,203. Exact counts on a precise question are a matter for a feasibility study, with cohorts.
Recompute catalogs computed before perturbation
The noise is derived from each cell’s patients, which the catalog computation records. A catalog computed with an earlier version of Linkr lacks that data: its noise is drawn from the cell itself. It stays stable from one publication to the next, but two publications made before and after a warehouse update could cancel it out. Recompute the catalog from the Configuration tab before publishing with a perturbation.
Seeing the impact of the settings
After changing the threshold, the mode or the perturbation, Run applies the settings and shows their effect, under the settings:
- Cells below the threshold and Cells masked to protect them — what each masking removes.
- Concepts shown as ”< threshold” (or Lost concepts, in Leave the concept out mode).
- Cells published and Patients in published cells — what stays visible.
- Impact per crossing — the same figures, table by table.
A crossing whose cells are mostly masked informs nobody: better widen its granularity — the quarter rather than the month, groups of units rather than each unit — or untick it in the Configuration tab.
The disclosure audit
The Disclosure audit, at the bottom of the tab, puts itself in the place of an outsider holding only the published files. Each published total gives an equation — the total is the sum of its cells — and the tables are linked through the totals they share. The audit solves these systems, taking the perturbation into account, and lists the masked cells it manages to pin down:
- their patients narrowed down between 1 and the threshold minus one — the disclosure the threshold was meant to prevent;
- their hospitalizations, unit stays or records recovered to within 2.
| Crossing | Cell | Count | Worked out |
|---|---|---|---|
| Concept × Period | Furosemide · 2022 | Patients | 2 – 7 |
| Concept × Period | Furosemide · 2022 | Records | 11 |
| Service × Age group | Paediatrics · 10–19 | Patients | 1 – 5 |
Audited in 47 s: 2,316 systems, 58,904 unknowns, noise ±0.
Run the audit starts it; like the computation, it keeps going if you switch tabs, and Stop interrupts it. The result is kept with the catalog’s results: a green banner when no cell is pinned down, red otherwise, with a table of examples — the crossing, the cell, the count at stake and what the audit worked out.
When the audit finds exposed cells, three levers close them: the perturbation (at ±3, a subtraction only gives an order of magnitude), removing the crossings involved, or raising the threshold.
Run the audit again after every change
An audit holds for the settings and results it checked. If you change an anonymization setting or recompute the catalog, it is flagged as out of date: run it again before publishing.
What the audit guarantees, and what it does not
The audit is sound: what it works out, an attacker can too. It is not exhaustive for all that — a tool solving all the tables in one go could sometimes pin down more. An empty audit, with a perturbation of ±3, is good assurance; it is not a mathematical guarantee.
Going further
- Building and publishing a catalog — variables, crossings, computation and publication.
- Cohorts — for exact counts on a precise question.
- Publishing content — sharing beyond your institution.