Linkr
Home Resources Tools Documentation Blog Demo
FR
  • What is Linkr?
  • Deployment modes
  • Quick start
  • Local install
  • With Docker
  • Manual install
  • Client-only
  • Your first project
  • Linkr in a clinical data warehouse
  • Workspaces and projects
  • The data pipeline
  • Entities and sharing
  • Versioning and collaboration
  • Overview
  • Projects
  • Wiki
  • Plugins
  • Members and roles
  • Settings
  • Schemas
  • Getting and exploring
  • Mapping
  • Databases
  • Derived sub-databases
  • Data quality
  • Data catalog
  • Build and publish
  • Anonymize
  • SQL script collections
  • ETL pipelines
  • Building and running
  • Generating the scripts
  • Overview
  • Mapping projects
  • Global view
  • Target concepts
  • Mapping editor
  • Suggestions
  • AI agent
  • Evaluation
  • Export
  • Overview
  • Databases
  • Concepts
  • Cohorts
  • Building
  • Results, SQL and report
  • Patient data
  • Pipeline
  • Datasets
  • IDE
  • Web apps
  • Versioning
  • Overview
  • Tabs and widgets
  • Built-in widgets
  • Analysis widgets
  • Control charts (SPC)
  • Surveys and eCRF
  • R and Python code
  • Filters, settings and export
  • Overview
  • Presentation mode
  • Exporting a report
  • Agents
  • MCP server
  • Skills
  • Import and export
  • Git versioning
  • Community catalog
  • Publishing content
  • Production install
  • Configuration
  • Authentication and permissions
  • Files on the server
  • Backup and restore
  • Contributing code
  • Glossary
  • Keyboard shortcuts
  • Release notes
Documentation Data warehouse Anonymize

Anonymizing a catalog

Anonymization threshold, secondary suppression, count perturbation and disclosure audit: what protects a data catalog before it is published.

Summary

Before a data catalog is published, the Anonymization tab protects it in three layers: a threshold masks the cells counting too few patients, secondary suppression keeps them from being recovered by subtraction, and an optional perturbation makes every count approximate. A disclosure audit then checks, from the published files alone, that no masked cell can be worked out.

Client Available in client-only mode — runs entirely in the browser, no backend. Backend Available with the FastAPI backend.

Why a threshold is not enough

A cell reading “2 patients, neonatal unit, March 2024” is not a statistic, it is an identification: someone who knows the unit will find those two children. The threshold removes it.

But a catalog also publishes totals — all units together, all periods together — and tables that overlap. If the March 2024 total is published, and every unit but the neonatal one is too, the masked cell comes back by one subtraction. The more crossings a catalog has, the more of these overlaps there are. The sections below describe how Linkr closes them.

The threshold

demo.linkr.interhop.org
Anonymization threshold
Threshold
10
Mode
Show “< threshold”
Perturbation (±)
3
Run
1,284
Cells below the threshold
7 % of cells
611
Cells masked to protect them
3 % of cells
2,140
Concepts shown as “< threshold”
38 % of concepts
90%
Cells published
The Anonymization tab: the threshold, the mode and the perturbation, then, once run, what the settings do — the masked cells, the concepts affected and the share of cells still published.

Threshold sets the minimum number of patients per cell — 10 is a common value. Any cell below it is masked: it carries no number in any published file.

The Mode handles the concept list:

  • Show ”< threshold” — concepts used by fewer patients than the threshold stay listed, with ”< 10” in place of their count. The reader knows the concept exists in the warehouse.
  • Leave the concept out — they disappear from the list.

In the crossings, cells below the threshold are always masked, whatever the mode.

Secondary suppression

So that a masked cell cannot be recovered by subtraction, Linkr masks other cells of the same group — this is called secondary suppression. On the published page, every masked cell is hatched: those below the threshold show ”< 10”, those masked to protect them show “masked” — they have enough patients, but are hidden to protect a small neighbouring cell.

View
demo.linkr.interhop.org
Patients by service × period
Service2021202220232024
Intensive care412438451397
Paediatrics9688masked102
Neonatology22masked< 1019
Cardiology205231219244
Neonatology has 4 patients in 2023: the cell is masked, and two neighbours are too, in its row and its column, so that no subtraction recovers it. The selector above plays the Publish tab's preview, which shows the hidden count struck through — the published files never do.

Two rules add to the classic “never a single masked cell in a row”:

  • The sum counts too. Two small cells masked together do not protect each other if their total can be computed: 3 patients hidden in two cells are still 3 patients. So the masked cells of a group must weigh, together, at least the threshold — otherwise one more cell is masked.
  • A masked total is not a hidden total. A total masked in one table can be recovered from another (the grand total, for instance). So Linkr also protects the groups whose total is masked.

Perturbation

Even well masked, overlapping tables sometimes let a cell be guessed to within a few units: the attacker does not know the count, but knows it lies between 1 and 3. Perturbation (±) closes that gap: every published count moves by at most that value. The setting is 0 by default, which publishes exact counts. ±3 is a good value.

The noise is not drawn at random on each publication. It is set by the cell’s patients — the cell key method, the one Eurostat recommends for censuses:

  • The same cell always gets the same noise. Republishing the catalog, or publishing another one on the same database, cannot average the noise out.
  • A subtraction only gives an order of magnitude. Each term carries its own noise; the difference of two noisy counts is only known to within a few units.
  • The threshold applies to the true counts. A cell is masked before any noise, and a published cell never shows less than the threshold.

The published page explains it in its Info tab, and a key figure that adds up several perturbed cells is preceded there by ”≈”.

Approximate counts are enough for a catalog

A catalog tells a team whether your warehouse is worth approaching: about 1,200 patients on mechanical ventilation in 2023 is the answer they are after, not 1,203. Exact counts on a precise question are a matter for a feasibility study, with cohorts.

Recompute catalogs computed before perturbation

The noise is derived from each cell’s patients, which the catalog computation records. A catalog computed with an earlier version of Linkr lacks that data: its noise is drawn from the cell itself. It stays stable from one publication to the next, but two publications made before and after a warehouse update could cancel it out. Recompute the catalog from the Configuration tab before publishing with a perturbation.

Seeing the impact of the settings

After changing the threshold, the mode or the perturbation, Run applies the settings and shows their effect, under the settings:

  • Cells below the threshold and Cells masked to protect them — what each masking removes.
  • Concepts shown as ”< threshold” (or Lost concepts, in Leave the concept out mode).
  • Cells published and Patients in published cells — what stays visible.
  • Impact per crossing — the same figures, table by table.

A crossing whose cells are mostly masked informs nobody: better widen its granularity — the quarter rather than the month, groups of units rather than each unit — or untick it in the Configuration tab.

The disclosure audit

The Disclosure audit, at the bottom of the tab, puts itself in the place of an outsider holding only the published files. Each published total gives an equation — the total is the sum of its cells — and the tables are linked through the totals they share. The audit solves these systems, taking the perturbation into account, and lists the masked cells it manages to pin down:

  • their patients narrowed down between 1 and the threshold minus one — the disclosure the threshold was meant to prevent;
  • their hospitalizations, unit stays or records recovered to within 2.
demo.linkr.interhop.org
Disclosure auditRun again
2 masked cells have their patients narrowed down to between 1 and 9.1 masked count of stays or records is recovered within 2.Perturbation closes these: at ±3, a subtraction gives only an order of magnitude. Leaving the crossings in question out, or a higher threshold, also does.
CrossingCellCountWorked out
Concept × PeriodFurosemide · 2022Patients2 – 7
Concept × PeriodFurosemide · 2022Records11
Service × Age groupPaediatrics · 10–19Patients1 – 5

Audited in 47 s: 2,316 systems, 58,904 unknowns, noise ±0.

The audit of a catalog published without perturbation: two masked cells have their patients pinned down, and one its records. For each, the crossing, the cell, the count at stake and what the audit worked out.

Run the audit starts it; like the computation, it keeps going if you switch tabs, and Stop interrupts it. The result is kept with the catalog’s results: a green banner when no cell is pinned down, red otherwise, with a table of examples — the crossing, the cell, the count at stake and what the audit worked out.

When the audit finds exposed cells, three levers close them: the perturbation (at ±3, a subtraction only gives an order of magnitude), removing the crossings involved, or raising the threshold.

Run the audit again after every change

An audit holds for the settings and results it checked. If you change an anonymization setting or recompute the catalog, it is flagged as out of date: run it again before publishing.

What the audit guarantees, and what it does not

The audit is sound: what it works out, an attacker can too. It is not exhaustive for all that — a tool solving all the tables in one go could sometimes pin down more. An empty audit, with a perturbation of ±3, is good assurance; it is not a mathematical guarantee.

Going further

  • Building and publishing a catalog — variables, crossings, computation and publication.
  • Cohorts — for exact counts on a precise question.
  • Publishing content — sharing beyond your institution.
PreviousBuild and publishNextSQL script collections

Product

  • Home
  • Demo

Resources

  • Documentation
  • Resources
  • Tools
  • Blog

Community

  • Framagit source code
  • Github source code

About

  • InterHop.org
  • Contact

2021–2026 InterHop — CC BY-NC-SA 4.0 (site) · GPLv3 (software)