Summary
A data catalog is built in the order of its tabs: the database described, the variables and crossings to count, the computation, anonymization, the Health-DCAT-AP metadata, then publication — an HTML page to download, a full archive, or a site deployed from git.
A catalog is created from the warehouse’s Data catalog page, with New catalog. It is then set up tab by tab: Overview, Configuration, Anonymization, Health-DCAT-AP and Publish.
1. Choose the database
A catalog describes one database, chosen when it is created and changeable in the Overview tab.
The database needs a schema
Without a schema mapping, Linkr does not know what a patient or a stay is: there is nothing to count. See Schemas.
2. Choose the variables and counts
In the Configuration tab, the Variables card sets what the catalog can be broken down by, and what each cell counts. Each variable is turned on with its switch; once on, it has its own totals and unfolds its options.
Counts
The Counts row sets what each cell counts. Patients are always counted — distinct patients: one seen in two periods counts once in each. You can add Hospitalizations and Unit stays; with a concept, a hospitalization counts when a record of the concept falls during that stay. Records — the event rows — are counted whenever the concept variable is part of a crossing.
Each added count takes time and memory on a large warehouse: only tick what the catalog’s readers will read.
Variables
Period
The start of the hospitalization or, crossed with concepts, the date of the event.
Granularity by month, quarter or year, and Every for wider steps — every 5 years, for instance.
Age group
The age at the hospitalization or at the event, from the birth date or failing that the birth year.
Preset Brackets — 5, 10 or 20 years, pediatric, clinical — or your own bounds: 18 and 65 give under 18, 18–64 and 65 and over.
Gender
From the gender column of the patients table, as mapped in the schema. No options.
Service
The Service level: the Care unit (unit stays) of each unit stay, or the Visit type.
Grouping keeps each unit as it is (All), keeps only the N largest, the rest becoming “Other services” (Top N), or gathers units into named groups (Groups) — the surest way to keep enough patients per cell.
Concept
Each clinical concept found in the event tables — measurements, drugs, diagnoses… The Category column and Subcategory column classify the concept list (in OMOP, typically the domain and the class): that is what makes the inventory browsable rather than a list of ten thousand codes.
Count by groups by Category or Subcategory instead of concept by concept, and Concepts crossed only crosses the N concepts with the most patients.
3. Choose the crossings
A crossing is a table that counts several variables together — concept × period, service × age group… The Crossings card offers them in three groups: One variable, Two variables — a grid where each box crosses the variable of its row with that of its column — and Three variables.
| Age group | Gender | Service | Concept | |
|---|---|---|---|---|
| Period | ||||
| Age group | ||||
| Gender | ||||
| Service |
The more variables a crossing has, the smaller its cells, and the more of them fall below the anonymization threshold. Estimate yields computes, for each crossing, the share of non-empty cells that would be published — green above 90%, amber between 60 and 90%, red below. A crossing where almost everything would be masked tells nobody anything, and is better left unticked.
A variable unticked under One variable is still counted on its own to order its values, but is not published on its own — gender, for instance, may only be of interest crossed with age.
4. Compute
Compute catalog goes through the database and counts, step by step — concepts, totals, ranking of units and concepts, then each crossing. On a large warehouse it takes a while: the computation keeps going if you switch tabs, and it is interruptible — Pause, Resume, or Start over from scratch. The configuration is locked while a computation is running or paused.
The result is kept with its date and duration. Once the catalog is computed, Recompute runs it again with the current configuration, and Discard deletes the results while keeping the configuration.
Changing the database invalidates the computation
The counts describe one specific database. If you switch to another, the existing results no longer mean anything: the computation must be run again.
5. Anonymize
This is the step that makes publication acceptable. In the Anonymization tab, you set a minimum number of patients per cell: below it, the cell is masked. Linkr also masks the cells that would let a masked cell be recovered by subtraction, can slightly perturb the published counts, and an audit checks that nothing can be worked out from the published files.
All of this is detailed in Anonymizing a catalog.
6. Describe the dataset
The Health-DCAT-AP tab gathers the catalog’s metadata: publisher, contacts, licence, health content, time coverage, access. Each section’s counter shows how many mandatory fields remain; Auto-fill fills the empty fields from what Linkr already knows (name, description, database…), and Mandatory only hides the rest. The Generated section is built from the database and the computed results — there is nothing to fill in, except where the published files will live.
Health-DCAT-AP, briefly
It is the European vocabulary for describing health datasets, set out by the European Health Data Space regulation. Using it means a data portal can harvest your catalog and display it alongside others, with no conversion.
Publishing
The Publish tab first shows a Preview of the page as it will be published. Show masked values displays there, for you alone, the count each masked cell hides — the downloaded and deployed files never contain those numbers.
The Export part offers two formats, in the chosen Page language:
Catalog page — HTML
A single file that opens in any browser, offline.
Its data is inside, nothing else to host. The Health-DCAT-AP metadata is embedded as JSON-LD, readable by a search engine or a data portal.
Full publication — ZIP
Everything a public site needs.
catalog.html, concepts.csv (one row per concept), one CSV per crossing under crossings/ — masked cells left empty with their status — and metadata.jsonld. The form to choose to feed an institutional catalog.
Deploy with GitLab / GitHub Pages
Instead of downloading, you can publish the page automatically from the catalog’s git repository. First link the catalog to a repository in the Versioning tab (see Git versioning), choose the Hosting, then Enable deployment: Linkr adds the site files to the repository, with a CI file that deploys them on every push. The site’s expected address can serve as the published catalog’s URL, in its metadata.
After a new computation, Update published site, then push from the Versioning tab.
Going further
- Anonymizing a catalog — threshold, secondary suppression, perturbation and audit, before any publication.
- Data catalog — what a catalog is for, and what a visitor of the published page reads.
- Schemas — required for the computation to know what to count.
- Git versioning — linking the catalog to a repository, to deploy it.