Summary
Linkr is not a warehouse management tool: it does not load the warehouse, manage access accounts or create extractions. It is installed inside a project’s secure environment (SPE), next to RStudio or Jupyter, and takes the project’s datamart from the silver level (cleaned data) to the gold level (analysis-ready data). A global instance, outside any project, can also centralise the institution’s mappings and scripts.
The layers of a clinical data warehouse
A clinical data warehouse (CDW) is not a single block. Between the care software and the result of a study, data crosses several layers, each run by a different team. The article Clinical data warehouses covers the hospital information system and ETL; this page focuses on what comes next, and on where Linkr fits.
1 · Hospital information system
CareElectronic health record, laboratory, pharmacy, billing… each application with its own database.
2 · Loading and warehouse — bronze, silver
CDW engineeringETL copies data from the hospital systems, pseudonymises and cleans it, into a data warehouse or a data lake.
3 · Project datamart
CDW platformFor an approved study, a subset of the warehouse: the patients, variables and period it needs.
4 · Project secure environment (SPE) — silver → gold
Linkr, RStudio, Jupyter…The research team works on its datamart: quality, standardisation, cohorts, variables, analyses. This is where Linkr sits.
The medallion architecture: bronze, silver, gold
Many recent warehouses, especially those built on a data lake, organise their data into three levels of refinement — the medallion architecture. The vocabulary comes from data engineering, but the logic applies to a classic warehouse too.
1 · Bronze
Raw copy
Data from the care software, as is, pseudonymised.
Warehouse · CDW team
2 · Silver
Cleaned and harmonised
Duplicates and aberrant values removed, sources brought together — sometimes already in OMOP format.
Warehouse · CDW team
3 · Gold
Reliable for one project
A study’s datamart, checked and standardised for its research question.
SPE · project team
- Bronze — the raw copy of the sources, as it arrives from the hospital systems, with no transformation other than pseudonymisation, usually required before anything enters the warehouse. It is kept so that everything can be replayed if a transformation rule changes.
- Silver — a first quality pass, for the whole warehouse: types fixed, duplicates removed, obviously aberrant values discarded (a birth date in 2050), tables from the different sources brought into one model. Some institutions convert the whole warehouse to the OMOP model at this level. It is reliable but generic data: the rules applied hold for every study, without being designed for any one in particular.
- Gold — data quality-checked for one specific project. It stays in the warehouse format — event tables, in long format —, but it has been checked and standardised with a research question in mind: a parameter’s plausibility bounds, the local codes to map, the units to harmonise depend on what the study will measure.
Bronze and silver are the CDW team’s work, for the whole institution. Gold is built project by project, in the SPE: that is where, close to the research question, the real quality work happens.
What about wide format?
Moving to a table with one row per patient and the study’s variables — going from long to wide format — comes after gold: it is variable preparation, described in The data pipeline. Some presentations of the medallion architecture put these analysis tables in gold; here, gold keeps the warehouse format.
The project datamart
A study never accesses the whole warehouse. Once the project is approved — protocol, scientific and ethics committee opinion, regulatory formalities —, the CDW platform extracts a datamart from it (also called an extraction or project dataset): only the patients, variables and periods needed.
Extracting a project’s regulatory datamart, managing access accounts to the warehouse: all of this belongs to a CDW management platform, not to Linkr. In server mode, however, Linkr can derive a sub-database from a cohort, restricted to its patients, from a database it already reads — useful to narrow a scope inside the SPE. See Derived sub-databases.
The project secure environment
The datamart is made available in a Secure Processing Environment, or SPE. It is an isolated space dedicated to the project:
- Cut off from the internet and from the rest of the information system.
- Reached through two-factor authentication, by the people authorised for the project only.
- With no direct way out: a result only leaves the SPE after approval, through a controlled export process.
The SPE provides analysis tools: RStudio, Jupyter, and Linkr. Linkr is therefore one tool among others in the SPE — it works on the same data as the team’s R or Python scripts, and does not replace them.
Linkr inherits the SPE's security
Two-factor authentication, network isolation, encryption and export control are provided by the SPE, not by Linkr. Linkr’s roles and permissions organise a team’s work — who edits what in a project —, they do not partition regulated data. Deploy one Linkr instance per SPE, rather than one instance shared between several research projects.
What Linkr does in the SPE: from silver to gold
In the SPE, Linkr connects to the project datamart. Its first job is to turn it into a gold datamart:
- Check quality — data quality rule sets check the datamart before any analysis: out-of-range values, inconsistent dates, missing data.
- Standardise — if the datamart is delivered in the institution’s local format, an ETL pipeline converts it to the OMOP model, so the community’s tools and analyses can be reused.
- Map terminologies — concept mapping links local codes (lab tests, drugs, monitoring parameters) to standard vocabularies.
- Reuse queries — SQL script collections gather the institution’s proven queries.
Then, from gold, it takes the team all the way to analysis:
- Build cohorts — select the study’s patients by criteria, in the datamart. See Cohorts.
- Prepare variables and analyse — go from long to wide format, then analyse and present in dashboards. See The data pipeline.
No identifier change for a cohort
A cohort built in Linkr, from an SPE’s datamart, keeps the datamart’s identifiers: the work already happens within one specific project, in its SPE.
A global instance for the institution
Alongside the instances installed in each SPE, an institution can run a global Linkr instance, outside any research project. It is not used to analyse patient data: it centralises the know-how built up on the warehouse.
- Concept mappings — the mapping of local codes to standard vocabularies, built once and reused by every project.
- SQL script collections and data quality rule sets — the queries and checks validated for the institution’s warehouse.
- ETL pipelines — the conversion from the local format to OMOP.
- The wiki — knowledge about the warehouse: what each table contains, the quirks of local coding, known pitfalls, the institution’s conventions. Centralised, it saves every project from rediscovering the same subtleties, and it grows with everyone’s experience.
These move from the global instance to the SPE instances through import and export or git versioning, depending on what the institution’s process allows. In the other direction, an improvement made during a project — a completed mapping, an added quality check — can flow back to the global instance, provided it contains no patient data and goes through the SPE’s export process.
The institution's global instance
No patient data: the know-how about the warehouse.
SPE — study A
Linkr, RStudio, Jupyter on the datamart
SPE — study B
Linkr, RStudio, Jupyter on the datamart
SPE — study C
Linkr, RStudio, Jupyter on the datamart
What Linkr does not do
To remove any ambiguity, here is which part of the infrastructure handles each need:
| Need | Infrastructure component |
|---|---|
| Load the warehouse from the hospital systems (bronze, silver) | CDW loading pipeline (ETL) |
| Pseudonymise the warehouse data | CDW loading pipeline (ETL) |
| Manage accounts and access authorisations to the CDW | CDW management platform |
| Extract a project’s regulatory datamart | CDW management platform |
| Two-factor authentication, network isolation, encryption | SPE |
| Approve and control exports out of the SPE | SPE export process |
| Quality, standardisation, mapping, cohorts, variables, analyses on the datamart | Linkr, with the SPE’s other tools |
Going further
- Clinical data warehouses — the hospital information system, ETL and common data models, explained from the start.
- Deployment modes — choosing between client-only and full-stack for an install in an SPE.
- Workspaces and projects — how Linkr organises work once installed.
- Concept mapping — the work the global instance lets you pool.