Linkr
Home Resources Tools Documentation Blog Demo
FR
  • What is Linkr?
  • Deployment modes
  • Quick start
  • Local install
  • With Docker
  • Manual install
  • Client-only
  • Your first project
  • Linkr in a clinical data warehouse
  • Workspaces and projects
  • The data pipeline
  • Entities and sharing
  • Versioning and collaboration
  • Overview
  • Projects
  • Wiki
  • Plugins
  • Members and roles
  • Settings
  • Schemas
  • Getting and exploring
  • Mapping
  • Databases
  • Derived sub-databases
  • Data quality
  • Data catalog
  • Build and publish
  • Anonymize
  • SQL script collections
  • ETL pipelines
  • Building and running
  • Generating the scripts
  • Overview
  • Mapping projects
  • Global view
  • Target concepts
  • Mapping editor
  • Suggestions
  • AI agent
  • Evaluation
  • Export
  • Overview
  • Databases
  • Concepts
  • Cohorts
  • Building
  • Results, SQL and report
  • Patient data
  • Pipeline
  • Datasets
  • IDE
  • Web apps
  • Versioning
  • Overview
  • Tabs and widgets
  • Built-in widgets
  • Analysis widgets
  • Control charts (SPC)
  • Surveys and eCRF
  • R and Python code
  • Filters, settings and export
  • Overview
  • Presentation mode
  • Exporting a report
  • Agents
  • MCP server
  • Skills
  • Import and export
  • Git versioning
  • Community catalog
  • Publishing content
  • Production install
  • Configuration
  • Authentication and permissions
  • Files on the server
  • Backup and restore
  • Contributing code
  • Glossary
  • Keyboard shortcuts
  • Release notes
Documentation Data warehouse Building and running

Building and running a pipeline

Building an ETL pipeline tab by tab: the tabs, writing the SQL scripts, the step-by-step run (pause, stop, errors), then the source / target quality check, statistics and concepts.

In short

An ETL pipeline is built in its tabs: you write or generate the SQL scripts, pick the source and target databases, run the chain in the Pipeline tab, then the Quality check tab compares both databases to tell whether the conversion lost anything.

Client Available in client-only mode — runs entirely in the browser, no backend. Backend Available with the FastAPI backend.

A pipeline is created from the warehouse’s ETL pipelines page. Its scripts refer to the databases by their role — source., target., vocab. — as the overview page explains; this page follows the rest, tab by tab.

The tabs

TabWhat you do there
PipelineThe run board: the scripts as cards, in the order they chain. This is where you pick the source and target databases, reorder, enable or disable a step, and launch.
ScriptsThe editor proper, with the query results — and the Generate from the schemas button.
Browse schemasSource and target tables side by side — essential while writing the correspondence.
VocabularyGenerate the vocabulary scripts from a concept mapping project.
Quality checkVerify what the run produced: both databases’ figures, and a concept-by-concept account of what was mapped.

Overview, Readme, License and Versioning sit alongside them, as for every other entity.

Only SQL files are steps

A pipeline’s folder can hold other files — a Markdown README, notes, helper code in Python or R. They travel with the pipeline, but the Pipeline tab only shows the .sql files, and a run chains those alone. A Markdown file opened in the Scripts tab is previewed rather than run.

Writing the scripts

An ETL script is ordinary SQL, which you can write entirely by hand. But most of a conversion to OMOP follows from what Linkr already knows, and two generators write it for you:

  • Generate from the schemas, in the Scripts tab — the scripts that load patients, stays and events, worked out from the source’s schema mapping and the target’s.
  • The Vocabulary tab — the scripts that load the translation of your local codes, worked out from a concept mapping project.

What they produce are ordinary scripts, which you review and edit; Linkr then spots the ones you touched before regenerating them. It is all in Generating the scripts.

Running, and checking

The Pipeline tab shows the chain as it will run: the source database at the top, the SQL scripts in their order, the target database at the bottom. There you pick the two databases, drag a card to move it, turn a step off with its switch, and Run pipeline launches the lot. Each card also has its own button, to run one script alone while working things out.

demo.linkr.interhop.org
7/8 scriptsICU hospital — exportICU hospital — OMOP 5.4
1
00_vocabulary.sql
2
10_person.sql
3
20_visit_occurrence.sql
4
30_visit_detail.sql
5
40_note.sqldisabled
6
50_measurement.sql
7
51_drug_exposure.sql
8
99_prune_vocabulary.sql

Click a node to view details

The Pipeline tab: the source (orange), eight scripts of which one is disabled, the target (green). Click ▶: the cards take their status one by one, the disabled script is skipped, and the run stops on 50_measurement.sql, whose detail opens on the right with the error. Try pausing and stopping along the way too.

Progress shows in the toolbar, query by query, with the name of the current script and the elapsed time — on a vocabulary script where one statement can take minutes, that is what distinguishes work in progress from a hang. A failing run stops at the first script in error, and Linkr opens the detail panel directly on the offending script.

Scripts run in the order of the cards, not of their names. When the two drift apart — a 35_… script dragged after a 50_… —, the toolbar’s alphabetical sort button realigns them; it is greyed out when the order already follows the names.

Pausing and stopping are not the same

Pause holds the run without ending it: resuming re-runs the interrupted script from its start, then carries on, within the same run.

Stop ends it. In both cases the statement already sent to the database runs to completion — so part of a script may have been applied. With a half-filled target database, recreating it empty beats re-running on top.

Then comes the Quality check tab, which answers the only question that matters: did the conversion lose anything? It offers two views.

demo.linkr.interhop.org
Filter:Computed 9/28/2026, 10:42 AM
StatusSource vocabularySource codeDescriptionSource patientsSource rowsTarget IDExpected rowsTarget rows
MissingREA_LOCALBIO_PCTProcalcitonin8122,140448171302,1400
FewerREA_LOCALBIO_KSerum potassium3,02148,902302310348,90246,115
MoreREA_LOCALBIO_NASerum sodium3,01948,877301955048,87749,012
OKREA_LOCALVS_FCHeart rate3,102912,4403027018912,440912,440
OKREA_LOCALVS_SPO2SpO23,098887,31040762499887,310887,310
OKREA_LOCALVS_TEMPTemperature3,087201,5443020891201,544201,544
The Quality check tab, Concepts view: each source code gets a verdict, and the chips in the bar act as filters — click “Missing” to keep only the concepts that arrived with no mapping. The Statistics view puts both databases side by side.

Statistics puts both databases side by side: patients, hospitalizations, unit stays, gender breakdown, lengths of stay, and each table’s row count. An unexpected gap — three thousand patients on one side, two thousand eight hundred on the other — points at a join that silently dropped rows.

Concepts is finer, and works differently from what you might assume: both columns are read from the target database. A source database in its original shape has no comparable columns; once converted to OMOP, however, each table carries both the original concept and the standard one it was mapped to. Comparing the two says exactly how many rows arrived with a source code, and how many actually found a match.

Each concept gets a verdict — Missing, Fewer, More or OK — usable as a filter, and the whole thing exports to CSV. “Missing” is the one to look at first: rows arrived, none were mapped.

The 'Expected rows' column is not redundant

When several source codes point at the same target concept, one code’s row count does not match the expected total. The column therefore shows the sum of every code feeding that concept — otherwise an “OK” verdict next to two different numbers would look like a bug.

Patient counts need a schema on both sides

With no schema mapping on a database, Linkr does not know what a patient is: it can only count rows per table. See Schemas.

Going further

  • Generating the scripts — the OMOP load and vocabulary scripts, written by Linkr.
  • ETL pipelines — the source, target and vocabulary roles, and sharing a pipeline.
  • Databases — creating the empty target database from a schema.
  • Schemas — required for the quality check’s patient counts.
  • Data quality — checking the target database once filled.
PreviousETL pipelinesNextGenerating the scripts

Product

  • Home
  • Demo

Resources

  • Documentation
  • Resources
  • Tools
  • Blog

Community

  • Framagit source code
  • Github source code

About

  • InterHop.org
  • Contact

2021–2026 InterHop — CC BY-NC-SA 4.0 (site) · GPLv3 (software)