Siebers 2026 — The Dutch Nationwide Pathology Databank (Palga)
A description of Palga, the Dutch nationwide pathology databank: 92 million records from 15 million people, built on a coded diagnosis line the pathologist writes at sign-out and a validator that refuses to let the report be authorised until it is well-formed.
Read status: full text, pre-proof. Received 5 March 2026, revised 9 July, accepted 19 July 2026. Figures 1–4 are referenced but are supplied as separate image files, so the data-flow diagrams themselves were not read — only their captions. Table 1 (an updated-vs-new matrix of features since Casparie 2007) extracted as an empty grid: the tick marks are glyphs that did not survive text extraction, so which features are “new” versus “updated” could not be read.
[unverified]
What it is
Not a model paper, so AGENTS.md §8’s four fields do not apply — there is no cohort, scanner, or validation strategy to extract. It is an infrastructure description, updating the widely cited Casparie 2007 account of the same system.
Palga began in 1971, reached all Dutch pathology departments in 1991, and now connects all 39 pathology departments in a country of 18 million people. As of mid-2026 it holds more than 92 million records on more than 15 million persons and has supported more than 750 publications. Those counts are cited to Palga’s own annual reports rather than to an independent audit.
The architecture splits cleanly in two, and the split is the most transferable idea in the paper:
- For patient care, each department’s LIMS connects to a local Palga Node, and the Nodes connect to a central Palga Hub. The Hub is an index only — all patient-identifiable data stays at the source department. A pathologist can pull a patient’s nationwide pathology history without any central store of identified records existing.
- For research, a separate central database — the Palga Scientific Database (PScDb) — receives pseudonymised excerpts, updated weekly.
The two mechanisms worth stealing
1. A coded diagnosis line, written by the pathologist, enforced at authorisation.
Every case gets a Palga One-line Diagnosis Summary (PODS): a series of asterisk-separated terms carrying at minimum the topography, the technique used to obtain the material, and at least one diagnosis term. The paper’s worked example: a conclusion reading “Cervical biopsy showing an hrHPV-positive adenocarcinoma” becomes
cervix*biopsy*adenocarcinoma*high-risk HPV type positive
which the system converts automatically to T83000*P11400*M81403*E33453M00021.
Terms come from the publicly accessible Palga Thesaurus, are based on SNOMED and ICD-O, and are now mapped to SNOMED CT. The codes — unlike the terms — are hierarchical, so querying T67 retrieves both T67000 (colon) and T67200 (ascending colon), and synonyms collapse onto one code. Groups of codes aggregate into higher-order retrieval terms such as “all primary carcinomas” or “all resections”.
The enforcement is the part that matters. The PODS Control Module sits on each Node and checks the PODS during the authorisation phase; a report it flags cannot be authorised or submitted until corrected. A second module (PF-CM) checks that mandatory report and personal fields are present and well-formed.
2. Two-step irreversible pseudonymisation with the key held outside the organisation.
Each Node sends data to the PScDb via an external Trusted Third Party (ZorgTTP). Identifiers are pseudonymised in two stages: a first-order pseudonym is generated at the source, then ZorgTTP converts it centrally into the final Palga pseudonym. The encryption key is stored outside Palga, so no one inside Palga can re-identify anyone, no link to the original is retained, and originals are not stored during the process.
Researchers never receive even the pseudonyms — they get randomly assigned study-specific identification numbers.
Standardised structured reporting: the numbers
| First SSR protocols (colorectal, breast resections) | 2009 |
| Palga Protocol Module (PPM) built, on the LogicNets platform | 2013 |
| PPM and SSR protocols CE-certified | since 2019 |
| Medical content approved by Dutch Society of Pathology expert groups | since 2022 |
| Operational SSR protocols, mid-2026 | 34 |
| SSR pathology reports recorded in 2025 | more than 900,000 |
| Share of all pathology reports generated using SSR | more than one third |
Protocols are built on ICCR datasets plus Dutch national treatment guidelines (FMS) and UICC TNM, defining mandatory core elements with optional non-core elements configurable per department. Each item within a protocol — not just the report — is bound to SNOMED CT.
Use is not mandatory in general, but it is mandatory for colon biopsies and cervical smears within the population screening programmes, and strongly recommended in national guidelines for lung, endometrial, breast, ovarian and prostate cancer.
The named barriers to further adoption are worth recording because they are not technical: clinician resistance to changing established reporting routines, incompatibility with existing workflows, and the perception of increased workload.
Linkage uncertainty, quantified
This is the most unusual thing in the paper and I have not seen it stated so plainly elsewhere.
The PScDb carries two different pseudonym constructions with an explicit, acknowledged trade-off:
- date of birth + sex + first four letters of surname at birth — less distinctive, so more prone to collision (two people sharing one pseudonym);
- date of birth + sex + first initial + first eight letters of surname at birth — more distinctive, so more prone to splitting (one person acquiring several pseudonyms, e.g. when the recorded name varies).
Because neither is exact, data stewards assemble a patient’s records by probabilistic linkage across pseudonymised characteristics, and every excerpt is shipped carrying a quality indicator, administrative_duplicate_probability, with three levels from “1. non” to “3. present – likely”. It is computed for both target and historical excerpts and reflects internal consistency across multiple target excerpts.
So the researcher receives a per-record measure of how confident the system is that these records belong to the same person — false-positive merges and false-negative splits are declared rather than hidden.
Access and governance, briefly
Departments are the GDPR data controllers; Palga is the data processor. Consent is opt-out by default (some departments are moving to opt-in). Secondary use without consent is permitted under four stated conditions: public-interest scientific or statistical research, pseudonymised data, consent being impracticable or disproportionate, and data minimisation.
150–200 data requests per year, generally free of charge, but always requiring a collaborating pathologist affiliated with a Dutch pathology department — which is what makes the PScDb unavailable to this group as a data source. Requests go through a portal and are reviewed by a scientific committee and a privacy committee; the privacy committee explicitly does not replace an accredited research ethics committee. Publication in a peer-reviewed journal is mandatory even for negative results. Data stewards query with SAS Enterprise Guide.
Three request types: exploratory (aggregate tables only; Palga disclaims validity and discourages publication from them), general, and linkage to outside cohorts. Linkage runs through the same TTP, with the rule that neither source may enrich its own database from the other — the researcher pulls from both and joins on a shared key.
A separate Palga Public Database gives anonymised aggregate counts by Palga code, stratifiable by year, age group, sex and material type, and suppresses any query returning 10 or fewer cases.
Archived FFPE material is reached through a different body, the Dutch National Tissue Portal (DNTP, since 2015), which owns the request-and-distribute process and acts as a track-and-trace system; Palga data stewards locate material and submit on the researcher’s behalf. Departments participate voluntarily.
What it says about images — the relevant gap
Whole-slide images are not part of the national infrastructure and never have been. Palga was built for structured and narrative report text. The paper is candid that integrating WSIs raises storage, heterogeneity, governance and funding problems it has not solved, and that Dutch initiatives toward a national image-sharing infrastructure are at feasibility stage with no formal roadmap or timeline, with architecture, governance and storage strategy all still open.
The authors position Palga against NHS Digital Pathology, TCGA and CPTAC and concede the comparison honestly: Palga leads on longitudinal nationwide text, and the others are further ahead on integrated imaging and image-based AI.
Future directions: LLMs on report text
AI is not currently used inside Palga. The paper names LLM applications it considers plausible: converting narrative reports to structured synoptic ones, extracting buried variables (ER/PR status, Bethesda, Gleason, Breslow thickness), assisting standardised coding, and redacting identifiers from free text. It also proposes recording which report elements were AI-generated, to make later auditing and post-market surveillance possible.
Academic projects supported by Palga are exploring a curated pseudonymised text corpus for training and benchmarking LLMs. The stated obstacle is not data volume but “highly specialized terminology, extensive use of abbreviations, and heterogeneity of reporting practices across institutions and time periods.”
Reading it critically
This is an infrastructure description written by the infrastructure’s own staff — nine of the ten authors are at the Palga Foundation. It is not a study and makes no measured claim about Palga itself, so the usual appraisal questions mostly do not apply. What is worth flagging:
- The one causal claim is second-hand. That SSR improved patient outcomes including survival in colorectal cancer is attributed to Sluijter 2019 (JCO Clinical Cancer Informatics), not demonstrated here. Treat it as a pointer to read, not an appraised finding — the design and confounding of a before-and-after national comparison is exactly where such a claim would be won or lost, and none of that is visible from this paper.
- Every headline count is self-reported, sourced to Palga’s own annual reports.
- No error rate anywhere. PODS coding is done by pathologists at sign-out under a structural validator — but the validator checks well-formedness, not correctness. Nothing here reports how often a syntactically valid PODS is clinically wrong, and that number bounds every study built on code-based retrieval. It is the single most useful missing figure in the paper.
administrative_duplicate_probabilityhas no published calibration. Three ordinal levels are shipped with each record, but the paper gives no sensitivity, specificity, or the similarity rule behind them. A researcher receives an uncertainty flag they cannot convert into a bias estimate.[unverified]- “More than one third” of reports are SSR — so approaching two thirds are still narrative, which the discussion concedes affects data uniformity. Any study drawing on structured fields is drawing on a non-random subset of departments and protocols, since adoption varies by both. That is a coverage-bias denominator problem and the paper does not quantify it.
- Declared: Mistral and ChatGPT were used for language editing. The competing-interests statement is present but rendered as an unfilled template in the pre-proof — both the “no competing interests” and the “the following interests” branches appear with nothing selected. Likely a pre-proof artefact.
[unverified]
Why it earns space here
Three reasons, in order of how much they change what to do:
- It is a working reference design for de-identification that is not a failure case. Every worked example on De-identification so far is one of this estate’s own slips. This is what the pattern looks like when it is built properly, at national scale: the key lives outside the organisation that holds the data.
- It reframes Report Text Extraction. The Dutch answer to “how do we get structured data out of pathology reports” is not a better parser — it is a mandatory coded field written at sign-out with a validator that blocks authorisation. That is an organisational intervention competing with a technical one, and it is the comparison the group’s own extraction work has never had.
- The linkage-uncertainty indicator is a rare example of a data provider shipping a measure of its own residual ambiguity rather than presenting a clean join.
Related: Synoptic Reporting — Palga is the largest national implementation of it, with the adoption numbers and the named barriers. Related: Record Linkage Under Pseudonymisation — the concept this paper is the worked example for. Related: imagebank — the local attempt at exactly the image layer Palga says it does not have.