Dunn 2025 — International Study of H&E Stain Variability

247 labs in the UK NEQAS programme stained the same circulated tissue; expert scores and quantitative colour deconvolution agreed less than expected, and stain ratio predicted quality better than stain intensity.
Author

Dunn C, Brettle D, Hodgson C, et al.

Doi

The same tissue sections were sent to 247 international labs, stained by each lab’s routine H&E protocol, then scored both by UK NEQAS expert assessors and by quantitative colour analysis — a rare case where subjective and objective stain quality can be compared at scale.

Scale and design

H&E accounts for over 80% of slides stained worldwide, and its quality control is “largely subjective”, assured by external quality assessment schemes resting on expert judgement. This study circulated tissue sections to 247 labs in the UK NEQAS Cellular Pathology Technique programme, had each stain them with its own routine protocol, and then analysed the returned slides two ways: independent expert assessor review, and digital quantification by H&E colour deconvolution plus colour difference (ΔE) in L*a*b* space.

Results, and the one that is actually interesting

  • 69% of labs scored good or excellent; assessors agreed well with each other (92.5% within one mark).
  • 60% of labs fell within 2ΔE of the mean — a difference “only perceptible through close observation”.
  • Little correlation between H&E intensity and assessor score. But the H&E intensity ratio trended with assessor score, suggesting an optimal relationship between the two stains rather than an optimal amount of either.

That last point is the finding to carry. The intuitive quantitative measure — how strongly the slide is stained — did not predict what experts considered good staining. The balance between haematoxylin and eosin did. Any automated H&E quality metric built on absolute intensity is therefore measuring the wrong thing, which is directly relevant to WSI Quality Control and to the group’s own stain-quality tooling.

Reservations

  • The relationship between intensity ratio and score is reported as a trend, not a fitted and validated threshold — so it identifies the right variable, not yet a usable cut-point. [unverified] what the effect size is; the abstract does not give one.
  • Expert score remains the reference standard, so the study measures agreement with expert opinion rather than with any downstream outcome. Whether “good” staining by this definition produces better diagnoses or better model performance is a separate question this design cannot answer.

Relation to the other EQA dataset here

This is a different study from the one behind Labquality EQA Staining Dataset (3 blocks, 66 labs, 11 countries, single scanner). Two independent multi-lab stain-variability datasets now sit in this wiki, with different scales and scoring schemes — worth knowing before either is cited as the evidence on inter-laboratory stain variation.

Related: Colour Calibration — the quantitative half of this study is colour deconvolution and ΔE, the measurement machinery that page describes. Related: Scanner and Stain Variability — this is the stain half of that variability, measured at 247-lab scale. Related: Interobserver Agreement — 92.5% of assessors within one mark is the ceiling any automated replacement is working against. Related: Labquality EQA Staining Dataset — the other multi-lab EQA staining dataset; compare before citing either.