Preparation Quality Versus Image Quality
Two different things a slide can be bad at — how the section was cut and mounted, versus what the scanner captured — which need separate checks because neither instrument sees the other’s failures.
What it is
“Slide quality” is routinely treated as one property. It is at least two, and they fail independently.
Preparation quality is a property of the physical slide, decided in the laboratory before anything is scanned: is the section centred on the glass, is it oblique, torn, folded, detached, overflowing the edge, cut too thick, are the levels parallel with no gaps between them, is there residue under the coverslip. These are microtomy, floating-out and mounting outcomes. A technician sees them by holding the slide up, and most are visible without a microscope at all.
Image quality is a property of the scan: out-of-focus regions, the artefacts that a segmentation model can find in pixels — folds where they darken the image, bubbles, pen marks, dark spots and foreign bodies, coverslip edges. This is what automated WSI QC tools such as GrandQC and HistoQC measure, and it is what WSI Quality Control is about.
The two overlap only partly. A fold is visible both ways. Off-centre placement is invisible to an artefact segmenter, because the tissue it can see is perfectly good tissue — it simply is not where it should be relative to a section boundary the model has no concept of. Conversely a subtle focus failure is invisible to the naked eye on glass and obvious in the scan.
The size of the gap has been measured here, once. GrandQC Quality Study mapped 554 quality problems recorded by Memorial technicians onto GrandQC’s five artefact classes, and 388 of them — 70% — had no corresponding class at all. Nine of the technicians’ thirteen criteria are preparation-geometry items with no image-artefact equivalent. Chance-corrected agreement between the technician’s suitability call and the model’s artefact percentage was Cohen’s κ = 0.10 — and not reliably above zero once the repeated slides in that dataset are accounted for. That is what near-independence looks like when two instruments score different constructs on the same objects.
That is the empirical content of this page: the low agreement was not a model failure and not a technician failure. It is what you get when a measurement of one axis is scored against a measurement of the other.
Why it matters for my work
It changes what a QC deployment is for. An artefact-segmentation model cannot replace the technician check, and a technician check cannot replace it. Running GrandQC and concluding “the slides are fine” leaves nine of thirteen recorded failure modes unexamined; running only the bench check leaves focus failures to be discovered by whoever reads the slide.
It decides which failures are recoverable and how expensively. A preparation failure usually requires going back to the block — recut, remount, restain — while an image failure often needs only a rescan. Those are very different costs in time and in tissue, and tissue is finite on a small biopsy. Conflating them makes the remediation decision unanswerable, and it feeds directly into Turnaround Time and rescan load in Scanning Time in Real Life.
It locates the defect in the process. Preparation faults point at microtomy, floating-out and mounting; image faults point at the scanner, its focus map, or the slide’s optical properties. A QC system that cannot separate them cannot tell a laboratory which of its steps to fix, which is most of the value of measuring quality at all.
It sets the threshold question. Neither axis has an established cut-off here, and they should not share one. [unverified] — no pre-specified exclusion rule is recorded for either.
How it connects
WSI Quality Control — the image-quality half, and the page that now carries the finding that it is only a half.
GrandQC Quality Study — the measurement this concept comes from, and the only local evidence of the size of the gap.
GrandQC-QuPath — the wrapper that made an image-quality model available locally; the reason its artefact vocabulary felt incomplete in routine use is on this page.
Interobserver Agreement — the statistical frame for comparing two raters, and the reason a low κ between instruments measuring different constructs is uninformative about either.
Laboratory Workload Measurement — preparation faults are attributable to a step and a cause, so a preparation-quality log is workload and process data, not only quality data.
Scanner and Stain Variability — a third axis again: colour and scanner differences are neither preparation geometry nor artefact presence, and need their own measurement. See Stain Quality.
Turnaround Time — recut versus rescan is the operational consequence of telling the two apart.
Model Abstention — whichever axis fails, the decision to exclude a slide before a model sees it is the same decision, and it should be pre-specified per axis.
Open questions
- Is there a published instrument that scores preparation quality from the scan? The technicians’ items are geometric and mostly measurable from a thumbnail — centring, overflow, obliquity, gaps between levels are all computable from a tissue mask. Nothing here has looked for such a tool, and it is a plausible gap in the open-source stack rather than an established absence.
[unverified] - Would the geometry items yield to a tissue mask plus simple shape statistics? If so this is a small piece of work with a clear payoff, since it would close the 70% rather than documenting it. GrandQC-QuPath already produces tissue masks as its first stage.
- Does preparation quality predict image quality? A thick or lifted section is a plausible cause of focus failure, so the two axes may be correlated even though they are conceptually distinct — and if they are, one measurement may partly stand in for the other. The data to test this exists in GrandQC Quality Study and has not been used for it.
- Which axis actually costs the laboratory more is unknown, because neither is currently counted in a way that reaches Laboratory Workload Measurement.