Rakha 2026 — Histological Grading of Invasive Breast Carcinoma in the Digital Era

Narrative review proposing a practical framework for applying Nottingham grading to WSI, with mitotic counting as the principal unsolved problem.
Author

Rakha EA (School of Medicine, University of Nottingham, UK; Department of Pathology, Cleveland Clinic Abu Dhabi and National Reference Laboratory, UAE)

Doi

A single-author narrative review arguing that Nottingham grading works on whole slide images but only after methodological adaptation — and that mitotic counting, not tubules or pleomorphism, is where the adaptation is actually needed.

Read status: abstract only. Wiley returned 402 on the article page; the abstract came from Europe PMC (PMID 42470156) and Crossref. The proposed framework’s operational detail — how hotspots should be selected step by step, what evidence sits behind the 2–3 mm² recommendation, and the magnitude of the WSI mitotic undercount — is not available. Claims that depend on that detail are marked [unverified].

What it is

A narrative review, published as Review + Journal Article. Method as stated: an “evidence-informed, methodological analysis … integrating published literature with expert practice.” No search strategy, inclusion criteria, or risk-of-bias assessment is described in the abstract. Expert opinion is an explicit input, not just a framing device.

There is no new cohort, no scanner, no magnification, no validation strategy, and no external validation — the fields the schema asks for on a model paper simply do not apply here. This is a proposal for a protocol, not a test of one. That is the single most important thing to hold onto when citing it.

The core claim, by grading component

The review’s central move is that WSI does not degrade the three Nottingham components equally, so a blanket “is digital grading safe?” question is the wrong question.

Component Effect of WSI, per the abstract
Tubule formation “Highly reproducible”, well suited to low-power digital assessment. Effectively a non-problem.
Nuclear pleomorphism “Moderate variability”, from intrinsic subjectivity and from display and perceptual factors.
Mitotic activity “The principal challenge” — hotspot identification, area calibration, and recognising mitotic figures without z-axis focus.

The practical recommendations

Three, all aimed at mitotic counting:

  1. Define the counting area in mm², optimally 2–3 mm².
  2. Prioritise regions of increased tumour cellularity when selecting the hotspot.
  3. Use calibrated high-power screen fields rather than assuming the microscope’s field equivalence carries over.

And one warning that matters more than the recommendations:

Systematic underestimation of mitotic counts on WSI compared with light microscopy is recognised, and is most relevant in cases near grading thresholds.

Direction of the bias is stated. Magnitude is not given in the abstract, and neither is the proportion of cases whose grade actually changes. [unverified]

The conclusion: grading “can be performed reliably using WSI but requires methodological adaptation rather than direct transfer of glass slide-based approaches.”

Why it is interesting here

Three concrete hooks into work already running, in descending order of immediacy.

1. It defines the human baseline that Aiforia Breast is being measured against. If a commercial breast AI is evaluated in routine sign-out and the human reference standard was read on glass while the algorithm reads WSI, the mitotic undercount is a bias built into the comparison design, not a property of the algorithm. Any disagreement on mitosis-driven readouts is then partly an artefact of the reading modality. That is worth resolving before, not after, the ECDP output.

2. Area calibration depends on scanner metadata the anonymisation step may delete. Counting in mm² requires µm/pixel resolution. The metadata-qupath tooling extracts scanner type and resolution from SVS, but the ecosystem notes state that anonymisation deletes this metadata — already flagged as an open question on Whole Slide Imaging. This paper turns that from a provenance nuisance into a measurement dependency: strip the resolution and you cannot calibrate a counting area at all.

3. It hands the department a cheap, well-posed agreement study. Grade is an ordinal three-level outcome read by multiple pathologists — the exact shape meddecide already handles. A glass-versus-WSI grading round on the same cases would measure the undercount locally instead of importing a number from a review that does not state one.

Statistical and methodological problems

Fewer than a data paper would have, but they change how the framework should be cited.

  1. Narrative, not systematic, and single-author. No search strategy, no inclusion criteria, no risk-of-bias assessment stated. The evidence tier for “grading can be performed reliably using WSI” is expert synthesis, not pooled data.
  2. The headline reassurance carries no effect size. “Can be performed reliably” appears without a pooled kappa, a concordance rate, or a confidence interval anywhere in the abstract. Reliability is a quantity; the abstract states it as a verdict. [unverified] whether the full text quantifies it.
  3. The one number a department needs is missing. The undercount’s direction is given, its magnitude is not — nor the fraction of cases crossing a grade boundary because of it. Without that, you cannot decide whether to adjust practice or ignore the effect.
  4. “Optimally 2–3 mm²” is stated without a visible derivation. Empirically optimised, consensus, or inherited from WHO convention? The abstract does not say. [unverified] A range given as “optimal” needs the criterion it was optimal against.
  5. Intellectual conflict of interest. The author leads the Nottingham group whose system the paper adapts. Not disqualifying — it is also why the expert practice is worth reading — but a framework’s originator is not a neutral evaluator of whether it transfers.
  6. The framework itself is unvalidated. It is a proposal. Nothing in the abstract reports testing it prospectively, in another department, or against glass. Recommending it and validating it are different acts, and only the first has happened.
  7. Display and perception are named as variance sources but not specified. Monitor, calibration, resolution, and viewing conditions are invoked for nuclear pleomorphism with no stated minimum. An unspecified confounder is not yet a controlled one.

What I would want before adopting the framework

The magnitude of the WSI-versus-glass mitotic difference and the grade-migration rate it causes; the evidence behind 2–3 mm²; the display specification implied by “calibrated high-power screen fields”; and whether any of it has been tested outside Nottingham.