Nottingham Grading

The three-component histological grade for invasive breast carcinoma — an ordinal score whose components behave very differently once the slide is read on a screen.

The three-component histological grade for invasive breast carcinoma — an ordinal score whose components behave very differently once the slide is read on a screen.

What it is

Three components — tubule formation, nuclear pleomorphism, mitotic count — are each scored 1–3 and summed, and the sum is banded into grade 1, 2 or 3. The precise mitotic score boundaries are area-dependent: they are defined for the size of the field being counted, so they must be read off the current table for the area actually used rather than carried over from a remembered “per 10 HPF” figure. That area dependence is the hinge on which everything below turns.

The structure matters statistically. Grade is ordinal, not categorical — grade 1 versus grade 3 is a worse disagreement than grade 1 versus grade 2 — and it is a thresholded sum, so a one-point shift in any single component can move the final grade only when the sum sits on a band boundary. Most cases are insensitive to small component errors. A minority are entirely determined by them.

What changes on whole slide images

Rakha 2026 (sources/papers/rakha-2026-wsi-breast-grading.md) makes the useful point that WSI does not affect the three components equally, so “is digital grading safe?” is the wrong question to ask as a single question.

  • Tubule formation — highly reproducible on WSI, and well suited to low-power digital assessment. Arguably easier than on glass, since the whole section is visible at once.
  • Nuclear pleomorphism — moderately variable, from intrinsic subjectivity plus display and perceptual factors. This is the component where the monitor becomes part of the measurement instrument.
  • Mitotic count — the principal problem, large enough that it gets its own page. See Mitotic Count.

The review’s conclusion is that grading transfers to WSI with methodological adaptation, not by direct transfer of glass-based habits. It is a proposal, not a validated protocol — no cohort, no concordance figure, and no external test are reported. Cite it as a framework, not as evidence that digital grading is equivalent.

Why it matters for my work

It defines the human baseline for Aiforia Breast. A commercial breast AI evaluated in routine sign-out is being compared against pathologists, and if those pathologists graded on glass while the algorithm reads WSI, part of any disagreement is the reading modality rather than the algorithm. The reference standard’s modality is a design variable that the project page does not currently record. [unverified]

It is a ready-made agreement study. An ordinal three-level outcome, multiple readers, two modalities — precisely what meddecide already implements. A glass-versus-WSI grading round on department cases would produce a local number where the literature offers a direction without a magnitude.

It sharpens an existing metadata gap. Grading on WSI needs calibrated area, which needs scanner resolution, which the anonymisation step may be deleting. See Whole Slide Imaging.

How it connects

Mitotic Count — the component that carries nearly all of the digital-era difficulty, and the reason the framework exists at all.

Interobserver Agreement — grade is ordinal, so weighted kappa is the right statistic and the linear-versus-quadratic choice must be stated; grading agreement also sets the ceiling for any AI trained against a single grader.

Biomarker Cut Points — grade is a thresholded sum, so the same near-boundary instability applies even though the threshold here is a guideline convention rather than a derived one.

Scanner and Stain Variability — display and perceptual factors are a third variance source alongside scanner and stain, and one this wiki had not previously named.

Whole Slide Imaging — the substrate, and the source of the µm/pixel metadata that area calibration depends on.

Synoptic Reporting — grade is a required element in structured breast reporting, which is what makes its reproducibility an operational and not merely academic concern.

Ki-67 Proliferation Index — the IHC proliferation index that parallels the mitotic component of grade; the two are the two ways breast proliferation gets quantified, and are easily conflated.

Aiforia Breast — the group’s breast AI evaluation, whose reference standard this concept directly governs.

Rakha 2026 — Histological Grading of Invasive Breast Carcinoma in the Digital Era — the review proposing how to apply this grade on a screen rather than down a microscope, and the source of the judgement that mitotic counting is the component that does not survive the move intact.

Open questions

  • Does the department currently grade breast cases on glass, on WSI, or a mixture? Nothing in the repo records this, and it determines whether the undercount applies here at all. [unverified]
  • Was the Aiforia Breast reference standard read on the same modality as the algorithm? Already listed as an open question on that project page; this paper makes it a bias question rather than a documentation one.
  • Is there a house display specification for primary digital reading — monitor, calibration, viewing conditions? The review names these as variance sources for nuclear pleomorphism without specifying a minimum. [unverified]
  • No page yet on intra-observer grading agreement, which is the companion measurement and the cheaper one to run.