Shaaban 2026 — UK Recommendations for Ki-67 Immunohistochemical Staining and Interpretation in Breast Cancer

National guidance from the UK breast pathology bodies on how to stain, score and report Ki-67, with digital pathology and AI positioned as the route to reproducibility.
Author

Shaaban AM, Dodson A, Rakha EA, Ellis IO, Rea D, Quinn C, Provenzano E, Pinder SE, on behalf of the UK National Coordinating Committee for Breast Pathology (NCCBP), the Association of Breast Pathology (ABP), and UK NEQAS ICC & ISH

Doi

A UK national recommendation, issued jointly by the NCCBP, the ABP and UK NEQAS ICC & ISH, on the technical staining, scoring and reporting of Ki-67 in breast cancer — and the first paper in this repo whose central argument is that reproducibility of a proliferation marker is now an EQA-and-AI problem, not only a scoring-rule problem.

Read status: abstract only. The Wiley article page returned 403 and the Europe PMC core record was rate-limited when this note was written. The high-level scope below is reliable (title, authorship, and the abstract’s own framing via Crossref). The operational specifics a laboratory would actually need — the recommended antibody clone(s), the scoring denominator (global average vs hotspot vs a fixed tumour-cell count), the number of cells/fields to count, any endorsed cut-point(s), and exactly which AI tools are considered validated — are not verified from the abstract. Every such specific below is marked [unverified] and must be confirmed against the full text before it is quoted or acted on.

What it is

A guideline / recommendation, not a primary study. There is no cohort, no scanner, no magnification and no validation experiment of the group’s own — the schema’s model-paper fields do not apply. What it carries instead is institutional weight: three UK bodies that between them set breast reporting practice, run the professional association, and operate the external quality assessment scheme for IHC/ISH have put their names to a single scoring standard. That is the thing to cite it for — a national reference point for “how Ki-67 should be done” — rather than as evidence that any particular method works.

The multi-body authorship is itself the content. UK NEQAS ICC & ISH being a named author means the recommendation is tied to an external quality assessment programme, so it is intended to be auditable across laboratories rather than adopted lab-by-lab. That is the same design logic as Labquality EQA Staining Dataset — a marker’s reproducibility is treated as a between-laboratory property to be measured, not a within-laboratory habit to be trusted.

What it covers (from the abstract)

Three strands, in the abstract’s own order:

  1. Technical immunohistochemistry. Pre-analytical and staining requirements for Ki-67 — the part most exposed to the between-laboratory variation that EQA exists to catch. Specific fixation, antibody and protocol requirements are [unverified].
  2. Scoring / interpretation methodology. How the stained slide is turned into a number. The denominator, the counting method, and whether a hotspot or a global estimate is recommended are the decisions that dominate Ki-67 irreproducibility, and they are [unverified] from the abstract.
  3. Digital pathology and AI. Quoted framing: these “have demonstrated improved reproducibility, reduced turnaround time, and prognostic performance superior to manual scoring.” This is the strongest and most citable claim in the abstract, and also the one that most needs its evidence base checked (see caveats).

The overall thrust, per the abstract, is standardisation and the use of validated approaches for clinical implementation.

Why it is interesting here

Four hooks into work already running, strongest first.

1. Ki-67 is already the running example on Biomarker Cut Points. That page uses “Ki-67 percentage” as its canonical continuous-marker-turned-categorical case, and the whole optimal-cutpoint hazard is precisely why a national recommendation on scoring matters: a guideline cut-point is the defensible alternative to a data-derived one. This paper is the concrete guideline that page was implicitly pointing at.

2. It is a reproducibility document, which is Interobserver Agreement territory. Ki-67 is the textbook low-agreement IHC marker; the reason it needs national guidance at all is between-reader and between-lab variability. The paper’s claim that digital/AI scoring improves reproducibility is a claim about lifting the agreement ceiling — directly the quantity that page is about, and directly relevant to any AI-versus-human comparison the group runs.

3. It sharpens the likely readout of Aiforia Breast. That project’s endpoint is undocumented, with Ki-67 / IHC quantification named as the most likely candidate given the vendor’s scope. If the Aiforia evaluation is in fact a Ki-67 readout, this recommendation is the reference standard the algorithm should be measured against — and the paper’s “AI beats manual scoring” claim is exactly what such an evaluation would be testing locally.

4. Technical staining variation ties to Scanner and Stain Variability and EQA. Ki-67 is an IHC stain, so between-laboratory staining spread is a first-order problem, and UK NEQAS authorship makes the EQA framing explicit. The abstract’s “reduced turnaround time” claim also touches Turnaround Time — automated scoring removing a manual counting step is a segment-level TAT effect, not a whole-process one.

Statistical and methodological caveats

Fewer than a data paper, but they change how the AI claim in particular should be cited.

  1. It is guidance, not evidence. “Validated approaches” and “superior prognostic performance” are summary verdicts. The abstract gives no pooled kappa, no concordance figure, no hazard ratio, and no citation trail for them. Whether the full text quantifies any of it is [unverified]. A recommendation inherits the strength of the evidence it rests on, and that evidence is not visible from the abstract.
  2. “Prognostic performance superior to manual scoring” is a strong claim on thin display. Superior on which cohort, against which manual protocol, with what external validation? None of this is in the abstract. This is the sentence most likely to be over-quoted and least supported by what can currently be read. [unverified]
  3. “Reduced turnaround time” is plausible but unmeasured here. It is asserted, not sized. Per Turnaround Time, TAT is composite and right-skewed; a plausible per-step saving is not the same as a measured whole-case improvement.
  4. No cut-point is confirmed. Whether the paper endorses a specific threshold (and if so how it was derived) is the single most consequential detail and is [unverified]. Per Biomarker Cut Points, a guideline threshold is only defensible if its derivation is stated.
  5. Generalisability outside the UK setting is a judgement call. EQA schemes, antibody availability and reporting conventions are national. The staining and audit specifics may not transfer directly to a Turkish laboratory even where the interpretive principles do.

What I would want before relying on it

The recommended antibody clone and staining protocol; the scoring denominator and counting method (hotspot vs global vs fixed cell count); any endorsed cut-point and its derivation; the actual evidence — cohorts, effect sizes, external validation — behind “reproducibility improved” and “prognostic performance superior to manual scoring”; and which AI tools the paper treats as validated. All of these are in the full text, none are in the abstract.