Unsupervised encoder agreement · within-donor test · A/D/TBI n=91

Do cancer-histology encoders see anything real in brain tissue? Yes, and it strengthens at scale.

Andy Grossberg · Waving Cat Learning Systems 2026-07-22 13,650 tiles · 91 donors · Allen Institute A/D/TBI

Three frozen H&E foundation encoders (Phikon-v2, UNI, Virchow2, all trained on colorectal-cancer pathology, none on brain) run over 13,650 tiles from 91 human donors. No fine-tuning, no supervised labels. We ask whether their top-10 nearest-neighbor lists agree, restricted to same-donor tiles so donor-level stain, scanner, and prep effects are held constant.

Fig 01 · Within-donor top-10 NN agreement · headline result Phikon-v2 ∩ UNI · same-donor tiles only

Two independent cancer encoders agree on 57.4% of same-donor neighbors, 8.56× above chance

Restricting the nearest-neighbor comparison to same-donor tiles holds donor-level stain, scanner, and prep constant, so the agreement cannot be cross-donor stain fingerprinting. The signal sits well above the within-donor random baseline of 1×.

Observed within-donor agreement is 8.56 times the random baseline of 1 time. 10× Within-donor random baseline 1.00× Observed · Phikon-v2 ∩ UNI top-10 NN agreement 57.4% agreement 8.56× bootstrap 95% CI [8.45× – 8.67×]
Within-donor NN agreement
57.4%
Phikon-v2 and UNI agree on 57.4% of their top-10 nearest-neighbor lists when the comparison is restricted to same-donor tiles.
Above within-donor chance
8.56×
Bootstrap 95% CI [8.45×, 8.67×]. Every one of the 91 donors clears the 7× threshold; 79 of 91 clear 8×.
This is the strict within-donor test, the one designed to catch encoders that fingerprint stain-batch instead of tissue structure. Held per donor, the signal does not collapse: it sits comfortably above chance across every donor in a sample nearly four times larger than the July 16 pilot.
Fig 02 · Per-donor threshold coverage · n=91 share of donors above each baseline multiple

Every donor clears 7×. No donor falls back to chance.

Share of the 91 donors whose within-donor agreement clears each multiple of the random baseline. There are no per-donor point estimates shown here, only how many donors sit above each floor.

0% 25% 50% 75% 100% Clear 7× baseline 91 of 91 donors 100% 91 / 91 Clear 8× baseline 79 of 91 donors 86.8% 79 / 91
Per-donor floor: all 91 donors clear 7× above chance; 79 of 91 also clear 8×. The headline 8.56× is the pooled estimate, not a per-donor value. These bars report coverage, not a fabricated distribution.
Fig 03 · Pilot → replication · scale-up donors per run

Nearly 4× the donors, signal holds

The July 16 pilot ran 24 donors, one section each. This run covers 91. Under the same strict within-donor test, the signal does not collapse at the larger scale.

0 25 50 75 100 24 Jul 16 pilot 24 donors · 1 section 91 Jul 22 run 91 donors · 1 section nearly 4× larger
July 16 pilot: 24 donors, one section each. July 22: 91 donors. The within-donor result replicates and strengthens across this ~4× scale-up.
Fig 04 · Cross-encoder-pair agreement · Spearman ρ 3 encoder pairs · all p < 1e-7

The same donors are jointly "clean" or "messy" across every encoder combination

Spearman rank correlation of per-donor agreement across the three encoder pairs falls in a single band. Only the range endpoints are known (0.53 to 0.81), so the band is shown as a range across all three pairs, not three fabricated point values.

Phikon-v2 / UNI Phikon-v2 / Virchow2 UNI / Virchow2
0.0 0.2 0.4 0.6 0.8 1.0 Spearman rank correlation of per-donor agreement (higher = more concordant across encoders) ρ = 0.53 ρ = 0.81 all 3 encoder pairs fall in this band
Spearman rank correlation across the three encoder pairs: ρ = 0.53 to 0.81, all p < 1e-7. That the same donors read as jointly clean or messy across every combination is consistent with the encoders resolving a shared underlying tissue structure, rather than each pair picking up its own noise.
Fig 05 · What this doesn't yet claim · honest limits the hedges are a feature

Two limits worth naming plainly

01
One section per donor

Section-level artifacts (staining batch, scanner focus, mounting day, fixation time) are still perfectly confounded with donor identity. The strict within-donor test rules out cross-donor stain artifact, but it cannot yet rule out a single-section-wide artifact that happens to persist across all 150 tiles of that section. The airtight follow-up is multi-section-per-donor; the download pattern is ready when we have a reason to run it.

02
No ground-truth labels

All agreement here is unsupervised. Nothing yet says the encoders are picking up cortical layer, myelin density, tangle burden, or any specific tissue feature a neuropathologist would care about. Only that whatever structure they resolve is shared, stable across three independent encoders, and not stain artifact alone.

The strict within-donor test is designed to catch stain-batch fingerprinting; it passes. What it does not do is assign meaning to the resolved structure. That is the next experiment, not this one.
Fig 06 · What we're doing with this · outreach labeled brain-tissue datasets wanted

The next experiment needs labeled brain-tissue data

With any labeled brain-tissue dataset (Braak stage, cortical layer, white-matter integrity), a linear probe or small MLP head on top of these frozen embeddings is hours-to-days of work with no re-encoding. If that overlaps with something you're working on, we'd like to hear.

Andy Grossberg · andy.grossberg [at] gmail.com · Waving Cat Learning Systems · 2026-07-22. Full technical report and analysis scripts available on request.