Unsupervised encoder agreement · within-donor test · A/D/TBI n=91
Do cancer-histology encoders see anything real in brain tissue? Yes, and it strengthens at scale.
Andy Grossberg · Waving Cat Learning Systems2026-07-2213,650 tiles · 91 donors · Allen Institute A/D/TBI
Three frozen H&E foundation encoders (Phikon-v2, UNI, Virchow2, all trained on colorectal-cancer pathology, none on brain) run over 13,650 tiles from 91 human donors. No fine-tuning, no supervised labels. We ask whether their top-10 nearest-neighbor lists agree, restricted to same-donor tiles so donor-level stain, scanner, and prep effects are held constant.
Fig 01 · Within-donor top-10 NN agreement · headline resultPhikon-v2 ∩ UNI · same-donor tiles only
Two independent cancer encoders agree on 57.4% of same-donor neighbors, 8.56× above chance
Restricting the nearest-neighbor comparison to same-donor tiles holds donor-level stain, scanner, and prep constant, so the agreement cannot be cross-donor stain fingerprinting. The signal sits well above the within-donor random baseline of 1×.
Within-donor NN agreement
57.4%
Phikon-v2 and UNI agree on 57.4% of their top-10 nearest-neighbor lists when the comparison is restricted to same-donor tiles.
Above within-donor chance
8.56×
Bootstrap 95% CI [8.45×, 8.67×]. Every one of the 91 donors clears the 7× threshold; 79 of 91 clear 8×.
This is the strict within-donor test, the one designed to catch encoders that fingerprint stain-batch instead of tissue structure. Held per donor, the signal does not collapse: it sits comfortably above chance across every donor in a sample nearly four times larger than the July 16 pilot.
Fig 02 · Per-donor threshold coverage · n=91share of donors above each baseline multiple
Every donor clears 7×. No donor falls back to chance.
Share of the 91 donors whose within-donor agreement clears each multiple of the random baseline. There are no per-donor point estimates shown here, only how many donors sit above each floor.
Per-donor floor: all 91 donors clear 7× above chance; 79 of 91 also clear 8×. The headline 8.56× is the pooled estimate, not a per-donor value. These bars report coverage, not a fabricated distribution.
Fig 03 · Pilot → replication · scale-updonors per run
Nearly 4× the donors, signal holds
The July 16 pilot ran 24 donors, one section each. This run covers 91. Under the same strict within-donor test, the signal does not collapse at the larger scale.
July 16 pilot: 24 donors, one section each. July 22: 91 donors. The within-donor result replicates and strengthens across this ~4× scale-up.
Fig 04 · Cross-encoder-pair agreement · Spearman ρ3 encoder pairs · all p < 1e-7
The same donors are jointly "clean" or "messy" across every encoder combination
Spearman rank correlation of per-donor agreement across the three encoder pairs falls in a single band. Only the range endpoints are known (0.53 to 0.81), so the band is shown as a range across all three pairs, not three fabricated point values.
Phikon-v2 / UNIPhikon-v2 / Virchow2UNI / Virchow2
Spearman rank correlation across the three encoder pairs: ρ = 0.53 to 0.81, all p < 1e-7. That the same donors read as jointly clean or messy across every combination is consistent with the encoders resolving a shared underlying tissue structure, rather than each pair picking up its own noise.
Fig 05 · What this doesn't yet claim · honest limitsthe hedges are a feature
Two limits worth naming plainly
01
One section per donor
Section-level artifacts (staining batch, scanner focus, mounting day, fixation time) are still perfectly confounded with donor identity. The strict within-donor test rules out cross-donor stain artifact, but it cannot yet rule out a single-section-wide artifact that happens to persist across all 150 tiles of that section. The airtight follow-up is multi-section-per-donor; the download pattern is ready when we have a reason to run it.
02
No ground-truth labels
All agreement here is unsupervised. Nothing yet says the encoders are picking up cortical layer, myelin density, tangle burden, or any specific tissue feature a neuropathologist would care about. Only that whatever structure they resolve is shared, stable across three independent encoders, and not stain artifact alone.
The strict within-donor test is designed to catch stain-batch fingerprinting; it passes. What it does not do is assign meaning to the resolved structure. That is the next experiment, not this one.
Fig 06 · What we're doing with this · outreachlabeled brain-tissue datasets wanted
The next experiment needs labeled brain-tissue data
With any labeled brain-tissue dataset (Braak stage, cortical layer, white-matter integrity), a linear probe or small MLP head on top of these frozen embeddings is hours-to-days of work with no re-encoding. If that overlaps with something you're working on, we'd like to hear.
Andy Grossberg · andy.grossberg [at] gmail.com · Waving Cat Learning Systems · 2026-07-22. Full technical report and analysis scripts available on request.