Spent both β the filter and the reseed β from cache: numpy over the cached document and query vectors and the centroids of all 41 partitions, no index built and no faiss search; the top-10 lists are numpy scoring inside the probed cells. The cached path reproduced the posted LEMUR counts, 175/177/188/210, exactly and NDCG at 10 within 4e-6. Your train-qrels point was right and went in: 919 rows over 809 queries, 1,258 pairs over 1,109 with test.
relevant-document mass shipped 40-draw mean Β± sd min..max rank
test 339, ranks 21β26 13 14.60 Β± 3.92 8..23 16/41
test 339, ranks 27β32 17 11.43 Β± 3.01 2..16 41/41
both 1,258, ranks 21β26 49 56.40 Β± 9.04 38..77 8/41
both 1,258, ranks 27β32 51 43.98 Β± 7.97 28..64 32/41
Neither of your branches holds at 21β26: 13 against 14.60 Β± 3.92, rank 16 of 41 β not thin, not a tail. But the filter did find the level story on both sides of that window: 27β32 is the heaviest of 41, and relevant-document recall is the lowest of 41 at every cell scored β 0.661 at 20 against 0.726 Β± 0.027, 0.699 at 26 against 0.769 Β± 0.026, 0.749 at 32 against 0.803 Β± 0.023. This partition places LEMUR's relevant documents late: short by rank 20, typical at 21β26, heaviest at 27β32.
My locus last round was wrong. The dozen were in the cells, 13 of them; what thinned was admission, and it is measurable: of the 13 at ranks 21β26, only 4 are their query's exact-top-10 members, against 8.05 Β± 3.16 over the draws, rank 2 of 41; the 17 at 27β32 carry 11 members, above every draw (5.47 Β± 2.10). A member is admitted the probe its cell is scanned; a non-member needs its betters unscanned. Your "cell content" stands, measured: which relevant documents the partition puts in each window, not how many.
reseed, forty partitions shipped 40-draw mean Β± sd min..max tail
admission 20β26 2 7.20 Β± 3.59 1..16 2 draws β€ 2
admission 26β32 11 4.55 Β± 2.29 0..10 0 draws β₯ 11
NDCG at 10, 20β26 +0.00448 +0.01786 Β± 0.00857 +0.00519..+0.03990 lowest of 41
NDCG at 10, 26β32 +0.02679 +0.01232 Β± 0.00601 +0.00211..+0.02931 rank 40 of 41
The prediction held: the shipped 20β26 NDCG step is the lowest of 41, the shipped 26β32 admission is above every draw, and no draw pairs the two tails. The two steps on another seed would have read about 7 and about 4.5, not 2 and 11. Over all 1,109 queries the pattern repeats: steps 16 and 30 against 30.1 Β± 7.7 and 21.5 Β± 6.1.
Against myself, and larger than a pair of steps: the shipped LEMUR partition is low on the whole grid β rank 1β3 of 41 at every probe count through 53, z β2.9 at 14 and at 26. MUVERA's cached encoding sits 2β4 documents off the posted cells, so it enters only as a control with its per-probe offset added back, +0.002 to +0.007. Cross-pairing the arms' 40 draws β the partitions are independent β the corrected gap is positive in the mean at every grid point; the median first-positive probe count over the 1,600 pairs is 8, six percent of cells; 90 percent are positive by 20; the posted pair's 48β50 crossing sits at the 97.7th percentile. Holding LEMUR at the shipped partition and redrawing only MUVERA puts the median crossing at 50, the posted point. The late crossing was the shipped LEMUR draw's price. Two edges cut the other way: 26 percent of pairs re-cross after first going positive, and the positive fraction dips at 26 β the flat spot is visible in the ensemble. The posted nprobe 8 cell reproduces to one document, 0.0011; every other cell is exact.
So the deployment line was mispriced, by me: "toward nlist/2" was this pair's price, not the encoders'. On a typical partition pair LEMUR-2048 is ahead of MUVERA-10240 from the default probe fraction, and the caveat that survives is variance: at 5,183 documents over 132 cells, single-partition numbers carry 0.01β0.02 of NDCG at 10 in partition luck, level and shape both. Pin IDMap,Flat where that matters.
Both instruments are spent, and the grid with them. The residual stays as named.