Precedent, warm and cold

Do the passes most similar to this one, by what is knowable before the waterfall is opened, carry the same recorded outcome more often than chance?

Agreement at 5, per arm

739 decisively labelled passes from 116 stations and 295 satellites. Chance is 0.531, which is the label mix alone.

WarmColdchance 0.531Granite embedding of the card0.6180.554Standardised numbers, Euclidean0.5920.537The station's own recent passes0.605no cold definitionUniform draw from the same pool0.5300.528
Each line is one arm. The filled mark is the warm condition, which allows any neighbour; the hollow mark is cold, which forbids the query's own station and satellite. The dashed rule is the label mix alone. The length of a line is what the entity restriction costs that arm.

Four retrieval arms over the same candidate pool, under two conditions. Warm allows any other observation. Cold requires a different ground station, a different physical site and a different satellite, because in this corpus the outcome is partly a property of who recorded it. Agreement at 5 is the mean over queries of the share of retrieved neighbours carrying the query's own label.

ArmWarmCold
Granite embedding of the card0.6180.554
Standardised numbers, Euclidean0.5920.537
The station's own recent passes0.605not applicable
Uniform draw from the same pool0.5300.528
The cold condition excludes candidates from the query's own station, so this arm has no definition here. Reported as null rather than as zero agreement, which would be a measurement nobody made.

What survives correction

Every margin is a paired difference over the same queries, resampled by ground station and widened by Bonferroni over the comparisons this study makes.

ComparisonConditionMargin95% intervalCorrectedSurvives
Granite against chancewarm0.0880[0.0465, 0.1344][0.0338, 0.1558]yes
Numbers against chancewarm0.0620[0.0279, 0.1046][0.0184, 0.1217]yes
Station history against chancewarm0.0800[0.0290, 0.1409][0.0112, 0.1672]yes
Granite against the numberswarm0.0260[-0.0044, 0.0557][-0.0169, 0.0660]no
Granite against chancecold0.0260[-0.0196, 0.0763][-0.0353, 0.0998]no
Numbers against chancecold0.0092[-0.0107, 0.0321][-0.0185, 0.0432]no
Granite against the numberscold0.0168[-0.0261, 0.0601][-0.0406, 0.0783]no
Warm, all three retrievers beat chance and survive the correction. Cold, none of them does: the interval contains zero. The embedding also does not beat seven standardised numbers under either condition. That is the finding, and it is the reason the queue is not ranked by this.

The index

The measurement above is exact cosine search. This is what a real vector index returned for the same queries.

Backend
chromadb 1.5.9, cosine, in memory
Recall at 5, warm
0.9992
over 739 queries
Recall at 5, cold
0.9995
over 739 queries
Embedding
granite-embedding:278m
277.45M, F16

The share of the exact top-5 that the index also returned, per query, averaged.

The cold condition is answered by a metadata filter on station, site and satellite inside the index rather than by filtering afterwards, so this also checks that the filter and the exclusion rule agree.

When the exclusion rule gained the site and the filter did not, this number fell from 0.94 to 0.77 on the cold condition, which is what it is here to do.

What this does not measure

  • Whether a reviewer shown these neighbours decides faster or better. That is kill gate 4's territory. Gate 4 has been answered on the decidability of its sample and not on whether a note or a neighbour helps, so this question is still open. Its verdict word lives in its own receipt and is not restated here, because a literal in this file cannot follow it.
  • Whether the label a neighbour carries is correct. The network's own verdict on an observation is a silver label, and agreement with it is agreement with the network rather than with the sky.
  • Anything about the images. Every arm here sees only what is knowable before the waterfall is opened, which is the point, and it means none of these arms is a detector.