The queue against the baselines

Random ordering is not the competition. These are: they cost nothing to implement, and if the queue cannot beat them then the model in front of it earned nothing. Every ordering is replayed on the same population, at the same budget, on the same resampled draws, so the comparison is paired rather than two separate measurements put side by side.

The conclusion

A baseline counts as beaten only when the Bonferroni-widened interval excludes zero under both the episode and the station resample, and both groupings agree on the direction.

Where they disagree the comparison is reported as not established with both directions named, because a disagreement between two defensible groupings is a real state rather than a reason to pick one.

Baselines compared
4
each replayed under two groupings
Beaten under both
0
none
Lost to
0
no baseline beat the queue

Oldest first

not established
Difference
+6 conflicts
By episode
[0.0, 13.0] does not survive
By station
[0.0, 14.0] does not survive

Not claimed. Grouped by pass episode the queue did better; grouped by ground station it did better. Neither survives the multiplicity correction, and a comparison is claimed only when both groupings survive it and agree.

Most uncertain image first

not established
Difference
+5 conflicts
By episode
[-1.0, 14.0] does not survive
By station
[-1.0, 13.0] does not survive

Not claimed. Grouped by pass episode the queue was not separated from the baseline; grouped by ground station it did better. Neither survives the multiplicity correction, and a comparison is claimed only when both groupings survive it and agree.

Largest frequency offset first

not established
Difference
0 conflicts
By episode
[-5.0, 5.0] does not survive
By station
[-4.0, 4.0] does not survive

Not claimed. Grouped by pass episode the queue was not separated from the baseline; grouped by ground station it was not separated from the baseline. Neither survives the multiplicity correction, and a comparison is claimed only when both groupings survive it and agree.

Physics score only

not established
Difference
+7 conflicts
By episode
[1.0, 15.0] survives
By station
[0.0, 15.0] does not survive

Not claimed. Grouped by pass episode the queue did better; grouped by ground station it did better. The episode grouping survives the multiplicity correction and the station grouping does not, and a comparison is claimed only when both groupings survive it and agree.

The measurements behind it

Conflicts found by the queue minus conflicts found by the baseline, at the same budget, on the same resampled population.

Null 0. The ratio beside it carries a +0.5 continuity correction on both terms in every draw, so a baseline that finds nothing does not produce an unbounded value and the estimator does not change between draws.

Grouped by pass episode

random 1.0The shipped queuethe composite score this project produces1.582Oldest firstwhat a reviewer working through a backlog does by default1.107Most uncertain image firstclassic active learning: review what the model is least sure about1.186Physics score onlypass geometry and corridor fit, with no image model at all1.028Largest frequency offset firstone line: sort on the offset that defines most of the conflicts, and see whether the other three score terms bought anything1.582
Each bar is a 95% interval and the dot inside it the point estimate. The dashed rule is random ordering at the same budget. An interval that contains 1.0 has not shown that its ordering beats chance, whatever its point estimate reads.
Budget 50, population 87, 22 conflicts in it, 87 groups. Random finds 12.6 on average.
OrderingConflictsLift95% intervalAgainst the queueBonferroni
The shipped queue
the composite score this project produces
201.582[1.353, 1.740]reference
Oldest first
what a reviewer working through a backlog does by default
141.107[0.797, 1.437]the queue beats it +6 conflicts[0.0, 13.0]
Most uncertain image first
classic active learning: review what the model is least sure about
151.186[0.791, 1.450]indistinguishable +5 conflicts[-1.0, 14.0]
Physics score only
pass geometry and corridor fit, with no image model at all
131.028[0.663, 1.353]the queue beats it +7 conflicts[1.0, 15.0]
Largest frequency offset first
one line: sort on the offset that defines most of the conflicts, and see whether the other three score terms bought anything
201.582[1.344, 1.740]indistinguishable 0 conflicts[-5.0, 5.0]

0 of the resamples were degenerate and are reported rather than dropped silently.

Grouped by ground station

random 1.0The shipped queuethe composite score this project produces1.582Oldest firstwhat a reviewer working through a backlog does by default1.107Most uncertain image firstclassic active learning: review what the model is least sure about1.186Physics score onlypass geometry and corridor fit, with no image model at all1.028Largest frequency offset firstone line: sort on the offset that defines most of the conflicts, and see whether the other three score terms bought anything1.582
Each bar is a 95% interval and the dot inside it the point estimate. The dashed rule is random ordering at the same budget. An interval that contains 1.0 has not shown that its ordering beats chance, whatever its point estimate reads.
Budget 50, population 87, 22 conflicts in it, 35 groups. Random finds 12.6 on average.
OrderingConflictsLift95% intervalAgainst the queueBonferroni
The shipped queue
the composite score this project produces
201.582[1.368, 1.735]reference
Oldest first
what a reviewer working through a backlog does by default
141.107[0.799, 1.502]the queue beats it +6 conflicts[0.0, 14.0]
Most uncertain image first
classic active learning: review what the model is least sure about
151.186[0.812, 1.402]the queue beats it +5 conflicts[-1.0, 13.0]
Physics score only
pass geometry and corridor fit, with no image model at all
131.028[0.642, 1.321]the queue beats it +7 conflicts[0.0, 15.0]
Largest frequency offset first
one line: sort on the offset that defines most of the conflicts, and see whether the other three score terms bought anything
201.582[1.369, 1.732]indistinguishable 0 conflicts[-4.0, 4.0]

0 of the resamples were degenerate and are reported rather than dropped silently.

The same numbers appear under both groupings for the point estimates and differ for the intervals, which is the whole reason both are shown.

A station contributes many passes and one receiver; treating its captures as independent would narrow every interval here by an amount that has nothing to do with the ranking being better.