The queue against the baselines
Random ordering is not the competition. These are: they cost nothing to implement, and if the queue cannot beat them then the model in front of it earned nothing. Every ordering is replayed on the same population, at the same budget, on the same resampled draws, so the comparison is paired rather than two separate measurements put side by side.
The conclusion
A baseline counts as beaten only when the Bonferroni-widened interval excludes zero under both the episode and the station resample, and both groupings agree on the direction.
Where they disagree the comparison is reported as not established with both directions named, because a disagreement between two defensible groupings is a real state rather than a reason to pick one.
Oldest first
not established- Difference
- +6 conflicts
- By episode
- [0.0, 13.0] does not survive
- By station
- [0.0, 14.0] does not survive
Not claimed. Grouped by pass episode the queue did better; grouped by ground station it did better. Neither survives the multiplicity correction, and a comparison is claimed only when both groupings survive it and agree.
Most uncertain image first
not established- Difference
- +5 conflicts
- By episode
- [-1.0, 14.0] does not survive
- By station
- [-1.0, 13.0] does not survive
Not claimed. Grouped by pass episode the queue was not separated from the baseline; grouped by ground station it did better. Neither survives the multiplicity correction, and a comparison is claimed only when both groupings survive it and agree.
Largest frequency offset first
not established- Difference
- 0 conflicts
- By episode
- [-5.0, 5.0] does not survive
- By station
- [-4.0, 4.0] does not survive
Not claimed. Grouped by pass episode the queue was not separated from the baseline; grouped by ground station it was not separated from the baseline. Neither survives the multiplicity correction, and a comparison is claimed only when both groupings survive it and agree.
Physics score only
not established- Difference
- +7 conflicts
- By episode
- [1.0, 15.0] survives
- By station
- [0.0, 15.0] does not survive
Not claimed. Grouped by pass episode the queue did better; grouped by ground station it did better. The episode grouping survives the multiplicity correction and the station grouping does not, and a comparison is claimed only when both groupings survive it and agree.
The measurements behind it
Conflicts found by the queue minus conflicts found by the baseline, at the same budget, on the same resampled population.
Null 0. The ratio beside it carries a +0.5 continuity correction on both terms in every draw, so a baseline that finds nothing does not produce an unbounded value and the estimator does not change between draws.
Grouped by pass episode
| Ordering | Conflicts | Lift | 95% interval | Against the queue | Bonferroni |
|---|---|---|---|---|---|
| The shipped queue the composite score this project produces | 20 | 1.582 | [1.353, 1.740] | reference | — |
| Oldest first what a reviewer working through a backlog does by default | 14 | 1.107 | [0.797, 1.437] | the queue beats it +6 conflicts | [0.0, 13.0] |
| Most uncertain image first classic active learning: review what the model is least sure about | 15 | 1.186 | [0.791, 1.450] | indistinguishable +5 conflicts | [-1.0, 14.0] |
| Physics score only pass geometry and corridor fit, with no image model at all | 13 | 1.028 | [0.663, 1.353] | the queue beats it +7 conflicts | [1.0, 15.0] |
| Largest frequency offset first one line: sort on the offset that defines most of the conflicts, and see whether the other three score terms bought anything | 20 | 1.582 | [1.344, 1.740] | indistinguishable 0 conflicts | [-5.0, 5.0] |
0 of the resamples were degenerate and are reported rather than dropped silently.
Grouped by ground station
| Ordering | Conflicts | Lift | 95% interval | Against the queue | Bonferroni |
|---|---|---|---|---|---|
| The shipped queue the composite score this project produces | 20 | 1.582 | [1.368, 1.735] | reference | — |
| Oldest first what a reviewer working through a backlog does by default | 14 | 1.107 | [0.799, 1.502] | the queue beats it +6 conflicts | [0.0, 14.0] |
| Most uncertain image first classic active learning: review what the model is least sure about | 15 | 1.186 | [0.812, 1.402] | the queue beats it +5 conflicts | [-1.0, 13.0] |
| Physics score only pass geometry and corridor fit, with no image model at all | 13 | 1.028 | [0.642, 1.321] | the queue beats it +7 conflicts | [0.0, 15.0] |
| Largest frequency offset first one line: sort on the offset that defines most of the conflicts, and see whether the other three score terms bought anything | 20 | 1.582 | [1.369, 1.732] | indistinguishable 0 conflicts | [-4.0, 4.0] |
0 of the resamples were degenerate and are reported rather than dropped silently.
The same numbers appear under both groupings for the point estimates and differ for the intervals, which is the whole reason both are shown.
A station contributes many passes and one receiver; treating its captures as independent would narrow every interval here by an amount that has nothing to do with the ranking being better.