Which satellite passes are worth a reviewer’s time, and the measurement that says so.
SatNOGS ground-station network · 50 observations examined · IBM Granite, on one machine
Volunteer ground stations record far more satellite passes than anyone reviews, and every recording is published as a waterfall image nobody has opened. TraceTriage reads the image and the orbital physics together and ranks 407 of them by how much a reviewer would learn from opening each, with the top 50 as the budget a volunteer actually has. It writes nothing back. Open the queue.
What holds. The evidence tools change what a local Granite model gets right, 22 of 24 against 2 of 24 with no tools at all, paired p = 0.000001. The grounding checker caught 525 of 525 planted falsehoods and refused 0 of 175 clean drafts. Read those two as a multiplication rather than as a count of attacks: 21 hand-written adversarial drafts and 7 clean ones, each against 25 packets. On real drafts, which are the ones nobody wrote to be caught, it refused 15 of 25. Every question in the agent study is answerable from the tools it was given, so that result measures whether the policy reaches for them, not whether it could have answered without.
IBM Granite does two jobs here and neither is the ranker: granite3.1-dense:8b answers questions over the evidence through this project’s own tools, and granite-embedding:278m retrieves a pass’s precedents, measured against a numeric baseline that ties with it. Both run on one machine over Ollama, with no hosted inference and no credential.
Chronologicalpre-registered, decides the gate
1.58×
Cold stationheld out, every station unseen in training
2.25×
The measurement is the point. A queue nobody tested is a preference.
The chronological split decides the gate because it was named in advance. One split passing is not the gate passing, and it is not nothing either.
0 of 2000 random orderings of the same 87 observations found as many conflicts inside the budget as this queue did (permutation p = 0.0005). The interval spans the threshold because the arithmetic caps it: 50 slots over 87 observations holding 22 conflicts put a perfect oracle at 1.740×, leaving 0.240 between the bar and perfection. How that bound was computed.
Behind this screen: all 407 ranked observations, rank 1 at the centre. Brightness is review value; 87 carry a fitted Doppler offset and drift at a rate set by its size.NO_REASONSTALE_CATALOGUE_FREQMODEL_LABEL_DISAGREEDISPLACED_STATION_CAP
Time, → one pass
Received frequency ↑
The corridor the orbit predicts, and 200 that were built from the same numbers.
Observation 14740031, a real SatNOGS capture. The white path is the Doppler shift computed from the satellite’s orbit, shifted by one fitted frequency offset bounded at 50 ppm. Each dark path is the same observation’s own Doppler values in a scrambled time order, given the identical offset search. They keep every frequency value and the whole swing, and lose only the shape. 6 of the 200 are drawn, including the closest one.
The plate is greyscale as the network published it, to within 1 part in 255, so every grey is a measured intensity and every coloured mark is something the pipeline computed. The colour is a map applied for legibility and changes no pixel’s rank against another.
- Fitted corridor
- 2.02 σ
- Best of 200 nulls
- 0.57 σ
- Nulls that reached it
- 0 of 200
- One-sided p
- 0.005
This is one observation. Gate 3 asked for the corridor to intersect a visible trace in 70% of reviewed positives. 224 of 289 scored observations discriminate, and the exact one-sided 95% lower bound on that is 0.731, which clears 70%. The gate is still not met, because the pre-registration counts groups rather than observations: over 68 (ground_station, UTC date) groups the rate is 0.471 with a lower bound of 0.366. So the gate is published as PASSED_UNGROUPED_ONLY: the observation-level pass is reported and not claimed. Open this observation, or read how every number here was generated.
Kill gate 6
NOT ESTABLISHED“Require the top review queue to find at least 1.5 times as many manually actionable conflicts as random ordering at the same budget.”
Lift 1.582, 95% interval [1.353, 1.740] over 4000 resamples grouped by pass episode and by ground station. The interval contains 1.5, so the gate is not met. It also sits entirely above 1.0, so the ranking is not nothing either.
What the measurement actually is
54 seconds, narrated: one pass measured against the curve its own orbit predicts.
scripts/render_explainer_narration.py. This clip shows what the measurement is and not that it is significant, and on this observation it is not: the matched filter reaches 0.396 sigma against the curved corridor and 0.352 against a vertical line, so the fit is the best of 408 offsets on an image whose trace the corridor cannot separate from noise. It was chosen for having the strongest corridor curvature in the shipped set, which is what makes the shape legible, and a decisive example would show a larger separation and a less readable curve. The gate 3 rate on the evaluation page is measured over observations that were null-calibrated; this one was not.What the queue found
Counts, not rates. Every number here carries the denominator it was measured over.
What counts as a conflict
Fixed before anything was measured. A criterion invented after seeing the ranking would measure the ranking against itself.
| Reason | What it means | Threshold |
|---|---|---|
| Model disagrees with label | Shipped arm (image_corridor) predicts with calibrated probability ≥ 0.75 that the observation is positive when the label is without-signal, or ≤ 0.25 (confident negative) when the label is with-signal. | prob positive floor 0.75, prob negative ceiling 0.25 |
| Stale catalogue frequency | Fitted frequency offset magnitude ≥ 20 ppm of the catalogue downlink frequency and offset_at_bound is false. Implies the SatNOGS transmitter catalogue frequency is at least 20 ppm stale. | abs offset ppm min 20 |
| Dead capture time | flat_row_frac ≥ 0.15: at least 15% of waterfall rows carry no luminance variation, indicating substantial dead capture time. | flat row frac min 0.15 |
- MODEL_LABEL_DISAGREE uses all decisively-labelled observations, not only single-vetted ones. The SatNOGS waterfall_status_user field is present in the snapshot but the single-vetter count per observation is not stored; restricting to single-vetted records would require a second API call not made at snapshot time.
- An uncorrected-Doppler outlier is not one of the criteria. The survey of correction status found only 3 uncorrected captures among the 24 it could vet, so applying the rule to the rest of the corpus would be a guess. The criterion keeps its name in the reason list and every split records it as not applicable, because saying the data do not support it is not the same as leaving it out.
- Corridor features are unavailable for 18 observations: 4 have no waterfall image and 14 have orbital elements too stale to propagate. Those score zero on the offset and flat-row signals and rank on model disagreement alone.
Concentration caps
A queue that spends its whole budget on one station has found one station's problem, not the corpus's.
| Cap | Share of budget | Entries at budget | Displaced | Binding |
|---|---|---|---|---|
| ground_station | 10% | 5 | 4 | bound |
| transmitter_uuid | 20% | 10 | 0 | inert |
Displaced nothing at budget 50 over 407 ranked observations: the per-transmitter cap.
Reported as inert on this split rather than as exercised.
A cap is credited with a displacement whenever it would have blocked the entry, even if another cap would have blocked it too, so an inert result here is a property of the data and not of the order the caps are checked in.
The queue
Fixed at 50 before any results were seen. The chronological test set has 87 decisively-labelled observations, so 50 is 57% coverage: large enough to measure lift with reasonable precision and small enough that 1.5x is non-trivial. On a split holding fewer than 50 decisively labelled observations, the budget is however many there are. 25 of these carry a waterfall you can open; the rest are listed with their measurements only, because shipping 2,500 images to prove a ranking is not evidence, it is weight. Where the data comes from.
Showing 60 of 407 matching rows, from 407 in the queue. Rank is the rank the pipeline assigned and does not change under a filter.
| Rank | Observation | Satellite | Score | Why it is here | Label | Model p(signal) | Offset ppm |
|---|---|---|---|---|---|---|---|
| 1● | 14746092 | SNUGLITE-II | 0.8796 | Model disagrees with label | without-signal | 1.000 | 15.8 |
| 2● | 14735140 | OBJECT BT | 0.8496 | Model disagrees with label | with-signal | 0.000 | 5.0 |
| 3● | 14732518 | QB50P1 | 0.5999 | Stale catalogue frequency | without-signal | 0.688 | 42.1 |
| 4● | 14745718 | SHINSEI (MS-F2) | 0.5962 | Stale catalogue frequency | without-signal | 0.688 | 31.2 |
| 5● | 14745990 | UND ROADS 1 | 0.5961 | Stale catalogue frequency | without-signal | 0.650 | 31.7 |
| 6● | 14744250 | GONETS M 18 | 0.5932 | Stale catalogue frequency | without-signal | 0.688 | 25.1 |
| 7● | 14746054 | TIROS 10 | 0.5846 | No conflict criterion | without-signal | 0.650 | -19.2 |
| 8● | 14744187 | MARINA | 0.5837 | Stale catalogue frequency | with-signal | 0.438 | 26.2 |
| 9● | 14733024 | KANYINI | 0.5830 | Stale catalogue frequency | without-signal | 0.688 | 22.7 |
| 10● | 14744248 | IPERDRONE.0 | 0.5816 | Stale catalogue frequency | without-signal | 0.500 | -25.3 |
| 11● | 14742034 | KNACKSAT-2 | 0.5791 | Stale catalogue frequency | without-signal | 0.500 | 24.9 |
| 12● | 14740027 | OBJECT AF | 0.5776 | Stale catalogue frequency | with-signal | 0.650 | 32.0 |
| 13● | 14740031 | OBJECT E | 0.5745 | Stale catalogue frequency | with-signal | 0.650 | 32.0 |
| 14● | 14741735 | PEARL-1B | 0.5723 | No conflict criterion | with-signal | 0.438 | 18.9 |
| 15● | 14742036 | OBJECT C | 0.5689 | Stale catalogue frequency | without-signal | 0.438 | 22.8 |
| 16● | 14745664 | OBJECT J | 0.5660 | No conflict criterion | with-signal | 0.438 | -16.4 |
| 17● | 14745929 | OBJECT H | 0.5643 | No conflict criterion | with-signal | 0.500 | -16.4 |
| 18● | 14735884 | CUTE-1.7+APD 2 | 0.5567 | No conflict criterion | with-signal | 0.500 | -14.2 |
| 19● | 14736746 | FRONTIERSAT | 0.5543 | No conflict criterion | with-signal | 0.438 | 11.7 |
| 20● | 14745985 | BRO-18 | 0.5513 | Stale catalogue frequency | without-signal | 0.250 | 36.2 |
| 21● | 14735176 | CONNECTA IOT-13 | 0.5500 | Stale catalogue frequency | with-signal | 0.688 | 22.7 |
| 22● | 14735743 | OBJECT H | 0.5461 | Stale catalogue frequency | with-signal | 0.778 | -39.2 |
| 23● | 14745984 | UND ROADS 2 | 0.5451 | No conflict criterion | with-signal | 0.650 | -13.0 |
| 24● | 14742229 | MARINA | 0.5441 | No conflict criterion | with-signal | 0.688 | 18.7 |
| 25● | 14744259 | OTTER PUP 2 | 0.5417 | Stale catalogue frequency | without-signal | 0.286 | -22.3 |
| 26● | 14745341 | OTP-2 | 0.5396 | No conflict criterion | with-signal | 0.688 | -16.9 |
| 27● | 14745823 | GONETS M 05 | 0.5395 | No conflict criterion | with-signal | 0.438 | 9.1 |
| 28● | 14740029 | ION SCV-008 | 0.5384 | No conflict criterion | with-signal | 0.650 | 11.0 |
| 29● | 14745983 | AISAT | 0.5378 | Model disagrees with label | without-signal | 0.947 | -7.9 |
| 30● | 14745342 | OTP-2 | 0.5352 | No conflict criterion | with-signal | 0.650 | 10.0 |
| 31● | 14745663 | OBJECT AD | 0.5350 | No conflict criterion | with-signal | 0.688 | -16.4 |
| 32● | 14745773 | LAPAN-TUBSAT | 0.5327 | No conflict criterion | without-signal | 0.650 | -2.2 |
| 33● | 14745601 | ISS (ZARYA) | 0.5307 | No conflict criterion | without-signal | 0.438 | 9.5 |
| 34● | 14740030 | OBJECT D | 0.5306 | No conflict criterion | with-signal | 0.778 | 18.7 |
| 35● | 14744188 | OBJECT AD | 0.5290 | No conflict criterion | with-signal | 0.778 | -18.6 |
| 36● | 14737233 | FRONTIERSAT | 0.5283 | No conflict criterion | without-signal | 0.438 | 8.7 |
| 37● | 14745911 | CUTE-1 | 0.5249 | No conflict criterion | with-signal | 0.778 | 17.2 |
| 38● | 14737313 | FRONTIERSAT | 0.5218 | No conflict criterion | with-signal | 0.650 | 6.6 |
| 39● | 14745611 | OBJECT AU | 0.5198 | Stale catalogue frequency | with-signal | 1.000 | 31.0 |
| 40● | 14737501 | FRONTIERSAT | 0.5188 | No conflict criterion | with-signal | 0.688 | -9.9 |
| 41● | 14743373 | OTP-2 | 0.5187 | No conflict criterion | with-signal | 0.438 | 1.4 |
| 42● | 14742035 | CUBESAT XI 4 | 0.5180 | No conflict criterion | without-signal | 0.438 | -3.6 |
| 43● | 14744185 | OTP-2 | 0.5173 | No conflict criterion | with-signal | 0.653 | -8.5 |
| 44● | 14746049 | HUBBLE 6 | 0.5143 | No conflict criterion | with-signal | 0.778 | -13.3 |
| 45● | 14734009 | OBJECT AY | 0.5132 | Stale catalogue frequency | with-signal | 1.000 | -24.6 |
| 46● | 14745839 | JAS 2 | 0.5126 | No conflict criterion | with-signal | 0.688 | -8.6 |
| 47● | 14744775 | NETSAT-2 | 0.5113 | No conflict criterion | without-signal | 0.438 | 0.9 |
| 48● | 14745821 | GONETS M 18 | 0.5113 | No conflict criterion | with-signal | 0.688 | 9.1 |
| 49● | 14746117 | GONETS M 17 | 0.5101 | No conflict criterion | with-signal | 0.688 | 9.1 |
| 50● | 14745661 | LEMUR 2 MJR092620 | 0.5080 | No conflict criterion | with-signal | 0.947 | 14.7 |
| 51 | 14740028 | BUGSAT 1 | 0.5260 | No conflict criterionDisplaced, station cap | without-signal | 0.438 | 7.7 |
| 52 | 14733022 | OTP-2 | 0.5206 | No conflict criterionDisplaced, station cap | with-signal | 0.650 | -5.2 |
| 53 | 14733025 | BOTSAT-1 | 0.5130 | No conflict criterionDisplaced, station cap | with-signal | 0.778 | -12.5 |
| 54 | 14738773 | KSM1-C | 0.5104 | No conflict criterionDisplaced, station cap | with-signal | 0.650 | 2.2 |
| 55 | 14745612 | INNOCUBE (TUBSAT-31) | 0.5074 | Stale catalogue frequency | without-signal | 0.000 | -28.7 |
| 56 | 14733021 | CUAVA-2 | 0.5054 | No conflict criterion | with-signal | 0.650 | 1.0 |
| 57 | 14745607 | OBJECT AE | 0.5026 | No conflict criterion | with-signal | 1.000 | -18.4 |
| 58 | 14733026 | BRO-6 | 0.5000 | No conflict criterion | with-signal | 0.947 | -11.5 |
| 59 | 14746047 | CONNECTA IOT-9 | 0.4998 | No conflict criterion | with-signal | 0.688 | 5.0 |
| 60 | 14745822 | GONETS-M 23 | 0.4977 | No conflict criterion | with-signal | 0.947 | 10.5 |