Which satellite passes are worth a reviewer’s time, and the measurement that says so.

SatNOGS ground-station network · 50 observations examined · IBM Granite, on one machine

Volunteer ground stations record far more satellite passes than anyone reviews, and every recording is published as a waterfall image nobody has opened. TraceTriage reads the image and the orbital physics together and ranks 407 of them by how much a reviewer would learn from opening each, with the top 50 as the budget a volunteer actually has. It writes nothing back. Open the queue.

What holds. The evidence tools change what a local Granite model gets right, 22 of 24 against 2 of 24 with no tools at all, paired p = 0.000001. The grounding checker caught 525 of 525 planted falsehoods and refused 0 of 175 clean drafts. Read those two as a multiplication rather than as a count of attacks: 21 hand-written adversarial drafts and 7 clean ones, each against 25 packets. On real drafts, which are the ones nobody wrote to be caught, it refused 15 of 25. Every question in the agent study is answerable from the tools it was given, so that result measures whether the policy reaches for them, not whether it could have answered without.

IBM Granite does two jobs here and neither is the ranker: granite3.1-dense:8b answers questions over the evidence through this project’s own tools, and granite-embedding:278m retrieves a pass’s precedents, measured against a numeric baseline that ties with it. Both run on one machine over Ollama, with no hosted inference and no credential.

Chronologicalpre-registered, decides the gate

1.58×

NOT ESTABLISHED
95% CI [1.353, 1.740]against a 1.5× threshold, so the gate is not met

Cold stationheld out, every station unseen in training

2.25×

PASSED
95% CI [1.920, 3.859]the whole interval clears 1.5× on stations the model never saw

The measurement is the point. A queue nobody tested is a preference.

The chronological split decides the gate because it was named in advance. One split passing is not the gate passing, and it is not nothing either.

0 of 2000 random orderings of the same 87 observations found as many conflicts inside the budget as this queue did (permutation p = 0.0005). The interval spans the threshold because the arithmetic caps it: 50 slots over 87 observations holding 22 conflicts put a perfect oracle at 1.740×, leaving 0.240 between the bar and perfection. How that bound was computed.

Behind this screen: all 407 ranked observations, rank 1 at the centre. Brightness is review value; 87 carry a fitted Doppler offset and drift at a rate set by its size.NO_REASONSTALE_CATALOGUE_FREQMODEL_LABEL_DISAGREEDISPLACED_STATION_CAP

Time, one pass

Received frequency

The corridor the orbit predicts, and 200 that were built from the same numbers.

Observation 14740031, a real SatNOGS capture. The white path is the Doppler shift computed from the satellite’s orbit, shifted by one fitted frequency offset bounded at 50 ppm. Each dark path is the same observation’s own Doppler values in a scrambled time order, given the identical offset search. They keep every frequency value and the whole swing, and lose only the shape. 6 of the 200 are drawn, including the closest one.

The plate is greyscale as the network published it, to within 1 part in 255, so every grey is a measured intensity and every coloured mark is something the pipeline computed. The colour is a map applied for legibility and changes no pixel’s rank against another.

Fitted corridor
2.02 σ
Best of 200 nulls
0.57 σ
Nulls that reached it
0 of 200
One-sided p
0.005

This is one observation. Gate 3 asked for the corridor to intersect a visible trace in 70% of reviewed positives. 224 of 289 scored observations discriminate, and the exact one-sided 95% lower bound on that is 0.731, which clears 70%. The gate is still not met, because the pre-registration counts groups rather than observations: over 68 (ground_station, UTC date) groups the rate is 0.471 with a lower bound of 0.366. So the gate is published as PASSED_UNGROUPED_ONLY: the observation-level pass is reported and not claimed. Open this observation, or read how every number here was generated.

Kill gate 6

NOT ESTABLISHED

Require the top review queue to find at least 1.5 times as many manually actionable conflicts as random ordering at the same budget.

0.8threshold 1.52.2

Lift 1.582, 95% interval [1.353, 1.740] over 4000 resamples grouped by pass episode and by ground station. The interval contains 1.5, so the gate is not met. It also sits entirely above 1.0, so the ranking is not nothing either.

What the measurement actually is

54 seconds, narrated: one pass measured against the curve its own orbit predicts.

A detector that assumes the trace is vertical looks in one column. The satellite is moving, so the received frequency sweeps, and the corridor is curved with its shape fixed by the pass geometry. Sliding that curve to its best match gives one number: how far off the capture was. For observation 14745984 that is 61 pixels, 5,648 Hz, 13.0 ppm. The frequency axis is cropped and exaggerated against the time axis so a 61 pixel shift on a 620 pixel image is visible at all, and the video says so on screen. Every figure spoken in it is read out of the scene that draws it, and a second model transcribed the track without seeing the script to check it was said: scripts/render_explainer_narration.py. This clip shows what the measurement is and not that it is significant, and on this observation it is not: the matched filter reaches 0.396 sigma against the curved corridor and 0.352 against a vertical line, so the fit is the best of 408 offsets on an image whose trace the corridor cannot separate from noise. It was chosen for having the strongest corridor curvature in the shipped set, which is what makes the shape legible, and a decisive example would show a larger separation and a less readable curve. The gate 3 rate on the evaluation page is measured over observations that were null-calibrated; this one was not.

What the queue found

Counts, not rates. Every number here carries the denominator it was measured over.

Conflicts found
20
in the top 50 the reviewer would reach
Random would find
12.6
expected count at the same budget, from the population rate
Queue length
407
from 410 test observations, one row per pass: one station, one satellite, one revolution, and where a station published several captures of the same pass the highest-scoring one is the row
Pass episodes
87
over 35 ground stations, the two groupings the interval is built on
The point estimate 1.582 is above the threshold and the interval is not, so this gate reads NOT_ESTABLISHED. The interval, and how it was resampled.

What counts as a conflict

Fixed before anything was measured. A criterion invented after seeing the ranking would measure the ranking against itself.

Written down in the pre-registration, before the ranking existed.
ReasonWhat it meansThreshold
Model disagrees with labelShipped arm (image_corridor) predicts with calibrated probability ≥ 0.75 that the observation is positive when the label is without-signal, or ≤ 0.25 (confident negative) when the label is with-signal.prob positive floor 0.75, prob negative ceiling 0.25
Stale catalogue frequencyFitted frequency offset magnitude ≥ 20 ppm of the catalogue downlink frequency and offset_at_bound is false. Implies the SatNOGS transmitter catalogue frequency is at least 20 ppm stale.abs offset ppm min 20
Dead capture timeflat_row_frac ≥ 0.15: at least 15% of waterfall rows carry no luminance variation, indicating substantial dead capture time.flat row frac min 0.15
  • MODEL_LABEL_DISAGREE uses all decisively-labelled observations, not only single-vetted ones. The SatNOGS waterfall_status_user field is present in the snapshot but the single-vetter count per observation is not stored; restricting to single-vetted records would require a second API call not made at snapshot time.
  • An uncorrected-Doppler outlier is not one of the criteria. The survey of correction status found only 3 uncorrected captures among the 24 it could vet, so applying the rule to the rest of the corpus would be a guess. The criterion keeps its name in the reason list and every split records it as not applicable, because saying the data do not support it is not the same as leaving it out.
  • Corridor features are unavailable for 18 observations: 4 have no waterfall image and 14 have orbital elements too stale to propagate. Those score zero on the offset and flat-row signals and rank on model disagreement alone.

Concentration caps

A queue that spends its whole budget on one station has found one station's problem, not the corpus's.

50 admitted, 4 displaced, budget 50.
CapShare of budgetEntries at budgetDisplacedBinding
ground_station10%54bound
transmitter_uuid20%100inert

Displaced nothing at budget 50 over 407 ranked observations: the per-transmitter cap.

Reported as inert on this split rather than as exercised.

A cap is credited with a displacement whenever it would have blocked the entry, even if another cap would have blocked it too, so an inert result here is a property of the data and not of the order the caps are checked in.

Without the caps the same queue would score 1.582 with interval [1.338, 1.740], which would be NOT_ESTABLISHED. The caps were fixed before measuring, so the capped queue is the one that counts. Reference only. The shipped queue is the capped queue and the verdict above is measured on it. This row exists so the price of entity-concentration control is on record instead of hidden by reporting only whichever queue scored better.

The queue

Fixed at 50 before any results were seen. The chronological test set has 87 decisively-labelled observations, so 50 is 57% coverage: large enough to measure lift with reasonable precision and small enough that 1.5x is non-trivial. On a split holding fewer than 50 decisively labelled observations, the budget is however many there are. 25 of these carry a waterfall you can open; the rest are listed with their measurements only, because shipping 2,500 images to prove a ranking is not evidence, it is weight. Where the data comes from.

Showing 60 of 407 matching rows, from 407 in the queue. Rank is the rank the pipeline assigned and does not change under a filter.

The shipped chronological queue. A row inside the review budget is one a reviewer would actually reach.
RankObservationSatelliteScoreWhy it is hereLabelModel p(signal)Offset ppm
114746092SNUGLITE-II0.8796Model disagrees with labelwithout-signal1.00015.8
214735140OBJECT BT0.8496Model disagrees with labelwith-signal0.0005.0
314732518QB50P10.5999Stale catalogue frequencywithout-signal0.68842.1
414745718SHINSEI (MS-F2)0.5962Stale catalogue frequencywithout-signal0.68831.2
514745990UND ROADS 10.5961Stale catalogue frequencywithout-signal0.65031.7
614744250GONETS M 180.5932Stale catalogue frequencywithout-signal0.68825.1
714746054TIROS 100.5846No conflict criterionwithout-signal0.650-19.2
814744187MARINA0.5837Stale catalogue frequencywith-signal0.43826.2
914733024KANYINI0.5830Stale catalogue frequencywithout-signal0.68822.7
1014744248IPERDRONE.00.5816Stale catalogue frequencywithout-signal0.500-25.3
1114742034KNACKSAT-20.5791Stale catalogue frequencywithout-signal0.50024.9
1214740027OBJECT AF0.5776Stale catalogue frequencywith-signal0.65032.0
1314740031OBJECT E0.5745Stale catalogue frequencywith-signal0.65032.0
1414741735PEARL-1B0.5723No conflict criterionwith-signal0.43818.9
1514742036OBJECT C0.5689Stale catalogue frequencywithout-signal0.43822.8
1614745664OBJECT J0.5660No conflict criterionwith-signal0.438-16.4
1714745929OBJECT H0.5643No conflict criterionwith-signal0.500-16.4
1814735884CUTE-1.7+APD 20.5567No conflict criterionwith-signal0.500-14.2
1914736746FRONTIERSAT0.5543No conflict criterionwith-signal0.43811.7
2014745985BRO-180.5513Stale catalogue frequencywithout-signal0.25036.2
2114735176CONNECTA IOT-130.5500Stale catalogue frequencywith-signal0.68822.7
2214735743OBJECT H0.5461Stale catalogue frequencywith-signal0.778-39.2
2314745984UND ROADS 20.5451No conflict criterionwith-signal0.650-13.0
2414742229MARINA0.5441No conflict criterionwith-signal0.68818.7
2514744259OTTER PUP 20.5417Stale catalogue frequencywithout-signal0.286-22.3
2614745341OTP-20.5396No conflict criterionwith-signal0.688-16.9
2714745823GONETS M 050.5395No conflict criterionwith-signal0.4389.1
2814740029ION SCV-0080.5384No conflict criterionwith-signal0.65011.0
2914745983AISAT0.5378Model disagrees with labelwithout-signal0.947-7.9
3014745342OTP-20.5352No conflict criterionwith-signal0.65010.0
3114745663OBJECT AD0.5350No conflict criterionwith-signal0.688-16.4
3214745773LAPAN-TUBSAT0.5327No conflict criterionwithout-signal0.650-2.2
3314745601ISS (ZARYA)0.5307No conflict criterionwithout-signal0.4389.5
3414740030OBJECT D0.5306No conflict criterionwith-signal0.77818.7
3514744188OBJECT AD0.5290No conflict criterionwith-signal0.778-18.6
3614737233FRONTIERSAT0.5283No conflict criterionwithout-signal0.4388.7
3714745911CUTE-10.5249No conflict criterionwith-signal0.77817.2
3814737313FRONTIERSAT0.5218No conflict criterionwith-signal0.6506.6
3914745611OBJECT AU0.5198Stale catalogue frequencywith-signal1.00031.0
4014737501FRONTIERSAT0.5188No conflict criterionwith-signal0.688-9.9
4114743373OTP-20.5187No conflict criterionwith-signal0.4381.4
4214742035CUBESAT XI 40.5180No conflict criterionwithout-signal0.438-3.6
4314744185OTP-20.5173No conflict criterionwith-signal0.653-8.5
4414746049HUBBLE 60.5143No conflict criterionwith-signal0.778-13.3
4514734009OBJECT AY0.5132Stale catalogue frequencywith-signal1.000-24.6
4614745839JAS 20.5126No conflict criterionwith-signal0.688-8.6
4714744775NETSAT-20.5113No conflict criterionwithout-signal0.4380.9
4814745821GONETS M 180.5113No conflict criterionwith-signal0.6889.1
4914746117GONETS M 170.5101No conflict criterionwith-signal0.6889.1
5014745661LEMUR 2 MJR0926200.5080No conflict criterionwith-signal0.94714.7
5114740028BUGSAT 10.5260No conflict criterionDisplaced, station capwithout-signal0.4387.7
5214733022OTP-20.5206No conflict criterionDisplaced, station capwith-signal0.650-5.2
5314733025BOTSAT-10.5130No conflict criterionDisplaced, station capwith-signal0.778-12.5
5414738773KSM1-C0.5104No conflict criterionDisplaced, station capwith-signal0.6502.2
5514745612INNOCUBE (TUBSAT-31)0.5074Stale catalogue frequencywithout-signal0.000-28.7
5614733021CUAVA-20.5054No conflict criterionwith-signal0.6501.0
5714745607OBJECT AE0.5026No conflict criterionwith-signal1.000-18.4
5814733026BRO-60.5000No conflict criterionwith-signal0.947-11.5
5914746047CONNECTA IOT-90.4998No conflict criterionwith-signal0.6885.0
6014745822GONETS-M 230.4977No conflict criterionwith-signal0.94710.5